Image processing system, image processing device, image processing method, program

JP2026143133APending Publication Date: 2026-09-08KYOCERA DOCUMENT SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025030575
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-09-08

AI Technical Summary

Benefits of technology

【0010】 本発明によれば、文字領域に含まれる文字が複数行に亘る場合であっても、各行に含まれる文字を高い精度で認識することが可能である画像処理システム、画像処理装置、画像処理方法、プログラムを提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026143133000001_ABST
    Figure 2026143133000001_ABST
Patent Text Reader

Abstract

This invention provides an image processing system, image processing device, image processing method, and program that can recognize characters in each line with high accuracy, even when the characters in a character area span multiple lines. [Solution] The image processing system (1) comprises a first processing unit (211), a second processing unit (211), and a third processing unit (211). The first processing unit (211) extracts character regions and characters contained in the image shown by the image data to be processed. The second processing unit (211) identifies the number of characters in a specific direction predetermined within the character region as the number of lines in the character region. The third processing unit (211) identifies the characters contained in each line of the number of lines in the character region identified by the second processing unit (211) within the character region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing system, an image processing apparatus, an image processing method, and a program.

Background Art

[0002] Generally, a technique is known for separating multi-line text from handwritten characters, which dilates and binarizes the handwritten characters, identifies a line with the lowest pixel frequency as a boundary, and separates multi-line text (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0004] By the way, in the case of handwritten characters freely written by a user, it is conceivable that a character string composed of the handwritten characters may be written in a state of being sloped upward to the right or downward to the right. However, when the character string is sloped, the above-described technique has a problem in that the accuracy of identifying line boundaries decreases.

[0005] An object of the present invention is to provide an image processing system, an image processing apparatus, an image processing method, and a program capable of recognizing characters included in each line with high accuracy even when characters included in a character region span a plurality of lines.

Means for Solving the Problem

[0006] An image processing system according to one aspect of the present invention comprises a first processing unit, a second processing unit, and a third processing unit. The first processing unit extracts character regions and characters contained in the image represented by the image data to be processed. The second processing unit identifies the number of characters in a predetermined specific direction within the character region as the number of lines in the character region. The third processing unit identifies the characters contained in each line of the number of lines in the character region identified by the second processing unit.

[0007] An image processing apparatus according to another aspect of the present invention comprises an image acquisition processing unit and an input processing unit. The image acquisition processing unit acquires the image data to be processed in the image processing system. The input processing unit accepts information of pre-set input items based on the extraction results output from the output processing unit of the image processing system.

[0008] Another aspect of the present invention relates to an image processing method in which one or more processors perform a first step, a second step, and a third step. The first step extracts character regions and characters contained in the image represented by the image data to be processed. The second step identifies the number of characters in the character region that are located in a predetermined specific direction as the number of lines in the character region. The third step identifies the characters contained in each line of the number of lines in the character region identified in the second step.

[0009] A program relating to another aspect of the present invention is a program that causes one or more processors to execute a first step, a second step, and a third step. The first step extracts a character region and the characters contained in the image shown by the image data to be processed. The second step identifies the number of characters in the character region that are located in a predetermined specific direction as the number of lines in the character region. The third step identifies the characters contained in each line of the number of lines in the character region identified in the second step. [Effects of the Invention]

[0010] According to the present invention, it is possible to provide an image processing system, image processing device, image processing method, and program that can recognize characters contained in each line with high accuracy, even when the characters contained in the character area span multiple lines. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a block diagram showing the configuration of an image processing system according to an embodiment of the present invention. [Figure 2] Figure 2 shows an example of candidate item information used in an image processing system according to an embodiment of the present invention. [Figure 3] Figure 3 shows an example of document-specific item information used in an image processing system according to an embodiment of the present invention. [Figure 4] Figure 4 is a flowchart illustrating an example of processing performed by an image processing system according to an embodiment of the present invention. [Figure 5] Figure 5 shows an example of an image that is subject to processing performed by the image processing system according to an embodiment of the present invention. [Figure 6] Figure 6 shows an example of first extraction result information used in an image processing system according to an embodiment of the present invention. [Figure 7] Figure 7 shows an example of recognition result information used in an image processing system according to an embodiment of the present invention. [Figure 8] Figure 8 is a flowchart illustrating an example of processing performed by an image processing system according to an embodiment of the present invention. [Figure 9] Figure 9 shows an example of an image that is subject to processing performed by the image processing system according to an embodiment of the present invention. [Figure 10] Figure 10 shows an example of second extraction result information used in the image processing system according to an embodiment of the present invention. [Figure 11]FIG. 11 is a diagram illustrating an example of second extraction result information used in the image processing system according to an embodiment of the present invention. [Figure 12] FIG. 12 is a diagram illustrating an example of second extraction result information used in the image processing system according to an embodiment of the present invention. [Figure 13] FIG. 13 is a diagram illustrating an example of second extraction result information used in the image processing system according to an embodiment of the present invention. [Figure 14] FIG. 14 is a diagram illustrating an example of recognition result information used in the image processing system according to an embodiment of the present invention. [Figure 15] FIG. 15 is a diagram illustrating an example of an image to be processed in the image processing system according to an embodiment of the present invention. [Figure 16] FIG. 16 is a flowchart for explaining an example of processing executed by the image processing system according to an embodiment of the present invention. [Figure 17] FIG. 17 is a diagram for explaining an example of a processing result executed by the image processing system according to an embodiment of the present invention. [Figure 18] FIG. 18 is a diagram for explaining an example of a processing result executed by the image processing system according to an embodiment of the present invention. [Figure 19] FIG. 19 is a flowchart for explaining an example of processing executed by the image processing system according to an embodiment of the present invention. [Figure 20] FIG. 20 is a diagram for explaining an example of a processing result executed by the image processing system according to an embodiment of the present invention. [Figure 21] FIG. 21 is a flowchart for explaining an example of processing executed by the image processing system according to an embodiment of the present invention. [Figure 22] FIG. 22 is a diagram illustrating an example of an image to be processed in the image processing system according to an embodiment of the present invention. [Figure 23] FIG. 23 is a diagram for explaining an example of a processing result executed by the image processing system according to an embodiment of the present invention. [Figure 24]Figure 24 is a diagram illustrating an example of the results of processing performed by an image processing system according to an embodiment of the present invention. [Figure 25] Figure 25 is a diagram illustrating an example of relationship information used in an image processing system according to an embodiment of the present invention. [Figure 26] Figure 26 is a diagram illustrating an example of the results of processing performed by an image processing system according to an embodiment of the present invention. [Figure 27] Figure 27 is a diagram illustrating an example of the results of processing performed by an image processing system according to an embodiment of the present invention. [Figure 28] Figure 28 is a diagram illustrating an example of the results of processing performed by an image processing system according to an embodiment of the present invention. [Figure 29] Figure 29 is a diagram illustrating an example of the results of processing performed by an image processing system according to an embodiment of the present invention. [Modes for carrying out the invention]

[0012] Embodiments of the present invention will be described below with reference to the attached drawings. The following embodiments are examples that embody the present invention and are not intended to limit the technical scope of the present invention.

[0013] [Image Processing System 1] As shown in Figure 1, the image processing system 1 according to this embodiment comprises one or more image processing devices 10, one or more servers 20, and one or more client terminals 30. The image processing devices 10, servers 20, and client terminals 30 are connected to each other via a communication network 100 such as the Internet or a LAN (Local Area Network).

[0014] In this embodiment, in the image processing system 1, image data to be processed for character recognition (hereinafter referred to as "OCR processing"), which recognizes characters contained in an image, is acquired by the image processing device 10. Then, the image data is transmitted from the image processing device 10 to the server 20. Subsequently, the server 20 performs OCR processing on the image data, and the result of the OCR processing is transmitted from the server 20 to the image processing device 10. In this embodiment, the execution entities of the various processes performed in the image processing system 1 are merely examples, and similar processing can be performed throughout the entire image processing system 1. For example, the OCR processing may be performed not only on the server 20, but also on the image processing device 10 or the client terminal 30.

[0015] [Image processing device 10] The image processing device 10 has an image reading function that reads an image of a document and generates image data. The image processing device 10 is, for example, a scanner, a facsimile machine, a multifunction printer, or a camera.

[0016] The image processing device 10 includes a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and an image acquisition unit 15, etc. The image processing device 10 may also include an image forming unit that forms an image on a sheet based on image data.

[0017] The control unit 11 comprises one or more processors and one or more memories, and executes various processes in the image processing device 10. The storage unit 12 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 12 stores image data acquired by the image acquisition unit 15, image data received from the client terminal 30, and recognition result information D12 corresponding to the results of the OCR processing performed on the server 20. The storage unit 12 also stores control programs and the like that cause the control unit 11 of the image processing device 10 to execute various processes, and the control unit 11 executes various processes according to the control programs.

[0018] The operation unit 13 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 14 is a liquid crystal display or organic EL display that displays various types of information.

[0019] The image acquisition unit 15 is controlled by the control unit 11 to perform image acquisition processing to acquire image data. Specifically, the image acquisition unit 15 has an image reading unit that reads an image from a document and outputs image data of the read image. For example, the image reading unit is an image reading unit that has an automatic document feeder (ADF), a document glass, a light source, multiple mirrors, optical lenses, and a CCD (Charge Coupled Device), etc.

[0020] In other embodiments, the image acquisition unit 15 may be a shooting unit such as a digital camera that captures images such as still images or videos and outputs image data corresponding to those images. Alternatively, the image acquisition unit 15 may be an interface for reading the image data from the shooting unit such as the digital camera. The object to be captured by the digital camera may be a display screen such as a tablet or display. Furthermore, the image acquisition unit 15 may acquire image data of an image displayed on a display screen by taking a screenshot of the display screen which includes handwritten characters using a stylus or the like via a touch panel. In addition, the image acquisition unit 15 may read image data stored in an external device such as the storage unit 12 of the image processing device 10, the server 20, or the storage unit 22 or storage unit 32 of the client terminal 30.

[0021] Furthermore, the control unit 11 can receive an OCR start operation via the operation unit 13, for example, at the start or end of image data acquisition by the image acquisition unit 15, to cause the server 20 to execute OCR processing on the image data. The control unit 11 then stores the recognition result information D12 (see Figure 7) of the OCR processing acquired from the server 20 in the storage unit 12. The control unit 11 can also display the processing result of the OCR processing on the display unit 14 based on the recognition result information D12, or output the recognition result information D12 to an external device such as a client terminal 30.

[0022] Furthermore, the control unit 11 can perform an automatic input process that automatically inputs characters corresponding to pre-set item groups from among the characters recognized by OCR processing as information corresponding to pre-set items in a predetermined database or the like. Specifically, the image processing device 10 has a document management function that manages information on one or more items contained in documents such as questionnaires, contracts, order forms, application forms, or receipts. The control unit 11 then performs an automatic input process that automatically inputs characters read from the document by OCR processing into a database stored in the storage unit 12 as information on specific items in the document. For example, if the document is a receipt, information on items such as "document name," "serial number," "recipient name," "transaction details," and "amount received" is entered into the database based on the results of the OCR processing. Note that the server 20 or client terminal 30 may also be equipped with the document management function and perform the automatic input process.

[0023] [Server 20] Server 20 is an information processing device comprising a control unit 21, a storage unit 22, an operation unit 23, a display unit 24, and an image acquisition unit 25, etc. Server 20 can be, for example, a personal computer, a smartphone, or a tablet terminal.

[0024] The control unit 21 comprises one or more processors and one or more memories, and executes various processes in the server 20. The storage unit 22 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 22 stores image data acquired by the image acquisition unit 25, image data acquired from the image processing device 10 or client terminal 30, and recognition result information D12 corresponding to the execution result of OCR processing performed on the server 20. The storage unit 22 also stores a control program that causes the control unit 21 of the server 20 to execute various processes, and the control unit 21 executes various processes according to the control program.

[0025] The operation unit 23 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 24 is a liquid crystal display or organic EL display that displays various types of information.

[0026] The image acquisition unit 25 performs processing to acquire image data that is subject to OCR processing performed by the control unit 21. Specifically, the image acquisition unit 25 is a communication unit that receives image data corresponding to images such as still images or videos from the image processing device 10 or client terminal 30. The image acquisition unit 25 may also read image data stored in an external device such as the storage unit 12 or storage unit 32 of the image processing device 10 or client terminal 30. Furthermore, the image acquisition unit 25 is a shooting unit of a digital camera or the like that captures images such as still images or videos and outputs image data corresponding to those images. The image acquisition unit 25 may also be an interface or the like that reads the image data from the shooting unit of the digital camera or the like.

[0027] [Client terminal 30] The client terminal 30 is an information processing device comprising a control unit 31, a storage unit 32, an operation unit 33, a display unit 34, and an image acquisition unit 35, etc. The client terminal 30 may be, for example, a personal computer, a smartphone, or a tablet terminal.

[0028] The control unit 31 comprises one or more processors and one or more memories, and executes various processes on the client terminal 30. The storage unit 32 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 32 stores image data acquired by the image acquisition unit 35, image data received from the image processing device 10, and recognition result information D12 corresponding to the results of OCR processing performed on the server 20. The storage unit 32 also stores a control program that causes the control unit 31 of the client terminal 30 to execute various processes, and the control unit 31 executes various processes according to the control program.

[0029] The operation unit 33 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 34 is a liquid crystal display or organic EL display that displays various types of information.

[0030] The image acquisition unit 35 is a shooting unit of a digital camera or the like that captures images such as still images or videos and outputs image data corresponding to those images. The image acquisition unit 35 may also be an interface or the like that reads the image data from the shooting unit of the digital camera or the like. The object to be captured by the digital camera or the like may be a display screen such as a tablet or display. The image acquisition unit 35 may also acquire image data of an image displayed on a display screen by taking a screenshot of the display screen which contains handwritten characters using a stylus or the like via a touch panel. Furthermore, the image acquisition unit 35 may also read image data stored in an external device such as the storage unit 12 or storage unit 22 of the image processing device 10 or server 20.

[0031] Incidentally, in the image processing system 1 configured in this way, the image data subject to OCR processing (hereinafter sometimes referred to as the "target image") may include printed or displayed typefaces and handwritten characters. In this embodiment, "typefaces" refers to characters expressed in a predetermined specific font such as Arial or Century, and excludes characters written by hand by humans or other means. For example, "typefaces" include printed characters printed on a sheet such as paper by a printer using the aforementioned specific font, and displayed characters shown on a display using the aforementioned specific font.

[0032] Furthermore, OCR processing may distinguish between handwritten characters and printed characters in a target image. In this case, if a series of characters, such as a single word or sentence, originally contains both handwritten and printed characters, the series of characters that should be recognized may be recognized separately as a handwritten character string and a printed character string. In contrast, there is a known technique that combines handwritten and printed characters into a series of characters according to the user's specified combination. However, such a configuration has the problem that specifying the handwritten and printed characters to be included in the series of characters places a burden on the user. In contrast, the image processing system 1 according to this embodiment can reduce the burden on the user in identifying handwritten and printed characters as a series of characters.

[0033] The detailed configuration of the image processing system 1 according to this embodiment will be described below.

[0034] [Storage unit 22 of server 20] The storage unit 22 of the server 20 stores candidate item information D1 and document-specific item information D2, etc. Various types of data such as candidate item information D1 and document-specific item information D2 may be set in the server 20 in response to user operations and stored in the storage unit 22, or they may be set by an external device such as the image processing device 10 or client terminal 30 and transmitted to the server 20 and stored in the storage unit 22.

[0035] <Candidate item information D1> As shown in Figure 2, the candidate item information D1 associates multiple pre-configured item groups with multiple pre-configured candidate item name strings. The candidate item information D1 is used by the control unit 21 to recognize the information of the candidate item name strings included as printed characters in the image to be processed by OCR as information corresponding to the item group. For example, the item group "Date" is associated with candidate item name strings such as "Application Date," "Entry Date," "Date," and "Date of Birth." Similarly, the item group "Name" is associated with candidate item name strings such as "Your Name," "Your Full Name," "Last Name," and "Given Name," and the item group "Address" is associated with candidate item name strings such as "Your Address," "Place," and "Address." In addition, as shown in Figure 2, the candidate item information D1 also includes item groups such as "Document Name," "Recipient Name," "Serial Number," and "Transaction Details."

[0036] <Document-specific item information D2> As shown in Figure 3, the document-specific item information D2 associates pre-set document types with the item groups. The control unit 21 uses the document-specific item information D2 to recognize the item groups contained in the target image according to the document type of the image to be processed by OCR. For example, in the example shown in Figure 3, the document type "receipt" stores item groups such as "document name," "serial number," "recipient name," "transaction details," "amount received," "date received," "document creator name," "document creator company name," and "document creator address." Similarly, in the example shown in Figure 3, the document type "order form" stores item groups such as "document name," "recipient name," "quantity ordered," "order amount," "order details," "order date," "orderer address," "orderer company name," and "orderer name." Various data such as candidate item information D1 and document-specific item information D2 may be set on the server 20 and stored in the storage unit 22, or they may be set on the image processing device 10 or client terminal 30, sent to the server 20, and stored in the storage unit 22. Furthermore, candidate item information D1 and document-specific item information D2 may be stored in the storage unit 12 or storage unit 32 of the image processing device 10 or client terminal 30.

[0037] [Control unit 21 of server 20] As shown in Figure 1, the control unit 21 of the server 20 includes an extraction processing unit 211, a specific processing unit 212, an output processing unit 213, and the like.

[0038] The extraction processing unit 211 performs OCR processing to recognize characters contained in the target image data. In particular, the extraction processing unit 211 is capable of distinguishing and recognizing handwritten characters and printed characters contained in the target image in a pre-set order. Furthermore, in the OCR processing of the target image, the extraction processing unit 211 can also perform named entity recognition processing to extract named entities such as names or place names from the characters contained in the target image.

[0039] The identification processing unit 212 performs a string identification process to identify one or more strings contained in the target image. Specifically, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit 211 and one or more pre-set item name candidate strings contained in the printed characters, the identification processing unit 212 identifies one or more handwritten characters and one or more printed characters that are consecutive in the extraction order by the extraction processing unit 211 as a series of strings. The identification processing unit 212 can also combine the one or more handwritten characters and one or more printed characters identified as a series of strings into a single string. In particular, the identification processing unit 212 may combine all of the handwritten characters and printed characters identified as a series of strings, or it may combine at least one of the handwritten characters and printed characters identified as a series of strings.

[0040] The output processing unit 213 generates recognition result information D12, D22, etc., based on the extraction results from the extraction processing unit 211 and the processing results from the identification processing unit 212, or both, and outputs the recognition result information D12, D22. The output methods of the output processing unit 213 include displaying, printing, transmitting, or storing the recognition result information D12, D22. For example, the recognition result information D12, D22 is transmitted to the image processing device 10 or the client terminal 30.

[0041] [Image processing method] Next, with reference to Figure 4, an example of the processing procedure in the image processing method executed by the image processing system 1 according to this embodiment will be described.

[0042] In the image processing system 1, the image processing method of the present invention is executed by one or more processors provided in the image processing system 1 performing various processes. Note that the various processes in the image processing method may be performed by one of the control unit 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30. Alternatively, the various processes in the image processing method may be divided and performed by two or three of the control unit 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30.

[0043] Specifically, steps S11, S12, etc. in Figure 4 represent the processing procedure (step) numbers for the processing executed by the control unit 11 of the image processing device 10 in the image processing method. Specifically, the control unit 11 of the image processing device 10 starts processing when it receives a request to execute an image reading process involving OCR processing from a user. For example, when the control unit 11 receives a request to execute the automatic input process from a user, it determines that a request to execute an image reading process involving OCR processing has been made.

[0044] Furthermore, steps S21, S22, etc. in Figure 4 represent the processing procedure (step) of the OCR processing executed by the control unit 21 of the server 20 in the image processing method. Specifically, the control unit 21 starts the OCR processing when the power to the server 20 is turned on.

[0045] <Image processing device 10: Step S11> In step S11, the control unit 11 controls the image acquisition unit 15 to perform an image acquisition process to acquire image data to be processed by OCR. Specifically, the control unit 11 controls the image acquisition unit 15 to perform the image reading process to acquire image data corresponding to the image of the document. Alternatively, in step S11, the control unit 11 may acquire image data to be processed by OCR by reading image data stored in the storage unit 12 in response to user operation.

[0046] <Image processing device 10: Step S12> In step S12, the control unit 11 sends the image data acquired in step S11 to the server 20 as the target for OCR processing.

[0047] <Server 20: Step S21> In step S21, the control unit 21 waits for the reception of image data to be processed by OCR (S21: No). Specifically, the control unit 21 receives image data to be processed by OCR from the image processing device 10 or the client terminal 30. When it is determined that the image data has been received (S21: Yes), the process proceeds to step S22.

[0048] <Server 20: Step S22> In step S22, the extraction processing unit 211 of the control unit 21 performs character extraction processing, which includes extracting and recognizing character regions and characters present in the target image indicated by the image data, based on the image data received as the target of OCR processing in step S21. In particular, the extraction processing unit 211 distinguishes and extracts handwritten characters and printed characters present in the character regions.

[0049] Specifically, the extraction processing unit 211 performs a separation process based on the image data to be processed, separating the target image represented by the image data into a handwritten image having only handwritten pixels that constitute handwritten characters and a printed image having only printed pixels that constitute printed characters or ruled lines. Conventional known techniques can be used for the separation process. For example, the separation process may use semantic segmentation using deep learning such as a deep neural network (DNN) with artificial intelligence. In the training of the deep learning, it is conceivable to generate multiple superimposed images by superimposing handwritten images and printed images and use these as training data. It is also conceivable to use handwritten binarized image data as training data such that the pixel values ​​of the binarized image data of the handwritten images corresponding to the superimposed images become label values ​​that represent handwritten pixels.

[0050] Next, the extraction processing unit 211 extracts characters such as handwritten characters and printed characters contained in the handwritten character image and the printed character image, respectively, in the target image, sequentially from a predetermined processing start point to a predetermined processing end point. Specifically, the extraction processing unit 211 generates first extraction result information D11 (see Figure 6), which associates the extracted characters with the classification of the extracted characters (handwritten characters or printed characters) and the position coordinates of the bounding rectangle of the extracted characters.

[0051] The extracted characters are a string containing one or more characters. The position coordinates are the position coordinates of the upper left corner and the lower right corner of the bounding rectangle in the target image, with the left-right direction as the X-axis and the up-down direction as the Y-axis. Furthermore, the X-coordinate of the position coordinates has a minimum value at the left edge and a maximum value at the right edge in the target image, and the Y-coordinate of the position coordinates has a minimum value at the top edge and a maximum value at the bottom edge in the target image. In this embodiment, the processing start point is the upper left corner and the processing end point is the lower right corner, and the extraction order of characters in the target image is described as proceeding with characters whose bounding rectangle's upper left corner is closer to the left edge and therefore closer to the top edge.

[0052] However, if multiple extracted characters that should be extracted in a predetermined order on the same line include handwritten characters, the position and size of the extracted characters may not be uniform. In this case, even on the same line, the upper left corners of each handwritten character may be shifted vertically relative to each other, or the upper left corners of the handwritten character and the type may be shifted vertically. As a result, multiple handwritten or typed characters that should be extracted in a predetermined order on the same line may be extracted in an order different from the predetermined order. Therefore, when the extraction processing unit 211 extracts characters sequentially from the upper left corner to the lower right corner, if the first bounding rectangle of the extracted character is detected on each line, it determines whether there are other bounding rectangles that interfere with the first bounding rectangle in the X-axis direction, at least in part. That is, the extraction processing unit 211 extracts other extracted characters whose range on the Y-axis overlaps with the bounding rectangle of the first detected extracted character at least in part. The extraction processing unit 211 then rearranges the extraction order of the multiple extraction characters corresponding to the multiple bounding rectangles whose ranges on the Y axis overlap, in ascending order of the X coordinate values ​​of the bounding rectangles, to generate the first extraction result information D1. This makes it possible to extract the extraction characters in their original extraction order even if the position and size of the extracted characters are not uniform.

[0053] Here, Figure 5 shows an example of a target image, and Figure 6 shows an example of the first extraction result information D11 extracted from the target image in Figure 5. Note that the dashed line in Figure 5 is not included in the target image, but indicates the character area within that target image.

[0054] As shown in Figure 6, the handwritten characters extracted sequentially from the top left of the target image in Figure 5 are "24", "7", "13", "Yamada Hanako", "Tokyo XX Ward", "△ Apartment No. 111", and "×× Company". Also, as shown in Figure 6, the printed characters extracted sequentially from the top left of the target image in Figure 5 are "Application Form", "Date of Entry", "20", "Year", "Month", "Day", "Name", and "Address". Below, we will explain the processing content of each step using the case where the first extraction result information D11 in Figure 6 is extracted from the target image in Figure 5 as an example.

[0055] <Server 20: Step S23> In step S23, the identification processing unit 212 of the control unit 21 extracts candidate item name strings included in candidate item information D1 (see Figure 2) from among the one or more characters extracted by the extraction processing unit 211, based on candidate item information D1 and first extraction result information D11. Specifically, in the target image shown in Figure 5, "Date of entry," "Your name," and "Your address" are extracted as candidate item name strings.

[0056] <Server 20: Step S24> In step S24, the identification processing unit 212 of the control unit 21 identifies one or more handwritten characters and one or more printed characters that are consecutively extracted by the extraction processing unit 211 as a series of strings, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit 211 and one or more candidate strings of item names extracted in step S23. The identification processing unit 212 then combines all of the identified series of strings. The identification processing unit 212 may combine at least one handwritten character and one printed character from the identified series of strings.

[0057] Specifically, the identification processing unit 212 identifies handwritten characters and printed characters that exist between two consecutive item name candidate strings based on the first extraction result information D11 as a series of strings corresponding to the item name candidate string that was extracted earlier among the two item name candidate strings. The identification processing unit 212 also identifies handwritten characters and printed characters that exist after the last item name candidate string as a series of strings corresponding to the last item name candidate string. The identification processing unit 212 identifies handwritten characters and printed characters that exist before the first item name candidate string as a series of strings, but does not associate this series of strings with the item name candidate strings. Furthermore, for the series of strings that exist before the first item name candidate string, "No Candidate" may be associated with the item name candidate string.

[0058] For example, the target image shown in Figure 5 contains the following as candidate strings for item names: "Date Entry," "Name," and "Address." The actual text corresponding to the candidate string "Date Entry" is "July 13, 2024." Similarly, the information corresponding to the candidate string "Name" is "Hanako Yamada," and the information corresponding to the candidate string "Address" is "111, △ Mansion, XX Ward, Tokyo." However, the target image shown in Figure 5 contains both handwritten and printed characters in the content of "Date Entry," which is "July 13, 2024." Therefore, in step S22, the printed "20," the handwritten "24," the printed "Year," the handwritten "7," the printed "Month," the handwritten "13," and the printed "Day" will be extracted as individual extracted characters.

[0059] In response, the identification processing unit 212 identifies the handwritten and printed characters that exist between the consecutive item name candidate strings "Entry Date" and "Name" in the first extraction result information D11 as a series of strings corresponding to the item name candidate string "Entry Date" which is extracted first. That is, "20", "24", "Year", "7", "Month", "13", and "Day" that exist between "Entry Date" and "Name" are identified as a series of strings "July 13, 2024" which corresponds to the item name candidate string "Entry Date". In particular, the identification processing unit 212 combines all the strings "20", "24", "Year", "7", "Month", "13", and "Day" into a single extracted character and associates that single extracted character with the item name candidate string "Entry Date". The identification processing unit 212 may combine at least one handwritten and printed character from the strings "20", "24", "Year", "7", "Month", "13", and "Day" into a single extracted character. For example, "20" and "24" may be combined, or "20", "24", and "year" may be combined. Also, the strings "20", "24", "year", "7", "month", "13", and "day" may be combined into "2024", "July", and "13th", respectively.

[0060] Similarly, the identification processing unit 212 identifies the handwritten and printed characters that exist between the consecutive item name candidate strings "Name" and "Address" in the first extraction result information D11 as a series of strings corresponding to the earlier item name candidate string "Name". That is, "Yamada Hanako" which exists between "Name" and "Address" is identified as the series of strings "Yamada Hanako" corresponding to the item name candidate string "Name".

[0061] Furthermore, the identification processing unit 212 identifies handwritten and printed characters that appear after "Your Address," which is the last candidate item name string in the extraction order of the first extraction result information D11, as a series of strings corresponding to "Your Address," which is the last candidate item name string in the extraction order. That is, "Tokyo XX Ward," "△ Apartment No. 111," and "×× Company," which appear after "Your Address," are identified as the series of strings corresponding to the candidate item name string "Your Address," namely "Tokyo XX Ward △ Apartment No. 111 ×× Company." Note that "×× Company" in the series of strings identified here, namely "Tokyo XX Ward △ Apartment No. 111 ×× Company," is not information corresponding to "Your Address." Therefore, in step S26 described later, "×× Company" is excluded from "Tokyo XX Ward △ Apartment No. 111 ×× Company." The specific processing unit 212 then combines the strings "Tokyo, XX Ward", "△ Apartment No. 111", and "×× Company" into a single extracted character, and associates this single extracted character with the item name candidate string "Your Address".

[0062] <Server 20: Step S25> In step S25, the identification processing unit 212 of the control unit 21 performs named entity recognition processing to extract named entities such as names (personal names), addresses (country, city), facilities, organizations, dates, times, amounts, email addresses, and telephone numbers for each of the series of strings identified in step S24.

[0063] The specific processing unit 212 may extract different named entities from multiple strings included in the series of strings. While a detailed explanation of the named entity extraction process is omitted as it is a well-known technique, this process may utilize, for example, a database, a machine learning model, or a database-type machine learning model combining these.

[0064] Specifically, as shown in Figure 7, the identification processing unit 212 determines that the named entity for the extracted text "July 13, 2024" is "Date". The identification processing unit 212 also determines that the named entity for the extracted text "Hanako Yamada" is "Name". Furthermore, as shown in Figure 8, the identification processing unit 212 determines that the named entities for the extracted text "Tokyo", "XX Ward", "△ Apartment", and "Room 111" are "Address", and that the named entity for the extracted text "XX Company" is "Organization".

[0065] <Server 20: Step S26> In step S26, the identification processing unit 212 of the control unit 21 generates recognition result information D12 (see Figure 7) indicating the result of the OCR processing based on the processing results in steps S22 to S25. In the recognition result information D12 shown in Figure 7, extracted characters representing a series of strings recognized by the OCR processing are associated with item names indicating the content of said extracted characters. The recognition result information D12 may also include the "classification" information from the first extraction result information D11. Furthermore, the recognition result information D12 may also include information on named entities extracted in the named entity recognition processing.

[0066] Specifically, the identification processing unit 212 identifies the candidate item name string corresponding to the series of strings as the item name of the series of strings if the candidate item name string corresponding to the series of strings is the same. The identification processing unit 212 then stores the extracted characters representing the series of strings in the recognition result information D12, associating them with the item name corresponding to the series of strings.

[0067] For example, as shown in Figure 7, the named entity corresponding to the series of strings "July 13, 2024" is "Date". Also, "Entry Date", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "July 13, 2024", is associated with the item group "Date" in the candidate item information D1. The named entity "Date" and the item group "Date" are the same. Therefore, as shown in Figure 7, the identification processing unit 212 stores the extracted string "July 13, 2024", identified as the series of strings, in the recognition result information D12, associating it with the item name "Entry Date", which is the candidate string for the item name corresponding to the series of strings.

[0068] Similarly, as shown in Figure 7, the named entity corresponding to the series of strings "Yamada Hanako" is "Name". Furthermore, "Name", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "Yamada Hanako", is associated with the item group "Name" in the candidate item information D1. The named entity "Name" and the item group "Name" are the same. Therefore, as shown in Figure 7, the identification processing unit 212 stores the extracted characters "Yamada Hanako", which was identified as the series of strings, in the recognition result information D12, associating them with the item name "Name", which is the candidate string for the item name corresponding to the series of strings.

[0069] Incidentally, the identification processing unit 212 excludes from the series of strings identified in step S24 any extracted characters whose named entity does not satisfy a first specific condition that has been set in advance in relation to the item group corresponding to the series of strings. Specifically, the first specific condition is that the named entity and the item group match, or that the named entity and the item group are similar.

[0070] For example, as shown in Figure 7, in the series of strings "Tokyo, XX Ward, △ Mansion No. 111, XX Company", the named entities corresponding to "Tokyo", "XX Ward", "△ Mansion", and "No. 111" are "Address", and the named entity corresponding to "XX Company" is "Organization". Furthermore, "Your Address", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "Tokyo, XX Ward, △ Mansion No. 111, XX Company", is associated with the item group "Address" in candidate item information D1. The named entity "Address" corresponding to "Tokyo", "XX Ward", "△ Mansion", and "No. 111" is the same as "Your Address" in the item group. On the other hand, the named entity "Organization" corresponding to "XX Company" is different from "Your Address" in the item group. Therefore, as shown in Figure 7, the identification processing unit 212 stores "Tokyo XX Ward △ Mansion No. 111 ×× Company", which is obtained by removing "×× Company" from the extracted characters identified as a series of strings, in the recognition result information D12, associating it with the item name candidate string "Your Address", which corresponds to the series of strings.

[0071] As shown in Figure 7, the extracted characters "Application Form" and "XX Company" extracted in step S22 are stored in the recognition result information D12 without being associated with an item name. In other embodiments, it is also conceivable that strings not associated with an item name are not stored in the recognition result information D12. Furthermore, it is also conceivable that strings not associated with the candidate item name strings are stored in the recognition result information D12 associated with an item name such as "No Item Name" to indicate that no corresponding item name exists.

[0072] <Server 20: Step S27> In step S27, the output processing unit 213 of the control unit 21 outputs recognition result information D12. Specifically, the output processing unit 213 transmits the recognition result information D12, which is the processing result of the OCR processing on the image data to be processed, to the image processing device 10, which is the source of the image data. The output processing unit 213 may also transmit the first extraction result information D11 along with the recognition result information D12 to the image processing device 10. In step S27, the output processing unit 213 may display, print, or store the first extraction result information D11 and the recognition result information D12, etc. The first extraction result information D11 and the recognition result information D12, etc. may also be transmitted to the client terminal 30.

[0073] <Image processing device 10: Step S13> In step S13, the control unit 11 waits for the recognition result information D12 to be received from the server 20 (S13: No). When it is determined that the recognition result information D12 has been received (S13: Yes), the process moves on to step S14.

[0074] <Image processing device 10: Step S14> In step S14, the control unit 11 outputs recognition result information D12. Specifically, if the request to execute the automatic input process has been executed, the control unit 11 executes an automatic input process that accepts input of information for pre-set input items in a predetermined database based on the contents of the recognition result information D12. That is, the control unit 11 when executing step S14 is an example of the reception processing unit according to the present invention. For example, in the automatic input process, as information for the input item that matches the item name in the recognition result information D12, the input of information for extracted characters associated with the item name in the recognition result information D12 is accepted. This simplifies the user's work of registering information about characters contained in the target image in the predetermined database.

[0075] In step S14, the image data that was the target of the OCR processing and the recognition result information D12 may be stored in the storage unit 12 in association with each other. In step S14, the recognition result information D12 may be combined with the image data that was the target of the OCR processing. In step S14, the image data that was the target of the OCR processing and the recognition result information D12 may be transmitted to the client terminal 30 either in association with each other or combined.

[0076] As explained above, in the image processing system 1, even if a series of characters, such as a single word or sentence, contains both handwritten and printed characters, those handwritten and printed characters are identified as a series of characters. Therefore, with the image processing system 1, the user does not need to perform an operation to specify the handwritten and printed characters that constitute a series of characters, thus reducing the burden on the user.

[0077] [Other examples of image processing methods] Incidentally, there is a known technique for identifying items that indicate the content of handwritten characters contained in an image based on printed characters present around the handwritten characters. However, if there are no printed characters around the handwritten characters, it is not possible to identify the item corresponding to those handwritten characters. On the other hand, among the handwritten characters, the one that contains one of the constituent characters such as numbers or symbols that have been set in advance to correspond to each item, and which has the fewest number of characters that are not included in those constituent characters, may be identified as the handwritten character corresponding to that item. However, with such a configuration, there is a problem that if the handwritten character contains characters other than numbers or symbols, it is not possible to identify the item corresponding to those handwritten characters. In contrast, the image processing system 1 according to this embodiment can identify items that indicate the content of handwritten characters for handwritten characters extracted from an image, as will be explained below.

[0078] The following describes other examples of the image processing method performed by the image processing system 1, with reference to the flowchart in Figure 8. In Figure 8, processing steps similar to those shown in Figure 4 are denoted by the same reference numerals and their explanations are omitted.

[0079] <Image processing device 10: Step S10> First, in the image processing apparatus 10, in step S10, the control unit 11 sets one or more item groups that indicate the item names of strings to be extracted by OCR processing from the strings included in the image data acquired in step S11. The control unit 11 when executing step S10 is an example of a setting processing unit according to the present invention.

[0080] Specifically, the control unit 11 displays a display screen for the user to select a document type stored in the document-specific item information D2, and accepts user input for selecting the document type on the display screen. The document type indicates the content of the image data acquired in step S11. Then, the control unit 11 sets one or more item groups stored in the document-specific item information D2, associated with the document type selected in response to the user input, as the item group corresponding to the image data acquired in step S11.

[0081] In other embodiments, the control unit 11 may display one or more item groups associated with the selected document type in the document-specific item information D2 on the display screen, and set one or more item groups selected by user operation from these item groups as the item group corresponding to the image data. In other embodiments, it is also conceivable that no document type is selected. For example, the control unit 11 may display multiple item groups as selection candidates on the display screen, and set one or more item groups selected by user operation from these item groups as the item group corresponding to the image data.

[0082] Furthermore, the control unit 11 may perform the processing in step S10 during or after the execution of step S11. In addition, the control unit 11 may automatically recognize the document type and item group corresponding to the image data by performing OCR processing or pattern matching processing on the image data acquired in step S11, based on information such as the document name contained in the image data.

[0083] The control unit 11 may set document-specific item information D2 according to user operation and send said document-specific item information D2 to the server 20. This allows the server 20 to perform OCR processing using the document-specific item information D2 received from the image processing device 10. The server 20 may also use the document-specific item information D2 received from the image processing device 10 only for OCR processing requested by the image processing device 10. The setting process of document-specific item information D2 by the control unit 11 is performed at any timing, such as before, during, or after the execution of step S10 or step S11.

[0084] <Image processing device 10: Step S12> When image data is acquired in step S11, in step S12, the control unit 11 associates the acquired image data with one or more item groups acquired in step S10 and transmits it to the server 20.

[0085] Furthermore, in step S12, the image data acquired in step S11 and the document type selected in step S10 may be associated with each other and sent to the server 20. As a result, if the document-specific item information D2 is stored in the storage unit 22 of the server 20, the control unit 21 can identify the item group corresponding to the image data based on the document-specific item information D2 and the document type received from the image processing device 10.

[0086] <Server 20: Step S22> As described above, when image data is received in step S21, in step S22, the extraction processing unit 211 of the control unit 21 executes processing for extracting a character region existing in the target image and characters existing in the character region. Then, the extraction processing unit 211 generates second extraction result information D21 in which extracted characters extracted from the target image, the classification of the extracted characters (handwritten characters or printed characters), and the position coordinates of the circumscribed rectangle of the extracted characters are associated with each other.

[0087] Here, FIG. 9 is a diagram showing an example of a target image, and FIG. 10 is a diagram showing an example of the second extraction result information D21 extracted from the target image of FIG. 9. It should be noted that the alternate long and short dash line in FIG. 9 is not included in the target image, but indicates a character region in the target image. Then, as shown in FIG. 10, the handwritten characters sequentially extracted from the upper left end to the lower right end of the target image in FIG. 9 are "00000000-01", "ABCD Co., Ltd.", "¥1,650-", "Product fee", "2024", "5", "24", "We hereby confirm receipt of the above amount", "¥1,500", ... "Chiyoda-ku, Tokyo ...", "2nd Floor, ABCDEF Building", "XYZ Corporation", "Hanako Yamada", and so on. Also, as shown in FIG. 10, the printed characters sequentially extracted from the upper left end to the lower right end of the target image in FIG. 9 are "Payee", "Receipt", "No.", "Dear", "Regarding", and so on. Hereinafter, the processing content of each step will be described by taking as an example a case where the second extraction result information D21 of FIG. 10 is extracted for the target image of FIG. 9.

[0088] <Server 20: Step S231> After the character region and characters are extracted in step S22, in the subsequent step S231, the identification processing unit 212 of the control unit 21 executes a first identification process for identifying an item group corresponding to each of the handwritten characters and printed characters extracted in step S22.

[0089] Specifically, as shown in Figure 11, the identification processing unit 212 identifies, for each character, the item group corresponding to the candidate string of item name that matches the character, based on the candidate item information D1 and the second extraction result information D21. For example, "Recipient" in the second extraction result information D21 is assigned the item group "Recipient Name" which corresponds to the candidate string of item name "Recipient" in the candidate item information D1. Similarly, "Receipt" in the second extraction result information D21 is assigned the item group "Document Name" which corresponds to the candidate string of item name "Receipt" in the candidate item information D1. In addition, "No." in the second extraction result information D21 is assigned the item group "Serial Number" which corresponds to the candidate string of item name "No." in the candidate item information D1, as shown in Figure 2. Furthermore, "However" in the second extraction result information D21 is assigned the item group "Transaction Details" which corresponds to the candidate string of item name "However" in the candidate item information D1, as shown in Figure 2.

[0090] Furthermore, based on the second extraction result information D21, the identification processing unit 212 identifies a type character whose positional relationship with the handwritten character satisfies a pre-set second specific condition, and identifies the item group corresponding to that type character as the item group for the handwritten character. For handwritten characters whose positional relationship with the type character does not satisfy the second specific condition based on the second extraction result information D21, the identification processing unit 212 identifies the corresponding item group in steps S232 to S233 described later.

[0091] Specifically, the second specific condition stipulates that the distance between the handwritten character and the type character closest to the handwritten character is less than or equal to a predetermined specific distance. This distance is, for example, the distance between the upper left corner of the handwritten character and the upper left corner of the type character. Alternatively, this distance may be the distance between the handwritten character and the type character at the point where they are closest to each other. Furthermore, the second specific condition may also stipulate that the type character is located in a predetermined direction relative to the handwritten character.

[0092] For example, in the examples shown in Figures 9 and 10, the typeface closest to the handwritten character "00000000-01" is "No.", so the item group "Serial Number" corresponding to "No." is assigned as the item group for the handwritten character "00000000-01". Similarly, the typeface closest to the handwritten character "Product Price" is "However", so the item group "Transaction Details" corresponding to "However" is assigned as the item group for the handwritten character "Product Price".

[0093] <Server 20: Step S232> In step S232, the identification processing unit 212 of the control unit 21 sets a determination group indicating the content of each handwritten character for which the corresponding item group was not identified in the first identification process in step S231.

[0094] Specifically, the identification processing unit 212 first performs a named entity extraction process for each handwritten character for which the corresponding item group was not identified in the first identification process in step S231, in order to extract named entities such as names, addresses, organization names, dates, times, amounts, email addresses, and telephone numbers. The identification processing unit 212 may extract different named entities from multiple characters contained in the handwritten characters. Although a detailed explanation of the named entity extraction process is omitted as it is a well-known technique, the named entity extraction process uses, for example, a database, a machine learning model, or a database-type machine learning model that combines these.

[0095] For example, in the examples shown in Figures 9 and 10, the handwritten characters for which the corresponding item group was not identified in the first identification process in step S231 are "¥1,650", "Tokyo, Chiyoda-ku...", "ABCDEF Building", "XYZ Corporation", and "Yamada Hanako". The identification processing unit 212 then determines that the named entity for one of the handwritten characters, "¥1,650", is "amount". The identification processing unit 212 also determines that for one of the handwritten characters, "Tokyo, Chiyoda-ku...", the named entities for "Tokyo", "Chiyoda-ku", and "Ichibancho" are "place names", and the named entity for "1-2-3" is "address". Furthermore, the identification processing unit 212 determines that for one of the handwritten characters, "ABCDEF Building 2nd floor", the named entity for "ABCDEF Building" is "facility name", and the named entity for "2nd floor" is "floor". Similarly, the specific processing unit 212 determines that the named entity of one of the handwritten characters, "XYZ Corporation," is "organization," and that the named entity of another handwritten character, "Hanako Yamada," is "name."

[0096] Next, the identification processing unit 212 sets the determination group for each handwritten character based on the named entity corresponding to each handwritten character. Specifically, the identification processing unit 212 sets the determination group for each handwritten character based on the named entity corresponding to each handwritten character and a pre-set determination rule.

[0097] For example, the determination rule may include a condition for identifying the determination group based solely on the named entity corresponding to the handwritten character. Alternatively, the determination rule may include a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the content of the handwritten character. Furthermore, the determination rule may include a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the item group associated with the printed characters surrounding the handwritten character. Note that the determination rule may include one or more of the above conditions. Now, a specific example of the determination rule will be described.

[0098] The determination rule includes that when the named entity of the handwritten character is "amount of money", a determination group of "amount of money" is set for the handwritten character. For example, the named entity of the handwritten character "¥1,650" is "amount of money". Therefore, as shown in FIG. 12, a determination group of "amount of money" is set for the handwritten character "¥1,650".

[0099] The determination rule includes that when the named entity of the handwritten character is "organization" and the handwritten character contains the character "company", a determination group of "company name" is set for the handwritten character. For example, the named entity of the handwritten character "XYZ Co., Ltd." is "organization", and "XYZ Co., Ltd." contains the character "company". Therefore, as shown in FIG. 12, a determination group of "company name" is set for the handwritten character "XYZ Co., Ltd.".

[0100] The determination rule includes that when the named entity of the handwritten character is "personal name", a determination group of "name" is set for the handwritten character. For example, the named entity of the handwritten character "Hanako Yamada" is "personal name". Therefore, as shown in FIG. 12, a determination group of "name" is set for the handwritten character "Hanako Yamada".

[0101] The determination rule includes that when two or more consecutive character strings (or three or more consecutive character strings) whose named entity is "place name" appear in the handwritten character, a determination group of "address" is set for the handwritten character. For example, regarding the handwritten character "Chiyoda Ward, Tokyo..., the named entities of "Tokyo", "Chiyoda Ward" and "Ichiban-cho" are all place names, and two or more consecutive character strings whose named entity is "place name" are present. Therefore, as shown in FIG. 12, a determination group of "address" is set for the handwritten character "Chiyoda Ward, Tokyo...".

[0102] The aforementioned determination rule includes the setting of the "address" determination group for handwritten characters if the named entity of the handwritten characters is "facility name" or "floor," and the item group or determination group corresponding to the closest handwritten character or printed character in any direction (up, down, left, or right) from the handwritten character is "address." For example, for the handwritten characters "ABCDEF Building 2nd Floor," the named entity of "ABCDEF Building" is "facility," and the named entity of "2nd Floor" is "floor." Furthermore, the determination group for the handwritten characters closest to "ABCDEF Building 2nd Floor," "Chiyoda-ku, Tokyo...," is "address." Therefore, as shown in Figure 12, the "address" determination group is set for the handwritten characters "ABCDEF Building 2nd Floor."

[0103] <Server 20: Step S233> In step S233, the identification processing unit 212 of the control unit 21 performs a second identification process to identify the item group corresponding to each of the handwritten characters for which the corresponding item group has not been identified in the first identification process, based on the determination group set for the handwritten character. For example, the identification processing unit 212 identifies the item group corresponding to the handwritten character based on the determination group corresponding to the handwritten character and the item group associated with the image data to be processed, using rule-based processing or the like.

[0104] Specifically, the identification processing unit 212 excludes from the identification target in the second identification process any item groups associated with the image data to be processed for which the corresponding handwritten character has already been identified in the second extraction result information D21. As a result, in the second identification process, only the remaining unspecified item groups, excluding the item groups associated with the image data to be processed for which the corresponding handwritten character has already been identified, are targeted for identification. Then, in the second identification process, the identification processing unit 212 identifies the item groups among the unspecified item groups that include the character of the judgment group corresponding to the handwritten character as the item group corresponding to the handwritten character. As a result, even if there are multiple item groups that include the character of the judgment group, in the second identification process, only the unspecified item groups are targeted for identification, increasing the likelihood that the item group can be identified based on the judgment group.

[0105] For example, if the determination group corresponding to the handwritten text "¥1,650" is "Amount," and the unspecified item group includes "Receipt Amount," then, as shown in Figure 13, "Receipt Amount," which includes the text "Amount," is identified as the item group corresponding to "¥1,650." Similarly, if the determination group corresponding to the handwritten text "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor" is "Address," and the unspecified item group includes "Document Creator Address," then, as shown in Figure 13, "Document Creator Address," which includes the text "Address," is identified as the item group corresponding to "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor," respectively. Furthermore, the identification processing unit 212 may combine "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor," which belong to the same item group, into a single extracted character.

[0106] Similarly, if the determination group corresponding to the handwritten text "XYZ Corporation" is "Company Name," and the aforementioned unspecified item group includes "Document Creator Company Name," then, as shown in Figure 13, "Document Creator Company Name," which includes the characters "Company Name," is identified as the item group corresponding to "XYZ Corporation." Also, if the determination group corresponding to the handwritten text "Hanako Yamada" is "Name," and the aforementioned unspecified item group includes "Document Creator Name," then, as shown in Figure 13, "Document Creator Name," which includes the characters "Name," is identified as the item group corresponding to "Hanako Yamada."

[0107] Furthermore, the identification processing unit 212 has described a case in which, in the rule-based processing, etc., the unspecified item group containing the character of the determination group corresponding to the handwritten character is identified as the item group corresponding to the handwritten character. On the other hand, the identification processing unit 212 may identify the item group corresponding to the handwritten character based on correspondence information that associates one or more of the determination groups with the item group.

[0108] Furthermore, in other embodiments, in the second identification process, all item groups associated with the image data to be processed may be identified. That is, the identification processing unit 212 may identify an item group that contains a character from a determination group corresponding to a handwritten character, among the item groups associated with the image data to be processed, for which the corresponding item group has not been identified in the first identification process, as the item group corresponding to the handwritten character. In this case as well, if there is only one item group containing a character from the determination group, that item group will be identified as the item group corresponding to the handwritten character.

[0109] <Server 20: Step S26> In step S26, the identification processing unit 212 of the control unit 21 generates recognition result information D22, which indicates the result of the OCR processing, based on the second extraction result information D21. Specifically, as shown in Figure 14, the recognition result information D22 associates the extracted characters extracted from the target image with the position coordinates of the extracted characters and the item group of the extracted characters. The recognition result information D22 may include either or both of the "classification" and "judgment group" information from the second extraction result information D21. The recognition result information D22 may not include "position coordinate" information. Subsequently, the recognition result information D22 is transmitted to the image processing device 10 in step S27. Then, in step S14, the control unit 11 of the image processing device 10 outputs the recognition result information D22 in the same way as the recognition result information D12. For example, if the automatic input processing execution request has been executed, the control unit 11 executes automatic input processing to accept input of information of pre-set input items in a predetermined database based on the contents of the recognition result information D22. In other words, the control unit 11 when step S14 is executed is an example of a reception processing unit according to the present invention.

[0110] As explained above, the image processing system 1 can identify the item group corresponding to handwritten characters, even if the corresponding item group could not be identified based on the surrounding printed characters, based on the named entity contained in the handwritten characters. In particular, even if the format of the document to be processed is not specified, the image processing system 1 can identify the item group corresponding to the handwritten characters based on the item group and the named entity, provided that the item groups contained in the document are identified.

[0111] [Multi-line text extraction function] Incidentally, there is a known technique for identifying the boundary between rows of handwritten characters written on multiple lines by expanding and binarizing the handwritten characters and determining the row with the lowest pixel frequency. However, in the case of handwritten characters freely written by a user, the string composed of handwritten characters may be written in a manner such as sloping upwards or downwards. However, when the string has a slope, the aforementioned technique has the problem of having low accuracy in identifying the row boundaries.

[0112] In contrast, the image processing system 1 according to this embodiment can identify the handwritten characters belonging to each line with high accuracy, even when a multi-line string of handwritten characters is tilted diagonally, as described below.

[0113] Specifically, in the character extraction process in step S22, the image processing system 1 has the specific processing unit 212 of the control unit 21 of the server 20 execute the multi-line extraction process described later. Note that the multi-line extraction process may be executed at a different timing than step S22. Furthermore, the multi-line extraction process is not limited to the control unit 21 of the server 20, but may also be executed by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc.

[0114] The following describes an example of the multi-line extraction process, referring to the flowchart in Figure 16. Furthermore, as shown in Figure 15, the explanation will use the example where the "Address" field in the target image, which is the image data subject to the multi-line extraction process, contains two lines of handwritten text: "Tokyo, XX Ward, △ Apartment, Room 111".

[0115] <Server 20: Step S221> In step S221, the extraction processing unit 211 of the control unit 21 extracts character regions present in the target image indicated by the image data, based on the image data to be processed. Note that step S221 is an example of the first step in the present invention, and the extraction processing unit 211 that executes step S221 is an example of the first processing unit according to the present invention.

[0116] Specifically, in the example shown in Figure 15, multiple character regions, including character region A1, are extracted. Note that the extraction process in step S221 can utilize conventional optical character recognition techniques, such as neural networks, so a detailed explanation is omitted here.

[0117] <Server 20: Step S222> In step S222, the extraction processing unit 211 of the control unit 21 extracts the characters present in each of the character areas extracted in step S221. The extraction processing unit 211 when executing step S222 is an example of the first processing unit according to the present invention.

[0118] Specifically, the extraction processing unit 211 detects the bounding rectangle (dotted line) of each character contained in character region A1, as shown in character region A1 of Figure 15. Note that, similar to step S221, the extraction process in step S222 can utilize conventional optical character recognition techniques, such as neural networks; therefore, a detailed explanation is omitted here.

[0119] <Server 20: Step S223> In step S223, the extraction processing unit 211 of the control unit 21 determines whether the string in each character area extracted in step S221 is written vertically or horizontally, based on the aspect ratio of each character area. This allows the image processing system 1 to automatically determine the writing direction of the characters contained in each character area, regardless of whether the writing direction is horizontal or vertical. The writing direction is perpendicular to the specific direction used to determine the number of lines of characters in the character area in step S224, described later. The extraction processing unit 211 determines the specific direction perpendicular to the writing direction. Note that the extraction processing unit 211 when executing step S223 is an example of the fourth processing unit according to the present invention.

[0120] Specifically, the control unit 21 determines that the text in the text area is written horizontally if the horizontal size w of the text area is greater than the vertical size h. On the other hand, the control unit 21 determines that the text in the text area is written vertically if the vertical size h of the text area is greater than the horizontal size w. For example, in the target image shown in Figure 15, the horizontal size w of text area A1 is greater than the vertical size h. Therefore, in step S223, it is determined that the text in text area A1 is written horizontally.

[0121] In other embodiments, the extraction processing unit 211 may accept a user operation to select whether the writing direction of the characters in the character area is horizontal or vertical, and determine the writing direction of the characters according to the user operation. For example, the extraction processing unit 211 may accept a selection operation for each of the character areas to determine whether the writing direction of the characters is horizontal or vertical individually. Alternatively, the extraction processing unit 211 may accept a selection operation to collectively select whether the writing direction of the characters in all the character areas in the target image is horizontal or vertical.

[0122] <Server 20: Step S224> In step S224, the extraction processing unit 211 of the control unit 21 performs a process to determine the number of lines of characters in each of the character areas. Specifically, the extraction processing unit 211 determines the number of lines of characters in the character area based on the writing direction of the characters in the character area (horizontal or vertical) and the number of characters in a specific direction orthogonal to the writing direction within the character area. Note that the extraction processing unit 211 when executing step S224 is an example of the second processing unit according to the present invention.

[0123] Specifically, for character regions where the characters are written horizontally, the extraction processing unit 211 counts the number of characters present in the vertical direction (an example of a specific direction) of the character region at multiple detection positions in the horizontal direction of the character region. For example, the extraction processing unit 211 counts the number of bounding rectangular areas of characters present in the vertical direction at each of the detection positions within the character region as the number of characters present in the vertical direction of the character region.

[0124] Here, as shown in Figure 16, the character area A1 includes a total of nine detection positions in the width direction. Specifically, it includes three first positions located at "2w / 10", "5w / 10", and "8w / 10" from the left edge in the width direction of the character area A1. In addition, the detection positions in the width direction of the character area A1 include a total of six second positions located at "±0.5w" from each of the first positions in the width direction of the character area A1.

[0125] In the example shown in Figure 17, the character counts for the characters present vertically in each of the first and second positions are "2, 0, 2, 0, 2, 0, 1, 0, 1" from left to right. The extraction processing unit 211 then identifies the maximum value among these counts as the number of lines of characters in the character area. In the example shown in Figure 17, since the maximum value of the count in character area A1 is "2", the extraction processing unit 211 identifies the number of horizontally written lines of characters in character area A1 as "2".

[0126] If, in step S223, it is determined that the writing direction is vertical, then the horizontal and vertical directions in the specific example for horizontal writing described here should be reversed, and therefore a detailed explanation of that case will be omitted.

[0127] <Server 20: Step S225> In step S225, if the character area identified in step S224 contains multiple lines of characters (S225:Yes), the extraction processing unit 211 proceeds to step S226. On the other hand, if the character area contains only one line of characters (S225:No), the processing from step S226 onward is not executed. The determination process in step S225 is executed sequentially for each character area until all character areas are targeted. That is, the execution of the processing in steps 226 to S228 is determined for each character area.

[0128] <Server 20: Step S226> In step S226, the extraction processing unit 211 performs a sampling process in which it samples a predetermined number of black pixels from the black pixels that constitute each character included in the character area. For example, the number of samples is four or five per character.

[0129] Specifically, Figures 18(A) and 18(B) show the sampling results when the sampling process is performed on the characters in character area A shown in Figure 15. In this way, by performing the sampling process on the characters, the influence of the shape of each character is suppressed, and the accuracy of the clustering process for each row in step S227 described later is improved. In other embodiments, the sampling process may be omitted, and the clustering process in step S227 described later may be performed on the black pixels that constitute the characters.

[0130] <Server 20: Step S227> In step S227, the extraction processing unit 211 identifies the characters belonging to each line in the character area based on the sampling results of step S226. In particular, in steps S225 to S227, the number of lines in the character area identified in step S224 is used as the number of lines of characters contained in the character area, and the characters contained in each line of that number of lines in the character area are identified. Note that the extraction processing unit 211 when executing steps S225 to S227 is an example of the third processing unit according to the present invention.

[0131] Specifically, the extraction processing unit 211 performs a clustering process that clusters the data using a cluster number k corresponding to the number of lines of characters in the character area. Specifically, as shown in Figure 15, if the number of lines of characters in character area A is "2", then the cluster number k in the clustering process is "2".

[0132] In this explanation, we will use the case where the clustering process is performed on the character region A shown in Figure 15, and the k-means method (k-means clustering), which is a type of non-hierarchical clustering, is used as an example. On the other hand, in other embodiments, the clustering process may be not limited to the k-means method, but may also use, for example, the x-means method, the k-medoids method, or spectral clustering.

[0133] First, if the number of clusters k is 2, the extraction processing unit 211 sets the centroid positions of the two clusters, the first cluster and the second cluster. Specifically, as shown in Figure 18(B), the upper end of the character area A, which is the center in the width direction of the character area A, is set as the initial value of the centroid position P11 of the first cluster (see reference). Similarly, as shown in Figure 18(B), the control unit 21 sets the lower end of the character area A, which is the center in the width direction of the character area A, as the initial value of the centroid position P12 of the second cluster.

[0134] Next, the extraction processing unit 211 performs a classification process to classify the sampled data close to the centroid position P11 into a first cluster and the sampled data close to the centroid position P12 into a second cluster. Subsequently, the extraction processing unit 211 performs an update process to update the centroid positions P11 and P12 corresponding to each of the clusters. Specifically, the extraction processing unit 211 calculates the centroid position of the sampled data belonging to the first cluster and sets the calculation result as the new centroid position P11. Similarly, the extraction processing unit 211 calculates the centroid position of the sampled data belonging to the second cluster and sets the calculation result as the new centroid position P12. The extraction processing unit 211 then repeatedly performs the classification process and the update process until the centroid positions P11 and P12 no longer change in the update process.

[0135] Subsequently, if the update process determines that the centroid positions P11 and P12 have not changed, the extraction processing unit 211 calculates the average value of the vertical (Y-axis) coordinates of the sampled data classified into each cluster. Then, the extraction processing unit 211 identifies the row number corresponding to each cluster, starting with the cluster with the smallest average value. Specifically, if the number of clusters k is "2", the unit identifies that the cluster with the smaller average value corresponds to the first row, and the cluster with the larger average value corresponds to the second row.

[0136] The extraction processing unit 211 then identifies each character corresponding to the sampled data as a string of characters in the row corresponding to the cluster to which the sampled data belongs. Specifically, in the example shown in Figure 18(B), "Tokyo XX Ward" is identified as the first row of data corresponding to the first cluster, and "△ Apartment No. 111" is identified as the second row of data corresponding to the second cluster. For example, for each character, the extraction processing unit 211 identifies it as a string of characters belonging to the first cluster if the sampled data corresponding to that character mostly belongs to the first cluster. Similarly, for each character, the extraction processing unit 211 identifies it as a string of characters belonging to the second cluster if the sampled data corresponding to that character mostly belongs to the second cluster.

[0137] It is also possible that the number of clusters k, which is the number of lines of characters included in the character area, is 3 or more. In this case, the extraction processing unit 211 divides the character area vertically into N equal parts (N = number of clusters k-1), and sets N positions, including the centers at the top and bottom of the character area and the centers of the imaginary lines that divide the character area into N equal parts, as centroid positions corresponding to the N clusters. Subsequently, the extraction processing unit 211 repeatedly executes the classification process and the update process as described above until the centroid positions no longer change. Then, the extraction processing unit 211 identifies each of the N clusters as the string of characters in the line closest to the first line in the character area, in ascending order of the average value of the vertical coordinates of the sampled data classified into that cluster.

[0138] <Server 20: Step S228> In step S227, when the characters belonging to each row in the character area are identified, in the subsequent step S228, the extraction processing unit 211 executes an image data generation process to generate image data of an image containing only the characters corresponding to each row of characters included in the character area.

[0139] Specifically, as shown in Figure 18(C), the extraction processing unit 211 converts other characters in the second line of character area A into white pixels and generates image data of an image containing only the characters in the first line. Similarly, as shown in Figure 18(D), the extraction processing unit 211 converts other characters in the first line of character area A into white pixels and generates image data of an image containing only the characters in the second line.

[0140] <Server 20: Step S229> In step S229, the extraction processing unit 211 determines whether the processing in steps S225 to S228 has been completed for all the character regions included in the target image. If it is determined that the processing in steps S225 to S228 has been completed for all the character regions included in the target image (S229: Yes), the processing proceeds to step S230. If it is determined that the processing in steps S225 to S228 has not been completed for all the character regions included in the target image (S229: No), the extraction processing unit 211 selects the next character region after the currently processed character region as the processing target and proceeds to step S225.

[0141] <Server 20: Step S230> In step S230, the extraction processing unit 211 recognizes the content of the characters contained in the target image for each character region in the target image, based on the character identification results in step S227. Specifically, the extraction processing unit 211 recognizes the content of the characters contained in each line based on the image data of each line generated in step S228. Note that, as with step S221, conventional optical character recognition technology using, for example, a neural network, can be used in step S230 as well, so its explanation is omitted here.

[0142] In particular, since the characters of each line are identified in step S227, the extraction processing unit 211 can recognize the order of the characters constituting the string in that line based on the X coordinates of the characters in that line. As a result, the recognition result information D11(D22) generated in step S26 reflects the character recognition results in the multi-line extraction process, and the recognition result information D11(D22) is output in step S27.

[0143] As explained above, the image processing system 1 can recognize characters in each line with high accuracy, even when the characters contained in each character area of ​​the target image span multiple lines. In particular, since the clustering process identifies the characters contained in each line of the character area in the image processing system 1, it is possible to recognize the characters contained in each line with high accuracy, even if there is a slant in the characters of each line contained in the character area.

[0144] In this embodiment, the example given was that the multi-line extraction process is performed for each character region included in the target image. On the other hand, the extraction processing unit 211 may perform the multi-line extraction process only for the character regions that include handwritten characters.

[0145] [Text identification function] Incidentally, when the image data to be processed contains text spanning multiple lines, general OCR processing may result in line breaks between each line of text being recognized, causing the text, which should be continuous, to be recognized in a fragmented manner. In contrast, the image processing system 1 according to this embodiment can identify that a multi-line string of text is a continuous sentence, as described below.

[0146] Specifically, in the character extraction process in step S22, the image processing system 1 has the identification processing unit 212 of the control unit 21 of the server 20 execute the text identification process described later. Note that the text identification process may be executed at a different timing than step S22. Furthermore, the text identification process is not limited to the control unit 21 of the server 20, but may also be executed by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc.

[0147] The following describes an example of the text identification process, referring to the flowchart in Figure 19. Here, we will explain using the example where the text contained in the target image A21 has two line breaks and contains three sets of text A211 to A213, as shown in Figure 20(A). The text identification process is executed when the control unit 21 of the server 20 receives image data to be processed from the image processing device 10 or the client terminal 30.

[0148] <Server 20: Step S31> In step S31, the extraction processing unit 211 of the control unit 21 extracts character regions present in the target image indicated by the image data, based on the image data to be processed. Note that the extraction process in step S31 can utilize conventional optical character recognition technologies, such as neural networks, so a detailed explanation of these technologies is omitted here.

[0149] <Server 20: Step S32> In step S32, the extraction processing unit 211 of the control unit 21 performs character extraction and recognition of the characters present in each of the character regions extracted in step S31. Note that, as with step S31, the extraction process in step S32 can utilize conventional optical character recognition technologies, such as neural networks; therefore, a detailed explanation is omitted here.

[0150] Specifically, as shown in Figure 20(B), the text area A22 extracted from the target image A21 (see Figure 20(A)) contains six line breaks in the entire text and includes seven sequences of sentences A221 to A227.

[0151] Furthermore, in step S32, the extraction processing unit 211 can distinguish and extract handwritten characters and printed characters present in the character area, similar to step S22 (see Figure 4). In this case, the extraction processing unit 211 may perform the processes in steps S33 to S38 described later only if handwritten characters are present in the character area.

[0152] <Server 20: Step S33> In step S33, the identification processing unit 212 of the control unit 21 determines whether the characters contained in the character area are in sentence form. If the characters contained in the character area are in sentence form (S33:Yes), the identification processing unit 212 proceeds to step S34; if the characters contained in the character area are not in sentence form (S33:No), the processing proceeds to step S36. The processing from step S33 onward is performed for each of the character areas extracted in step S31, and is performed sequentially for each character area until all of the character areas are targeted (S39:No).

[0153] Specifically, in step S33, the specific processing unit 212 determines whether the characters in the character area constitute a sentence based on the content of the characters contained within that character area. For example, the specific processing unit 212 decomposes the characters in the character area into parts of speech and determines whether nouns and verbs are included in the character area at least a predetermined number of times. The predetermined number of times is a pre-set number, such as 1 to 3 times. The specific processing unit 212 determines that the characters in the character area constitute a sentence if nouns and verbs are included at least a predetermined number of times. On the other hand, the specific processing unit 212 determines that the characters in the character area do not constitute a sentence if the number of nouns or verbs in the character area is less than the predetermined number of times. The method for determining whether the characters in the character area constitute a sentence is not limited to this, and various conventionally known techniques may be used.

[0154] Furthermore, in step S33, the specific processing unit 212 may determine whether the text-formatted characters contained in the character area span multiple lines. In this case, the specific processing unit 212 proceeds to step S34 if the characters contained in the character area are text and the text spans multiple lines, and proceeds to step S36 if the characters do not span multiple lines. For example, the specific processing unit 212 can determine whether the characters span multiple lines by performing the same processing as in steps S223 to S224 (see Figure 16) in the multi-line extraction process described above.

[0155] <Server 20: Step S34> In step S34, the identification processing unit 212 of the control unit 21 determines whether there is a relationship between two consecutive lines of text extracted by the extraction processing unit 211, and if it is determined that there is a relationship, it performs an identification process to identify the two lines of text as a series of texts.

[0156] Specifically, in step S34, the specific processing unit 212 determines whether there is a contextual connection between any two consecutive lines of text contained in the character area. For example, if there are three lines of text contained in the character area, the control unit 21 determines whether there is a contextual connection between the first and second lines of text contained in the character area, and also determines whether there is a contextual connection between the second and third lines of text contained in the character area. In other embodiments, the determination of whether there is a contextual connection between any two consecutive lines of text may be made based on more lines of text than the two lines of text contained in the character area.

[0157] More specifically, the specific processing unit 212 determines whether or not there is a relationship based on the meaning of the sentence when the two lines of text are joined together. For example, the specific processing unit 212 may determine whether or not there is a contextual connection between the two lines of text depending on whether or not the two lines of text become a meaningful sentence when joined together. Alternatively, the specific processing unit 212 may compare the case where the two lines of text are joined with the case where the two lines of text are not joined and determine whether or not there is a contextual connection between the two lines of text depending on which becomes a meaningful sentence.

[0158] For example, in step S34, the specific processing unit 212 uses BERT (Bidirectional Encoder Representations from Transformers) NSP (Next Sentence Prediction) to determine whether two consecutive lines of text are connected in context. In particular, it is desirable to use a dataset that includes both contextually connected positive examples and contextually unconnected negative examples for pre-training of the NSP. This makes the NSP specialized in determining whether two consecutive lines of text are connected in context. Therefore, the accuracy of determining whether two consecutive lines of text are connected in context is improved compared to predicting whether the second line of two consecutive lines of text is logically or semantically connected to the first line using a general NSP. Note that other methods may be used to determine whether two consecutive lines of text are connected in context, not just the NSP.

[0159] Then, if the identification processing unit 212 determines that there is a contextual connection between two consecutive lines of text, it identifies those two consecutive lines of text as a single sentence and combines them. For example, if it is determined in step S34 that the strings in the first and second lines have a contextual connection, the identification processing unit 212 identifies the strings in the first and second lines as a single sentence and combines the strings in the first and second lines as a single sentence. Also, if it is determined in step S34 that the strings in the second and third lines have a contextual connection, the control unit 212 identifies the strings in the second and third lines as a single sentence and combines the strings in the second and third lines as a single sentence. On the other hand, if it is determined in step S34 that there is no contextual connection between the strings in the third and fourth lines, the strings in the third and fourth lines are not combined.

[0160] For example, in Figure 20(B), the first line of sentence A221, "I found this product very interesting," ends with the word "interesting," while the second line of sentence A222, "I made a discovery," begins with the word "discovery." Therefore, if a method were to combine sentences A221 and A222 by determining whether or not words are split before and after a line break, sentences A221 and A222, which do not have words split before and after a line break, would not be combined. In contrast, the identification processing unit 212 determines whether or not there is a contextual connection between two consecutive lines of sentences. Therefore, even though sentences A221 and A222 in the first and second lines do not have words split before and after a line break, they are determined to have a contextual connection and are identified and combined as a series of sentences, as shown in Figure 20(C). Similarly, in the text area A22 shown in Figure 20(B), the sentence A232, which is formed by combining the sentences from lines 3 to 6, is identified and combined as a series of sentences, and the sentence A233, which is the sentence from line 7, is identified as a series of sentences.

[0161] <Server 20: Step S35> In step S35, the control unit 21 performs punctuation correction processing to correct the text within the character area identified in step S34, such as adding or deleting punctuation marks. For example, a pre-trained model using BERT (Bidirectional Encoder Representations from Transformers) MLM (Masked Language Modeling) may be used for the punctuation correction processing.

[0162] For example, in step S35, a comma is added to the phrase "I like the tap function" in sentence A232 of text area A22 in Figure 20(C). As a result, as shown in Figure 20(D), sentence A232 is corrected to the same as target image A21, "I like the tap function. Function". Also, a comma is added to the phrase "It is also well-equipped and very useful in daily life" in sentence A232 of text area A22 in Figure 20(C). As a result, as shown in Figure 20(D), sentence A232 is corrected to the same as target image A21, "It is also well-equipped and very useful in daily life".

[0163] <Server 20: Step S36> In step S36, the specific processing unit 212 performs morphological analysis on the characters in the character area that were determined not to be in sentence form in step S33, and subdivides the characters into morphological strings of the smallest units (morphemes) that have meaning in language.

[0164] <Server 20: Step S37> In step S37, the identification processing unit 212 extracts a named entity from each of the morphological strings obtained in the morphological analysis in step S36, and associates the named entity label corresponding to the named entity with the morphological string. For example, from a morphological string representing a prefecture, the named entity "country, city" is extracted, and the named entity label "GPE" corresponding to the named entity "country, city" is associated with the morphological string. Similarly, from a morphological string representing a name, the named entity "name" is extracted, and the named entity label "PERSON" corresponding to the named entity "name" is associated with the morphological string.

[0165] <Server 20: Step S38> In step S38, the identification processing unit 212 identifies and combines a series of morphological strings, where the named entity labels extracted in step S37 satisfy the pre-set joining conditions, into a single string.

[0166] For example, in the aforementioned join condition, it is conceivable that the named entity labels "GPE" corresponding to the morphological string indicating a prefecture, "LOC" corresponding to the morphological string indicating a city or town name, and "FAC" corresponding to the morphological string indicating a building would be set as the objects to be joined. In this case, if the string on the first line is "Tokyo XX Ward" and the string on the second line is "△ Apartment No. 111", the named entity labels of the strings on the first and second lines will be "GPE" and "FAC", and the join condition will be satisfied. Therefore, the strings on the first and second lines will be identified as a series of strings and joined as the string "Tokyo XX Ward △ Apartment No. 111". On the other hand, if the string on the third line is "Yamada Hanako", the named entity labels of the strings on the second and third lines will be "FAC" and "PERSON", and the join condition will not be satisfied. Therefore, the strings on the second and third lines will not be identified as a series of strings and will not be joined.

[0167] In this embodiment, the case in which steps S34 and S35 to S38 are performed has been described. On the other hand, in other embodiments, steps S34 or S35 to S38 may be omitted.

[0168] <Server 20: Step S39> In step S39, the extraction processing unit 211 determines whether the processing from step S33 onwards has been completed for all the character regions included in the target image. If it is determined that the processing from step S33 onwards has been completed for all the character regions included in the target image (S39: Yes), the processing proceeds to step S40. If it is determined that the processing from step S33 onwards has not been completed for all the character regions included in the target image (S39: No), the extraction processing unit 211 selects the next character region after the currently processed character region as the processing target and proceeds to step S33.

[0169] <Server 20: Step S40> In step S40, the output processing unit 213 of the control unit 21 performs output processing to output each of the sentences in the character area after the execution of steps S34 to S35, or each of the strings after the execution of steps S36 to S38, as recognition result information indicating the result of the sentence identification process. The output methods of the recognition result information include display, printing, transmission, or storage. For example, if the image data that was the subject of the recognition result information was transmitted from the image processing device 10 or the client terminal 30, the recognition result information is transmitted to the image processing device 10 or the client terminal 30, which is the source of the image data. In the image processing device 10, the control unit 11 performs processing to output the recognition result information. Specifically, the control unit 11 performs automatic input processing to accept input of information of pre-set input items in a predetermined database based on the content of the recognition result information. That is, the control unit 11 when performing such processing is an example of the reception processing unit according to the present invention. For example, in the automatic input processing, the sentence shown in the recognition result information is input as the information of the input item. This simplifies the user's task of registering information about the characters contained in the target image in the predetermined database. The control unit 11 may also store the image data that was the target of OCR processing and the recognition result information in the storage unit 12 in association with each other. The control unit 11 may also combine the recognition result information with the image data that was the target of OCR processing. The control unit 11 may also transmit the image data that was the target of OCR processing and the recognition result information, either in association with each other or combined, to the client terminal 30.

[0170] As explained above, the image processing system 1 can identify that a multi-line string of characters is actually a continuous sentence, and can obtain useful recognition result information.

[0171] [Summary output function] Incidentally, images of various documents may contain not only printed text but also diagrams such as tables or graphs. Furthermore, handwritten annotations, such as handwritten annotation characters or handwritten annotation figures, may be added to these diagrams. In response to this, the image processing system 1 according to this embodiment can generate summaries that take into account the content of the diagrams, as will be explained below.

[0172] Specifically, in the image processing system 1, the control unit 21 of the server 20 performs the summary output processing described later. For example, the extraction processing unit 211 extracts handwritten annotations, which include either or both handwritten annotation characters and handwritten annotation figures, and the components of the charts and graphs included in the target image, as indicated by the image data to be processed. The output processing unit 213 generates and outputs summary information D32 based on the handwritten annotations extracted by the extraction processing unit 211 and the target components among the components that are the subject of the handwritten annotations. Furthermore, the output processing unit 213 generates and outputs summary information D32 based on the handwritten annotations, the target components, and related components among the components that are associated with the target components.

[0173] Furthermore, the summary output processing may be performed not only by the control unit 21 of the server 20, but also by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc. Also, the summary output processing may be shared and performed by two or three of the control units 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30.

[0174] Hereinafter, with reference to Figure 21, an example of the processing procedure in the summary output method, which is executed as one of the image processing methods in the image processing system 1 according to this embodiment, will be described.

[0175] In Figure 21, steps S111, S112, etc., represent the processing procedure (step) numbers of the processing performed by the control unit 11 of the image processing device 10 in the summary output method. Specifically, the control unit 11 of the image processing device 10 starts processing when it receives a request from the user to execute the image reading process that includes the summary output process.

[0176] Furthermore, steps S41, S42, etc. in Figure 21 represent the processing procedure (step) of the summary output process executed by the control unit 21 of the server 20 in the summary output method. Specifically, the control unit 21 starts the summary output process when the power to the server 20 is turned on.

[0177] <Image processing device 10: Step S111> In step S111, the control unit 11 of the image processing device 10 executes an image acquisition process to acquire image data to be processed by the summary output process by controlling the image acquisition unit 15. Specifically, the control unit 11 acquires image data corresponding to the image of the original document by controlling the image acquisition unit 15 to execute the image reading process. Alternatively, in step S111, the control unit 11 may acquire image data to be processed by the summary output process by reading image data stored in the storage unit 12 in response to user operation.

[0178] <Image processing device 10: Step S112> In step S112, the control unit 11 of the image processing device 10 transmits the image data acquired in step S111 to the server 20 as the target of the summary output process.

[0179] <Server 20: Step S41> In step S41, the control unit 21 of the server device 20 waits for the reception of image data to be processed by the summary output process (S41: No). Specifically, the control unit 21 receives the image data to be processed by the summary output process from the image processing device 10 or the client terminal 30. When it is determined that the image data has been received (S41: Yes), the process proceeds to step S42.

[0180] The following describes the case where the image data corresponding to the target image P10 shown in Figure 22 is acquired as the target of the summary output process. Specifically, the target image P10 shown in Figure 22 includes a string of characters written in type and a table, which is an example of a diagram. The components of the table shown in Figure 22 include row names such as "notebook," "ballpoint pen," "eraser," and "pencil," and column names such as "product name," "sales price [yen]," "quantity [units]," "cost [yen]," "sales [yen]," "sales area," and "gross profit [yen]," as well as non-elemental data. The components of the table shown in Figure 22 also include elemental data entered into each cell of the table. Then, in the summary output process, the output processing unit 213 generates summary information D32 that includes the elemental data and the non-elemental data.

[0181] Furthermore, the target image P10 includes a handwritten annotation figure C1 that annotates the text "Expand sales area" from the printed text, and a handwritten annotation character C2 that indicates the content of the annotation by the handwritten annotation figure C1, which reads "Southeast Asia, etc."

[0182] Furthermore, the target image P10 includes handwritten annotation shapes C11 and C12, which are used to annotate the row name "Pencil," which is non-elemental data in the figure and table, and handwritten annotation text C13, which is "New Product," indicating the content of the annotations made by the handwritten annotation shapes C11 and C12. In addition, the target image P10 includes a handwritten annotation shape C21, which is used to annotate the column name "Cost [yen]," which is non-elemental data in the figure and table.

[0183] Similarly, the target image P10 includes a handwritten annotation figure C31 that annotates "150," which is the element data corresponding to the combination of "eraser" and "gross profit [yen]" among the element data of the chart, and handwritten annotation text C32, which indicates the content of the annotation by the handwritten annotation figure C31, and "minimum." In addition, the target image P10 also includes a handwritten annotation figure C41 that annotates "1," which is the element data corresponding to the combination of "ballpoint pen" and "quantity [yen]" among the element data of the chart, and handwritten annotation text C42, which indicates the content of the annotation by the handwritten annotation figure C41, and "not very popular."

[0184] Furthermore, the target image P10 includes a handwritten annotation figure C51 that annotates "800," which is the element data corresponding to the combination of "pencil" and "profit [yen]" among the element data of the chart. In addition, the target image P10 includes handwritten annotation figures C43 and C44 that annotate "Netherlands," which is part of the element data corresponding to the combination of "ballpoint pen" and "sales area" among the element data of the chart, and handwritten annotation text C45, which indicates the content of the annotations made by handwritten annotation figures C43 and C44, and "economic boom."

[0185] Furthermore, the target image P10 includes handwritten annotation shapes C61, which are annotation targets for the column name "Sales [yen]", which is non-element data of the chart, and the element data "450", "1,000", "250", and "1,400", which are the element data corresponding to each combination of "notebook", "ballpoint pen", "eraser", and "pencil" and "Sales [yen]" from the element data of the chart.

[0186] <Server 20: Step S42> In step S42, the extraction processing unit 211 of the control unit 21 performs a first separation process based on the image data to be processed, separating the target image represented by the image data into a character image and a chart image. The chart image is an image of the target image that includes at least a chart such as a table or graph and the printed characters within the chart. The character image is an image of the target image that includes at least a string of characters other than the printed characters in the chart. The character image may also include lines and the like. Conventional techniques may be used for the first separation process, but for example, a trained model of a neural network that detects the shape of a chart may be used.

[0187] In particular, if the target image contains handwritten annotations such as handwritten annotation characters or handwritten annotation figures that are the target of the chart, the extraction processing unit 211 separates the image including the handwritten annotations as the chart image. The handwritten annotation figures include arrows, leader lines, borders, or underlines. Similarly, if the target image contains handwritten annotations that are the target of the chart, excluding the typefaces in the chart, the extraction processing unit 211 separates the image including the handwritten annotations as the character image.

[0188] For example, Figure 23(A) shows a character image P20, which is an example of a character image separated from the target image P10 shown in Figure 22 by the first separation process. Also, Figure 23(B) shows a chart image P30, which is an example of a chart image separated from the target image P10 shown in Figure 22 by the first separation process.

[0189] <Server 20: Step S43> In step S43, the extraction processing unit 211 of the control unit 21 performs a second separation process, similar to step S22, to separate the chart image into a handwritten image containing only the handwritten annotations included in the chart image and a typeset image containing only the chart and the typeset text within the chart.

[0190] For example, Figure 24(A) shows a handwritten image P31, which is an example of the handwritten image separated by the second separation process from the chart image P30 shown in Figure 23(B). Also, Figure 24(B) shows a typeface image P32, which is an example of the typeface image separated by the second separation process from the chart image P30 shown in Figure 23(B).

[0191] <Server 20: Step S44> In step S44, the extraction processing unit 211 of the control unit 21 extracts the components of the chart and the typefaces included in the typeface image separated in step S43.

[0192] Specifically, the extraction processing unit 211 extracts the components of a figure or table from the printed image using a trained neural network model, such as a graph neural network. For example, if the printed image contains a table as a figure or table, the extraction processing unit 211 extracts non-elemental data including at least one of the table title, column names, or row names of the table, and elemental data indicating the elements of the table. Also, if the printed image contains a graph as a figure or table, the extraction processing unit 211 extracts non-elemental data including at least one of the graph title, legend, vertical (value) axis, horizontal (item) axis, vertical axis label, horizontal axis label, vertical axis tick, horizontal axis tick, vertical axis tick label, horizontal axis tick label, or data label of the graph, and elemental data indicating the elements of the graph.

[0193] Furthermore, the extraction processing unit 211 performs OCR processing to recognize the typefaces included in the components of the diagram. The extraction processing unit 211 then associates the recognized components of the diagram with the typefaces included in those components and stores them in the memory of the control unit 21 or the storage unit 22. For example, if the diagram is a table, the components such as the table title, column names, row names, and element data are associated with the recognition results of the typefaces that represent the content of those components.

[0194] <Server 20: Step S45> In step S45, the output processing unit 213 of the control unit 21 associates the handwritten annotations included in the chart image separated in step S42 with the target components, which are type characters in the chart extracted in step S44, with related components associated with those type characters. Here, the target components are the components extracted in step S44 that are the subject of the handwritten annotations. The related components are the components extracted in step S44 that are related to the target components. For example, if the target components are element data, the related components are non-element data such as row names or column names corresponding to the target components. Also, if the target components are non-element data such as row names or column names, the related components are element data corresponding to the target components.

[0195] Specifically, the output processing unit 213 identifies which typeface in the diagram the handwritten annotation refers to, and stores the typeface as the target component, associating it with the handwritten annotation, in the memory or storage unit 22 of the control unit 21. The output processing unit 213 also stores the typeface identified as the target component, along with non-elemental data such as row names or column names related to the typeface, as related components, associating it with the typeface, in the memory or storage unit 22 of the control unit 21.

[0196] For example, in the example shown in Figure 22, the handwritten annotation shape C11 with a surrounding line and the handwritten annotation shape C12 with a leader line are associated with the handwritten annotation character C13, "New Product," which is located at the end of the handwritten annotation shape C12, and the element data "Pencil," which is surrounded by the handwritten annotation shape C11. In this case, "Pencil" is an example of a target component. Also, the handwritten annotation shape C21 with an underline is associated with the column name "Cost [yen]," to which the handwritten annotation shape C21 is attached. In this case, "Pencil" is an example of a target component, and "Cost [yen]" is an example of a related component. Furthermore, in the example shown in Figure 22, the handwritten annotation shape C31 with an arrow is associated with the handwritten annotation character C32, "Minimum," which is located at the beginning of the handwritten annotation shape C31, and the element data "150," which is one of the components in the chart pointed to by the end of the handwritten annotation shape C31, and the column name "Gross Profit [yen]," which corresponds to "150." In this case, "150" is an example of the target component, and "Gross Profit [yen]" is an example of the related component.

[0197] Similarly, in the example shown in Figure 22, the handwritten annotation shape C41 of the arrow is associated with the handwritten annotation text C42, which is "not very popular," the element data "1," which is one of the components in the chart pointed to by the end of the handwritten annotation shape C41 of the arrow, and the related "quantity [pieces]." In this case, "1" is an example of the target component, and "quantity [pieces]" is an example of the related component. Also, in the example shown in Figure 22, the handwritten annotation shape C43 of the surrounding line and the handwritten annotation shape C44 of the leader line are associated with the handwritten annotation text C45, which is "booming economy," the element data "Netherlands" surrounded by the handwritten annotation shape C43, and the column name "sales area" related to "Netherlands." In this case, "Netherlands" is an example of the target component, and "sales area" is an example of the related component. Furthermore, the underlined handwritten annotation figure C51 is associated with the element data "800" to which the handwritten annotation figure C51 is attached, and the column name "Gross Profit [yen]" which is related to "800". In this case, "800" is an example of the target component, and "Gross Profit [yen]" is an example of the related component. Also, in the example shown in Figure 22, the bounding handwritten annotation figure C61 is associated with the column name "Cost [yen]" and the element data "450", "1,000", "250", and "1,400" which are surrounded by the handwritten annotation figure C61, and the row names "Notebook", "Ballpoint pen", "Eraser", and "Pencil" which are related to "450", "1,000", "250", and "1,400". In this case, "450", "1,000", "250", and "1,400" are examples of the target components, and "Notebook", "Ballpoint pen", "Eraser", and "Pencil" are examples of the related components.

[0198] <Server 20: Step S46> In step S46, the output processing unit 213 of the control unit 21 estimates the relationship between the handwritten annotation associated in step S45 and the typefaces in the diagram, and stores the estimated result of this relationship in the memory or storage unit 22 of the control unit 21.

[0199] Specifically, as shown in Figure 25, the memory unit 22 stores relationship information D31 in which a corresponding relationship R1 to R5 is set for each combination of the type of handwritten annotation and the type of type in the chart associated with the handwritten annotation. The output processing unit 213 then identifies which of the relationship R1 to R5 the relationship between the handwritten annotation and the type in the chart associated with the handwritten annotation is based on the relationship information D31.

[0200] As shown in Figure 25, in relational information D31, the type of type in the figure or table associated with the handwritten annotation includes non-elemental data and elemental data of the figure or table. In addition, in relational information D31, the type of handwritten annotation includes annotations with text, annotations without text, annotations with text that include statistical indicators, and annotations with text that do not include statistical indicators. For example, in relational information D31, if the type of type in the figure or table to which the handwritten annotation is annotated is elemental data of the figure or table, and the type of handwritten annotation is an annotation with text that includes statistical indicators, then relational R3 is associated.

[0201] <Server 20: Step S47> In step S47, the output processing unit 213 of the control unit 21 generates summary information D32 that shows a summary of the target image based on the extraction results in step S44 and the estimation results in step S46.

[0202] The output processing unit 213 uses the extraction results from step S44 and the estimation results from step S46 as input information to generate summary information D32 containing content corresponding to relationships R1 to R5, using a conditional text generation model such as "Text-to-Text Transfer Transformer (T5)". According to the conditional model generation model, the intended text information is generated by providing arbitrary contextual information and conditions for what kind of text data to generate. In particular, the conditional model generation model has been trained to use non-elemental data from the extraction results in step S44 to generate summary information D32 as needed.

[0203] In particular, when generating summary information D32, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the chart image P30. For example, as shown in Figure 22, if the character image P20 (see Figure 23) contains typefaces and handwritten annotations of handwritten annotation figures C1 and handwritten annotation characters C2 that are the target of the annotations, the output processing unit 213 generates summary information D32 that supplements the typefaces based on the handwritten annotations. For example, in the example shown in Figure 22, based on the handwritten annotation figures C1 and handwritten annotation characters C2, the handwritten annotation character C2 "Southeast Asia, etc." is added to the typeface "Expand sales area" contained in the character image P20, and summary information D32 containing the content "Expand sales area to Southeast Asia, etc." is generated.

[0204] In other embodiments, the storage unit 22 may store generation rules that, in association with each of the relationships R1 to R5 that are the estimation results in step S46, generate summary information D32 of the content described below based on the handwritten annotations and the typefaces in the charts associated with the handwritten annotations. In this case, the output processing unit 213 generates the summary information D32 according to the generation rules, using non-elemental data from the extraction results in step S44 as needed.

[0205] The following describes an example of summary information D32, which includes a summary generated based on relationships R1 to R5, figures and tables, and handwritten annotations on those figures and tables.

[0206] Figure 26(A) shows an example of summary information D32 generated by the conventional technology in which the summary output processing is not performed on the target image P10 shown in Figure 22. Figure 26(B) shows an example of summary information D32 generated when the summary output processing is performed on the target image P10 shown in Figure 22. As shown in Figure 26(A), in the conventional technology, the content of handwritten annotations corresponding to the figures and tables included in the target image P10 was not reflected in the summary information D32. In contrast, the summary information D32 shown in Figure 26(B) reflects the content of handwritten annotations corresponding to the figures and tables included in the target image P10.

[0207] Relationship R1 indicates that the handwritten annotation supplements the typeset text in the figure or table associated with the handwritten annotation. The conditional text generation model has learned that for relationship R1, it generates text that supplements the typeset text in the figure or table associated with the handwritten annotation with the handwritten annotation text included in the handwritten annotation. For example, in the example shown in Figure 22, the type of handwritten annotation, which includes handwritten annotation shapes C11 and C12 and the handwritten annotation text C13, "New Product," is an annotation with text, and the typeset text in the figure or table associated with the handwritten annotation shapes C11 and C12 is the non-element data "Pencil." In this case, the output processing unit 213 generates summary information D32 that includes the text "Pencil is a new product," as an example, as shown in Figure 26(B), where "Pencil is a new product" is used to supplement "Pencil" with the handwritten annotation text C13, "New Product."

[0208] Relationship R2 indicates that the handwritten annotation emphasizes the typeface in the figure or table associated with the handwritten annotation. The conditional text generation model has been trained to generate text that emphasizes the content of the typeface in the figure or table associated with the handwritten annotation for relationship R2. For example, in the example shown in Figure 22, the handwritten annotation figure C21 is an annotation without text, and the typeface in the figure or table associated with the handwritten annotation figure C21 is the non-element data "Cost [yen]". In this case, the generation of summary information D32 uses the content of the element data corresponding to the typeface "Cost [yen]" in the figure or table associated with the handwritten annotation figure C21, and the row name of the non-element data "Cost [yen]" corresponding to each of the element data. For example, the output processing unit 213 generates summary information D32, which includes the sentence "The cost [yen] is 30 for the notebook, 400 for the ballpoint pen, 20 for the eraser, and 30 for the pencil," as shown in Figure 26(B), as a sentence that emphasizes the row name "Cost [yen]" in the chart associated with the handwritten annotation figure C21.

[0209] Relationship R3 indicates that the handwritten annotation corresponds to a statistical indicator in the chart associated with the handwritten annotation. For example, the statistical indicator may include maximum, minimum, mean, mode, or median. The conditional text generation model has learned that for relationship R3, it generates text that explains the content of the text in the chart associated with the handwritten annotation using the statistical indicator indicated by the handwritten annotation and the non-elemental data of the chart. For example, in the example shown in Figure 22, the type of handwritten annotation that includes the statistical indicator "minimum" in the handwritten annotation shape C31 and the handwritten annotation character C32 is an annotation with text that includes a statistical indicator, and the text in the chart associated with the handwritten annotation shape C31 is the elemental data "150". In this case, the generation of summary information D32 uses the text "150" in the chart associated with the handwritten annotation shape C31, the statistical indicator "minimum" in the handwritten annotation character C32, and the column name "Gross Profit [yen]", which is the non-elemental data corresponding to the text "150". For example, the output processing unit 213 generates summary information D32, which includes the sentence "The minimum value of the gross profit [yen] is 150," as the summary content corresponding to the handwritten annotation character C31 and the printed character "150" in the figure, as shown in Figure 26(B).

[0210] Relationship R4 indicates that the handwritten annotation supplements the typeset text in the chart associated with the handwritten annotation. The conditional text generation model has been trained to generate text that supplements the content of the typeset text in the chart associated with the handwritten annotation based on the handwritten annotation characters and the non-elemental data corresponding to the typeset text in the chart, for relationship R4. For example, handwritten annotation figure C41 and handwritten annotation character C42 are annotations with text that do not contain statistical indicators, the elemental data in the chart associated with handwritten annotation figure C41 is "1", and the non-elemental data corresponding to this elemental data is the row name "ballpoint pen" and the column name "cost [yen]". In this case, the generation of summary information D32 uses the typeset "1" in the chart associated with handwritten annotation figure C41 and the column name "quantity [pieces]" and row name "ballpoint pen" of the non-elemental data corresponding to the typeset "1". Also, in the example shown in Figure 22, the character image P20 contains the typeset text "Meanwhile, one ballpoint pen was sold." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "The number of ballpoint pens [pieces] sold was 1, and they were not very popular," as shown in Figure 26(B), based on the type information contained in the character image P20 and the handwritten annotation figure C41 and handwritten annotation characters C42. That is, if both the character image P20 and the diagram image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the diagram image P30. In other embodiments, the output processing unit 213 may not combine the summary of the character image P20 and the summary of the diagram image P30, but may generate the summary of the character image P20 and the summary of the diagram image P30 individually.

[0211] Similarly, handwritten annotation shapes C43, C44, and handwritten annotation text C45 are annotations with text that do not include statistical indicators. The element data within the chart associated with handwritten annotation shapes C43 and C44 is "Netherlands," and the non-element data corresponding to this element data is the row name "Ballpoint pen" and the column name "Sales area." That is, the relationship between handwritten annotation shapes C43, C44, and handwritten annotation text C45 and the typefaces within the chart is relationship R4. In this case, the generation of summary information D32 uses the element data "Netherlands" within the chart associated with handwritten annotation shapes C43 and C44, and the non-element data column name "Sales area" and row name "Ballpoint pen" corresponding to the typeface "Netherlands." Also, in the example shown in Figure 22, the text image P20 contains the typeface "Ballpoint pens have the second highest sales after pencils." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "Ballpoint pens have a sales area that includes the booming Netherlands, and are the second highest-selling product after pencils," as shown in Figure 26(B), based on the type information contained in the character image P20 and the handwritten annotation figures C43, C44 and handwritten annotation characters C45. In other words, if both the character image P20 and the chart image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the chart image P30.

[0212] Relationship R5 indicates that the handwritten annotation emphasizes the typeset text within the chart associated with the handwritten annotation. The conditional text generation model has learned that for relationship R5, it generates text that emphasizes the content of the typeset text within the chart associated with the handwritten annotation, taking into account the constituent elements of the chart. For example, in the example shown in Figure 22, the handwritten annotation figure C52 is an annotation without text, the element data within the chart associated with the handwritten annotation figure C52 is "800", and the non-element data corresponding to this element data is "Gross Profit [yen]". In this case, the generation of summary information D32 uses the typeset text "800" from the element data associated with the handwritten annotation figure C52, and the column name "Gross Profit [yen]" and row name "Pencil" from the non-element data corresponding to the typeset text "800". Also, in the example shown in Figure 22, the text image P20 contains the typeset text "Pencil sales were particularly strong, with 20 sold." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "The pencil is a new product, the gross profit [yen] was 800, and 20 units were sold," based on the type information contained in the character image P20 and the handwritten annotation figure C52. In other words, if both the character image P20 and the figure image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the figure image P30.

[0213] Furthermore, multiple relationships among relationships R1 to R5 may be valid simultaneously. For example, relationships R2 and R5 may be valid at the same time. The conditional text generation model has learned to generate summary information D32 based on the multiple relationships when multiple relationships among relationships R1 to R5 are valid simultaneously. For example, in the example shown in Figure 22, the handwritten annotation figure C61 is an annotation without text, and the handwritten annotation figure C61 is associated with the column name "Sales [yen]", which is non-element data, and the element data "450", "1,000", "250", and "1,400". In this case, the generation of summary information D32 uses the non-element data and element data associated with the handwritten annotation figure C61, and the row name, which is non-element data corresponding to the element data. For example, the output processing unit 213 generates summary information D32 that includes the sentence, "Sales [yen] were 450 for notebooks, 1,000 for ballpoint pens, 250 for erasers, and 1,400 for pencils," based on the handwritten annotation figure C61. In other words, if multiple relationships among relationships R1 to R5 are simultaneously established, the output processing unit 213 generates summary information D32 based on those multiple relationships.

[0214] <Server 20: Step S48> In step S48, the output processing unit 213 of the control unit 21 outputs the summary information D32 generated in step S47. Specifically, the output processing unit 213 transmits the summary information D32 to the image processing device 10, which is the source of the image data that was the target of the summary output processing. The output processing unit 213 may also transmit the extraction results from step S44 to the image processing device 10 along with the summary information D32. In step S48, the output processing unit 213 may display, print, or store the summary information D32 and the extraction results from step S44. The summary information D32 and the extraction results from step S44 may also be transmitted to the client terminal 30.

[0215] <Image processing device 10: Step S113> In step S113, the control unit 11 of the image processing device 10 waits for the reception of summary information D32 from the server 20 (S113: No). When it is determined that the summary information D32 has been received (S113: Yes), the process proceeds to step S114.

[0216] <Image processing device 10: Step S114> In step S114, the control unit 11 of the image processing device 10 performs a result output process to output summary information D32. Specifically, in the result output process, the control unit 11 performs an image formation process to form an image on a sheet based on the image data that was the target of the summary output process. In the result output process, the image data that was the target of the summary output process and the summary information D32 may be stored in the storage unit 12 in association. In the result output process, the summary information D32 may be combined with the image data that was the target of the summary output process. In the result output process, the image data that was the target of the summary output process and the summary information D32 may be transmitted to the client terminal 30 in association or combined.

[0217] As explained above, the image processing system 1 can generate summary information D32 while taking into account handwritten annotations added to tables or graphs. This reduces the effort required from the user compared to when the user generates similar summary information D32 themselves.

[0218] [Other examples of charts and images] Furthermore, in this embodiment, the case in which the figure / table included in figure / table image P30 is a table was used as an example for explanation. On the other hand, it is also possible that the figure / table included in figure / table image P30 is various graphs, and in this case as well, the output processing unit 213 generates summary information D32 based on handwritten annotations such as handwritten annotation characters and handwritten annotation figures, and the target components of the graph that are the subject of the handwritten annotations. Below, the summary information D32 generated when figure / table images P40, P50, or P60 are included in target image P10 (see Figure 22) instead of figure / table image P30 will be explained.

[0219] The output processing unit 213 generates summary information D32 based on the handwritten annotations contained in the figure image P40, P50, or P60, and the elemental data and non-elemental data that are components of the bar graph contained in the figure image P40, P50, or P60. Specifically, the output processing unit 213 converts the bar graph data contained in the figure image P40, P50, or P60 into tabular data such as the aforementioned figure image P30, and then generates summary information D32 based on the figure image P30.

[0220] [Figure / Table Image P40] Figure 27(A) shows an example of figure image P40, which includes a bar graph, as another example of figure image P30 included in target image P10 (see Figure 22). In figure image P40, element data is indicated by vertical bar line data markers. Figure image P40 includes the graph title "2024 / 4 / 10 Product Sales", the legend "Sales [yen], Gross Profit [yen]", vertical axis tick labels "0", "500", "1000", "1500", and horizontal axis tick labels "Notebook", "Ballpoint pen", "Eraser", and "Pencil" as non-element data. Figure image P40 also includes the vertical axis tick as non-element data. Furthermore, figure image P40 includes element data corresponding to the horizontal axis tick labels "Notebook", "Ballpoint pen", "Eraser", and "Pencil". In addition, figure image P40 includes handwritten annotations: handwritten annotation shapes C71 and C72, and handwritten annotation text C73. Figure 27(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes the chart image P40.

[0221] Handwritten annotation shape C71 is a bounding line that annotates the horizontal axis tick label "eraser," which is non-elemental data in the bar graph of figure image P40. Handwritten annotation shape C72 is an arrow that annotates the elemental data "pencil sales [yen]" in the bar graph of figure image P40. Handwritten annotation text C73 is the annotation content for the subject of handwritten annotation shape C72, which is "up from the previous day."

[0222] In this case, the output processing unit 213 first converts the data from the bar graph into tabular data by inputting the element data from the bar graph into cells, using "Sales [yen]" and "Gross Profit [yen]" from the legend of the bar graph included in the chart image P40 as column names, and the horizontal axis scale labels "Notebook," "Ballpoint pen," "Eraser," and "Pencil" as row names.

[0223] For example, the output processing unit 213 calculates the value of each element data based on the non-element data of the bar graph and the Y-coordinate of the vertical axis of the data marker representing each element data in the bar graph. The Y-coordinate is assumed to be 0 at the top of the target image P10 or chart image P40 and the maximum value at the bottom. Furthermore, y1 is the value of the vertical axis tick label corresponding to the vertical axis tick closest to the Y-coordinate, which is less than or equal to the Y-coordinate of the upper end of the element data's data marker; y2 is the Y-coordinate of the vertical axis tick corresponding to that vertical axis tick label; y3 is the Y-coordinate of the lower end of the element data's data marker; and y4 is the Y-coordinate of the upper end of the element data's data marker. In this case, the value of the element data is calculated using the formula y1 × ((y3-y4) / (y3-y2)). Specifically, for data marker M1 corresponding to the sales [yen] of the horizontal axis tick label "pencil" in Figure 27(A), the value of the vertical axis tick label of vertical axis tick M2 is y1, the Y-axis coordinate of vertical axis tick M2 is y2, the Y-axis coordinate of the lower end of vertical axis tick M3 of data marker M1 is y3, and the Y-coordinate of the upper end of data marker M4 is y4. For example, for data marker M1 of pencil sales [yen], if y1=1500, y2=220, y3=290, and y4=227, then using the above calculation formula, the solution to 1500 × ((290-227) / (290-220)) is 1,350, which is the value of the element data for pencil sales [yen].

[0224] In the example shown in Figure 27(A), the type of handwritten annotation, including the handwritten annotation figure C71, is a textless annotation, and the non-element data associated with the handwritten annotation figure C71 is "eraser" (relationship R2). Therefore, for example, the output processing unit 213 generates summary information D32 that includes the sentence, "Erasers generated sales of 250 yen and gross profit of 150 yen," as a sentence emphasizing the handwritten annotation character C71, "eraser," as shown in Figure 27(B).

[0225] Furthermore, in the example shown in Figure 27(A), the type of handwritten annotation, which includes the handwritten annotation shape C72 and the handwritten annotation text C73 "up from the previous day," is a text-based annotation that does not include statistical indicators, and the handwritten annotation shape C72 is associated with the element data "1,350" (relationship R4). Also, as mentioned above, the text image P20 shown in Figure 22 contains the text "Pencil sales were particularly strong, with 20 sold." Therefore, for example, the output processing unit 213 generates summary information D32 that includes the text "20 pencils were sold, and sales [yen] were 1,350, up from the previous day," as a sentence that supplements the element data pencil sales [yen] of "1,350," using the handwritten annotation text C73 "up from the previous day."

[0226] [Figure / Table Image P50] Figure 28(A) shows an example of a chart image P50, which includes a line graph, as another example of a chart image P30 included in the target image P10 (see Figure 22). In chart image P50, element data is indicated by line data markers. Chart image P50 includes the graph title "Product Sales," the legend "Notebook, Ballpoint Pen, Eraser, Pencil," the vertical axis tick labels "0," "500," "1000," "1500," the horizontal axis tick labels "April 7th," "April 8th," "April 9th," "April 10th," the vertical axis label "Sales [yen]," and the horizontal axis label "Date" as non-element data. Note that the vertical and horizontal axis ticks are also included as non-element data in chart image P50. Furthermore, chart image P50 includes element data corresponding to the legend "Notebook," "Ballpoint Pen," "Eraser," and "Pencil." Furthermore, the figure image P50 includes handwritten annotations: handwritten annotation figures C81 and C82, and handwritten annotation characters C83. Figure 28(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes the figure image P50.

[0227] Handwritten annotation shape C81 is a bounding line that annotates the horizontal axis tick label "April 8th," which is non-elemental data of the line graph in Figure Image P50. Handwritten annotation shape C82 is an arrow that annotates the elemental data "Sales [yen]" for "April 10th" of the "Pencil" line in the line graph in Figure Image P50. Handwritten annotation text C83 is the handwritten text "Doing well," which indicates the content of the annotation for the subject of handwritten annotation shape C82.

[0228] In this case, the output processing unit 213 first converts the data of the line graph into tabular data by inputting the element data from the line graph into cells, using the legend entries "notebook," "ballpoint pen," "eraser," and "pencil" from the line graph included in the chart image P50 as column names, and the horizontal axis scale labels "April 7th," "April 8th," "April 9th," and "April 10th" as row names.

[0229] For example, the output processing unit 213 calculates the value of each element data based on the non-element data of the line graph and the Y-coordinate of the vertical axis of the data marker representing each element data in the line graph. The Y-coordinate value is assumed to be 0 at the top end of the target image P10 or chart image P50 and the maximum value at the bottom end. Furthermore, y1 is the value of the vertical axis tick label that corresponds to the vertical axis tick closest to the Y-coordinate, which is less than or equal to the Y-coordinate of the intersection of the data marker of the element data and the horizontal axis tick; y2 is the Y-coordinate of the vertical axis tick corresponding to that vertical axis tick label; y3 is the Y-coordinate of the vertical axis tick with the smallest Y-coordinate; and y4 is the Y-coordinate of the intersection of the data marker of the element data and the horizontal axis tick. In this case, the value of the element data is calculated using the formula y1 × ((y3-y4) / (y3-y2)). Specifically, for data marker M11 corresponding to the sales [yen] for the horizontal axis tick label "April 8th" in Figure 28(A), y1 is the value of the vertical axis tick label M12 at the vertical axis tick M12, y2 is the Y-axis coordinate of the vertical axis tick M12, y3 is the Y-axis coordinate of the vertical axis tick M13 at the lower end of the vertical axis tick label, and y4 is the Y-coordinate of the intersection of data marker M11 and the horizontal axis tick of "April 8th". For example, for data marker M11 for pencil sales [yen], if y1=1500, y2=220, y3=290, and y4=227, then using the above calculation formula, the solution to 1500 × ((290-227) / (290-220)) is 1,350, which is the value of the element data for pencil sales [yen].

[0230] In the example shown in Figure 28(A), the type of handwritten annotation, including the handwritten annotation figure C81, is a textless annotation, and the non-elemental data associated with the handwritten annotation figure C81 is "April 8th" (relationship R2). Therefore, for example, the output processing unit 213 generates summary information D32 based on the annotation image C81, which includes the sentence "On April 8th, there were 450 notebooks, 2,000 ballpoint pens, 350 erasers, and 770 pencils," as a sentence that emphasizes the non-elemental data "April 8th," as shown in Figure 28(B).

[0231] Furthermore, in the example shown in Figure 28(A), the type of handwritten annotation, which includes the handwritten annotation shape C82 and the handwritten annotation text C83 "doing well", is an annotation with text, and the non-element data associated with the handwritten annotation shape C82 is "April 10th" (relationship R1). Therefore, for example, the output processing unit 213 generates summary information D32 that includes the sentence, "Pencils are doing well, with sales of 630 on April 7th, 770 on April 8th, 1,050 on April 9th, and 1,400 on April 10th," as a sentence to supplement the element data "pencils" sales [yen] of "April 10th," which is "1,350," as shown in Figure 28(B).

[0232] [Figure / Table Image P60] Figure 29(A) shows an example of a chart image P60, which includes a pie chart, as another example of a chart image P30 included in the target image P10 (see Figure 22). In chart image P60, element data is indicated by circular data markers. Chart image P60 includes the graph title "Sales Composition Ratio" and the legend "Notebook, Ballpoint Pen, Eraser, Pencil" as non-element data. In addition, chart image P60 includes element data "15%", "32%", "8%", and "45%" corresponding to the legend "Notebook", "Ballpoint Pen", "Eraser", and "Pencil", respectively. Furthermore, chart image P60 includes handwritten annotations of handwritten annotation shapes C91 and handwritten annotation characters C92. Finally, Figure 29(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes chart image P60.

[0233] Handwritten annotation shape C91 is an arrow that annotates the element data "8%" of "Sales Composition Ratio" for "Erasers" in the pie chart image P60. Handwritten annotation text C92 is handwritten text that indicates the content of the annotation for the subject of handwritten annotation shape C91, saying "A small number, but 5 were sold."

[0234] In this case, the output processing unit 213 first converts the data from the pie chart into tabular data by inputting the element data from the pie chart into cells, using the legend entries "notebook," "ballpoint pen," "eraser," and "pencil" from the pie chart included in the chart image P60 as column names and the title of the pie chart, "sales composition ratio," as row names.

[0235] In the example shown in Figure 29(A), the type of handwritten annotation, including the handwritten annotation figure C91 and the annotation image C92, is an annotation with text, and the element data associated with the handwritten annotation figure C91 is "8%", which corresponds to "pencil" in the legend (relationship R4). Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "Erasers account for a small percentage of sales at 8%, but 5 were sold," as a sentence supplementing the element data "8%", as shown in Figure 29(B).

[0236] [Notes on the invention] The following is an overview of the invention extracted from the above-described embodiments. Note that each configuration and processing function described below can be selected and combined as desired.

[0237] <Note A1> An extraction processing unit that extracts handwritten characters and printed characters contained in the image data to be processed in a predetermined order, An identification processing unit identifies one or more handwritten characters and one or more printed characters that are consecutively extracted by the extraction processing unit as a series of strings, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit and one or more pre-set item name candidate strings contained in the printed characters. An image processing system equipped with the following features.

[0238] <Appendix A2> The identification processing unit identifies the handwritten characters and printed characters that exist between two consecutive item name candidate strings as a series of strings. The image processing system described in Appendix A1.

[0239] <Note A3> The identification processing unit identifies the handwritten characters and printed characters that are extracted after the last candidate string of item names in the extraction order as a series of strings. The image processing system described in Appendix A1 or A2.

[0240] <Note A4> The specified processing unit extracts named entities from the series of strings and excludes strings from the series of strings that do not satisfy the pre-set conditions for the named entities and the item groups associated with the candidate item name strings. The image processing system described in any of the appendices A1 to A3.

[0241] <Note A5> The specified processing unit combines at least one of the handwritten characters and the printed characters included in the series of strings into a single string. The image processing system described in one of the appendices A1 to A4.

[0242] <Note A6> The specified processing unit combines all the handwritten characters and printed characters included in the series of strings into a single string. The image processing system described in Appendix A5.

[0243] <Note A7> The system includes an output processing unit that outputs recognition result information generated based on the extraction results from the extraction processing unit and the processing results from the identification processing unit. An image processing system as described in any of the appendices A1 to A6.

[0244] <Note A8> An image processing device provided in the image processing system described in Appendix A7, An image acquisition unit that acquires the image data to be processed, a reception processing unit that receives input of information on preset input items based on the extraction result output from said output processing unit; An image processing apparatus comprising:

[0245] <Supplementary Note A9> An image processing method executed by one or more processors, the method comprising: an extraction step of extracting handwritten characters and printed characters included in an image represented by image data to be processed in a preset order; a specifying step of specifying, as a series of character strings, one or more of said handwritten characters and one or more of said printed characters that are consecutive in the extraction order obtained by said extraction step, based on a positional relationship between said handwritten characters and said printed characters extracted by said extraction step and one or more preset item name candidate character strings included in said printed characters; An image processing method for executing said steps.

[0246] <Supplementary Note A10> A program for causing one or more processors to execute: an extraction step of extracting handwritten characters and printed characters included in an image represented by image data to be processed in a preset order; a specifying step of specifying, as a series of character strings, one or more of said handwritten characters and one or more of said printed characters that are consecutive in the extraction order obtained by said extraction step, based on a positional relationship between said handwritten characters and said printed characters extracted by said extraction step and one or more preset item name candidate character strings included in said printed characters; A program for causing said steps to be executed.

[0247] <Supplementary Note B1> an extraction processing unit that extracts handwritten characters included in an image represented by image data to be processed; a specifying processing unit that specifies said item group corresponding to said handwritten characters based on an execution result of named entity extraction processing for extracting a named entity from said handwritten characters extracted by said extraction processing unit and a plurality of item groups preset in association with said image data; An image processing system comprising:

[0248] <Note B2> The system includes a setting processing unit that sets the item group corresponding to the image data in response to user operation. The image processing system described in Appendix B1.

[0249] <Note B3> The setting processing unit sets the item group according to the selection of the document type corresponding to the image data. The image processing system described in Appendix B1 or B2.

[0250] <Note B4> The extraction processing unit extracts handwritten characters and printed characters contained in the image shown by the image data to be processed. The identification processing unit performs a first identification process to identify the item group corresponding to the handwritten character based on the positional relationship between the handwritten character extracted by the extraction processing unit and the printed character extracted by the extraction processing unit, and a second identification process to identify the item group corresponding to the handwritten character based on the unidentified item groups among a plurality of item groups corresponding to the image data for which the handwritten character corresponding to the handwritten character was not identified by the first identification process, and the result of the named entity recognition process. An image processing system as described in any of the appendices B1 to B3.

[0251] <Note B5> The identification processing unit sets a determination group indicating the content of the handwritten character based on the named entity extracted from the handwritten character and a predetermined determination rule, and identifies the item group corresponding to the handwritten character based on the determination group and the item group. An image processing system as described in one of the appendices B1 to B4.

[0252] <Note B6> The determination rules include one or more of the following conditions for identifying the determination group based solely on the content of the named entity extracted from the handwritten character, conditions for identifying the determination group based on the named entity corresponding to the handwritten character and the content of the handwritten character, and conditions for identifying the determination group based on the named entity corresponding to the handwritten character and the item group associated with the printed characters surrounding the handwritten character. The image processing system described in Appendix B5.

[0253] <Note B7> The system includes an output processing unit that outputs recognition result information generated based on the identification result by the identification processing unit and the extraction result by the extraction processing unit. An image processing system as described in any of the appendices B1 to B6.

[0254] <Note B8> An image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix B7, A receiving processing unit accepts input of information for pre-set input items based on the extraction results output from the output processing unit, An image processing device equipped with the following features.

[0255] <Note B9> One or more processors An extraction step to extract handwritten characters contained in the image data to be processed, A selection step to identify the item group corresponding to the handwritten character, based on the result of executing a named entity recognition process to extract named entities from the handwritten characters extracted in the extraction step, and a plurality of item groups that are set in advance in association with the image data, An image processing method that performs this task.

[0256] <Note B10> One or more processors, An extraction step to extract handwritten characters contained in the image data to be processed, a specifying step of specifying the item group corresponding to the handwritten characters based on an execution result of a named entity extraction process that extracts named entities from the handwritten characters extracted by said extraction step, and a plurality of item groups set in advance in association with said image data; A program for causing execution of the above steps.

[0257] <Supplementary Note C1> a first processing unit that extracts a character region included in an image represented by image data to be processed and a character included in said character region; a second processing unit that specifies, as the number of lines in the character region, the number of characters present in a predetermined direction specified in advance in said character region; a third processing unit that specifies, in said character region, characters included in each line of the number of lines in said character region specified by said second processing unit; An image processing system comprising:

[0258] <Supplementary Note C2> said third processing unit specifies characters included in each line in said character region based on a result obtained by clustering the character string with a number corresponding to the number of lines in said character region specified by said second processing unit, The image processing system according to Supplementary Note C1.

[0259] <Supplementary Note C3> said third processing unit executes said clustering by a k-means method, The image processing system according to Supplementary Note C2.

[0260] <Supplementary Note C4> said third processing unit samples each character in said character region by a preset number, and executes said clustering based on sampled data after the sampling, The image processing system according to Supplementary Note C2 or C3.

[0261] <Supplementary Note C5> The system includes a fourth processing unit that determines the specific direction in the character area based on the aspect ratio of the character area. An image processing system as described in any of the appendices C1 to C4.

[0262] <Appendix C6> The first processing unit distinguishes and extracts handwritten characters and printed characters included in the character area, The second processing unit, the third processing unit, and the fourth processing unit perform processing only on the character area including the handwritten characters. An image processing system as described in any of the appendices C1 to C5.

[0263] <Note C7> The third processing unit performs clustering only on the character regions where the number of lines identified by the second processing unit is multiple. An image processing system as described in any of the appendices C1 to C6.

[0264] <Note C8> The system includes an output processing unit that outputs recognition result information generated based on the identification result by the third processing unit, An image processing system as described in any of the appendices C1 to C7.

[0265] <Note C9> An image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix C8, A receiving processing unit accepts input of information for pre-set input items based on the extraction results output from the output processing unit, An image processing device equipped with the following features.

[0266] <Note C10> One or more processors The first step involves extracting the character regions and the characters contained within the image data to be processed, A second step involves determining the number of characters present in a predetermined specific direction within the character area as the number of lines in that character area. In the character area, the third step is to identify the characters contained in each line of the number of lines in the character area identified in the second step, An image processing method that performs this task.

[0267] <Note C11> One or more processors, The first step involves extracting the character regions and the characters contained within the image data to be processed, A second step involves determining the number of characters present in a predetermined specific direction within the character area as the number of lines in that character area. In the character area, the third step is to identify the characters contained in each line of the number of lines in the character area identified in the second step, An image processing program to execute this.

[0268] <Note D1> An extraction processing unit that extracts characters contained in the image data to be processed, An identification processing unit that determines whether two consecutive lines of text extracted by the extraction processing unit are related, and if it is determined that they are related, performs an identification process to identify the two lines of text as a series of texts, An image processing system equipped with the following features.

[0269] <Note D2> The specified processing unit determines whether or not there is a relationship based on the meaning of the sentence when the two lines of text are joined together. The image processing system described in Appendix D1.

[0270] <Note D3> The aforementioned specific processing unit uses BERT (Bidirectional Encoder Representations from Transformers) NSP (Next Sentence Prediction) to determine whether or not the relationship exists. The image processing system described in Appendix D1 or D2.

[0271] <Note D4> The specified processing unit executes the specified processing on the character area only if the characters in the character area extracted by the extraction processing unit are in sentence form. An image processing system as described in any of the appendices D1 to D3.

[0272] <Note D5> The extraction processing unit distinguishes and extracts handwritten characters and printed characters contained in the image, The specified processing unit executes the specified processing for the character area only if the character in the character area extracted by the extraction processing unit is a handwritten character. An image processing system as described in any of the appendices D1 to D4.

[0273] <Note D6> The identification processing unit combines the two lines of text when the two lines of text are identified as a series of texts. An image processing system as described in any of the appendices D1 to D5.

[0274] <Note D7> The system includes an output processing unit that outputs recognition result information generated based on the identification result by the identification processing unit and the extraction result by the extraction processing unit. An image processing system as described in any of the appendices D1 to D6.

[0275] <Note D8> An image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix D7, A receiving processing unit accepts input of information for pre-set input items based on the extraction results output from the output processing unit, An image processing device equipped with the following features.

[0276] <Note D9> One or more processors An extraction step to extract characters contained in the image data to be processed, A determination step is made to determine whether the two consecutive lines of text extracted by the extraction step are related, and if it is determined that they are related, to identify the two lines of text as a series of texts. An image processing method that performs this task.

[0277] <Note D10> One or more processors, An extraction step to extract characters contained in the image data to be processed, A determination step is made to determine whether the two consecutive lines of text extracted by the extraction step are related, and if it is determined that they are related, to identify the two lines of text as a series of texts. A program to execute.

[0278] <Note E1> An extraction processing unit that extracts handwritten annotations, including either or both handwritten annotation characters and handwritten annotation figures, contained in the image data to be processed, and components of diagrams and charts contained in the image, An output processing unit generates summary information based on the handwritten annotation extracted by the extraction processing unit and the target component among the components that is the subject of the handwritten annotation; An image processing system equipped with the following features.

[0279] <Note E2> The output processing unit generates the summary information based on the handwritten annotation, the target component, and related components among the components that are related to the target component. The image processing system described in Appendix E1.

[0280] <Note E3> The above figures and tables are tables, The components of the aforementioned figure include element data and non-element data excluding said element data. The aforementioned non-elemental data includes at least one of the table title, column name, or row name, The output processing unit generates the summary information including the element data and the non-element data. The image processing system described in Appendix E1 or E2.

[0281] <Note E4> The output processing unit generates summary information that includes the element data and either or both of the column name and row name corresponding to the element data, when the target component is the element data. An image processing system as described in any of the appendices E1 to E3.

[0282] <Note E5> The aforementioned figure is a graph, The components of the aforementioned figure include element data and non-element data excluding said element data. The aforementioned non-elemental data includes at least one of the following: graph title, legend, vertical (value) axis, horizontal (item) axis, vertical axis label, horizontal axis label, vertical axis tick, horizontal axis tick, vertical axis tick label, horizontal axis tick label, or data label. The output processing unit generates the summary information including the element data and the non-element data. An image processing system as described in any of the appendices E1 to E4.

[0283] <Note E6> The output processing unit separates the image into a character image containing at least printed characters and a chart image containing at least the chart, and generates the summary information by combining a summary of the printed characters contained in the character image and a summary of the chart contained in the chart image. An image processing system as described in any of the appendices E1 to E5.

[0284] <Note E7> The output processing unit outputs the generated summary information. An image processing system as described in any of the appendices E1 to E6.

[0285] <Note E8> An image processing apparatus comprising an image acquisition processing unit that acquires the image data to be processed in the image processing system described in any of the appendices E1 to E7.

[0286] <Note E9> One or more processors Extraction step of extracting handwritten annotations, including either or both handwritten annotation characters and handwritten annotation figures, and components of diagrams and charts included in the image data to be processed. An output step that generates summary information based on the handwritten annotations extracted by the extraction step and the target component among the components that is the subject of the handwritten annotations, An image processing method that performs this task.

[0287] <Note E10> One or more processors, Extraction step of extracting handwritten annotations, including either or both handwritten annotation characters and handwritten annotation figures, and components of diagrams and charts included in the image data to be processed. An output step that generates summary information based on the handwritten annotations extracted by the extraction step and the target component among the components that is the subject of the handwritten annotations, A program to execute. [Explanation of Symbols]

[0288] 1. Image Processing System 10 Image Processing Device 11 Control Unit 15 Image acquisition unit 20 servers 21 Control Unit 211 Extraction Processing Unit 212 Specific Processing Unit 213 Output Processing Unit 30 client terminals 31 Control Unit 35 Image acquisition unit

Claims

1. A first processing unit that extracts the character region contained in the image indicated by the image data to be processed and the characters contained in said character region, A second processing unit that determines the number of characters present in a predetermined specific direction within the character area as the number of lines in the character area, In the character area, a third processing unit identifies the characters contained in each line of the number of lines in the character area identified by the second processing unit, An image processing system equipped with the following features.

2. The third processing unit identifies the characters contained in each line of the character region based on the result of clustering the string in a number corresponding to the number of lines in the character region identified by the second processing unit. The image processing system according to claim 1.

3. The third processing unit performs the clustering using the k-means method. The image processing system according to claim 2.

4. The third processing unit samples each character in the character area by a predetermined number of times, and performs the clustering based on the sampled data after sampling. The image processing system according to claim 2 or 3.

5. The system includes a fourth processing unit that determines the specific direction in the character area based on the aspect ratio of the character area. The image processing system according to claim 1.

6. The first processing unit distinguishes and extracts handwritten characters and printed characters included in the character area, The second processing unit, the third processing unit, and the fourth processing unit perform processing only on the character area including the handwritten characters. The image processing system according to claim 1.

7. The third processing unit performs clustering only on the character regions where the number of lines identified by the second processing unit is multiple. The image processing system according to claim 1.

8. The system includes an output processing unit that outputs recognition result information generated based on the identification result by the third processing unit, The image processing system according to claim 1.

9. An image acquisition processing unit that acquires the image data to be processed in the image processing system according to claim 8, A receiving processing unit accepts input of information for pre-set input items based on the extraction results output from the output processing unit, An image processing device equipped with the following features.

10. One or more processors The first step involves extracting the character regions and the characters contained within the image data to be processed, A second step is to determine the number of characters in a specific direction that is predetermined in the character area as the number of lines in that character area, In the character area, the third step is to identify the characters contained in each line of the number of lines in the character area identified in the second step, An image processing method that performs the following.

11. One or more processors, The first step involves extracting the character regions and the characters contained within the image data to be processed, A second step is to determine the number of characters in a specific direction that is predetermined in the character area as the number of lines in that character area, In the character area, the third step is to identify the characters contained in each line of the number of lines in the character area identified in the second step, An image processing program to execute this.

Citation Information

Patent Citations

  • Image processing system and image processing method

    JP2023014964A