Image processing system, image processing device, and image processing method

WO2026181943A1PCT designated stage Publication Date: 2026-09-03KYOCERA DOCUMENT SOLUTIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/006348
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-27
Filing Date
2026-02-20
Publication Date
2026-09-03

Smart Images

  • Figure JP2026006348_03092026_PF_FP_ABST
    Figure JP2026006348_03092026_PF_FP_ABST
Patent Text Reader

Abstract

This image processing system comprises an extraction processing unit and an identification processing unit. The extraction processing unit extracts handwritten characters included in an image represented by image data to be processed. The identification processing unit identifies an item group corresponding to the handwritten characters, on the basis of: the results of executing named entity extraction processing for extracting a named entity from the handwritten characters extracted by the extraction processing unit; and a plurality of item groups preset in association with the image data.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing system, image processing apparatus, image processing method

[0001] The present invention relates to an image processing system, an image processing apparatus, and an image processing method.

[0002] In general, there is known a technique for specifying an item indicating the content of a handwritten character from which an image is extracted, based on printed characters existing around the handwritten character (see, for example, Patent Document 1).

[0003] Japanese Unexamined Patent Application Publication No. 2023-22573

[0004] Incidentally, when there is no printed character around a handwritten character, an item corresponding to the handwritten character cannot be specified. On the other hand, among handwritten characters, one that includes any of constituent characters such as numbers or symbols preset in association with each item and has the smallest number of characters not included in the constituent characters may be specified as the handwritten character corresponding to the item. However, with such a configuration, when a handwritten character includes another character that is not a number or a symbol, the item corresponding to the handwritten character cannot be specified.

[0005] An object of the present invention is to provide an image processing system, an image processing apparatus, and an image processing method capable of specifying an item indicating the content of a handwritten character extracted from an image for the handwritten character.

[0006] An image processing system according to one aspect of the present invention includes an extraction processing unit and a specification processing unit. The extraction processing unit extracts a handwritten character included in an image represented by image data to be processed. The specification processing unit specifies the item group corresponding to the handwritten character based on an execution result of named entity extraction processing for extracting a named entity from the handwritten character extracted by the extraction processing unit and a plurality of item groups preset in association with the image data.

[0007] An image processing apparatus according to another aspect of the present invention is an image processing apparatus provided in the image processing system, comprising an image acquisition unit and a reception processing unit. The image acquisition unit acquires the image data to be processed. The reception processing unit accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

[0008] Another aspect of the present invention relates to an image processing method in which one or more processors perform an extraction step and a identification step. The extraction step extracts handwritten characters contained in an image represented by the image data to be processed. The identification step identifies the item group corresponding to the handwritten characters based on the result of a named entity recognition process that extracts named entities from the handwritten characters extracted by the extraction step and a plurality of item groups that are set in advance in association with the image data.

[0009] According to the present invention, it is possible to provide an image processing system, an image processing device, and an image processing method that can identify items indicating the content of handwritten characters extracted from an image.

[0010] Figure 1 is a block diagram showing the configuration of an image processing system according to an embodiment of the present invention. Figure 2 is a diagram showing an example of candidate item information used in the image processing system according to an embodiment of the present invention. Figure 3 is a diagram showing an example of document-specific item information used in the image processing system according to an embodiment of the present invention. Figure 4 is a flowchart illustrating an example of processing performed by the image processing system according to an embodiment of the present invention. Figure 5 is a diagram showing an example of an image that is the target of processing performed by the image processing system according to an embodiment of the present invention. Figure 6 is a diagram showing an example of first extraction result information used in the image processing system according to an embodiment of the present invention. Figure 7 is a diagram showing an example of recognition result information used in the image processing system according to an embodiment of the present invention. Figure 8 is a flowchart illustrating an example of processing performed by the image processing system according to an embodiment of the present invention. Figure 9 is a diagram showing an example of an image that is the target of processing performed by the image processing system according to an embodiment of the present invention. Figure 10 is a diagram showing an example of second extraction result information used in the image processing system according to an embodiment of the present invention. Figure 11 is a diagram showing an example of second extraction result information used in the image processing system according to an embodiment of the present invention. Figure 12 is a diagram showing an example of second extraction result information used in the image processing system according to an embodiment of the present invention. Figure 13 is a diagram showing an example of second extraction result information used in the image processing system according to an embodiment of the present invention. Figure 14 is a diagram showing an example of recognition result information used in the image processing system according to an embodiment of the present invention. Figure 15 is a diagram showing an example of an image to be processed by the image processing system according to an embodiment of the present invention. Figure 16 is a flowchart illustrating an example of processing performed by the image processing system according to an embodiment of the present invention. Figure 17 is a diagram illustrating an example of the result of processing performed by the image processing system according to an embodiment of the present invention. Figure 18 is a diagram illustrating an example of the result of processing performed by the image processing system according to an embodiment of the present invention. Figure 19 is a flowchart illustrating an example of processing performed by the image processing system according to an embodiment of the present invention. Figure 20 is a diagram illustrating an example of the result of processing performed by the image processing system according to an embodiment of the present invention.Figure 21 is a flowchart illustrating an example of processing performed by an image processing system according to an embodiment of the present invention. Figure 22 is a diagram illustrating an example of an image that is the target of processing performed by an image processing system according to an embodiment of the present invention. Figure 23 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention. Figure 24 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention. Figure 25 is a diagram illustrating an example of relationship information used in an image processing system according to an embodiment of the present invention. Figure 26 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention. Figure 27 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention. Figure 28 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention. Figure 29 is a diagram illustrating an example of the result of processing performed by an image processing system according to an embodiment of the present invention.

[0011] The embodiments of the present invention will be described below with reference to the attached drawings. The following embodiments are examples that embody the present invention and are not intended to limit the technical scope of the present invention.

[0012] [Image Processing System 1] As shown in Figure 1, the image processing system 1 according to this embodiment comprises one or more image processing devices 10, one or more servers 20, and one or more client terminals 30. The image processing devices 10, servers 20, and client terminals 30 are connected to each other via a communication network 100 such as the Internet or a LAN (Local Area Network).

[0013] In this embodiment, in the image processing system 1, image data to be processed by character recognition processing (hereinafter referred to as "OCR processing") to recognize characters contained in the image is acquired by the image processing device 10. Then, the image data is transmitted from the image processing device 10 to the server 20. Subsequently, the server 20 performs OCR processing on the image data, and the result of the OCR processing is transmitted from the server 20 to the image processing device 10. In this embodiment, the execution entities of the various processes performed in the image processing system 1 are merely examples, and similar processing can be performed throughout the image processing system 1. For example, the OCR processing may be performed not only on the server 20, but also on the image processing device 10 or the client terminal 30.

[0014] [Image Processing Device 10] The image processing device 10 has an image reading function that reads an image of a document and generates image data. The image processing device 10 is, for example, a scanner, a facsimile machine, a multifunction printer, or a camera.

[0015] The image processing device 10 includes a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and an image acquisition unit 15, etc. The image processing device 10 may also include an image forming unit that forms an image on a sheet based on image data.

[0016] The control unit 11 comprises one or more processors and one or more memories, and executes various processes in the image processing device 10. The storage unit 12 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 12 stores image data acquired by the image acquisition unit 15, image data received from the client terminal 30, and recognition result information D12 corresponding to the execution result of OCR processing performed on the server 20. The storage unit 12 also stores control programs and the like that cause the control unit 11 of the image processing device 10 to execute various processes, and the control unit 11 executes various processes according to the control programs.

[0017] The operation unit 13 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 14 is a liquid crystal display or organic EL display that displays various types of information.

[0018] The image acquisition unit 15 is controlled by the control unit 11 to perform an image acquisition process to acquire image data. Specifically, the image acquisition unit 15 has an image reading unit that reads an image from a document and outputs image data of the read image. For example, the image reading unit is an image reading unit that has an automatic document feeder (ADF), a document tray, a light source, multiple mirrors, optical lenses, and a CCD (Charge Coupled Device), etc.

[0019] In other embodiments, the image acquisition unit 15 may be a shooting unit such as a digital camera that captures images such as still images or videos and outputs image data corresponding to those images. Alternatively, the image acquisition unit 15 may be an interface for reading the image data from the shooting unit such as the digital camera. The object to be photographed by the digital camera may be a display screen such as a tablet or display. Furthermore, the image acquisition unit 15 may acquire image data of an image displayed on a display screen by taking a screenshot of the display screen which includes handwritten characters using a stylus or the like via a touch panel. In addition, the image acquisition unit 15 may read image data stored in an external device such as the storage unit 12 of the image processing device 10, the server 20, or the storage unit 22 or storage unit 32 of the client terminal 30.

[0020] Furthermore, the control unit 11 can receive an OCR start operation via the operation unit 13, for example, when the image acquisition unit 15 starts acquiring image data or after the acquisition is completed, to cause the server 20 to execute OCR processing on the image data as the processing target. The control unit 11 then stores the recognition result information D12 (see Figure 7) of the OCR processing acquired from the server 20 in the storage unit 12. The control unit 11 can also display the processing result of the OCR processing on the display unit 14 based on the recognition result information D12, or output the recognition result information D12 to an external device such as a client terminal 30.

[0021] Furthermore, the control unit 11 can perform an automatic input process that automatically inputs characters corresponding to a pre-set item group from among the characters recognized by OCR processing as information corresponding to a pre-set item in a predetermined database or the like. Specifically, the image processing device 10 has a document management function that manages information on one or more items contained in documents such as questionnaires, contracts, order forms, application forms, or receipts. The control unit 11 then performs an automatic input process that automatically inputs characters read from the document by OCR processing into a database stored in the storage unit 12 as information on a specific item in the document. For example, if the document is a receipt, information on items such as "document name," "serial number," "recipient name," "transaction details," and "amount received" is entered into the database based on the results of the OCR processing. Note that the server 20 or client terminal 30 may also be equipped with the document management function and perform the automatic input process.

[0022] [Server 20] Server 20 is an information processing device comprising a control unit 21, a storage unit 22, an operation unit 23, a display unit 24, and an image acquisition unit 25, etc. Server 20 is, for example, a personal computer, a smartphone, or a tablet terminal.

[0023] The control unit 21 comprises one or more processors and one or more memories, and executes various processes in the server 20. The storage unit 22 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 22 stores image data acquired by the image acquisition unit 25, image data acquired from the image processing device 10 or client terminal 30, and recognition result information D12 corresponding to the execution result of OCR processing performed in the server 20. The storage unit 22 also stores a control program that causes the control unit 21 of the server 20 to execute various processes, and the control unit 21 executes various processes according to the control program.

[0024] The operation unit 23 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 24 is a liquid crystal display or organic EL display that displays various types of information.

[0025] The image acquisition unit 25 performs processing to acquire image data that is subject to OCR processing performed by the control unit 21. Specifically, the image acquisition unit 25 is a communication unit that receives image data corresponding to images such as still images or videos from the image processing device 10 or client terminal 30. The image acquisition unit 25 may also read image data stored in an external device such as the storage unit 12 or storage unit 32 of the image processing device 10 or client terminal 30. Furthermore, the image acquisition unit 25 is a shooting unit of a digital camera or the like that captures images such as still images or videos and outputs image data corresponding to those images. The image acquisition unit 25 may also be an interface or the like that reads the image data from the shooting unit of the digital camera or the like.

[0026] [Client Terminal 30] The client terminal 30 is an information processing device comprising a control unit 31, a storage unit 32, an operation unit 33, a display unit 34, and an image acquisition unit 35, etc. The client terminal 30 is, for example, a personal computer, a smartphone, or a tablet terminal.

[0027] The control unit 31 comprises one or more processors and one or more memories, and executes various processes on the client terminal 30. The storage unit 32 is a hard disk drive (HD) or solid-state drive (SSD) or the like that stores various data non-temporarily. For example, the storage unit 32 stores image data acquired by the image acquisition unit 35, image data received from the image processing device 10, and recognition result information D12 corresponding to the execution result of OCR processing performed on the server 20. The storage unit 32 also stores a control program that causes the control unit 31 of the client terminal 30 to execute various processes, and the control unit 31 executes various processes according to the control program.

[0028] The operation unit 33 includes a keyboard, mouse, touch panel, and voice input device that accept user input. The display unit 34 is a liquid crystal display or organic EL display that displays various types of information.

[0029] The image acquisition unit 35 is a shooting unit of a digital camera or the like that captures images such as still images or videos and outputs image data corresponding to those images. The image acquisition unit 35 may also be an interface or the like that reads the image data from the shooting unit of the digital camera or the like. The object to be captured by the digital camera or the like may be a display screen such as a tablet or display. The image acquisition unit 35 may also acquire image data of an image displayed on a display screen by taking a screenshot of the display screen which contains handwritten characters using a stylus or the like via a touch panel. Furthermore, the image acquisition unit 35 may also read image data stored in an external device such as the storage unit 12 or storage unit 22 of the image processing device 10 or server 20.

[0030] Incidentally, in the image processing system 1 configured in this way, the image shown by the image data subject to OCR processing (hereinafter sometimes referred to as the "target image") may include printed or displayed typefaces and handwritten characters. In this embodiment, "typefaces" refers to characters expressed in a predetermined specific font such as Arial or Century, and excludes characters written by hand by humans or other means. For example, "typefaces" include printed characters printed on a sheet such as paper by a printer using the aforementioned specific font, and displayed characters shown on a display using the aforementioned specific font.

[0031] Furthermore, OCR processing may distinguish between handwritten characters and printed characters in a target image. In this case, if a series of characters, such as a single word or sentence, originally contains both handwritten and printed characters, the series of characters that should be recognized may be recognized separately as a handwritten character string and a printed character string. In contrast, there is a known technique that combines handwritten and printed characters as a series of characters according to the user's specification operation for the characters to be combined. However, such a configuration has the problem that the operation to specify the handwritten and printed characters to be included in the series of characters becomes burdensome for the user. In contrast, the image processing system 1 according to this embodiment makes it possible to reduce the burden on the user in order to identify handwritten and printed characters as a series of characters.

[0032] The detailed configuration of the image processing system 1 according to this embodiment will be described below.

[0033] [Storage Unit 22 of Server 20] The storage unit 22 of Server 20 stores candidate item information D1 and document-specific item information D2, etc. Various types of data such as candidate item information D1 and document-specific item information D2 may be set in Server 20 according to user operation and stored in Storage Unit 22, or they may be set by an external device such as an image processing device 10 or a client terminal 30 and transmitted to Server 20 and stored in Storage Unit 22.

[0034] <Candidate Item Information D1> As shown in Figure 2, Candidate Item Information D1 associates multiple pre-configured item groups with multiple pre-configured candidate item name strings. Candidate Item Information D1 is used by the control unit 21 to recognize the information of the candidate item name strings included as printed characters in the image to be processed by OCR as information corresponding to the item group. For example, the item group "Date" is associated with candidate item name strings such as "Application Date," "Entry Date," "Date," and "Date of Birth." Similarly, the item group "Name" is associated with candidate item name strings such as "Your Name," "Your Full Name," "Last Name," and "Given Name," and the item group "Address" is associated with candidate item name strings such as "Your Address," "Place," and "Address." Also, as shown in Figure 2, Candidate Item Information D1 includes item groups such as "Document Name," "Recipient Name," "Serial Number," and "Transaction Details."

[0035] <Document-Specific Item Information D2> As shown in Figure 3, in Document-Specific Item Information D2, pre-set document types are associated with the item groups. Document-Specific Item Information D2 is used by the control unit 21 to recognize the item groups included in the target image according to the document type of the target image for OCR processing. For example, in the example shown in Figure 3, the document type "Receipt" stores item groups such as "Document Name," "Serial Number," "Recipient Name," "Transaction Details," "Receipt Amount," "Receipt Date," "Document Creator Name," "Document Creator Company Name," and "Document Creator Address." Similarly, in the example shown in Figure 3, the document type "Order Form" stores item groups such as "Document Name," "Recipient Name," "Order Quantity," "Order Amount," "Order Details," "Order Date," "Orderer Address," "Orderer Company Name," and "Orderer Name." Furthermore, various data such as candidate item information D1 and document-specific item information D2 may be set on the server 20 and stored in the storage unit 22, or they may be set on the image processing device 10 or client terminal 30, transmitted to the server 20, and stored in the storage unit 22. In addition, candidate item information D1 and document-specific item information D2 may be stored in the storage unit 12 or storage unit 32 of the image processing device 10 or client terminal 30.

[0036] [Control Unit 21 of Server 20] As shown in Figure 1, the control unit 21 of Server 20 includes an extraction processing unit 211, a specific processing unit 212, an output processing unit 213, and the like.

[0037] The extraction processing unit 211 performs OCR processing to recognize characters contained in the target image data. In particular, the extraction processing unit 211 is capable of distinguishing and recognizing handwritten characters and printed characters contained in the target image in a predetermined order. Furthermore, in the OCR processing of the target image, the extraction processing unit 211 can also perform named entity recognition processing to extract named entities such as names or place names from the characters contained in the target image.

[0038] The identification processing unit 212 performs a string identification process to identify one or more strings contained in the target image. Specifically, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit 211 and one or more pre-set item name candidate strings contained in the printed characters, the identification processing unit 212 identifies one or more handwritten characters and one or more printed characters that are consecutive in the extraction order by the extraction processing unit 211 as a series of strings. The identification processing unit 212 can also combine the one or more handwritten characters and one or more printed characters identified as a series of strings into a single string. In particular, the identification processing unit 212 may combine all of the handwritten characters and printed characters identified as a series of strings, or it may combine at least one of the handwritten characters and printed characters identified as a series of strings.

[0039] The output processing unit 213 generates recognition result information D12, D22, etc., based on either or both of the extraction results from the extraction processing unit 211 and the processing results from the identification processing unit 212, and outputs the recognition result information D12, D22. The output methods of the output processing unit 213 include displaying, printing, transmitting, or storing the recognition result information D12, D22. For example, the recognition result information D12, D22 is transmitted to the image processing device 10 or the client terminal 30.

[0040] [Image Processing Method] Next, with reference to Figure 4, an example of the processing procedure in the image processing method executed by the image processing system 1 according to this embodiment will be described.

[0041] In the image processing system 1, the image processing method of the present invention is executed by one or more processors provided in the image processing system 1 performing various processes. Note that the various processes in the image processing method may be performed by one of the control unit 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30. Alternatively, the various processes in the image processing method may be divided and performed by two or three of the control unit 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30.

[0042] Specifically, steps S11, S12, etc. in Figure 4 represent the processing procedure (step) numbers for the processing executed by the control unit 11 of the image processing device 10 in the image processing method. Specifically, the control unit 11 of the image processing device 10 starts processing when it receives a request to execute an image reading process involving OCR processing from a user. For example, when the control unit 11 receives a request to execute the automatic input process from a user, it determines that a request to execute an image reading process involving OCR processing has been made.

[0043] Furthermore, steps S21, S22, etc. in Figure 4 represent the processing procedure (step) of the OCR processing executed by the control unit 21 of the server 20 in the image processing method. Specifically, the control unit 21 starts the OCR processing when the power to the server 20 is turned on.

[0044] <Image Processing Apparatus 10: Step S11> In step S11, the control unit 11 controls the image acquisition unit 15 to execute an image acquisition process for acquiring image data to be processed by OCR processing. Specifically, the control unit 11 controls the image acquisition unit 15 to execute the image reading process, thereby acquiring image data corresponding to the image of the document. In addition, in step S11, the control unit 11 may acquire image data to be processed by OCR processing by reading out the image data stored in the storage unit 12 in response to a user operation.

[0045] <Image Processing Apparatus 10: Step S12> In step S12, the control unit 11 transmits the image data acquired in step S11 to the server 20 as a processing target of OCR processing.

[0046] <Server 20: Step S21> In step S21, the control unit 21 waits to receive image data to be processed by OCR processing (S21: No). Specifically, the control unit 21 receives image data to be processed by OCR processing from the image processing apparatus 10 or the client terminal 30. When it is determined that the image data has been received (S21: Yes), the process proceeds to step S22.

[0047] <Server 20: Step S22> In step S22, the extraction processing unit 211 of the control unit 21 executes character extraction processing including extracting and recognizing a character region existing in the target image represented by the image data and characters existing in the character region, based on the image data received as the processing target of OCR processing in step S21. In particular, the extraction processing unit 211 distinguishes and extracts handwritten characters and printed characters existing in the character region.

[0048] Specifically, the extraction processing unit 211 executes separation processing based on image data to be processed to separate a target image represented by the image data into a handwritten image having only handwritten pixels constituting handwritten characters, and a print image having only print pixels constituting printed characters or ruled lines. It should be noted that any conventionally known technology may be employed for the separation processing. For example, in the separation processing, semantic segmentation using deep learning such as a deep neural network (DNN) based on artificial intelligence is used. In the deep learning training, it is conceivable to use, as training data, a plurality of generated superimposed images obtained by superimposing a handwritten image and a print image. It is also conceivable to use handwritten binarized image data as teacher data such that pixel values of image data obtained by binarizing the handwritten image corresponding to the superimposed image serve as label values indicating handwritten pixels.

[0049] Next, the extraction processing unit 211 sequentially extracts characters such as handwritten characters and printed characters included in the handwritten character image and the print image respectively, from a preset processing start end to a preset processing end end in the target image. Specifically, the extraction processing unit 211 generates first extraction result information D11 (see FIG. 6) in which an extracted character extracted from the target image, a classification of the extracted character (handwritten character or printed character), and position coordinates of a circumscribed rectangle of the extracted character are associated with each other.

[0050] The extracted character is a character string including one or a plurality of characters. The position coordinates are the position coordinates of the upper-left end and the position coordinates of the lower-right end of the circumscribed rectangle when the horizontal direction in the target image is taken as the X axis and the vertical direction is taken as the Y axis. In addition, for the X coordinate in the position coordinates, the left end of the target image is a minimum value and the right end is a maximum value, and for the Y coordinate in the position coordinates, the upper end of the target image is a minimum value and the lower end is a maximum value. In the present embodiment, description is given on the assumption that the processing start end is the upper-left end, the processing end end is the lower-right end, and the closer the position coordinate of the upper-left end of the circumscribed rectangle of a character is to the left end and the upper end, the earlier the character is extracted in the target image.

[0051] However, if multiple extracted characters that should be extracted in a predetermined order on the same line include handwritten characters, the position and size of the extracted characters may not be uniform. In this case, even on the same line, the upper left corners of each handwritten character may be shifted vertically relative to each other, or the upper left corners of the handwritten character and the type may be shifted vertically. As a result, multiple handwritten or typed characters that should be extracted in a predetermined order on the same line may be extracted in an order different from the predetermined order. Therefore, when the extraction processing unit 211 extracts characters sequentially from the upper left corner to the lower right corner, if the first bounding rectangle of the extracted character is detected on each line, it determines whether there are other bounding rectangles that interfere with the first bounding rectangle in the X-axis direction, at least in part. That is, the extraction processing unit 211 extracts other extracted characters whose range on the Y-axis overlaps with the bounding rectangle of the first detected extracted character at least in part. The extraction processing unit 211 then rearranges the extraction order of the multiple extraction characters corresponding to the multiple bounding rectangles whose ranges on the Y axis overlap, in ascending order of the X coordinate values ​​of the bounding rectangles, to generate the first extraction result information D11. This makes it possible to extract the extraction characters in their original extraction order even if the position and size of the characters are not uniform.

[0052] Here, Figure 5 shows an example of a target image, and Figure 6 shows an example of the first extraction result information D11 extracted from the target image in Figure 5. Note that the dashed line in Figure 5 is not included in the target image, but indicates the character area in the target image.

[0053] As shown in Figure 6, the handwritten characters extracted sequentially from the top left of the target image in Figure 5 are "24", "7", "13", "Yamada Hanako", "Tokyo XX Ward", "△ Mansion No. 111", and "×× Company". Also, as shown in Figure 6, the printed characters extracted sequentially from the top left of the target image in Figure 5 are "Application Form", "Date of Entry", "20", "Year", "Month", "Day", "Name", and "Address". Below, we will explain the processing content of each step using the case where the first extraction result information D11 in Figure 6 is extracted from the target image in Figure 5 as an example.

[0054] <Server 20: Step S23> In step S23, the identification processing unit 212 of the control unit 21 extracts candidate item name strings included in candidate item information D1 (see Figure 2) from among the one or more type characters extracted by the extraction processing unit 211, based on candidate item information D1 and first extraction result information D11. Specifically, in the target image shown in Figure 5, "Date of entry", "Name", and "Address" are extracted as candidate item name strings.

[0055] <Server 20: Step S24> In step S24, the identification processing unit 212 of the control unit 21 identifies one or more handwritten characters and one or more printed characters that are consecutively extracted by the extraction processing unit 211 as a series of strings, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit 211 and one or more candidate strings of item names extracted in step S23. The identification processing unit 212 also combines all of the identified series of strings. The identification processing unit 212 may combine at least one handwritten character and one printed character from the identified series of strings.

[0056] Specifically, the identification processing unit 212 identifies handwritten characters and printed characters that exist between two consecutive item name candidate strings based on the first extraction result information D11 as a series of strings corresponding to the item name candidate string that was extracted earlier among the two item name candidate strings. The identification processing unit 212 also identifies handwritten characters and printed characters that exist after the last item name candidate string as a series of strings corresponding to the last item name candidate string. The identification processing unit 212 identifies handwritten characters and printed characters that exist before the first item name candidate string as a series of strings, but does not associate this series of strings with the item name candidate strings. Furthermore, for the series of strings that exist before the first item name candidate string, "No Candidate" may be associated with the item name candidate string.

[0057] For example, the target image shown in Figure 5 contains the following as candidate strings for item names: "Date of Entry," "Name," and "Address." The actual text corresponding to the candidate string "Date of Entry" is "July 13, 2024." Similarly, the information corresponding to the candidate string "Name" is "Hanako Yamada," and the information corresponding to the candidate string "Address" is "111, △ Mansion, XX Ward, Tokyo." However, the target image shown in Figure 5 contains both handwritten and printed characters in the content of "Date of Entry," which is "July 13, 2024." Therefore, in step S22, the printed "20," the handwritten "24," the printed "Year," the handwritten "7," the printed "Month," the handwritten "13," and the printed "Day" are extracted as individual extracted characters.

[0058] In response, the identification processing unit 212 identifies the handwritten characters and printed characters that exist between "Entry Date" and "Name," which are consecutive item name candidate strings in the first extraction result information D11, as a series of strings corresponding to "Entry Date," which is the earlier item name candidate string in the extraction order. That is, "20," "24," "Year," "7," "Month," "13," and "Day" that exist between "Entry Date" and "Name" are identified as a series of strings corresponding to "Entry Date" in the item name candidate string, "July 13, 2024." In particular, the identification processing unit 212 combines all the strings "20," "24," "Year," "7," "Month," "13," and "Day" into a single extracted character and associates that single extracted character with "Entry Date" in the item name candidate string. The identification processing unit 212 may also combine at least one handwritten character and printed character from the strings "20," "24," "Year," "7," "Month," "13," and "Day" into a single extracted character. For example, "20" and "24" may be combined, or "20", "24", and "year" may be combined. Also, the strings "20", "24", "year", "7", "month", "13", and "day" may be combined into "2024", "July", and "13th", respectively.

[0059] Similarly, the identification processing unit 212 identifies the handwritten and printed characters that exist between the consecutive item name candidate strings "Name" and "Address" in the first extraction result information D11 as a series of strings corresponding to the earlier item name candidate string "Name". That is, "Yamada Hanako" which exists between "Name" and "Address" is identified as the series of strings "Yamada Hanako" corresponding to the item name candidate string "Name".

[0060] Furthermore, the identification processing unit 212 identifies handwritten and printed characters that appear after "Your Address," which is the last candidate item name string in the extraction order of the first extraction result information D11, as a series of strings corresponding to "Your Address," which is the last candidate item name string in the extraction order. That is, "Tokyo XX Ward," "△ Apartment No. 111," and "×× Company," which appear after "Your Address," are identified as a series of strings corresponding to the candidate item name string "Your Address," namely "Tokyo XX Ward △ Apartment No. 111 ×× Company." Note that "×× Company" is not information corresponding to "Your Address" among the series of strings identified here, namely "Tokyo XX Ward △ Apartment No. 111 ×× Company." Therefore, in step S26 described later, "×× Company" is excluded from "Tokyo XX Ward △ Apartment No. 111 ×× Company." Then, the specific processing unit 212 combines the strings "Tokyo, XX Ward", "△ Apartment No. 111", and "×× Company" into a single extracted character, and associates this single extracted character with the item name candidate string "Your Address".

[0061] <Server 20: Step S25> In step S25, the identification processing unit 212 of the control unit 21 performs named entity recognition processing to extract named entities such as names (personal names), addresses (country, city), facilities, organizations, dates, times, amounts, email addresses, and telephone numbers for each of the series of strings identified in step S24.

[0062] The specific processing unit 212 may extract different named entities from multiple strings included in the series of strings. While a detailed explanation of the named entity extraction process is omitted as it is a well-known technique, this process may utilize, for example, a database, a machine learning model, or a database-type machine learning model combining these.

[0063] Specifically, as shown in Figure 7, the specific processing unit 212 determines that the named entity for the extracted text "July 13, 2024" is "Date". The specific processing unit 212 also determines that the named entity for the extracted text "Hanako Yamada" is "Name". Furthermore, as shown in Figure 8, the specific processing unit 212 determines that the named entities for the extracted text "Tokyo", "XX Ward", "△ Mansion", and "No. 111" are "Address", and that the named entity for the extracted text "XX Company" is "Organization".

[0064] <Server 20: Step S26> In step S26, the identification processing unit 212 of the control unit 21 generates recognition result information D12 (see Figure 7) indicating the result of the OCR processing based on the processing results in steps S22 to S25. In the recognition result information D12 shown in Figure 7, extracted characters representing a series of strings recognized by the OCR processing are associated with item names indicating the content of said extracted characters. The recognition result information D12 may also include the "classification" information from the first extraction result information D11. Furthermore, the recognition result information D12 may also include named entity information extracted in the named entity recognition processing.

[0065] Specifically, the identification processing unit 212 identifies the candidate item name string corresponding to the series of strings as the item name of the series of strings if the candidate item name string corresponding to the series of strings is the same. The identification processing unit 212 then stores the extracted characters representing the series of strings in the recognition result information D12, associating them with the item name corresponding to the series of strings.

[0066] For example, as shown in Figure 7, the named entity corresponding to the series of strings "July 13, 2024" is "Date". Also, "Entry Date", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "July 13, 2024", is associated with the item group "Date" in the candidate item information D1. The named entity "Date" and the item group "Date" are the same. Therefore, as shown in Figure 7, the identification processing unit 212 stores the extracted characters "July 13, 2024", which was identified as the series of strings, in the recognition result information D12, associating them with the item name "Entry Date", which is the candidate string for the item name corresponding to the series of strings.

[0067] Similarly, as shown in Figure 7, the named entity corresponding to the series of strings "Yamada Hanako" is "Name". Furthermore, "Name", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "Yamada Hanako", is associated with the item group "Name" in the candidate item information D1. The named entity "Name" and the item group "Name" are the same. Therefore, as shown in Figure 7, the identification processing unit 212 stores the extracted characters "Yamada Hanako", which was identified as the series of strings, in the recognition result information D12, associating them with "Name", the candidate string for the item name corresponding to the series of strings, as the item name.

[0068] Incidentally, the identification processing unit 212 excludes from the series of strings identified in step S24 any extracted characters whose named entity does not satisfy a first specific condition that has been set in advance in relation to the item group corresponding to the series of strings. Specifically, the first specific condition is that the named entity and the item group match, or that the named entity and the item group are similar.

[0069] For example, as shown in Figure 7, in the series of strings "Tokyo, XX Ward, △ Mansion No. 111, XX Company", the named entities corresponding to "Tokyo", "XX Ward", "△ Mansion", and "No. 111" are "Address", and the named entity corresponding to "XX Company" is "Organization". Furthermore, "Your Address", which was identified in step S23 as the candidate string for the item name corresponding to the series of strings "Tokyo, XX Ward, △ Mansion No. 111, XX Company", is associated with the item group "Address" in candidate item information D1. The named entities "Address" corresponding to "Tokyo", "XX Ward", "△ Mansion", and "No. 111" are the same as the "Address" in the item group. On the other hand, the named entity "Organization" corresponding to "XX Company" is different from the "Address" in the item group. Therefore, as shown in Figure 7, the identification processing unit 212 stores "Tokyo XX Ward △ Mansion No. 111 ×× Company", which is obtained by removing "×× Company" from the extracted characters identified as a series of strings, in the recognition result information D12, associating it with the item name candidate string "Your Address", which corresponds to the series of strings.

[0070] As shown in Figure 7, the extracted characters "Application Form" and "XX Company" extracted in step S22 are stored in the recognition result information D12 without being associated with an item name. In other embodiments, it is also conceivable that strings not associated with an item name are not stored in the recognition result information D12. Furthermore, it is also conceivable that strings not associated with the candidate item name strings are stored in the recognition result information D12 associated with an item name such as "No Item Name" to indicate that no corresponding item name exists.

[0071] <Server 20: Step S27> In step S27, the output processing unit 213 of the control unit 21 outputs recognition result information D12. Specifically, the output processing unit 213 transmits the recognition result information D12, which is the processing result of the OCR processing of the image data to be processed, to the image processing device 10, which is the source of the image data. The output processing unit 213 may also transmit the first extraction result information D11 to the image processing device 10 along with the recognition result information D12. In step S27, the output processing unit 213 may display, print, or store the first extraction result information D11 and the recognition result information D12, etc. The first extraction result information D11 and the recognition result information D12, etc. may also be transmitted to the client terminal 30.

[0072] <Image Processing Device 10: Step S13> In step S13, the control unit 11 waits for the recognition result information D12 to be received from the server 20 (S13: No). When it is determined that the recognition result information D12 has been received (S13: Yes), the process moves to step S14.

[0073] <Image Processing Device 10: Step S14> In step S14, the control unit 11 outputs recognition result information D12. Specifically, if the request to execute the automatic input process has been executed, the control unit 11 executes an automatic input process that accepts input of information of pre-set input items in a predetermined database based on the contents of the recognition result information D12. That is, the control unit 11 when executing step S14 is an example of the reception processing unit according to the present invention. For example, in the automatic input process, as information of the input item that matches the item name in the recognition result information D12, the input of information of extracted characters associated with the item name in the recognition result information D12 is accepted. This simplifies the user's work of registering information of characters contained in the target image in the predetermined database.

[0074] In step S14, the image data that was the target of the OCR processing and the recognition result information D12 may be stored in the storage unit 12 in association with each other. Alternatively, in step S14, the recognition result information D12 may be combined with the image data that was the target of the OCR processing. Alternatively, in step S14, the image data that was the target of the OCR processing and the recognition result information D12 may be transmitted to the client terminal 30 either in association with each other or combined.

[0075] As explained above, in the image processing system 1, even if a series of characters, such as a single word or sentence, contains both handwritten and printed characters, the handwritten and printed characters are identified as a single series of characters. Therefore, with the image processing system 1, the user does not need to perform an operation to specify the handwritten and printed characters that constitute the series of characters, thus reducing the burden on the user.

[0076] [Other Examples of Image Processing Methods] Incidentally, there is a known technique for identifying items that indicate the content of handwritten characters contained in an image based on typefaces present around the handwritten characters. However, if there are no typefaces around the handwritten characters, it is not possible to identify the item corresponding to those handwritten characters. On the other hand, among the handwritten characters, the one that contains one of the constituent characters such as numbers or symbols that have been set in advance to correspond to each item, and which has the fewest number of characters that are not included in those constituent characters, may be identified as the handwritten character corresponding to that item. However, with such a configuration, there is a problem that if the handwritten character contains characters other than numbers or symbols, it is not possible to identify the item corresponding to those handwritten characters. In contrast to this, the image processing system 1 according to this embodiment makes it possible to identify items that indicate the content of handwritten characters for handwritten characters extracted from an image, as will be explained below.

[0077] The following describes other examples of the image processing method performed by the image processing system 1, with reference to the flowchart in Figure 8. In Figure 8, processing steps similar to those shown in Figure 4 are denoted by the same reference numerals and their explanations are omitted.

[0078] <Image Processing Device 10: Step S10> First, in step S10, the control unit 11 of the image processing device 10 sets one or more item groups that indicate the item names of the strings to be extracted by OCR processing from the strings included in the image data acquired in step S11. The control unit 11 when executing step S10 is an example of a setting processing unit according to the present invention.

[0079] Specifically, the control unit 11 displays a screen for the user to select a document type stored in the document-specific item information D2, and accepts user input for selecting the document type on the display screen. The document type indicates the content of the image data acquired in step S11. Then, the control unit 11 sets one or more item groups stored in the document-specific item information D2, associated with the document type selected in response to the user input, as item groups corresponding to the image data acquired in step S11.

[0080] In other embodiments, the control unit 11 may display one or more item groups associated with the selected document type in the document-specific item information D2 on the display screen, and set one or more item groups selected by user operation from these item groups as the item group corresponding to the image data. In other embodiments, it is also conceivable that no document type is selected. For example, the control unit 11 may display multiple item groups as selection candidates on the display screen, and set one or more item groups selected by user operation from these item groups as the item group corresponding to the image data.

[0081] Furthermore, the control unit 11 may perform the processing in step S10 during or after the execution of step S11. In addition, the control unit 11 may automatically recognize the document type and item group corresponding to the image data by performing OCR processing or pattern matching processing on the image data acquired in step S11, based on information such as the document name contained in the image data.

[0082] The control unit 11 may set document-specific item information D2 in response to user operation and transmit said document-specific item information D2 to the server 20. This allows the server 20 to perform OCR processing using the document-specific item information D2 received from the image processing device 10. The server 20 may also use the document-specific item information D2 received from the image processing device 10 only in the OCR processing requested by the image processing device 10. The setting process of document-specific item information D2 by the control unit 11 is performed at any timing, such as before, during, or after the execution of step S10 or step S11.

[0083] <Image Processing Device 10: Step S12> When image data is acquired in step S11, in step S12, the control unit 11 associates the acquired image data with one or more item groups set in step S10 and transmits it to the server 20.

[0084] In addition, in step S12, the image data acquired in step S11 and the document type selected in step S10 may be associated with each other and sent to the server 20. As a result, if the document-specific item information D2 is stored in the storage unit 22 of the server 20, the control unit 21 can identify the item group corresponding to the image data based on the document-specific item information D2 and the document type received from the image processing device 10.

[0085] <Server 20: Step S22> As described above, when image data is received in step S21, in step S22, the extraction processing unit 211 of the control unit 21 performs a process to extract character regions and characters present in the target image. The extraction processing unit 211 then generates second extraction result information D21, which associates the extracted characters from the target image with the classification of the extracted characters (handwritten or printed) and the position coordinates of the bounding rectangle of the extracted characters.

[0086] Here, Figure 9 is a diagram showing an example of a target image, and Figure 10 is a diagram showing an example of the second extraction result information D21 extracted from the target image in Figure 9. Note that the dashed line in Figure 9 is not included in the target image, but indicates the character area in the target image. As shown in Figure 10, the handwritten characters extracted sequentially from the upper left to the lower right of the target image in Figure 9 are "00000000-01", "ABCD (Co., Ltd.)", "¥1,650-", "Product price", "2024", "5", "24", "We have received the above correctly", "¥1,500", ... "Chiyoda-ku, Tokyo...", "ABCDEF Building 2nd floor", "XYZ Corporation", "Yamada Hanako", etc. Also, as shown in Figure 10, the printed characters extracted sequentially from the upper left to the lower right of the target image in Figure 9 are "Recipient", "Receipt", "No.", "Mr. / Ms.", "However", etc. In the following, we will explain the processing content of each step, using the case where the second extraction result information D21 in Figure 10 is extracted from the target image in Figure 9 as an example.

[0087] <Server 20: Step S231> After the character area and characters are extracted in step S22, in the following step S231, the identification processing unit 212 of the control unit 21 performs a first identification process to identify the item group corresponding to each of the handwritten characters and printed characters extracted in step S22.

[0088] Specifically, as shown in Figure 11, the identification processing unit 212 identifies, for each type character, the item group corresponding to the candidate string of item name that matches the type character, based on the candidate item information D1 and the second extraction result information D21. For example, "Recipient" in the second extraction result information D21 is assigned the item group "Recipient Name" which corresponds to the candidate string of item name "Recipient" in the candidate item information D1. Similarly, "Receipt" in the second extraction result information D21 is assigned the item group "Document Name" which corresponds to the candidate string of item name "Receipt" in the candidate item information D1. Furthermore, "No." in the second extraction result information D21 is assigned the item group "Serial Number" which corresponds to the candidate string of item name "No." in the candidate item information D1, as shown in Figure 2. In addition, "However" in the second extraction result information D21 is assigned the item group "Transaction Details" which corresponds to the candidate string of item name "However" in the candidate item information D1, as shown in Figure 2.

[0089] Furthermore, based on the second extraction result information D21, the identification processing unit 212 identifies a type character whose positional relationship with the handwritten character satisfies a pre-set second specific condition, and identifies the item group corresponding to that type character as the item group for the handwritten character. For handwritten characters whose positional relationship with the type character does not satisfy the second specific condition based on the second extraction result information D21, the identification processing unit 212 identifies the corresponding item group in steps S232 to S233 described later.

[0090] Specifically, the second specific condition stipulates that the distance between the handwritten character and the type character closest to the handwritten character is less than or equal to a predetermined specific distance. This distance is, for example, the distance between the upper left corner of the handwritten character and the upper left corner of the type character. Alternatively, this distance may be the distance between the handwritten character and the type character at the point where they are closest to each other. Furthermore, the second specific condition may also stipulate that the type character is located in a predetermined direction relative to the handwritten character.

[0091] For example, in the examples shown in Figures 9 and 10, the typeface closest to the handwritten character "00000000-01" is "No.", so the item group "Serial Number" corresponding to "No." is assigned as the item group for the handwritten character "00000000-01". Similarly, the typeface closest to the handwritten character "Product Price" is "However", so the item group "Transaction Details" corresponding to "However" is assigned as the item group for the handwritten character "Product Price".

[0092] <Server 20: Step S232> In step S232, the identification processing unit 212 of the control unit 21 sets a determination group indicating the content of each handwritten character for which the corresponding item group was not identified in the first identification process in step S231.

[0093] Specifically, the identification processing unit 212 first performs a named entity extraction process for each handwritten character for which the corresponding item group was not identified in the first identification process in step S231, in order to extract named entities such as names, addresses, organization names, dates, times, amounts, email addresses, and telephone numbers. The identification processing unit 212 may extract different named entities from multiple characters contained in the handwritten characters. Although a detailed explanation of the named entity extraction process is omitted as it is a well-known technique, the named entity extraction process uses, for example, a database, a machine learning model, or a database-type machine learning model that combines these.

[0094] For example, in the examples shown in Figures 9 and 10, the handwritten characters for which the corresponding item group was not identified in the first identification process in step S231 are "¥1,650", "Chiyoda-ku, Tokyo...", "ABCDEF Building", "XYZ Corporation", and "Hanako Yamada". The identification processing unit 212 then determines that the named entity for one of the handwritten characters, "¥1,650", is "amount". The identification processing unit 212 also determines that for one of the handwritten characters, "Chiyoda-ku, Tokyo...", the named entities for "Tokyo", "Chiyoda-ku", and "Ichibancho" are "place names", and the named entity for "1-2-3" is "address". Furthermore, the identification processing unit 212 determines that for one of the handwritten characters, "ABCDEF Building 2nd floor", the named entity for "ABCDEF Building" is "facility name", and the named entity for "2nd floor" is "floor". Similarly, the specific processing unit 212 determines that the unique identifier of one of the handwritten characters, "XYZ Corporation," is "organization," and that the unique identifier of one of the handwritten characters, "Hanako Yamada," is "name."

[0095] Next, the specific processing unit 212 sets the determination group for each of the handwritten characters based on the named entity corresponding to each of the handwritten characters. Specifically, the specific processing unit 212 sets the determination group for each of the handwritten characters based on the named entity corresponding to each of the handwritten characters and a predetermined determination rule.

[0096] For example, the determination rule may include a condition for identifying the determination group based solely on the named entity corresponding to the handwritten character. Alternatively, the determination rule may include a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the content of the handwritten character. Furthermore, the determination rule may include a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the item group associated with the printed characters surrounding the handwritten character. Note that the determination rule may include one or more of the above conditions. Now, a specific example of the determination rule will be described.

[0097] The determination rule includes that when the named entity of the handwritten character is "amount of money", a determination group of "amount of money" is set for the handwritten character. For example, the named entity of the handwritten character "¥1,650" is "amount of money". Therefore, as shown in FIG. 12, a determination group of "amount of money" is set for the handwritten character "¥1,650".

[0098] The determination rule includes that when the named entity of the handwritten character is "organization" and the handwritten character includes the character "company", a determination group of "company name" is set for the handwritten character. For example, the named entity of the handwritten character "XYZ Inc." is "organization", and "XYZ Inc." includes the character "company". Therefore, as shown in FIG. 12, a determination group of "company name" is set for the handwritten character "XYZ Inc.".

[0099] The determination rule includes that when the named entity of the handwritten character is "personal name", a determination group of "name" is set for the handwritten character. For example, the named entity of the handwritten character "Yamada Hanako" is "personal name". Therefore, as shown in FIG. 12, a determination group of "name" is set for the handwritten character "Yamada Hanako".

[0100] The determination rule includes that when two or more consecutive character strings (or three or more consecutive character strings) whose named entity is "place name" appear in the handwritten character, a determination group of "address" is set for the handwritten character. For example, for the handwritten character "Chiyoda-ku, Tokyo ...", the named entities of "Tokyo", "Chiyoda-ku" and "Ichiban-cho" are all "place name", and two or more consecutive character strings whose named entity is "place name" are present. Therefore, as shown in FIG. 12, a determination group of "address" is set for the handwritten character "Chiyoda-ku, Tokyo ...".

[0101] The aforementioned determination rule includes the setting of the "address" determination group for handwritten characters if the named entity of the handwritten characters is "facility name" or "floor," and the item group or determination group corresponding to the closest handwritten character or printed character in any direction (up, down, left, or right) from the handwritten character is "address." For example, for the handwritten characters "ABCDEF Building 2nd Floor," the named entity of "ABCDEF Building" is "facility," and the named entity of "2nd Floor" is "floor." Furthermore, the determination group for the handwritten characters closest to "ABCDEF Building 2nd Floor," which is "Chiyoda-ku, Tokyo...", is "address." Therefore, as shown in Figure 12, the "address" determination group is set for the handwritten characters "ABCDEF Building 2nd Floor."

[0102] <Server 20: Step S233> In step S233, the identification processing unit 212 of the control unit 21 performs a second identification process to identify the item group corresponding to each of the handwritten characters for which the corresponding item group has not been identified in the first identification process, based on the determination group set for the handwritten character. For example, the identification processing unit 212 identifies the item group corresponding to the handwritten character based on the determination group corresponding to the handwritten character and the item group associated with the image data to be processed, using rule-based processing or the like.

[0103] Specifically, the identification processing unit 212 excludes from the identification target in the second identification process any item groups associated with the image data to be processed for which the corresponding handwritten character has already been identified in the second extraction result information D21. As a result, in the second identification process, only the remaining unspecified item groups, excluding the item groups associated with the image data to be processed for which the corresponding handwritten character has already been identified, are targeted for identification. Then, in the second identification process, the identification processing unit 212 identifies the item groups among the unspecified item groups that include the character of the determination group corresponding to the handwritten character as the item group corresponding to the handwritten character. As a result, even if there are multiple item groups that include the character of the determination group, in the second identification process, only the unspecified item groups are targeted for identification, increasing the likelihood that the item group can be identified based on the determination group.

[0104] For example, if the determination group corresponding to the handwritten text "¥1,650" is "Amount," and the unspecified item group includes "Receipt Amount," then, as shown in Figure 13, "Receipt Amount," which includes the word "Amount," is identified as the item group corresponding to "¥1,650." Also, if the determination group corresponding to the handwritten text "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor" is "Address," and the unspecified item group includes "Document Creator Address," then, as shown in Figure 13, "Document Creator Address," which includes the word "Address," is identified as the item group corresponding to "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor," respectively. Furthermore, the identification processing unit 212 may combine "Chiyoda-ku, Tokyo..." and "ABCDEF Building 2nd Floor," which belong to the same item group, into a single extracted character.

[0105] Similarly, if the determination group corresponding to the handwritten text "XYZ Corporation" is "Company Name," and the aforementioned unspecified item group includes "Document Creator Company Name," then, as shown in Figure 13, "Document Creator Company Name," which includes the characters "Company Name," is identified as the item group corresponding to "XYZ Corporation." Also, if the determination group corresponding to the handwritten text "Hanako Yamada" is "Name," and the aforementioned unspecified item group includes "Document Creator Name," then, as shown in Figure 13, "Document Creator Name," which includes the characters "Name," is identified as the item group corresponding to "Hanako Yamada."

[0106] Furthermore, the identification processing unit 212 has explained the case in which, in the rule-based processing, etc., the unspecified item group containing the character of the determination group corresponding to the handwritten character is identified as the item group corresponding to the handwritten character. On the other hand, the identification processing unit 212 may identify the item group corresponding to the handwritten character based on correspondence information that associates one or more of the determination groups with the item group.

[0107] Furthermore, in other embodiments, in the second identification process, all item groups associated with the image data to be processed may be identified. That is, the identification processing unit 212 may identify an item group that contains a character from a determination group corresponding to a handwritten character, among the item groups associated with the image data to be processed, for which the corresponding item group has not been identified in the first identification process, as the item group corresponding to the handwritten character. In this case as well, if there is only one item group containing a character from the determination group, that item group will be identified as the item group corresponding to the handwritten character.

[0108] <Server 20: Step S26> In step S26, the identification processing unit 212 of the control unit 21 generates recognition result information D22, which indicates the result of the OCR processing, based on the second extraction result information D21. Specifically, as shown in Figure 14, the recognition result information D22 associates the extracted characters extracted from the target image with the position coordinates of the extracted characters and the item group of the extracted characters. The recognition result information D22 may include either or both of the "classification" and "judgment group" information from the second extraction result information D21. The recognition result information D22 may not include "position coordinate" information. Subsequently, the recognition result information D22 is transmitted to the image processing device 10 in step S27. Then, the control unit 11 of the image processing device 10 outputs the recognition result information D22 in step S14, similar to the recognition result information D12. For example, if the request to execute the automatic input process has been made, the control unit 11 executes an automatic input process that accepts the input of information for pre-set input items in a predetermined database based on the contents of the recognition result information D22. That is, the control unit 11 when step S14 is executed is an example of the reception processing unit according to the present invention.

[0109] As explained above, the image processing system 1 can identify the item group corresponding to handwritten characters, even if the corresponding item group could not be identified based on the surrounding printed characters, based on the named entity contained in the handwritten characters. In particular, even if the format of the document to be processed is not specified, the image processing system 1 can identify the item group corresponding to the handwritten characters based on the item group and the named entity, provided that the item groups contained in the document are identified.

[0110] [Multi-line character extraction function] Incidentally, there is a known technique for identifying the line with the lowest pixel frequency as the boundary for handwritten characters written on multiple lines by expanding and binarizing the handwritten characters. However, in the case of handwritten characters freely written by the user, it is conceivable that the string composed of handwritten characters may be written in a manner such as sloping upwards to the right or downwards to the right. However, when there is a slope in the string, the aforementioned technique has the problem of having low accuracy in identifying the line boundaries.

[0111] In contrast, the image processing system 1 according to this embodiment can identify handwritten characters belonging to each line with high accuracy, even when a multi-line string of handwritten characters is tilted diagonally, as described below.

[0112] Specifically, in the character extraction process in step S22, the image processing system 1 has the specific processing unit 212 of the control unit 21 of the server 20 execute the multi-line extraction process described later. Note that the multi-line extraction process may be executed at a different timing than step S22. Furthermore, the multi-line extraction process is not limited to the control unit 21 of the server 20, but may also be executed by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc.

[0113] The following describes an example of the multi-line extraction process, referring to the flowchart in Figure 16. Furthermore, the following explanation will use the example shown in Figure 15, where the "Address" field in the target image, which is the image data subject to the multi-line extraction process, contains two lines of handwritten text: "Tokyo, XX Ward, △ Apartment, Room 111".

[0114] <Server 20: Step S221> In step S221, the extraction processing unit 211 of the control unit 21 extracts character regions present in the target image indicated by the image data, based on the image data to be processed. Note that step S221 is an example of the first step in the present invention, and the extraction processing unit 211 that executes step S221 is an example of the first processing unit according to the present invention.

[0115] Specifically, in the example shown in Figure 15, multiple character regions, including character region A1, are extracted. Note that the extraction process in step S221 can utilize conventional optical character recognition technologies, such as neural networks; therefore, a detailed explanation of these technologies is omitted here.

[0116] <Server 20: Step S222> In step S222, the extraction processing unit 211 of the control unit 21 extracts the characters present in each of the character areas extracted in step S221. Note that the extraction processing unit 211 when executing step S222 is an example of the first processing unit according to the present invention.

[0117] Specifically, the extraction processing unit 211 detects the circumscribing rectangle (dotted line) of each character contained in character area A1, as shown in character area A1 of Figure 15. Note that, similar to step S221, the extraction process in step S222 can utilize conventional optical character recognition technology, such as a neural network; therefore, its explanation is omitted here.

[0118] <Server 20: Step S223> In step S223, the extraction processing unit 211 of the control unit 21 determines whether the string in each character area extracted in step S221 is written vertically or horizontally, based on the aspect ratio of each character area. This makes it possible for the image processing system 1 to automatically determine the writing direction of the characters contained in each character area, regardless of whether the writing direction is horizontal or vertical. The writing direction is perpendicular to the specific direction used to determine the number of lines of characters in the character area in step S224, described later, and the extraction processing unit 211 determines the specific direction perpendicular to the writing direction. Note that the extraction processing unit 211 when executing step S223 is an example of the fourth processing unit according to the present invention.

[0119] Specifically, the control unit 21 determines that the text in the text area is written horizontally if the horizontal size w of the text area is greater than the vertical size h. On the other hand, the control unit 21 determines that the text in the text area is written vertically if the vertical size h of the text area is greater than the horizontal size w. For example, in the target image shown in Figure 15, the horizontal size w of text area A1 is greater than the vertical size h. Therefore, in step S223, it is determined that the text in text area A1 is written horizontally.

[0120] In other embodiments, the extraction processing unit 211 may accept a user operation to select whether the writing direction of the characters in the character area is horizontal or vertical, and determine the writing direction of the characters according to the user operation. For example, the extraction processing unit 211 may accept a selection operation for each of the character areas to determine whether the writing direction of the characters is horizontal or vertical individually. Alternatively, the extraction processing unit 211 may accept a selection operation to select whether the writing direction of the characters in all the character areas in the target image is horizontal or vertical at once.

[0121] <Server 20: Step S224> In step S224, the extraction processing unit 211 of the control unit 21 performs a process to determine the number of lines of characters in each of the character areas. Specifically, the extraction processing unit 211 determines the number of lines of characters in the character area based on the writing direction of the characters in the character area (horizontal or vertical) and the number of characters in a specific direction orthogonal to the writing direction within the character area. Note that the extraction processing unit 211 when executing step S224 is an example of the second processing unit according to the present invention.

[0122] Specifically, for character regions where the characters are written horizontally, the extraction processing unit 211 counts the number of characters present in the vertical direction (an example of a specific direction) of the character region at multiple detection positions in the horizontal direction of the character region. For example, the extraction processing unit 211 counts the number of bounding rectangular areas of characters present in the vertical direction at each of the detection positions within the character region as the number of characters present in the vertical direction of the character region.

[0123] Here, as shown in Figure 16, the character area A1 includes a total of nine detection positions in the width direction. Specifically, the character area A1 includes three first positions located at "2w / 10", "5w / 10", and "8w / 10" from the left edge in the width direction. In addition, the detection positions in the width direction of the character area A1 include a total of six second positions located at "±0.5w" from each of the first positions in the width direction of the character area A1.

[0124] In the example shown in Figure 17, the character counts for the characters present vertically in each of the first and second positions are "2, 0, 2, 0, 2, 0, 1, 0, 1" from left to right. The extraction processing unit 211 then identifies the maximum value among these counts as the number of lines of characters in the character area. In the example shown in Figure 17, since the maximum value of the count in character area A1 is "2", the extraction processing unit 211 identifies the number of lines of horizontally written characters in character area A1 as "2".

[0125] If it is determined in step S223 that the writing direction is vertical, then the horizontal and vertical directions in the specific example for horizontal writing described here should be swapped, and a detailed explanation of that case will be omitted.

[0126] <Server 20: Step S225> In step S225, if the number of lines of characters in the character area identified in step S224 is multiple (S225: Yes), the extraction processing unit 211 proceeds to step S226. On the other hand, if the number of lines of characters in the character area is not multiple (S225: No), the processing from step S226 onwards is not executed. The determination process in step S225 is executed sequentially for each of the character areas until all of the character areas are targeted. That is, whether or not to execute the processing in steps 226 to S228 is determined for each character area.

[0127] <Server 20: Step S226> In step S226, the extraction processing unit 211 performs a sampling process in which it samples a predetermined number of black pixels from the black pixels that constitute each character included in the character area. For example, the number of samples is four or five per character.

[0128] Specifically, Figures 18(A) and 18(B) show the sampling results when the sampling process is performed on the characters in character area A shown in Figure 15. In this way, by performing the sampling process on the characters, the influence of the shape of each character is suppressed, and the accuracy of the clustering process for each row in step S227 described later is improved. In other embodiments, the sampling process may be omitted, and the clustering process in step S227 described later may be performed on the black pixels that constitute the characters.

[0129] <Server 20: Step S227> In step S227, the extraction processing unit 211 identifies the characters belonging to each line in the character area based on the sampling results of step S226. In particular, in steps S225 to S227, the number of lines in the character area identified in step S224 is used as the number of lines of characters contained in the character area, and the characters contained in each line of that number of lines in the character area are identified. Note that the extraction processing unit 211 when executing steps S225 to S227 is an example of the third processing unit according to the present invention.

[0130] Specifically, the extraction processing unit 211 performs a clustering process that clusters the data using a cluster number k corresponding to the number of lines of characters in the character area. Specifically, as shown in Figure 15, if the number of lines of characters in character area A is "2", then the cluster number k in the clustering process is "2".

[0131] In this explanation, we will use the case where the clustering process is performed on the character area A shown in Figure 15, and the k-means method (k-average method), which is a type of non-hierarchical clustering, is used for the clustering process. On the other hand, in other embodiments, the clustering process is not limited to the k-means method; for example, the x-means method, k-medoids method, spectral clustering method, etc., may be used.

[0132] First, if the number of clusters k is 2, the extraction processing unit 211 sets the centroid positions of the two clusters, the first cluster and the second cluster. Specifically, as shown in Figure 18(B), the upper end of the character area A, which is the center in the width direction of the character area A, is set as the initial value of the centroid position P11 of the first cluster (see reference). Similarly, as shown in Figure 18(B), the control unit 21 sets the lower end of the character area A, which is the center in the width direction of the character area A, as the initial value of the centroid position P12 of the second cluster.

[0133] Next, the extraction processing unit 211 performs a classification process to classify the sampled data close to the centroid position P11 into a first cluster and the sampled data close to the centroid position P12 into a second cluster. Subsequently, the extraction processing unit 211 performs an update process to update the centroid positions P11 and P12 corresponding to each of the clusters. Specifically, the extraction processing unit 211 calculates the centroid position of the sampled data belonging to the first cluster and sets the calculation result as the new centroid position P11. Similarly, the extraction processing unit 211 calculates the centroid position of the sampled data belonging to the second cluster and sets the calculation result as the new centroid position P12. The extraction processing unit 211 then repeatedly performs the classification process and the update process until the centroid positions P11 and P12 no longer change in the update process.

[0134] Subsequently, when it is determined that the centroid positions P11 and P12 have not changed in the update process, the extraction processing unit 211 calculates the average value of the vertical (Y-axis) coordinates of the sampled data classified into each cluster. Then, the extraction processing unit 211 identifies the row number corresponding to each cluster, starting with the cluster with the smallest average value. Specifically, if the number of clusters k is "2", it identifies that the cluster with the smaller average value corresponds to the first row, and the cluster with the larger average value corresponds to the second row.

[0135] The extraction processing unit 211 then identifies each character corresponding to the sampled data as a string of characters in the row corresponding to the cluster to which the sampled data belongs. Specifically, in the example shown in Figure 18(B), "Tokyo XX Ward" is identified as the first row of data corresponding to the first cluster, and "△ Apartment No. 111" is identified as the second row of data corresponding to the second cluster. For example, for each character, the extraction processing unit 211 identifies it as a string of characters belonging to the first cluster if the sampled data corresponding to that character mostly belongs to the first cluster. Similarly, for each character, the extraction processing unit 211 identifies it as a string of characters belonging to the second cluster if the sampled data corresponding to that character mostly belongs to the second cluster.

[0136] It is also possible that the number of clusters k, which is the number of lines of characters included in the character area, is 3 or more. In this case, the extraction processing unit 211 divides the character area vertically into N equal parts (N = number of clusters k - 1), and sets N positions, including the centers at the top and bottom ends of the character area and the centers of the imaginary lines that divide the character area into N equal parts, as centroid positions corresponding to the N clusters. Subsequently, the extraction processing unit 211 repeatedly executes the classification process and the update process as described above until the centroid positions no longer change. Then, the extraction processing unit 211 identifies each of the N clusters as the string of characters in the line closest to the first line in the character area, in ascending order of the average value of the vertical coordinates of the sampled data classified into that cluster.

[0137] <Server 20: Step S228> Once the characters belonging to each row in the character area are identified in step S227, in the following step S228, the extraction processing unit 211 executes an image data generation process to generate image data containing only the characters corresponding to each row of characters in the character area.

[0138] Specifically, as shown in Figure 18(C), the extraction processing unit 211 converts other characters in the second line of character area A into white pixels and generates image data of an image containing only the characters in the first line. Similarly, as shown in Figure 18(D), the extraction processing unit 211 converts other characters in the first line of character area A into white pixels and generates image data of an image containing only the characters in the second line.

[0139] <Server 20: Step S229> In step S229, the extraction processing unit 211 determines whether the processing in steps S225 to S228 has been completed for all of the character regions included in the target image. If it is determined that the processing in steps S225 to S228 has been completed for all of the character regions included in the target image (S229: Yes), the processing moves to step S230. If it is determined that the processing in steps S225 to S228 has not been completed for all of the character regions included in the target image (S229: No), the extraction processing unit 211 takes the next character region after the currently processed character region as the processing target and moves the processing to step S225.

[0140] <Server 20: Step S230> In step S230, the extraction processing unit 211 recognizes the content of the characters contained in the target image for each character region in the target image, based on the results of identifying the characters contained in each line in step S227. Specifically, the extraction processing unit 211 recognizes the content of the characters contained in each line based on the image data of each line generated in step S228. Note that, as with step S221, etc., conventional optical character recognition technology using, for example, a neural network can be used in step S230, so its explanation is omitted here.

[0141] In particular, since the characters of each line are identified in step S227, the extraction processing unit 211 can recognize the order of the characters constituting the string in that line based on the X coordinates of the characters in that line. As a result, the recognition result information D12 (D22) generated in step S26 reflects the character recognition results in the multi-line extraction process, and the recognition result information D12 (D22) is output in step S27.

[0142] As explained above, the image processing system 1 can recognize characters in each line with high accuracy, even when the characters contained in each character area of ​​the target image span multiple lines. In particular, since the clustering process identifies the characters contained in each line of the character area in the image processing system 1, it is possible to recognize the characters contained in each line with high accuracy, even if there is a slant in the characters of each line contained in the character area.

[0143] In this embodiment, the example given was that the multi-line extraction process is performed for each character region included in the target image. On the other hand, the extraction processing unit 211 may perform the multi-line extraction process only for the character regions that include handwritten characters.

[0144] [Text Identification Function] However, when the text contained in the image data to be processed is a multi-line sentence, in general OCR processing, line breaks may exist between each line of the recognized sentence, and the sentence, which should be continuous, may be recognized in a split manner. In contrast, the image processing system 1 according to this embodiment can identify that a multi-line string of characters is a continuous sentence, as will be explained below.

[0145] Specifically, in the character extraction process in step S22, the image processing system 1 has the identification processing unit 212 of the control unit 21 of the server 20 execute the text identification process described later. Note that the text identification process may be executed at a different timing than step S22. Furthermore, the text identification process is not limited to the control unit 21 of the server 20, but may also be executed by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc.

[0146] The following describes an example of the text identification process, referring to the flowchart in Figure 19. Here, as shown in Figure 20(A), the text contained in the target image A21 has two line breaks and contains three sets of text A211 to A213. The text identification process is executed when the control unit 21 of the server 20 receives image data to be processed from the image processing device 10 or the client terminal 30.

[0147] <Server 20: Step S31> In step S31, the extraction processing unit 211 of the control unit 21 extracts character regions present in the target image indicated by the image data, based on the image data to be processed. Note that the extraction process in step S31 can be performed using conventional optical character recognition technology, such as a neural network, so its explanation is omitted here.

[0148] <Server 20: Step S32> In step S32, the extraction processing unit 211 of the control unit 21 performs character extraction and recognition of the characters present in each of the character regions extracted in step S31. Note that, as with step S31, the extraction process in step S32 can utilize conventional optical character recognition technology, such as a neural network, so its explanation is omitted here.

[0149] Specifically, as shown in Figure 20(B), the text area A22 extracted from the target image A21 (see Figure 20(A)) contains six line breaks in the entire text and includes seven sets of sentences A221 to A227.

[0150] Furthermore, in step S32, the extraction processing unit 211 can distinguish and extract handwritten characters and printed characters present in the character area, similar to step S22 (see Figure 4). In this case, the extraction processing unit 211 may perform the processes of steps S33 to S38 described later only if handwritten characters are present in the character area.

[0151] <Server 20: Step S33> In step S33, the identification processing unit 212 of the control unit 21 determines whether the characters contained in the character area are in sentence form. If the characters contained in the character area are in sentence form (S33: Yes), the identification processing unit 212 moves the process to step S34, and if the characters contained in the character area are not in sentence form (S33: No), the process moves to step S36. The processing from step S33 onward is performed for each of the character areas extracted in step S31, and is performed sequentially for each character area until all of the character areas are targeted (S39: No).

[0152] Specifically, in step S33, the specific processing unit 212 determines whether the characters in the character area constitute a sentence based on the content of the characters contained therein. For example, the specific processing unit 212 decomposes the characters in the character area into parts of speech and determines whether nouns and verbs are included in the character area at least a predetermined number of times. The predetermined number of times is a number set in advance, such as 1 to 3 times. The specific processing unit 212 determines that the characters in the character area constitute a sentence if nouns and verbs are included at least a predetermined number of times each. On the other hand, the specific processing unit 212 determines that the characters in the character area do not constitute a sentence if the number of nouns and verbs in the character area is less than the predetermined number of times. The method for determining whether the characters in the character area constitute a sentence is not limited to this, and various conventionally known techniques may be used.

[0153] Furthermore, in step S33, the specific processing unit 212 may determine whether the text-formatted characters included in the character area span multiple lines. In this case, the specific processing unit 212 proceeds to step S34 if the characters included in the character area are text and the text spans multiple lines, and proceeds to step S36 if the characters do not span multiple lines. For example, the specific processing unit 212 can determine whether the characters span multiple lines by performing the same processing as in steps S223 to S224 (see Figure 16) in the multi-line extraction process described above.

[0154] <Server 20: Step S34> In step S34, the identification processing unit 212 of the control unit 21 determines whether there is a relationship between two consecutive lines of text extracted by the extraction processing unit 211, and if it is determined that there is a relationship, it performs an identification process to identify the two lines of text as a series of texts.

[0155] Specifically, in step S34, the specific processing unit 212 determines whether there is a contextual connection between two consecutive lines of text contained in the character area. For example, if there are three lines of text contained in the character area, the control unit 21 determines whether there is a contextual connection between the first and second lines of text contained in the character area, and also determines whether there is a contextual connection between the second and third lines of text contained in the character area. In other embodiments, the determination of whether there is a contextual connection between two lines of text may be made based on more lines of text than the two lines of text contained in the character area.

[0156] More specifically, the specific processing unit 212 determines whether or not there is a relationship based on the meaning of the sentence when the two lines of text are joined together. For example, the specific processing unit 212 may determine whether or not there is a contextual connection between the two lines of text depending on whether or not the two lines of text become a meaningful sentence when joined together. Alternatively, the specific processing unit 212 may compare the case where the two lines of text are joined with the case where the two lines of text are not joined and determine whether or not there is a contextual connection between the two lines of text depending on which becomes a meaningful sentence.

[0157] For example, in step S34, the specific processing unit 212 uses the Next Sentence Prediction (NSP) of BERT (Bidirectional Encoder Representations from Transformers) to determine whether two consecutive lines of text have a contextual connection. In particular, it is desirable that the NSP be pre-trained using a dataset that includes both contextually connected positive examples and contextually unconnected negative examples. This makes the NSP specialized in determining whether two consecutive lines of text have a contextual connection. Therefore, the accuracy of determining whether two consecutive lines of text have a contextual connection is higher compared to predicting whether the second line of two consecutive lines of text is logically or semantically connected to the first line using a general NSP. Note that other methods may be used to determine whether two consecutive lines of text have a contextual connection, not just the NSP.

[0158] Then, if the specific processing unit 212 determines that there is a contextual connection between two consecutive lines of text, it identifies the two consecutive lines of text as a single sentence and combines them. For example, if it is determined in step S34 that the strings of the first and second lines have a contextual connection, the specific processing unit 212 identifies the strings of the first and second lines as a single sentence and combines the strings of the first and second lines as a single sentence. Also, if it is determined in step S34 that the strings of the second and third lines have a contextual connection, the control unit 212 identifies the strings of the second and third lines as a single sentence and combines the strings of the second and third lines as a single sentence. On the other hand, if it is determined in step S34 that there is no contextual connection between the strings of the third and fourth lines, the strings of the third and fourth lines are not combined.

[0159] For example, in Figure 20(B), the first line of sentence A221, "I found this product very interesting," ends with the word "interesting," and the second line of sentence A222, "I made a discovery," begins with the word "discovery." Therefore, if a method were used to determine whether or not words are split before and after a line break and then combine sentences A221 and A222, sentences A221 and A222, which do not have words split before and after a line break, would not be combined. In contrast, the identification processing unit 212 determines whether or not there is a contextual connection between two consecutive lines of sentences. Therefore, even though sentences A221 and A222, which do not have words split before and after a line break, are determined to have a contextual connection, they are identified as a series of sentences and combined, as shown in Figure 20(C). Similarly, in the text area A22 shown in Figure 20(B), the sentence A232, which is formed by combining the sentences from the 3rd to the 6th lines, is identified and combined as a series of sentences, and the sentence A233, which is the sentence from the 7th line, is identified as a series of sentences.

[0160] <Server 20: Step S35> In step S35, the control unit 21 performs punctuation correction processing to correct the text within the character area identified in step S34, such as adding or deleting punctuation marks. For example, a pre-trained model using BERT (Bidirectional Encoder Representations from Transformers) MLM (Masked Language Modeling) is used for the punctuation correction processing.

[0161] For example, in step S35, a comma is added to the phrase "I like the tap function" in the sentence A232 of the text area A22 in Figure 20(C). As a result, as shown in Figure 20(D), sentence A232 is corrected to the same as the target image A21, "I like the tap function. Function". Also, a comma is added to the phrase "It is also well-equipped and very useful in daily life" in the sentence A232 of the text area A22 in Figure 20(C). As a result, as shown in Figure 20(D), sentence A232 is corrected to the same as the target image A21, "It is also well-equipped and very useful in daily life".

[0162] <Server 20: Step S36> In step S36, the specific processing unit 212 performs morphological analysis on the characters in the character area that were determined not to be in sentence form in step S33, and subdivides the characters into morphological strings of the smallest units (morphemes) that have meaning in language.

[0163] <Server 20: Step S37> In step S37, the specific processing unit 212 extracts a named entity from each of the morphological strings obtained in the morphological analysis in step S36, and associates the named entity label corresponding to the named entity with the morphological string. For example, the named entity "country, city" is extracted from the morphological string representing a prefecture, and the named entity label "GPE" corresponding to the named entity "country, city" is associated with the morphological string. Also, the named entity "name" is extracted from the morphological string representing a name, and the named entity label "PERSON" corresponding to the named entity "name" is associated with the morphological string.

[0164] <Server 20: Step S38> In step S38, the identification processing unit 212 identifies and combines a plurality of morphological strings, which are consecutive morphological strings, and whose named entity labels extracted in step S37 satisfy the pre-set combination conditions, as a series of strings.

[0165] For example, in the aforementioned combination condition, it is conceivable that the named entity labels "GPE" corresponding to the morphological string indicating a prefecture, "LOC" corresponding to the morphological string indicating a city or town name, and "FAC" corresponding to the morphological string indicating a building are set as the combination targets. In this case, if the string on the first line is "Tokyo XX Ward" and the string on the second line is "△ Apartment No. 111", the named entity labels of the strings on the first and second lines become "GPE" and "FAC", and the combination condition is satisfied. Therefore, the strings on the first and second lines are identified as a series of strings and combined into the string "Tokyo XX Ward △ Apartment No. 111". On the other hand, if the string on the third line is "Yamada Hanako", the named entity labels of the strings on the second and third lines become "FAC" and "PERSON", and the combination condition is not satisfied. Therefore, the strings on the second and third lines are not identified as a series of strings and are not combined.

[0166] In this embodiment, the case in which steps S34 and S35 to S38 are performed has been described. On the other hand, in other embodiments, steps S34 or S35 to S38 may be omitted.

[0167] <Server 20: Step S39> In step S39, the extraction processing unit 211 determines whether the processing from step S33 onwards has been completed for all of the character regions included in the target image. If it is determined that the processing from step S33 onwards has been completed for all of the character regions included in the target image (S39: Yes), the processing moves to step S40. If it is determined that the processing from step S33 onwards has not been completed for all of the character regions included in the target image (S39: No), the extraction processing unit 211 takes the next character region after the currently processed character region as the processing target and moves the processing to step S33.

[0168] <Server 20: Step S40> In step S40, the output processing unit 213 of the control unit 21 performs output processing to output each of the sentences in the character area after the execution of steps S34 to S35, or each of the strings after the execution of steps S36 to S38, as recognition result information indicating the result of the sentence identification process. The output methods of the recognition result information include display, printing, transmission, or storage. For example, if the image data that was the target of the recognition result information was transmitted from the image processing device 10 or the client terminal 30, the recognition result information is transmitted to the image processing device 10 or the client terminal 30, which is the source of the image data. In the image processing device 10, the control unit 11 performs processing to output the recognition result information. Specifically, the control unit 11 performs automatic input processing to accept input of information of pre-set input items in a predetermined database based on the content of the recognition result information. That is, the control unit 11 when performing such processing is an example of the reception processing unit according to the present invention. For example, in the automatic input processing, the sentence shown in the recognition result information is input as the information of the input item. This simplifies the user's task of registering information about the characters contained in the target image in the predetermined database. The control unit 11 may also store the image data that was the target of OCR processing and the recognition result information in the storage unit 12 in association with each other. The control unit 11 may also combine the recognition result information with the image data that was the target of OCR processing. The control unit 11 may also transmit the image data that was the target of OCR processing and the recognition result information, either in association with each other or combined, to the client terminal 30.

[0169] As explained above, the image processing system 1 can identify that a multi-line string of characters is actually a continuous sentence, and can obtain useful recognition result information.

[0170] [Summary Output Function] Incidentally, images of various documents may contain not only printed text but also diagrams such as tables or graphs. Furthermore, handwritten annotations such as handwritten annotation characters or handwritten annotation figures may be added to these diagrams. In response to this, the image processing system 1 according to this embodiment can generate summaries that take into account the content of the diagrams, as will be explained below.

[0171] Specifically, in the image processing system 1, the control unit 21 of the server 20 executes the summary output processing described later. For example, the extraction processing unit 211 extracts handwritten annotations, including either or both handwritten annotation characters and handwritten annotation figures, and components of charts and graphs included in the target image indicated by the image data to be processed. The output processing unit 213 generates and outputs summary information D32 based on the handwritten annotations extracted by the extraction processing unit 211 and the target components among the components that are the subject of the handwritten annotations. Furthermore, the output processing unit 213 generates and outputs summary information D32 based on the handwritten annotations, the target components, and related components among the components that are associated with the target components.

[0172] Furthermore, the summary output processing may be performed not only by the control unit 21 of the server 20, but also by the control unit 11 of the image processing device 10 or the control unit 31 of the client terminal 30, etc. Also, the summary output processing may be divided and performed by two or three of the control units 11 of the image processing device 10, the control unit 21 of the server 20, and the control unit 31 of the client terminal 30.

[0173] Hereinafter, with reference to Figure 21, an example of the processing procedure in the summary output method, which is executed as one of the image processing methods in the image processing system 1 according to this embodiment, will be described.

[0174] In Figure 21, steps S111, S112, etc., represent the processing procedure (step) numbers of the processing performed by the control unit 11 of the image processing device 10 in the summary output method. Specifically, the control unit 11 of the image processing device 10 starts processing when it receives a request from the user to execute the image reading process that includes the summary output process.

[0175] Furthermore, steps S41, S42, etc. in Figure 21 represent the processing procedure (step) of the summary output process executed by the control unit 21 of the server 20 in the summary output method. Specifically, the control unit 21 starts the summary output process when the power to the server 20 is turned on.

[0176] <Image Processing Device 10: Step S111> In step S111, the control unit 11 of the image processing device 10 controls the image acquisition unit 15 to perform an image acquisition process to acquire image data to be processed by the summary output process. Specifically, the control unit 11 controls the image acquisition unit 15 to perform the image reading process to acquire image data corresponding to the image of the original document. Alternatively, in step S111, the control unit 11 may acquire image data to be processed by the summary output process by reading image data stored in the storage unit 12 in response to user operation.

[0177] <Image Processing Device 10: Step S112> In step S112, the control unit 11 of the image processing device 10 transmits the image data acquired in step S111 to the server 20 as the target of the summary output process.

[0178] <Server 20: Step S41> In step S41, the control unit 21 of the server 20 waits for the reception of image data to be processed by the summary output process (S41: No). Specifically, the control unit 21 receives the image data to be processed by the summary output process from the image processing device 10 or the client terminal 30. When it is determined that the image data has been received (S41: Yes), the process moves to step S42.

[0179] The following describes the case where the image data corresponding to the target image P10 shown in Figure 22 is acquired as the target of the summary output process. Specifically, the target image P10 shown in Figure 22 includes a string of characters written in type and a table, which is an example of a diagram. The components of the table shown in Figure 22 include row names such as "notebook," "ballpoint pen," "eraser," and "pencil," and column names such as "product name," "sales price [yen]," "quantity [units]," "cost [yen]," "sales [yen]," "sales area," and "gross profit [yen]," as well as non-elemental data. The components of the table shown in Figure 22 also include elemental data entered into each cell of the table. Then, in the summary output process, the output processing unit 213 generates summary information D32 that includes the elemental data and the non-elemental data.

[0180] Furthermore, the target image P10 includes a handwritten annotation figure C1 that annotates the string of characters shown in print, "Expand sales area," and a handwritten annotation character C2 that indicates the content of the annotation by the handwritten annotation figure C1, which is "Southeast Asia, etc."

[0181] Furthermore, the target image P10 includes handwritten annotation shapes C11 and C12 that annotate the row name "Pencil," which is non-elemental data of the figure and table, and handwritten annotation text C13, "New Product," which indicates the content of the annotations made by the handwritten annotation shapes C11 and C12. In addition, the target image P10 includes a handwritten annotation shape C21 that annotates the column name "Cost [yen]," which is non-elemental data of the figure and table.

[0182] Similarly, the target image P10 includes a handwritten annotation figure C31 that annotates "150," which is the element data corresponding to the combination of "eraser" and "gross profit [yen]" among the element data of the chart, and handwritten annotation text C32, which indicates the content of the annotation by the handwritten annotation figure C31, and "minimum." In addition, the target image P10 includes a handwritten annotation figure C41 that annotates "1," which is the element data corresponding to the combination of "ballpoint pen" and "quantity [yen]" among the element data of the chart, and handwritten annotation text C42, which indicates the content of the annotation by the handwritten annotation figure C41, and "not very popular."

[0183] Furthermore, the target image P10 includes a handwritten annotation figure C51 that annotates "800," which is the element data corresponding to the combination of "pencil" and "gross profit [yen]" among the element data of the chart. In addition, the target image P10 includes handwritten annotation figures C43 and C44 that annotate "Netherlands," which is part of the element data corresponding to the combination of "ballpoint pen" and "sales area" among the element data of the chart, and handwritten annotation text C45, which indicates the content of the annotations made by the handwritten annotation figures C43 and C44, and "economic boom."

[0184] Furthermore, the target image P10 includes a handwritten annotation figure C61 that annotates the column name "Sales [yen]", which is non-elemental data of the chart, and the elemental data "450", "1,000", "250", and "1,400", which correspond to each combination of "notebook", "ballpoint pen", "eraser", and "pencil" and "Sales [yen]" from the elemental data of the chart.

[0185] <Server 20: Step S42> In step S42, the extraction processing unit 211 of the control unit 21 performs a first separation process based on the image data to be processed, separating the target image represented by the image data into a character image and a chart image. The chart image is an image of the target image that includes at least a chart such as a table or graph and the typefaces within the chart. The character image is an image of the target image that includes at least a string of typefaces other than the typefaces within the chart. The character image may also include ruled lines, etc. Conventional techniques may be used for the first separation process, but for example, a trained model of a neural network that detects the shape of a chart may be used.

[0186] In particular, if the target image contains handwritten annotations such as handwritten annotation characters or handwritten annotation figures that are the target of the chart, the extraction processing unit 211 separates the image including the handwritten annotations as the chart image. The handwritten annotation figures include arrows, leader lines, borders, or underlines. Similarly, if the target image contains handwritten annotations that are the target of the chart, other than the typefaces in the chart, the extraction processing unit 211 separates the image including the handwritten annotations as the character image.

[0187] For example, Figure 23(A) shows a character image P20, which is an example of a character image separated from the target image P10 shown in Figure 22 by the first separation process. Also, Figure 23(B) shows a chart image P30, which is an example of a chart image separated from the target image P10 shown in Figure 22 by the first separation process.

[0188] <Server 20: Step S43> In step S43, the extraction processing unit 211 of the control unit 21 performs a second separation process, similar to step S22, to separate the chart image into a handwritten image containing only the handwritten annotations included in the chart image and a type image containing only the chart and typefaces within the chart included in the chart image.

[0189] For example, Figure 24(A) shows a handwritten image P31, which is an example of the handwritten image separated by the second separation process from the chart image P30 shown in Figure 23(B). Also, Figure 24(B) shows a typeset image P32, which is an example of the typeset image separated by the second separation process from the chart image P30 shown in Figure 23(B).

[0190] <Server 20: Step S44> In step S44, the extraction processing unit 211 of the control unit 21 extracts the components of the chart and the typefaces included in the typeface image separated in step S43.

[0191] Specifically, the extraction processing unit 211 extracts the elements of a figure or table from the printed image using a trained neural network model, such as a graph neural network. For example, if the printed image contains a table as a figure or table, the extraction processing unit 211 extracts non-elemental data including at least one of the table title, column names, or row names of the table, and elemental data indicating the elements of the table. Also, if the printed image contains a graph as a figure or table, the extraction processing unit 211 extracts non-elemental data including at least one of the graph title, legend, vertical (value) axis, horizontal (item) axis, vertical axis label, horizontal axis label, vertical axis tick, horizontal axis tick, vertical axis tick label, horizontal axis tick label, or data label of the graph, and elemental data indicating the elements of the graph.

[0192] Furthermore, the extraction processing unit 211 performs OCR processing to recognize the typefaces included in the components of the diagram. The extraction processing unit 211 then associates the recognized components of the diagram with the typefaces included in those components and stores them in the memory or storage unit 22 of the control unit 21. For example, if the diagram is a table, the components such as the table title, column names, row names, and element data are associated with the recognition results of the typefaces that represent the content of those components.

[0193] <Server 20: Step S45> In step S45, the output processing unit 213 of the control unit 21 associates the handwritten annotations included in the chart image separated in step S42 with the target components, which are type characters in the chart extracted in step S44, with the related components associated with those type characters. Here, the target components are the components extracted in step S44 that are the subject of the handwritten annotations. The related components are the components extracted in step S44 that are related to the target components. For example, if the target components are element data, the related components are non-element data such as row names or column names corresponding to the target components. Also, if the target components are non-element data such as row names or column names, the related components are element data corresponding to the target components.

[0194] Specifically, the output processing unit 213 identifies which typeface in the diagram the handwritten annotation refers to, and stores the typeface as the target component, associating it with the handwritten annotation, in the memory or storage unit 22 of the control unit 21. The output processing unit 213 also stores the typeface identified as the target component, along with non-elemental data such as row names or column names related to the typeface, as related components, associating them with the typeface, in the memory or storage unit 22 of the control unit 21.

[0195] For example, in the example shown in Figure 22, the handwritten annotation figure C11 with a surrounding line and the handwritten annotation figure C12 with a leader line are associated with the handwritten annotation character C13 "New Product" located at the end of the handwritten annotation figure C12, and the element data "Pencil" surrounded by the handwritten annotation figure C11. In this case, "Pencil" is an example of a target component. Also, the handwritten annotation figure C21 with an underline is associated with the column name "Cost [yen]" to which the handwritten annotation figure C21 is attached. In this case, "Pencil" is an example of a target component, and "Cost [yen]" is an example of a related component. Furthermore, in the example shown in Figure 22, the handwritten annotation figure C31 with an arrow is associated with the handwritten annotation character C32 "Minimum" located at the beginning of the handwritten annotation figure C31, the element data "150" which is one of the components in the chart pointed to by the end of the handwritten annotation figure C31, and the column name "Gross Profit [yen]" corresponding to "150". In this case, "150" is an example of the target component, and "Gross Profit [yen]" is an example of the related component.

[0196] Similarly, in the example shown in Figure 22, the handwritten annotation shape C41 of the arrow is associated with the handwritten annotation text C42, which is "not very popular," the element data "1," which is one of the components in the chart pointed to by the end of the handwritten annotation shape C41 of the arrow, and the "quantity [pieces]" associated with "1." In this case, "1" is an example of the target component, and "quantity [pieces]" is an example of the associated component. Also, in the example shown in Figure 22, the handwritten annotation shape C43 of the surrounding line and the handwritten annotation shape C44 of the leader line are associated with the handwritten annotation text C45, which is "booming economy," the element data "Netherlands" surrounded by the handwritten annotation shape C43, and the column name "sales area" associated with "Netherlands." In this case, "Netherlands" is an example of the target component, and "sales area" is an example of the associated component. Furthermore, the underlined handwritten annotation figure C51 is associated with the element data "800" to which the handwritten annotation figure C51 is attached, and the column name "Gross Profit [yen]" which is related to "800". In this case, "800" is an example of the target component, and "Gross Profit [yen]" is an example of the related component. Also, in the example shown in Figure 22, the enclosed handwritten annotation figure C61 is associated with the column name "Cost [yen]" and the element data "450", "1,000", "250", and "1,400" which are enclosed by the handwritten annotation figure C61, and the row names "Notebook", "Ballpoint pen", "Eraser", and "Pencil" which are related to "450", "1,000", "250", and "1,400". In this case, "450", "1,000", "250", and "1,400" are examples of the target components, and "Notebook", "Ballpoint pen", "Eraser", and "Pencil" are examples of the related components.

[0197] <Server 20: Step S46> In step S46, the output processing unit 213 of the control unit 21 estimates the relationship between the handwritten annotation associated in step S45 and the typefaces in the chart, and stores the estimated result of the relationship in the memory or storage unit 22 of the control unit 21.

[0198] Specifically, as shown in Figure 25, the memory unit 22 stores relationship information D31 in which a corresponding relationship R1 to R5 is set for each combination of the type of handwritten annotation and the type of type in the chart associated with the handwritten annotation. The output processing unit 213 then identifies which of the relationship R1 to R5 the relationship between the handwritten annotation and the type in the chart associated with the handwritten annotation is based on the relationship information D31.

[0199] As shown in Figure 25, in relational information D31, the type of type in the chart associated with the handwritten annotation includes non-elemental data and elemental data of the chart. In addition, in relational information D31, the type of handwritten annotation includes annotations with text, annotations without text, annotations with text that include statistical indicators, and annotations with text that do not include statistical indicators. For example, in relational information D31, if the type of type in the chart to which the handwritten annotation is annotated is elemental data of the chart, and the type of handwritten annotation is an annotation with text that includes statistical indicators, then relational R3 is associated.

[0200] <Server 20: Step S47> In step S47, the output processing unit 213 of the control unit 21 generates summary information D32 showing a summary of the target image based on the extraction results in step S44 and the estimation results in step S46.

[0201] The output processing unit 213 uses the extraction results from step S44 and the estimation results from step S46 as input information to generate summary information D32 containing content corresponding to relationships R1 to R5, using a conditional text generation model such as "Text-to-Text Transfer Transformer (T5)". According to the conditional model generation model, the intended text information is generated by providing arbitrary contextual information and conditions for what kind of text data to generate. In particular, the conditional model generation model has been learned to use non-elemental data from the extraction results in step S44 to generate summary information D32 as needed.

[0202] In particular, when the output processing unit 213 generates summary information D32, it generates summary information D32 that combines the summary of the character image P20 and the summary of the chart image P30. For example, as shown in Figure 22, if the character image P20 (see Figure 23) contains typefaces and handwritten annotations of handwritten annotation figures C1 and handwritten annotation characters C2 that are the target of the annotations, the output processing unit 213 generates summary information D32 that supplements the typefaces based on the handwritten annotations. For example, in the example shown in Figure 22, based on the handwritten annotation figure C1 and handwritten annotation characters C2, the handwritten annotation character C2 "Southeast Asia, etc." is added to the typeface "Expand sales area" contained in the character image P20, and summary information D32 that includes the content "Expand sales area to Southeast Asia, etc." is generated.

[0203] In other embodiments, the storage unit 22 may store generation rules that, in association with each of the relationships R1 to R5 that are the estimation results in step S46, generate summary information D32 of the content described below based on the handwritten annotations and the typefaces in the charts associated with the handwritten annotations. In this case, the output processing unit 213 generates the summary information D32 according to the generation rules, using non-elemental data from the extraction results in step S44 as needed.

[0204] The following describes an example of summary information D32, which includes a summary generated based on relationships R1 to R5, figures and tables, and handwritten annotations on those figures and tables.

[0205] Figure 26(A) shows an example of summary information D32 generated by the conventional technology in which the summary output processing is not performed on the target image P10 shown in Figure 22. Figure 26(B) shows an example of summary information D32 generated when the summary output processing is performed on the target image P10 shown in Figure 22. As shown in Figure 26(A), in the conventional technology, the content of handwritten annotations corresponding to the figures and tables included in the target image P10 was not reflected in the summary information D32. In contrast, the summary information D32 shown in Figure 26(B) reflects the content of handwritten annotations corresponding to the figures and tables included in the target image P10.

[0206] Relationship R1 indicates that the handwritten annotation supplements the typeset text in the chart associated with the handwritten annotation. The conditional text generation model has learned that for relationship R1, it generates text that supplements the typeset text in the chart associated with the handwritten annotation with the handwritten annotation text included in the handwritten annotation. For example, in the example shown in Figure 22, the type of the handwritten annotation, which includes handwritten annotation shapes C11 and C12 and the handwritten annotation text C13, which is "new product", is an annotation with text, and the typeset text in the chart associated with the handwritten annotation shapes C11 and C12 is the non-element data "pencil". In this case, the output processing unit 213 generates summary information D32 that includes the text "Pencil is a new product" as text that supplements "pencil" with the handwritten annotation text C13, which is "new product", as shown in Figure 26(B).

[0207] Relationship R2 indicates that the handwritten annotation emphasizes the typeset text within the chart associated with the handwritten annotation. The conditional text generation model has learned that for relationship R2, it generates text that emphasizes the content of the typeset text within the chart associated with the handwritten annotation. For example, in the example shown in Figure 22, the handwritten annotation figure C21 is an annotation without text, and the typeset text within the chart associated with the handwritten annotation figure C21 is the non-element data "Cost [yen]". In this case, the generation of summary information D32 uses the content of the element data corresponding to the typeset text "Cost [yen]" within the chart associated with the handwritten annotation figure C21, and the column name of the non-element data "Cost [yen]" corresponding to each of the element data. For example, the output processing unit 213 generates summary information D32, which includes the sentence "The cost [yen] is 30 for the notebook, 400 for the ballpoint pen, 20 for the eraser, and 30 for the pencil," as shown in Figure 26(B), as a sentence that emphasizes the column name "Cost [yen]" in the chart associated with the handwritten annotation figure C21.

[0208] Relationship R3 indicates that the handwritten annotation is a statistical indicator corresponding to the printed text in the chart associated with the handwritten annotation. For example, the statistical indicator may include maximum, minimum, mean, mode, or median. The conditional text generation model has learned that for relationship R3, it generates text that explains the content of the printed text in the chart associated with the handwritten annotation using the statistical indicator indicated by the handwritten annotation and the non-elemental data of the chart. For example, in the example shown in Figure 22, the type of handwritten annotation that includes the statistical indicator "minimum," which is the handwritten annotation figure C31 and the handwritten annotation character C32, is an annotation with text that includes a statistical indicator, and the printed text in the chart associated with the handwritten annotation figure C31 is the elemental data "150." In this case, the generation of summary information D32 uses the printed text "150" in the chart associated with the handwritten annotation figure C31, the statistical indicator "minimum" which is the handwritten annotation character C32, and the column name "Gross Profit [yen]," which is the non-elemental data corresponding to the printed text "150." For example, the output processing unit 213 generates summary information D32, which includes the sentence "The minimum gross profit [yen] is 150," as the content of the summary corresponding to the handwritten annotation character C32 and the typeset "150" in the chart, as shown in Figure 26(B).

[0209] Relationship R4 indicates that the handwritten annotation supplements the typeset text in the chart associated with the handwritten annotation. The conditional text generation model has been trained to generate text that supplements the content of the typeset text in the chart associated with the handwritten annotation based on the handwritten annotation characters and the non-elemental data corresponding to the typeset text in the chart, for relationship R4. For example, handwritten annotation figure C41 and handwritten annotation character C42 are annotations with text that do not contain statistical indicators, the elemental data in the chart associated with handwritten annotation figure C41 is "1", and the non-elemental data corresponding to this elemental data is the row name "ballpoint pen" and the column name "quantity [pieces]". In this case, the generation of summary information D32 uses the typeset "1" in the chart associated with handwritten annotation figure C41 and the column name "quantity [pieces]" and row name "ballpoint pen" of the non-elemental data corresponding to the typeset "1". Also, in the example shown in Figure 22, the character image P20 contains the typeset text "Meanwhile, only one ballpoint pen was sold." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "The number of ballpoint pens sold was [number] units, and they were not very popular," as shown in Figure 26(B), based on the type information contained in the character image P20 and the handwritten annotation figures C41 and handwritten annotation characters C42. That is, if both the character image P20 and the diagram image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the diagram image P30. In other embodiments, the output processing unit 213 may not combine the summary of the character image P20 and the summary of the diagram image P30, but may generate the summary of the character image P20 and the summary of the diagram image P30 individually.

[0210] Similarly, handwritten annotation figures C43, C44, and handwritten annotation text C45 are annotations with text that do not include statistical indicators. The element data within the chart associated with handwritten annotation figures C43 and C44 is "Netherlands," and the non-element data corresponding to this element data are the row name "Ballpoint pen" and the column name "Sales area." That is, the relationship between handwritten annotation figures C43, C44, and handwritten annotation text C45 and the typefaces within the chart is relationship R4. In this case, the generation of summary information D32 uses the element data "Netherlands" within the chart associated with handwritten annotation figures C43 and C44, and the non-element data column name "Sales area" and row name "Ballpoint pen" corresponding to the typeface "Netherlands." Also, in the example shown in Figure 22, the text image P20 contains the typeface "Ballpoint pens have the second highest sales after pencils." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "Ballpoint pens have a sales area that includes the booming Netherlands, and are the second highest-selling product after pencils," as shown in Figure 26(B), based on the information of the typefaces contained in the character image P20 and the handwritten annotation figures C43, C44 and handwritten annotation characters C45. In other words, if both the character image P20 and the chart image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the chart image P30.

[0211] Relationship R5 indicates that the handwritten annotation emphasizes the typeset text within the chart associated with the handwritten annotation. The conditional text generation model has learned that, for relationship R5, it generates text that emphasizes the content of the typeset text within the chart associated with the handwritten annotation, taking into account the constituent elements of the chart. For example, in the example shown in Figure 22, the handwritten annotation figure C51 is an annotation without text, the element data within the chart associated with the handwritten annotation figure C51 is "800", and the non-element data corresponding to this element data is "Gross Profit [yen]". In this case, the generation of summary information D32 uses the typeset text "800" from the element data associated with the handwritten annotation figure C51, and the column name "Gross Profit [yen]" and row name "Pencil" from the non-element data corresponding to the typeset text "800". Also, in the example shown in Figure 22, the text image P20 contains the typeset text "Pencil sales were particularly strong, with 20 sold." Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "The pencil is a new product, the gross profit [yen] was 800, and 20 units were sold," based on the type information contained in the character image P20 and the handwritten annotation figure C51. In other words, if both the character image P20 and the figure image P30 contain similar content, the output processing unit 213 generates summary information D32 by combining the summary of the character image P20 and the summary of the figure image P30.

[0212] Furthermore, multiple relationships among relationships R1 to R5 may be valid simultaneously. For example, relationships R2 and R5 may be valid at the same time. The conditional text generation model has learned that when multiple relationships among relationships R1 to R5 are valid simultaneously, it generates summary information D32 based on those multiple relationships. For example, in the example shown in Figure 22, the handwritten annotation figure C61 is an annotation without text, and the handwritten annotation figure C61 is associated with the column name "Sales [yen]", which is non-element data, and the element data "450", "1,000", "250", and "1,400". In this case, the generation of summary information D32 uses the non-element data and element data associated with the handwritten annotation figure C61, and the row name, which is non-element data corresponding to the element data. For example, the output processing unit 213 generates summary information D32 that includes the sentence, "Sales [yen] were 450 for notebooks, 1,000 for ballpoint pens, 250 for erasers, and 1,400 for pencils," based on the handwritten annotation figure C61. That is, if multiple relationships among relationships R1 to R5 are simultaneously established, the output processing unit 213 generates summary information D32 based on those multiple relationships.

[0213] <Server 20: Step S48> In step S48, the output processing unit 213 of the control unit 21 outputs the summary information D32 generated in step S47. Specifically, the output processing unit 213 transmits the summary information D32 to the image processing device 10, which is the source of the image data that was the target of the summary output processing. The output processing unit 213 may also transmit the extraction results from step S44 to the image processing device 10 along with the summary information D32. In step S48, the output processing unit 213 may display, print, or store the summary information D32 and the extraction results from step S44. The summary information D32 and the extraction results from step S44 may also be transmitted to the client terminal 30.

[0214] <Image Processing Device 10: Step S113> In step S113, the control unit 11 of the image processing device 10 waits for the reception of summary information D32 from the server 20 (S113: No). When it is determined that the summary information D32 has been received (S113: Yes), the process moves to step S114.

[0215] <Image Processing Device 10: Step S114> In step S114, the control unit 11 of the image processing device 10 performs a result output process to output summary information D32. Specifically, in the result output process, the control unit 11 performs an image formation process to form an image on a sheet based on the image data that was the target of the summary output process. In addition, in the result output process, the image data that was the target of the summary output process and the summary information D32 may be stored in the storage unit 12 in association. In addition, in the result output process, the summary information D32 may be combined with the image data that was the target of the summary output process. In addition, in the result output process, the image data that was the target of the summary output process and the summary information D32 may be transmitted to the client terminal 30 in association or combined.

[0216] As explained above, the image processing system 1 can generate summary information D32 while taking into account handwritten annotations added to tables or graphs. This reduces the effort required from the user compared to when the user generates similar summary information D32 themselves.

[0217] [Other Examples of Figures and Charts] In this embodiment, the example of a figure or chart in figure / chart image P30 has been used for explanation. On the other hand, it is also possible that the figure / chart in figure / chart image P30 is various types of graphs. In this case as well, the output processing unit 213 generates summary information D32 based on handwritten annotations such as handwritten annotation characters and handwritten annotation figures, and the target components of the graph that are the subject of the handwritten annotations. Below, the summary information D32 generated when figure / chart images P40, P50, or P60 are included in the target image P10 (see Figure 22) instead of figure / chart image P30 will be explained.

[0218] The output processing unit 213 generates summary information D32 based on the handwritten annotations contained in the chart image P40, P50, or P60, and the elemental data and non-elemental data that are components of the bar graph contained in the chart image P40, P50, or P60. Specifically, the output processing unit 213 converts the bar graph data contained in the chart image P40, P50, or P60 into tabular data such as the aforementioned chart image P30, and then generates summary information D32 based on the chart image P30.

[0219] [Figure Image P40] Figure 27(A) shows an example of Figure Image P40, which includes a bar graph, as another example of Figure Image P30 included in the target image P10 (see Figure 22). In Figure Image P40, element data is indicated by vertical bar line data markers. Figure Image P40 includes the graph title "2024 / 4 / 10 Product Sales", the legend "Sales [yen], Gross Profit [yen]", the vertical axis tick labels "0", "500", "1000", "1500", and the horizontal axis tick labels "Notebook", "Ballpoint Pen", "Eraser", and "Pencil" as non-element data. Note that Figure Image P40 also includes the vertical axis tick as non-element data. Furthermore, Figure Image P40 includes element data corresponding to the horizontal axis tick labels "Notebook", "Ballpoint Pen", "Eraser", and "Pencil". Furthermore, the figure image P40 includes handwritten annotations: handwritten annotation figures C71 and C72, and handwritten annotation characters C73. Figure 27(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes the figure image P40.

[0220] Handwritten annotation shape C71 is a bounding line that annotates the horizontal axis tick label "eraser," which is non-elemental data in the bar graph of figure image P40. Handwritten annotation shape C72 is an arrow that annotates the elemental data "pencil sales [yen]" in the bar graph of figure image P40. Handwritten annotation text C73 is the annotation content for the subject of handwritten annotation shape C72, which is "up from the previous day."

[0221] In this case, the output processing unit 213 first converts the data from the bar graph into tabular data by inputting the element data from the bar graph into cells, using the legend "Sales [yen]" and "Gross Profit [yen]" from the bar graph included in the chart image P40 as column names, and the horizontal axis scale labels "Notebook," "Ballpoint pen," "Eraser," and "Pencil" as row names.

[0222] For example, the output processing unit 213 calculates the value of each element data based on the non-element data of the bar graph and the Y-coordinate of the vertical axis of the data marker representing each element data in the bar graph. The Y-coordinate is assumed to be 0 at the top of the target image P10 or chart image P40 and the maximum value at the bottom. Furthermore, y1 is the value of the vertical axis tick label that corresponds to the vertical axis tick closest to the Y-coordinate, which is less than or equal to the Y-coordinate of the top of the data marker of the element data; y2 is the Y-coordinate of the vertical axis tick corresponding to the vertical axis tick label; y3 is the Y-coordinate of the bottom of the data marker of the element data; and y4 is the Y-coordinate of the top of the data marker of the element data. In this case, the value of the element data is calculated using the formula y1 × ((y3 - y4) / (y3 - y2)). Specifically, for data marker M1 corresponding to the sales [yen] of the horizontal axis scale label "pencil" in Figure 27(A), the value of the vertical axis scale label of the vertical axis scale M2 is y1, the Y-axis coordinate of the vertical axis scale M2 is y2, the Y-axis coordinate of the vertical axis scale M3 at the lower end of data marker M1 is y3, and the Y-axis coordinate of the upper end M4 of data marker M1 is y4. For example, for data marker M1 of pencil sales [yen], if y1 = 1500, y2 = 220, y3 = 290, and y4 = 227, then using the above calculation formula, the solution to 1500 × ((290 - 227) / (290 - 220)) is 1,350, which is the value of the element data for pencil sales [yen].

[0223] In the example shown in Figure 27(A), the type of handwritten annotation, including the handwritten annotation figure C71, is a textless annotation, and the non-elemental data associated with the handwritten annotation figure C71 is "eraser" (relationship R2). Therefore, for example, the output processing unit 213 generates summary information D32 that includes the sentence, "Erasers generated sales of 250 yen and gross profit of 150 yen," as a sentence emphasizing the handwritten annotation character C73, "eraser," as shown in Figure 27(B).

[0224] Furthermore, in the example shown in Figure 27(A), the type of handwritten annotation, which includes the handwritten annotation figure C72 and the handwritten annotation text C73 "up from the previous day," is a text-based annotation that does not include statistical indicators, and the handwritten annotation figure C72 is associated with the element data "1,350" (relationship R4). Also, as mentioned above, the text image P20 shown in Figure 22 contains the text "Pencil sales were particularly strong, with 20 sold." Therefore, for example, the output processing unit 213 generates summary information D32 that includes the text "20 pencils were sold, and sales [yen] were 1,350, up from the previous day," as a sentence that supplements the element data pencil sales [yen] of "1,350," using the handwritten annotation text C73 "up from the previous day."

[0225] [Figure Image P50] Figure 28(A) shows an example of Figure Image P50, which includes a line graph, as another example of Figure Image P30 included in the target image P10 (see Figure 22). In Figure Image P50, element data is indicated by line data markers. Figure Image P50 includes the graph title "Product Sales", the legend "Notebook, Ballpoint Pen, Eraser, Pencil", vertical axis tick labels "0", "500", "1000", "1500", horizontal axis tick labels "April 7th", "April 8th", "April 9th", "April 10th", vertical axis label "Sales [yen]", and horizontal axis label "Date" as non-element data. In addition, Figure Image P50 includes element data corresponding to the legend "Notebook", "Ballpoint Pen", "Eraser", and "Pencil". Furthermore, the figure image P50 includes handwritten annotations: handwritten annotation figures C81 and C82, and handwritten annotation characters C83. Figure 28(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes the figure image P50.

[0226] Handwritten annotation shape C81 is a bounding line that annotates the horizontal axis tick label "April 8th," which is non-elemental data of the line graph in Figure Image P50. Handwritten annotation shape C82 is an arrow that annotates the elemental data "Sales [yen]" for "April 10th" of the "Pencil" line in the line graph in Figure Image P50. Handwritten annotation text C83 is the handwritten text "Doing well," which indicates the content of the annotation for the subject of handwritten annotation shape C82.

[0227] In this case, the output processing unit 213 first converts the data of the line graph into tabular data by inputting the element data of the line graph into cells, using the legend "notebook," "ballpoint pen," "eraser," and "pencil" of the line graph included in the chart image P50 as column names, and the horizontal axis scale labels "April 7th," "April 8th," "April 9th," and "April 10th" as row names.

[0228] For example, the output processing unit 213 calculates the value of each element data based on the non-element data of the line graph and the Y-coordinate of the vertical axis of the data marker representing each element data in the line graph. The Y-coordinate is assumed to be 0 at the top end of the target image P10 or chart image P50 and the maximum value at the bottom end. Furthermore, y1 is the value of the vertical axis tick label that corresponds to the vertical axis tick closest to the Y-coordinate, which is less than or equal to the Y-coordinate of the intersection of the data marker of the element data and the horizontal axis tick; y2 is the Y-coordinate of the vertical axis tick corresponding to that vertical axis tick label; y3 is the Y-coordinate of the vertical axis tick with the smallest Y-coordinate; and y4 is the Y-coordinate of the intersection of the data marker of the element data and the horizontal axis tick. In this case, the value of the element data is calculated using the formula y1 × ((y3 - y4) / (y3 - y2)). Specifically, for data marker M11 corresponding to the sales [yen] of the horizontal axis scale label "April 8th" in Figure 28(A), y1 is the value of the vertical axis scale label of vertical axis scale M12, y2 is the Y-axis coordinate of vertical axis scale M12, y3 is the Y-axis coordinate of vertical axis scale M13 at the lower end of the vertical axis scale label, and y4 is the Y-coordinate of the intersection of data marker M11 and the horizontal axis scale of "April 8th". For example, for data marker M11 of pencil sales [yen], if y1 = 1500, y2 = 220, y3 = 290, and y4 = 227, then using the above calculation formula, the solution to 1500 × ((290 - 227) / (290 - 220)) is 1,350, which is the value of the element data for pencil sales [yen].

[0229] In the example shown in Figure 28(A), the type of handwritten annotation, including the handwritten annotation figure C81, is a textless annotation, and the non-elemental data associated with the handwritten annotation figure C81 is "April 8th" (relationship R2). Therefore, for example, the output processing unit 213 generates summary information D32 that includes the sentence, "On April 8th, there were 450 notebooks, 2,000 ballpoint pens, 350 erasers, and 770 pencils," as a sentence that emphasizes the non-elemental data "April 8th," as shown in Figure 28(B).

[0230] Furthermore, in the example shown in Figure 28(A), the type of handwritten annotation, which includes the handwritten annotation figure C82 and the handwritten annotation character C83 "doing well", is an annotation with text, and the non-element data associated with the handwritten annotation figure C82 is "April 10th" (relationship R1). Therefore, for example, the output processing unit 213 generates summary information D32 that includes the sentence, "Pencils are doing well, with sales of 630 on April 7th, 770 on April 8th, 1,050 on April 9th, and 1,400 on April 10th," as a sentence to supplement the element data "pencils" sales [yen] of "April 10th", which is "1,350", as shown in Figure 28(B).

[0231] [Figure P60] Figure 29(A) shows an example of Figure P60, which includes a pie chart, as another example of Figure P30 included in the target image P10 (see Figure 22). In Figure P60, element data is indicated by circular data markers. Figure P60 includes the graph title "Sales Composition Ratio" and the legend "Notebook, Ballpoint Pen, Eraser, Pencil" as non-element data. Figure P60 also includes element data "15%", "32%", "8%", and "45%" corresponding to the legend "Notebook", "Ballpoint Pen", "Eraser", and "Pencil", respectively. Furthermore, Figure P60 includes handwritten annotations of handwritten annotation shapes C91 and handwritten annotation characters C92. Figure 29(B) shows an example of summary information D32 generated by the output processing unit 213 for the target image P10, which includes Figure P60.

[0232] The handwritten annotation shape C91 is an arrow that annotates the element data "8%" of "Sales Composition Ratio" for "Erasers" in the pie chart on the chart image P60. The handwritten annotation text C92 is the handwritten text "A small number, but 5 were sold" which indicates the content of the annotation for the subject of the handwritten annotation shape C91.

[0233] In this case, the output processing unit 213 first converts the data from the pie chart into tabular data by inputting the element data from the pie chart into cells, using the legend "notebook," "ballpoint pen," "eraser," and "pencil" of the pie chart included in the chart image P60 as column names and the title of the pie chart "sales composition ratio" as row names.

[0234] In the example shown in Figure 29(A), the type of handwritten annotation, which includes the handwritten annotation figure C91 and the annotation text C92, is an annotation with text, and the element data associated with the handwritten annotation figure C91 is "8%", which corresponds to "eraser" in the legend (relationship R4). Therefore, for example, the output processing unit 213 generates summary information D32, which includes the sentence "Erasers account for a small percentage of sales at 8%, but 5 were sold," as a sentence that supplements the element data "8%", as shown in Figure 29(B).

[0235] [Addendum to the Invention] The following is an outline of the invention extracted from the above-described embodiments. Note that each configuration and processing function described in the following addendum can be selected and combined as desired.

[0236] <Note A1> An image processing system comprising: an extraction processing unit that extracts handwritten characters and printed characters contained in an image shown by an image data to be processed in a predetermined order; and an identification processing unit that identifies one or more handwritten characters and one or more printed characters whose extraction order by the extraction processing unit is consecutive, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction processing unit and one or more predetermined item name candidate strings contained in the printed characters, as a series of strings.

[0237] <Appendix A2> The image processing system according to Appendix A1, wherein the identification processing unit identifies the handwritten characters and printed characters that exist between two consecutive item name candidate strings as a series of strings.

[0238] <Appendix A3> The image processing system according to Appendix A1 or A2, wherein the identification processing unit identifies the handwritten characters and printed characters that are extracted after the last candidate string of item names in the extraction order as a series of strings.

[0239] <Appendix A4> The image processing system according to any one of Appendix A1 to A3, wherein the specified processing unit extracts named entities included in the series of strings, and excludes from the series of strings any strings in which the named entities and the item group associated with the candidate item name strings do not satisfy the pre-set conditions.

[0240] <Appendix A5> The image processing system according to any one of Appendices A1 to A4, wherein the specific processing unit combines at least one of the handwritten characters and the printed characters included in the series of strings into a single string.

[0241] <Appendix A6> The image processing system described in Appendix A5, wherein the specified processing unit combines all the handwritten characters and printed characters included in the series of strings into a single string.

[0242] <Appendix A7> An image processing system according to any one of Appendix A1 to A6, comprising an output processing unit that outputs recognition result information generated based on the extraction result by the extraction processing unit and the processing result by the identification processing unit.

[0243] <Appendix A8> An image processing device provided in the image processing system described in Appendix A7, comprising: an image acquisition unit that acquires the image data to be processed; and a reception processing unit that accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

[0244] <Note A9> An image processing method comprising: an extraction step in which one or more processors extract handwritten characters and printed characters contained in an image shown by an image data to be processed in a predetermined order; and an identification step in which, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction step and one or more predetermined item name candidate strings contained in the printed characters, one or more of the handwritten characters and one or more of the printed characters whose extraction order by the extraction step is consecutive are identified as a series of strings.

[0245] <Note A10> A program for causing one or more processors to execute: an extraction step of extracting handwritten characters and printed characters contained in an image shown by an image data to be processed in a predetermined order; and an identification step of identifying one or more handwritten characters and one or more printed characters whose extraction order by the extraction step is consecutive, based on the positional relationship between the handwritten characters and printed characters extracted by the extraction step and one or more predetermined item name candidate strings contained in the printed characters, as a series of strings.

[0246] <Note B1> An image processing system comprising: an extraction processing unit that extracts handwritten characters contained in an image shown by an image data to be processed; and an identification processing unit that identifies the item group corresponding to the handwritten characters based on the result of executing a named entity recognition process that extracts named entities from the handwritten characters extracted by the extraction processing unit and a plurality of item groups that are set in advance in association with the image data.

[0247] <Appendix B2> The image processing system according to Appendix B1, further comprising a setting processing unit that sets the item group corresponding to the image data in response to user operation.

[0248] <Appendix B3> The image processing system described in Appendix B1 or B2, wherein the setting processing unit sets the item group according to the selection of the document type corresponding to the image data.

[0249] <Appendix B4> The image processing system according to any one of the appendices B1 to B3, wherein the extraction processing unit extracts handwritten characters and printed characters contained in the image shown by the image data to be processed, and the identification processing unit performs a first identification process to identify the item group corresponding to the handwritten characters based on the positional relationship between the handwritten characters extracted by the extraction processing unit and the printed characters extracted by the extraction processing unit, and a second identification process to identify the item group corresponding to the handwritten characters based on the unidentified item groups among a plurality of item groups corresponding to the image data for which the handwritten characters corresponding to the handwritten characters were not identified by the first identification process, and the result of the named entity recognition process.

[0250] <Appendix B5> The image processing system according to any one of the appendices B1 to B4, wherein the specific processing unit sets a determination group indicating the content of the handwritten character based on a named entity extracted from the handwritten character and a predetermined determination rule, and identifies the item group corresponding to the handwritten character based on the determination group and the item group.

[0251] <Appendix B6> The image processing system described in Appendix B5, wherein the judgment rule includes one or more of the following conditions: a condition for identifying the judgment group based solely on the content of the named entity extracted from the handwritten character; a condition for identifying the judgment group based on the named entity corresponding to the handwritten character and the content of the handwritten character; and a condition for identifying the judgment group based on the named entity corresponding to the handwritten character and the item group associated with the printed characters present around the handwritten character.

[0252] <Appendix B7> An image processing system according to any one of Appendix B1 to B6, comprising an output processing unit that outputs recognition result information generated based on the identification result by the identification processing unit and the extraction result by the extraction processing unit.

[0253] <Appendix B8> An image processing apparatus comprising: an image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix B7; and a reception processing unit that accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

[0254] <Note B9> An image processing method comprising: an extraction step in which one or more processors extract handwritten characters contained in an image shown by an image data to be processed; and an identification step in which, based on the result of executing a named entity recognition process to extract named entities from the handwritten characters extracted by the extraction step and a plurality of item groups set in advance in association with the image data, the item group corresponding to the handwritten characters.

[0255] <Note B10> A program for causing one or more processors to execute: an extraction step of extracting handwritten characters contained in an image shown by an image data to be processed; an identification step of identifying an item group corresponding to a handwritten character based on the result of executing a named entity recognition process that extracts named entities from the handwritten characters extracted by the extraction step and a plurality of item groups set in advance in association with the image data.

[0256] <Appendix C1> An image processing system comprising: a first processing unit that extracts character regions and characters contained in an image indicated by image data to be processed; a second processing unit that identifies the number of characters in a specific direction predetermined in the character region as the number of lines in the character region; and a third processing unit that identifies the characters contained in each line of the number of lines in the character region identified by the second processing unit.

[0257] <Appendix C2> The image processing system according to Appendix C1, wherein the third processing unit identifies the characters contained in each line of the character area based on the result of clustering the strings in a number corresponding to the number of lines in the character area identified by the second processing unit.

[0258] <Appendix C3> The third processing unit is the image processing system described in Appendix C2, which performs the clustering using the k-means method.

[0259] <Appendix C4> The image processing system according to Appendix C2 or C3, wherein the third processing unit samples each character in the character area by a predetermined number of times and performs the clustering based on the sampled data after sampling.

[0260] <Appendix C5> An image processing system according to any one of the appendices C1 to C4, comprising a fourth processing unit that determines the specific direction in the character area based on the aspect ratio of the character area.

[0261] <Appendix C6> The image processing system according to any one of the appendices C1 to C5, wherein the first processing unit distinguishes and extracts handwritten characters and printed characters included in the character area, and the second processing unit, the third processing unit, and the fourth processing unit perform processing only on the character area including the handwritten characters.

[0262] <Appendix C7> The image processing system according to any one of the appendices C1 to C6, wherein the third processing unit performs clustering only on the character regions where the number of lines identified by the second processing unit is multiple.

[0263] <Appendix C8> An image processing system according to any one of the appendices C1 to C7, comprising an output processing unit that outputs recognition result information generated based on the specific result of the third processing unit.

[0264] <Appendix C9> An image processing apparatus comprising: an image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix C8; and a reception processing unit that accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

[0265] <Appendix C10> An image processing method comprising: a first step of extracting character regions and characters contained in an image indicated by image data to be processed by one or more processors; a second step of identifying the number of characters in a specific direction predetermined in the character region as the number of lines in the character region; and a third step of identifying the characters contained in each line of the number of lines in the character region identified in the second step.

[0266] <Note C11> An image processing program that causes one or more processors to perform the following steps: a first step of extracting character regions and characters contained in an image indicated by image data to be processed; a second step of identifying the number of characters in a character region that are located in a specific direction that has been identified in advance as the number of lines in the character region; and a third step of identifying the characters contained in each line of the number of lines in the character region that has been identified in the second step.

[0267] <Note D1> An image processing system comprising: an extraction processing unit that extracts characters contained in an image shown by an image data to be processed; and an identification processing unit that determines whether there is a relationship between two consecutive lines of text extracted by the extraction processing unit, and if it is determined that there is a relationship, performs an identification process to identify the two lines of text as a series of texts.

[0268] <Appendix D2> The image processing system described in Appendix D1, wherein the specific processing unit determines whether or not there is a relationship based on the meaning of the sentence when the two lines of text are joined together.

[0269] <Note D3> The image processing system described in Note D1 or D2, wherein the specified processing unit determines the presence or absence of the relationship using the NSP (Next Sentence Prediction) of BERT (Bidirectional Encoder Representations from Transformers).

[0270] <Appendix D4> The image processing system according to any one of Appendix D1 to D3, wherein the specific processing unit performs the specific processing on the character area only when the characters in the character area extracted by the extraction processing unit are in sentence form.

[0271] <Note D5> The image processing system according to any one of Notes D1 to D4, wherein the extraction processing unit distinguishes and extracts handwritten characters and printed characters contained in the image, and the identification processing unit performs the identification process on the character area only when the characters in the character area extracted by the extraction processing unit are handwritten characters.

[0272] <Note D6> The image processing system according to any one of Notes D1 to D5, wherein the specified processing unit combines the two lines of text when they are identified as a series of texts.

[0273] <Appendix D7> An image processing system according to any one of the appendices D1 to D6, comprising an output processing unit that outputs recognition result information generated based on the identification result by the identification processing unit and the extraction result by the extraction processing unit.

[0274] <Appendix D8> An image processing apparatus comprising: an image acquisition processing unit that acquires the image data to be processed in the image processing system described in Appendix D7; and a reception processing unit that accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

[0275] <Note D9> An image processing method comprising: an extraction step in which one or more processors extract characters contained in an image represented by an image data to be processed; and an identification step in which one determines whether there is a relationship between two consecutive lines of text extracted by the extraction step, and if it is determined that there is a relationship, identifies the two lines of text as a series of texts.

[0276] <Note D10> A program that causes one or more processors to execute: an extraction step of extracting characters contained in an image represented by an image data to be processed; and an identification step of determining whether two consecutive lines of text extracted by the extraction step are related, and if it is determined that they are related, identifying the two lines of text as a series of texts.

[0277] <Appendix E1> An image processing system comprising: an extraction processing unit that extracts handwritten annotations, which include either or both handwritten annotation characters and handwritten annotation figures, and components of a chart included in the image data to be processed; and an output processing unit that generates summary information based on the handwritten annotations extracted by the extraction processing unit and the target components among the components that are the subject of the handwritten annotations.

[0278] <Appendix E2> The image processing system according to Appendix E1, wherein the output processing unit generates the summary information based on the handwritten annotation, the target component, and related components among the components that are related to the target component.

[0279] <Appendix E3> The image processing system according to Appendix E1 or E2, wherein the figure is a table, the components of the figure include element data and non-element data excluding the element data, the non-element data includes at least one of a table title, column names, or row names, and the output processing unit generates the summary information including the element data and the non-element data.

[0280] <Appendix E4> The image processing system according to any one of Appendix E1 to E3, wherein the output processing unit generates summary information including the element data and either or both of the column name and row name corresponding to the element data when the target component is the element data.

[0281] <Appendix E5> The image processing system according to any one of the appendices E1 to E4, wherein the figure is a graph, the components of the figure include element data and non-element data excluding the element data, the non-element data includes at least one of a graph title, legend, vertical (value) axis, horizontal (item) axis, vertical axis label, horizontal axis label, vertical axis tick, horizontal axis tick, vertical axis tick label, horizontal axis tick label, or data label, and the output processing unit generates the summary information including the element data and the non-element data.

[0282] <Appendix E6> The image processing system according to any one of the appendices E1 to E5, wherein the output processing unit separates the image into a character image containing at least printed characters and a chart image containing at least the chart, and generates the summary information by combining a summary of the printed characters contained in the character image and a summary of the chart contained in the chart image.

[0283] <Appendix E7> The output processing unit outputs the generated summary information, an image processing system according to any one of the appendices E1 to E6.

[0284] <Appendix E8> An image processing apparatus comprising an image acquisition processing unit that acquires the image data to be processed in the image processing system described in any of the appendices E1 to E7.

[0285] <Appendix E9> An image processing method comprising: an extraction step in which one or more processors extract handwritten annotations, which include either or both handwritten annotation characters and handwritten annotation figures, and components of a chart included in the image shown by the image data to be processed; and an output step in which summary information is generated based on the handwritten annotations extracted by the extraction step and the target components among the components that are the subject of the handwritten annotations.

[0286] <Appendix E10> A program for causing one or more processors to execute: an extraction step of extracting handwritten annotations, which include either or both handwritten annotation characters and handwritten annotation figures, and components of charts and graphs included in the image data to be processed; and an output step of generating summary information based on the handwritten annotations extracted by the extraction step and the target components among the components that are the subject of the handwritten annotations.

Claims

1. An image processing system comprising: an extraction processing unit that extracts handwritten characters contained in an image shown by an image data to be processed; and an identification processing unit that identifies an item group corresponding to the handwritten characters based on the result of executing a named entity recognition process that extracts named entities from the handwritten characters extracted by the extraction processing unit and a plurality of item groups that are set in advance in association with the image data.

2. The image processing system according to claim 1, further comprising a setting processing unit that sets the item group corresponding to the image data in response to user operation.

3. The image processing system according to claim 2, wherein the setting processing unit sets the item group according to the selection of the document type corresponding to the image data.

4. The image processing system according to claim 1, wherein the extraction processing unit extracts handwritten characters and printed characters contained in the image shown by the image data to be processed, and the identification processing unit performs a first identification process to identify the item group corresponding to the handwritten characters based on the positional relationship between the handwritten characters extracted by the extraction processing unit and the printed characters extracted by the extraction processing unit, and a second identification process to identify the item group corresponding to the handwritten characters based on the unidentified item groups among a plurality of item groups corresponding to the image data for which the handwritten characters corresponding to the handwritten characters were not identified by the first identification process, and the result of the named entity recognition process.

5. The image processing system according to claim 1, wherein the identification processing unit sets a determination group indicating the content of the handwritten character based on a named entity extracted from the handwritten character and a predetermined determination rule, and identifies the item group corresponding to the handwritten character based on the determination group and the item group.

6. The image processing system according to claim 5, wherein the determination rule includes one or more of the following: a condition for identifying the determination group based solely on the content of a named entity extracted from the handwritten character; a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the content of the handwritten character; and a condition for identifying the determination group based on the named entity corresponding to the handwritten character and the item group associated with printed characters present around the handwritten character.

7. The image processing system according to claim 1, further comprising an output processing unit that outputs recognition result information generated based on the identification result by the identification processing unit and the extraction result by the extraction processing unit.

8. An image processing apparatus comprising: an image acquisition processing unit that acquires the image data to be processed in the image processing system according to claim 7; and a reception processing unit that accepts input of information of pre-set input items based on the extraction results output from the output processing unit.

9. An image processing method comprising: an extraction step in which one or more processors extract handwritten characters contained in an image represented by an image data to be processed; and an identification step in which, based on the result of executing a named entity recognition process to extract named entities from the handwritten characters extracted by the extraction step and a plurality of item groups set in advance in association with the image data, the item group corresponding to the handwritten characters.