Information Processing System, Method, and Program

The information processing system addresses the cognitive burden of character recognition corrections by generating natural sentence replacement candidates for multi-character corrections, enhancing the efficiency of the correction process.

JP7693410B2Active Publication Date: 2025-06-17PFU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021103109
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-22
Publication Date
2025-06-17
Estimated Expiration
2041-06-22

AI Technical Summary

Technical Problem

Conventional character recognition correction methods impose a significant cognitive burden on users due to the need to distinguish between similar candidate characters for each character in a string, and often require multiple selections and corrections.

Method used

An information processing system that includes a character recognition result acquisition means, a speculation target designation means, a speculation means, and a replacement candidate generation means to generate replacement candidate character strings that include the correct characters for one or more misrecognized characters, presented as natural sentences for user selection.

Benefits of technology

This approach reduces the user's burden in correcting character recognition errors by allowing for the selection of natural sentence replacement candidates, thereby simplifying the correction process and reducing the need for multiple selections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693410000001
    Figure 0007693410000001
  • Figure 0007693410000002
    Figure 0007693410000002
  • Figure 0007693410000003
    Figure 0007693410000003
Patent Text Reader

Abstract

To provide an information processing system, method and program, capable of relieving a user's burden in correcting character strings obtained by character recognition.SOLUTION: An information processor 1 as an information processing system includes: a character recognition result acquisition section 22 which acquires a recognized character string obtained by character recognition of a character string image; a prediction object specification section 25 which specifies one or more prescribed positions of the recognized character string as one or more prediction object positions; a prediction section 26 which predicts a correct answer character to enter each prediction object position specified by the prediction object specification section 25; and a replacement candidate character generation section 27 which generates one or more replacement candidate character strings including the at least one correct answer character as a character string for replacement with at least a part of the recognized character string.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to character recognition technology.

Background Art

[0002] Conventionally, various technologies have been proposed to reduce the user's workload when correcting the results of character recognition (see Patent Documents 1 to 4).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventionally, the correction of a character string obtained by character recognition has been performed on a character-by-character basis. Also, when candidate characters for correction are presented, they are for correcting misrecognition in character recognition, so the glyphs of the presented candidate characters are similar to each other. For this reason, it is a great cognitive burden on the user to distinguish the correct character from the presented candidate characters.

[0005] In view of the above problems, an object of the present disclosure is to reduce the burden on the user related to the correction of a character string obtained by character recognition.

Means for Solving the Problems

[0006] An example of the present disclosure includes a character recognition result acquisition means for acquiring a recognition character string obtained by performing character recognition on a character string image, a speculation target designation means for designating a predetermined position among the recognition character strings as a speculation target position, a speculation means for speculating a correct character to be entered in the speculation target position designated by the speculation target designation means, and a replacement candidate generation means for generating a replacement candidate character string including the correct character as a character string for replacing at least a part of the recognition character string. A natural sentence determination means for determining whether or not the replacement candidate string is a natural sentence It is an information processing system including the above.

[0007] The present disclosure can be grasped as an information processing apparatus, a system, a method executed by a computer, or a program to be executed by a computer. Further, the present disclosure can also be grasped as a recording medium in which such a program is recorded and can be read by a computer or other devices, machines, etc. Here, a recording medium readable by a computer or the like refers to a recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read from a computer or the like.

Effects of the Invention

[0008] According to the present disclosure, it is possible to reduce the burden on the user related to the correction of the character string obtained by character recognition.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments of an information processing apparatus, a system, a method, and a program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are examples of embodiments, and do not limit the information processing apparatus, system, method, and program according to the present disclosure to the specific configurations described below. In practice, specific configurations according to the implementation mode may be appropriately adopted, and various improvements and modifications may be made.

[0011] In the present embodiment, embodiments when the information processing apparatus, system, method, and program according to the present disclosure are implemented in an electronic form document quality confirmation work support system will be described. However, the information processing apparatus, system, method, and program according to the present disclosure can be widely used for character recognition technology, and the application target of the present disclosure is not limited to the examples shown in the embodiments.

[0012] Conventionally, to confirm and correct the results of optical character recognition (OCR), the user visually compares the string image that is the object of character recognition with the recognized string obtained by OCR. When a misrecognized character is found in the recognized string, the user performs the following operations for each misrecognized character: move the cursor to the position of the misrecognized character, delete the misrecognized character, and input the correct character. Furthermore, when the misrecognized character is a single-byte character, the input of the correct character can be simply performed by an operation such as pressing the key of the correct character. However, when inputting multi-byte characters such as Chinese characters and kana, it is necessary to perform a conversion process (Japanese conversion process in the case of Japanese). The conversion process includes operations such as inputting conversion elements, searching for and selecting the correct character from the conversion candidates presented by the system based on the conversion elements, and finalizing. The work of the user when confirming and correcting a recognized string including multi-byte characters is more complicated than that when confirming and correcting a recognized string including only single-byte characters.

[0013] Also, in response to the above problems, conventionally, when a misrecognized character is selected by the user, a method has been proposed in which a list of candidate characters that were not adopted during character recognition is displayed, and the user is allowed to select the correct character from the list. However, even when such a method is adopted, since the correction of misrecognized characters is in units of characters, when there are multiple misrecognized characters, it is necessary to display the list each time and have the user select. Furthermore, since the candidate characters are for correcting misrecognitions in character recognition, characters with similar glyphs are arranged in the list of candidate characters, imposing a large cognitive burden on the user to distinguish the correct character. In addition, since the content of the candidate character list depends on the recognition accuracy of the OCR system employed, there may be a case where the correct character is not present in the candidate character list in the first place.

[0014] Therefore, in the system described in this embodiment, it is possible to obtain the correct character from outside the candidate characters in character recognition, and present a replacement candidate string that includes the correct characters for one or more misrecognized characters and is a natural phrase for the user to select, thereby reducing the burden on the user for correcting the character string obtained by character recognition. However, when implementing the technology according to the present disclosure, it is not necessary to solve all the problems described above by adopting all the configurations described below. When implementing the technology according to the present disclosure, it is also possible to solve some of the problems described above by adopting a part of the configurations described below.

[0015] <Configuration of the System> FIG. 1 is a schematic diagram showing the configuration of the system according to this embodiment. The system according to this embodiment includes a scanner 3 and an information processing apparatus 1 that are communicably connected to each other via a network or other communication means.

[0016] The information processing apparatus 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage device 14 such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or an HDD (Hard Disk Drive), an input device 15 such as a keyboard, a mouse, or a touch panel, an output device 16 such as a display, and a communication unit 17, etc. However, regarding the specific hardware configuration of the information processing apparatus 1, appropriate omission, replacement, or addition is possible according to the implementation mode. Further, the information processing apparatus 1 is not limited to a device composed of a single housing. The information processing apparatus 1 may be realized by a plurality of devices using so-called cloud or distributed computing technologies, etc.

[0017] The scanner 3 is a device that acquires image data by imaging a document, business card, receipt, photo / illustration, or other original set by the user. In this embodiment, the scanner 3 is exemplified as a device for acquiring a target image, but the device used for acquiring an image is not limited to a so-called scanner. For example, it is also possible to image the target using a digital camera or a camera sensor built into a smartphone / tablet to obtain an image.

[0018] The scanner 3 according to this embodiment has a function of transmitting the image data obtained by imaging to the information processing device 1 via a network. Further, the scanner 3 may further have a user interface for enabling character input / output and item selection, such as a touch panel display and a keyboard, and a web browsing function and a server function. The communication means and hardware configuration, etc. of the scanner that can adopt the method according to this embodiment are not limited to the examples in this embodiment.

[0019] FIG. 2 is a diagram showing an outline of the functional configuration of the information processing device 1 according to this embodiment. In the information processing device 1, the program recorded in the storage device 14 is read into the RAM 13 and executed by the CPU 11, and each hardware provided in the information processing device 1 is controlled, whereby a character recognition unit 21, a character recognition result acquisition unit 22, an output unit 23, a selection reception unit 24, a speculation target designation unit 25, a speculation unit 26, a replacement candidate generation unit 27, a natural sentence determination unit 28, a correction unit 29, and a replacement unit 30 are provided. In this embodiment and other embodiments described later, each function provided in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but a part or all of these functions may be executed by one or more dedicated processors.

[0020] The character recognition unit 21 performs character recognition (OCR) processing on the character string image included in the input image data and outputs the recognized character string. Here, the image data may be image data obtained by imaging with a scanner 3 or the like, or may be an image generated as image data from the beginning. In the present embodiment, the character recognition unit 21 obtains one or more candidate characters for each character included in the character string image, and creates and outputs a recognized character string by adopting the character with the highest accuracy. However, the character recognition unit 21 can also output the candidate characters that were not adopted in the recognized character string because their accuracy is second or lower as reference data. In the present embodiment, an example in which the information processing apparatus 1 includes the character recognition unit 21 will be described. However, the information processing apparatus 1 only needs to be able to acquire the recognized character string that is the result of character recognition, and does not necessarily need to include the character recognition unit 21. The character recognition unit 21 may be provided in a device other than the information processing apparatus 1, such as a scanner 3, an external device, or a server.

[0021] The character recognition result acquisition unit 22 acquires the recognized character string (the character string obtained by character recognition) obtained by performing character recognition on the character string image. Also, in the present embodiment, the character recognition result acquisition unit 22 further acquires, in addition to the recognized character string, the candidate characters that were not adopted in the recognized character string during character recognition.

[0022] The output unit 23 displays the character string image side by side with the recognized character string of the character string image.

[0023] FIG. 3 is a diagram showing an example of the confirmation / correction screen 5 displayed on the display of the information processing apparatus 1 in the present embodiment. On the confirmation / correction screen 5, the character string image 51 to be read and the recognized character string 52 recognized from the character string image 51 are displayed side by side (vertically when written horizontally). At this time, the recognized character string 52 may be displayed in an arbitrary font.

[0024] The selection reception unit 24 receives the selection of the characters to be corrected from the user. The user who views the confirmation / correction screen 5 compares the string images 51 and the recognized characters 52 displayed side by side to check whether the character recognition result by the character recognition unit 21 is correct. If there is a character misrecognized in the character recognition result, the user designates the character as a character to be corrected (hereinafter referred to as "character to be corrected") by aligning the cursor 53 with the character, thereby setting the character to be corrected in a selected state. At this time, the operation of the cursor 53 may be performed by an input operation using a direction key or the like, or may be performed by an input operation using a pointing device.

[0025] The speculation target designating unit 25 designates the speculation target location by generating a query string in which the characters at a predetermined location (hereinafter referred to as "speculation target location") in the recognized characters 52 are replaced with predetermined mask characters. Here, the speculation target designating unit 25 generates a query string in which at least the characters to be corrected designated by the user are replaced with mask characters as the speculation target location.

[0026] In the present embodiment, in order to further reduce the operation burden on the user and obtain more natural correction candidates as a whole, the speculation target designating unit 25 generates a query string in which the characters to be corrected and any characters after the characters to be corrected are replaced with mask characters. By doing so, even when there are misrecognized characters other than the characters to be corrected explicitly designated by the user, without waiting for the user to select the misrecognized characters, a natural replacement candidate string including the correct characters for a plurality of characters is proposed, and a plurality of characters can be corrected together. The speculation target designating unit 25 generates a plurality of query strings obtained by combinations of the characters to be corrected replaced by the mask characters and any characters after the characters to be corrected replaced by the mask characters.

[0027] Furthermore, in addition to replacement with mask characters, the speculation target designating unit 25 may generate a query string by replacing the characters corresponding to the candidate characters in the recognition character string 52 with those characters. In this case, the speculation target designating unit 25 generates a plurality of query strings obtained by a combination of replacement with mask characters and replacement with candidate characters. As described above, the candidate characters are the characters not adopted in the recognition character string 52 during character recognition by the character recognition result acquisition unit 22. By doing so, it becomes possible to speculate the correct character in consideration of the characters that were candidates but not adopted in the process by the character recognition unit 21.

[0028] The speculation unit 26 speculates the correct character to be entered at the speculation target location designated by the speculation target designating unit 25. In the present embodiment, the speculation unit 26 speculates the correct character by speculating the character to be entered at the speculation target location designated by the mask characters in the generated query string.

[0029] The replacement candidate generation unit 27 generates a replacement candidate character string in which the mask characters in the query string are replaced with the speculation results (correct characters) by the speculation unit 26 as a character string for replacing at least a part of the recognition character string 52. Here, when a plurality of query strings are generated, the replacement candidate generation unit 27 generates a plurality of replacement candidate character strings corresponding to each of the plurality of query strings.

[0030] The natural sentence determination unit 28 determines whether the replacement candidate character string is a natural sentence. Further, the natural sentence determination unit 28 further calculates, as a priority used when displaying a plurality of replacement candidate character strings, an index indicating the degree to which the replacement candidate character string is a natural sentence.

[0031] When it is determined that the replacement candidate character string is not a natural sentence, the correction unit 29 performs a correction of deleting the last character on the replacement candidate character string. Then, when the replacement candidate character string is corrected by the correction unit 29, the natural sentence determination unit 28 further determines whether the corrected replacement candidate character string is a natural sentence.

[0032] The replacement unit 30 replaces the corresponding character in the recognition character string 52 with the replacement candidate character string selected by the user.

[0033] <Flow of processing> Next, the flow of processing executed by the information processing apparatus 1 according to the present embodiment will be described. Note that the specific content and processing order of the processing described below are examples for implementing the present disclosure. The specific processing content and processing order may be appropriately selected according to the embodiment of the present disclosure.

[0034] In the present embodiment, first, character recognition by the character recognition unit 21 is executed. Then, the character recognition result acquisition unit 22 acquires the recognition character string 52 output by the character recognition unit 21 and candidate characters (characters that were not adopted in the recognition character string 52 because their accuracy is after the second). The output unit 23 outputs the character string image 51 that was the target of character recognition and the recognition character string 52 acquired from the character recognition unit 21 side by side for the user's confirmation and correction work (see FIG. 3). The user visually compares the character string image 51 and the recognition character string 52, and if a misrecognized character is found in the recognition character string 52, moves the cursor 53 to the position of the misrecognized character. The selection reception unit 24 receives the movement operation of the cursor 53 by the user as the selection of the character to be corrected, and sets the misrecognized character in a selected state.

[0035] FIG. 4 is a flowchart showing an outline of the flow of candidate presentation processing according to the present embodiment. The processing shown in this flowchart is executed when an arbitrary character (misrecognized character) in the recognition character string 52 is selected by a user operation.

[0036] In step S101, a query string for inference processing is generated. The inference target specifying unit 25 generates one or more query strings for input to the character inference learning model in the correct character inference process using the character inference learning model by the inference unit 26 described later. In the present embodiment, the inference target specifying unit 25 generates a query string by replacing one or more characters after the character to be corrected selected by the user with a combination of a mask character and a candidate character. Here, the mask character is a special character for indicating the inference target location to the character inference learning model, and in the present embodiment, "@" is used. However, other symbols may be used as the mask character, and a method other than the mask character may be adopted for specifying the inference target. Also, as described above, the candidate character is a character that was not adopted in the recognition character string 52 because it was determined that the accuracy in the character recognition process was the second or lower by the character recognition result acquisition unit 22.

[0037] In the present embodiment, the inference target specifying unit 25 generates one or more query strings by replacing one or more characters after the character to be corrected with all possible combinations of a mask character and a candidate character. However, an upper limit may be set on the number of query strings generated. By doing so, the load of the candidate presentation process shown in this flowchart can be made a load according to the processing capacity of the information processing apparatus 1. Also, the upper limit on the number of characters included in the query string after the character to be corrected, the range of candidate characters used for generating the query string (up to which candidate characters with what accuracy are used), the upper limit on the number of mask characters included in the query string, etc. may also be arbitrarily settable.

[0038] Fig. 5 is a diagram showing an example of a query string generated by the information processing device 1 in this embodiment. According to the example shown in Fig. 5, it can be seen that one or more characters following the character to be corrected are replaced with multiple combinations of a mask character (the special character "@" in the figure) and a candidate character (the bold characters "雲" or "雲" and "パ" in the figure), and multiple query strings (1 to 625 in the figure) are generated. When the generation of the query string for the inference process is completed, the process proceeds to step S102.

[0039] In step S102, a correct character to be inserted in the position indicated by the masked character in the query string is guessed, and a replacement candidate string including the correct character is generated. The guess unit 26 gives one or more query strings generated in step S101 as input to a character guess learning model created by performing learning in advance, and obtains a guess result of a correct character to be inserted in the position indicated by the masked character in the query string as an output from the character guess learning model.

[0040] The guess result may be output as the correct character, or may be output as a character string (replacement candidate character string) corresponding to the query string including the correct character. When the correct character is obtained as the guess result, the replacement candidate generating unit 27 generates a replacement candidate character string in which the mask character in the query string is replaced with the correct character. Here, when multiple query strings are generated in step S101, multiple replacement candidate character strings corresponding to the multiple query strings are obtained. Then, the process proceeds to step S103. However, when only one replacement candidate character string is obtained, the processes from step S103 to step S108 may be skipped.

[0041] From step S103 to step S107, it is determined whether the replacement candidate string is a natural phrase. The natural phrase determination unit 28 gives each of the plurality of replacement candidate strings obtained in step S102 as an input to a natural phrase determination learning model created by performing learning in advance, and as an output from the natural phrase determination learning model, for each of the plurality of replacement candidate strings, a determination result as to whether the replacement candidate string is a natural phrase and an index indicating the degree to which the replacement candidate string is a natural phrase are obtained (step S103). In the present embodiment, as an index indicating the degree to which the replacement candidate string is a natural phrase, a probability that the replacement candidate string is a natural phrase (for example, a value between 0.0 and 1.0) is obtained. However, other aspects such as scores and ranks may be adopted as the index. Further, the determination result as to whether the replacement candidate string is a natural phrase may be such that a determination result that the replacement candidate string is a natural phrase is obtained when an index indicating the degree to which the replacement candidate string is a natural phrase is equal to or greater than a predetermined threshold (for example, 0.6) set in advance. However, the threshold for determining whether it is a natural phrase may be arbitrarily set.

[0042] When a determination result that the replacement candidate string is not a natural phrase is obtained (NO in step S104), the correction unit 29 corrects the replacement candidate string by deleting one character from the end (step S105). Here, if as a result of the correction of deleting one character from the end, the corrected replacement candidate string satisfies any of the determination skip conditions (YES in step S106), the process proceeds to step S107. Here, the determination skip condition means that the last character of the corrected replacement candidate string is a character corresponding to the character to be corrected selected by the user as a misrecognized character (in other words, the last character of the corrected replacement candidate string is a character corresponding to the mask character that first appears in the query string), or the corrected replacement candidate string matches a replacement candidate string for which natural phrase determination has already been performed. On the other hand, if the corrected replacement candidate string does not satisfy any of the determination skip conditions (NO in step S106), the process returns to step S103, and the natural phrase determination unit 28 further determines whether the corrected replacement candidate string is a natural phrase (step S103). That is, the process from step S103 to step S106 is repeatedly executed while deleting one character from the end until the target replacement candidate string is determined to be a natural phrase or any of the determination skip conditions is satisfied.

[0043] When it is determined that the target replacement candidate string is a natural phrase or any of the determination skip conditions is satisfied, it is determined whether the natural phrase determination for all the replacement candidate strings generated in step S102 has been completed (step S107). If there is an unjudged replacement candidate string (NO in step S107), the natural phrase determination unit 28 executes the process from step S103 to step S106 for the next unjudged replacement candidate string. That is, the process from step S103 to step S106 is repeatedly executed while changing the target replacement candidate string until the natural phrase determination for all the replacement candidate strings generated in step S102 is completed. When the natural phrase determination for all the generated replacement candidate strings is completed (YES in step S107), the process proceeds to step S108.

[0044] In steps S108 and S109, replacement candidate strings are output according to a predetermined priority order. The output unit 23 sorts the plurality of replacement candidate strings obtained by the processing up to step S107 in descending order of the degree / probability of being a natural sentence obtained from the natural sentence determination learning model, and further sorts them in descending order of the string length (step S108). By doing so, the plurality of replacement candidate strings are arranged in descending order of the string length, and for candidates with the same string length, they are arranged in descending order of the degree / probability of being a natural sentence.

[0045] Then, the output unit 23 causes the display device to display a replacement candidate string list including the generated one or more replacement candidate strings side by side with the character string image 51 that is the target of character recognition and the recognition string 52 obtained from the character recognition unit 21 (step S109). At this time, since the replacement candidate strings are sorted as described above in step S108, the output unit 23 can preferentially display longer replacement candidate strings by displaying the plurality of replacement candidate strings in the sorted order, and can also preferentially display replacement candidate strings with a higher degree of being a natural sentence based on the index calculated by the natural sentence determination unit 28. Therefore, the user can select a longer and more natural sentence with fewer operations and correct a larger number of misrecognized characters at once. After that, the processing shown in this flowchart ends.

[0046] FIG. 6 is a diagram showing an example of the replacement candidate string list 54 displayed on the confirmation / correction screen 5 in the present embodiment. As described with reference to FIG. 3, on the confirmation / correction screen 5, the character string image 51 to be read and the recognized character string 52 recognized from the character string image 51 are displayed side by side. The replacement candidate string list 54 is displayed in the order sorted in step S108 below the recognized character string 52 including the misrecognized character that the user has selected by moving the cursor 53 thereto. In the present embodiment, when a character in the recognized character string 52 is selected due to the movement of the cursor 53, the candidate presentation process is executed, and the replacement candidate string list 54 is automatically displayed. Note that the length of the waiting time from when the selection state is entered until the replacement candidate string list 54 is displayed can be changed by setting. Here, the output unit 23 highlights (displayed in bold in FIG. 6) the characters (the characters to be corrected) that are the difference between the recognized character string 52 and the replacement candidate strings in the list 54.

[0047] FIG. 7 is a diagram showing an example of a state in which a replacement candidate string is selected from the replacement candidate string list 54 in the present embodiment. The output unit 23 highlights the portion of the recognized character string 52 corresponding to the replacement candidate string in the selected state. In the example shown in FIG. 7, it can be seen that when the replacement candidate string "Comprehensive Exhibition" with index number 4 is in the selected state, the corresponding portion "Comprehensive Exhibition Book Fair" in the recognized character string 52 is highlighted as the range to be replaced.

[0048] In this embodiment, the user can select a replacement candidate string desired by the user using various operation methods, such as (1) an operation method of changing a replacement candidate string in a selected state by operating a mouse wheel and confirming the selection of a desired replacement candidate string by clicking, (2) an operation method of changing a replacement candidate string in a selected state by operating a direction key on a keyboard and confirming the selection of a desired replacement candidate string by operating an Enter key, and (3) an operation method of confirming the selection of a replacement candidate string by operating a number key corresponding to an index number displayed near the replacement candidate string (in the examples of Figs. 6 and 7, the numbers 1 to 5 displayed to the left of the replacement candidate string). However, the specific operation method for selecting a replacement candidate string desired by the user from the replacement candidate string list 54 is not limited to the example in this embodiment. For example, the replacement candidate string desired by the user may be selected using a pointing device.

[0049] When the user selects and confirms a desired replacement candidate string from the displayed replacement candidate string list 54, the replacement unit 30 replaces (corrects) the corresponding portion of the recognized string 52 with the replacement candidate string selected by the user. Using the example shown in FIG. 7, when the user selects the replacement candidate string with index number 1, the replacement unit 30 replaces the entire recognized string "Sougou Tenhonkai Haneru" with the replacement candidate string "Sougou Tenkai Panel". That is, according to the system according to the present embodiment, the two erroneously recognized characters "Hon" and "Ha" included in the recognized string "Sougou Tenhonkai Haneru" can be corrected at the same time by presenting "Sougou Tenkai Panel", which is a natural phrase with a meaning as a whole, to the user as the highest priority candidate and having the user select it. On the other hand, when the user selects the replacement candidate string with index number 4, the replacement unit 30 replaces the portion "Sougou Tenhonkai" corresponding to the replacement candidate string with index number 4 in the recognized string "Sougou Tenhonkai Haneru" with the replacement candidate string "Sougou Tenkai". In this case, too, by presenting the user with "Sogo Tenshikai," which is a natural phrase that has overall meaning, and allowing the user to select it, the erroneously recognized character "hon" contained in the recognized character string "Sogo Tenshikai" can be corrected.

[0050] When the replacement (correction) of the corresponding part of the recognition string 52 using the replacement candidate string is completed, the confirmation operation of the recognition string 52 by the user is resumed. When a misrecognized character is discovered again by the user and the misrecognized character is set to the selected state, the candidate presentation process is executed again. For example, in the example shown in FIG. 7, when the replacement candidate string with index number 4 is selected by the user, the misrecognized character "ha" remains, but the user can execute the candidate presentation process again by moving the cursor 53 to make the misrecognized character "ha" in the selected state. The user performs the operation of confirming and correcting the recognition string 52 by repeating a series of operations until no misrecognized character is found.

[0051] <Learning model> In the technology according to the present disclosure, the type of the specific learning model that can be adopted for the character speculation learning model is not limited, and any model that can take a query string as an input and output a speculation result of the correct character to be entered at the position indicated by the mask character may be used. For example, the correct character is output as a result of multi-class classification targeting all characters in the first and second levels of Shift_JIS. Hereinafter, an example of the pre-learning process for creating a character speculation learning model that can be used in the present embodiment will be described.

[0052] In pre-training, a large number of phrases (nouns or compound words) are collected, and any proportion of characters (at least one character) in these phrases are randomly selected and replaced with masked characters. Using a learning model, the characters at the positions indicated by the masked characters are inferred, the error between the characters before replacement with the masked characters and the inferred characters is calculated, and the parameters of the learning model are corrected by the error backpropagation method. By executing a series of such processes any number of times, a learning model for character inference is created. For example, one-hot vectors or the like may be used for calculating the error. Note that the proportion of characters replaced with masked characters can be arbitrarily set. For example, by masking a proportion of characters according to the misrecognition rate of the OCR system used in the implementation, the prediction accuracy can be improved. Also, the masked characters may be randomly selected as described above, but may be selected such that the masking rate of the characters at the rear is higher than that of the characters at the front of the character string. This is because when confirming and correcting the character recognition results, the character string is corrected in order from the front to the rear, and in view of the entire correction work, there is a situation where the characters at the rear of the character string are masked more frequently. This is to create a learning model that can accurately infer the characters at the rear in accordance with such a situation.

[0053] In the technology according to the present disclosure, the types of specific learning models that can be adopted for the natural sentence determination learning model are not limited, and any model that can output a determination result as to whether an input is a meaningful natural sentence, the degree to which it is a natural sentence, or the probability that it is a natural sentence, etc., as long as it can output such information for an arbitrary character string as input. An example of the pre-training process for creating a natural sentence determination learning model that can be used in the present embodiment will be described below.

[0054] In pre-training, a large number of phrases (nouns or compound words) are collected, and negative examples are created by randomly selecting multiple characters in these phrases and replacing them with random other characters. Also, phrases for which replacement is not performed are used as positive examples as they are. The positive and negative examples may be assigned by dividing the collected phrases into positive or negative examples with a 50% probability. Then, using these phrases prepared as positive and negative examples as input, the learning model is used to infer whether the input phrase is a natural phrase or not, and the error between the probability that the input phrase is a natural phrase (a value between 0.0 and 1.0) as the inference result and the correct answer (1.0 if the input phrase is a positive example, 0.0 if it is a negative example) is calculated, and the parameters of the learning model are corrected by the error backpropagation method. By executing this series of processes any number of times, a learning model for natural language sentence determination is created.

[0055] <Variation> Note that in the above-described embodiment, an example of specifying the inference target location by replacement with a mask character has been described, but the method of specifying the inference target location is not limited to the examples in this embodiment. For example, the inference target location may be specified by the character number (the number of characters from the beginning of the character string) in the recognized character string.

Description of Signs

[0056] 1 Information processing device

Claims

1. A character recognition result acquisition means for acquiring a recognition character string obtained by performing character recognition on a character string image; A speculation target designation means for designating a predetermined location in the recognition character string as a speculation target location; A speculation means for speculating a correct character to be entered in the speculation target location designated by the speculation target designation means; A replacement candidate generation means for generating a replacement candidate character string including the correct character as a character string for replacing at least a part of the recognition character string; A natural sentence determination means for determining whether the replacement candidate character string is a natural sentence; An information processing system comprising the above.

2. The speculation target designation means designates the speculation target location by generating a query character string in which the character at the speculation target location in the recognition character string is replaced with a predetermined mask character, The speculation means speculates the character to be entered in the speculation target location designated by the mask character in the generated query character string, The replacement candidate generation means generates the replacement candidate character string in which the mask character in the query character string is replaced with the speculation result by the speculation means. The information processing system according to claim 1.

3. Further comprising a selection reception means for receiving a selection of a character to be corrected from a user, The speculation target designation means generates a query character string in which at least the character to be corrected is replaced with the mask character. The information processing system according to claim 2.

4. The speculation target designation means generates a query character string in which the character to be corrected and any character after the character to be corrected are replaced with the mask character. The information processing system according to claim 3.

5. The speculation target specifying means generates a plurality of query character strings obtained by combinations of the character to be corrected replaced by the mask character and any characters after the character to be corrected replaced by the mask character, The replacement candidate generating means generates a plurality of replacement candidate character strings corresponding to the plurality of query character strings, The information processing system according to claim 4.

6. The character recognition result acquisition means further acquires candidate characters not adopted for the recognition character string during the character recognition, in addition to the recognition character string, The speculation target specifying means generates the query character string by replacing the character corresponding to the candidate character in the recognition character string with the candidate character, in addition to the replacement by the mask character, The information processing system according to any one of claims 2 to 5.

7. The speculation target specifying means generates a plurality of query character strings obtained by a combination of replacement by the mask character and replacement by the candidate character, The replacement candidate generating means generates a plurality of replacement candidate character strings corresponding to the plurality of query character strings, The information processing system according to claim 6.

8. When it is determined that the replacement candidate character string is not a natural phrase, correction means for performing correction to delete the last character from the replacement candidate character string is further provided, When the replacement candidate character string is corrected by the correction means, the natural phrase determination means further determines whether or not the corrected replacement candidate character string is a natural phrase, The information processing system according to claim 1.

9. The natural phrase determination means further calculates an index indicating the degree to which the replacement candidate character string is a natural phrase as a priority used when displaying the plurality of replacement candidate character strings, The information processing system according to claim 1 or 8.

10. Further comprising output means for causing the display device to display the generated replacement candidate string. The information processing system according to any one of claims 1 to 9.

11. When the output means displays a plurality of the replacement candidate strings, the output means preferentially displays a longer replacement candidate string. The information processing system according to claim 10.

12. The output means highlights characters that are the difference from the recognition string among the replacement candidate strings. The information processing system according to claim 10 or 11.

13. The output means further displays the recognition string and highlights a portion of the recognition string corresponding to the replacement candidate string selected by the user. The information processing system according to any one of claims 10 to 12.

14. The output means displays the character string image side by side with the recognition string of the character string image. The information processing system according to any one of claims 10 to 13.

15. Further comprising replacement means for replacing a corresponding character string in the recognition string with a replacement candidate string selected by the user. The information processing system according to any one of claims 1 to 14.

16. A computer A character recognition result acquisition step of acquiring a recognition string obtained by character recognition of a character string image, A speculation target designation step of designating a predetermined location in the recognition string as a speculation target location, A speculation step of speculating a correct character to enter the speculation target location designated in the speculation target designation step, A replacement candidate generation step of generating a replacement candidate string including the correct character as a string for replacing at least a part of the recognition string; A natural sentence determination step of determining whether the replacement candidate string is a natural sentence; A method for executing.

17. A computer, A character recognition result acquisition means for acquiring a recognition string obtained by character recognition of a string image; A speculation target designation means for designating a predetermined position in the recognition string as a speculation target position; A speculation means for speculating the correct character to be entered at the speculation target position designated by the speculation target designation means; A replacement candidate generation means for generating a replacement candidate string including the correct character as a string for replacing at least a part of the recognition string; A natural sentence determination means for determining whether the replacement candidate string is a natural sentence; A program for causing it to function as.

Citation Information

Patent Citations

  • Correcting device for japanese document

    JP1984116882A

  • Character correcting method of character recognizing device

    JP1985254387A

  • Chinese character ocr

    JP1993135199A

  • Character recognizing device, erroneous recognition character correcting method and occidental document processor

    JP1993298495A

  • Japanese character reader

    JP1993342402A