Estimation support device, estimation support method, and program
The estimation support device addresses the inefficiencies in OCR-based data entry by converting OCR character strings into wildcards, searching databases, and calculating similarities, resulting in improved accuracy and reduced labor in data entry processing.
Patent Information
- Application Number
- JP2023199876
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
Existing systems that assist with data entry using OCR are unable to reduce labor requirements, leading to inefficiencies and potential typing errors.
An estimation support device and method that convert parts of character strings generated by OCR into wildcards, search a database for corresponding strings, and calculate similarities between OCR and database strings to improve data entry accuracy and reduce labor.
The solution enables more accurate and efficient data entry processing by reducing the need for manual correction and minimizing labor requirements, thereby improving the overall accuracy and speed of data entry tasks.
Smart Images

Figure 2025086069000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an inference support device and an inference support method for supporting input processing in data input using optical character recognition / reader (OCR), and further relates to a program for realizing these. [Background technology]
[0002] In recent years, with the development of computer systems, various data are processed and stored on computers. For this reason, there is a need to convert information written on paper into digital data that can be used by computers using OCR.
[0003] An example of the need to convert paper information into digital data is currency exchange processing by financial institutions. To be more specific, financial institutions have traditionally converted paper-based information into digital data by applying OCR to paper currency transfer request forms filled out by customers.
[0004] However, since it is difficult for OCR to completely recognize all characters written on paper, the operator must supplement the information that is not fully recognized by OCR by typing. Furthermore, typing by operators is not always perfect, and there is a possibility of typing errors.
[0005] For this reason, systems that assist with input have been proposed (see, for example, Patent Document 1). Such systems have the function of storing information used in past transactions, such as sender information and recipient information, and searching for information to be supplemented from the stored information using information with a high recognition rate and typed information as keys. Such systems are expected to reduce typing errors by operators. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] JP 2003-6441 A Summary of the Invention [Problem to be solved by the invention]
[0007] In recent years, there has been a growing demand for labor-saving by reducing the number of workers required for simple tasks such as data entry. However, the above-mentioned system only has a function of presenting candidates when an operator enters data, and the introduction of the above-mentioned system does not reduce the number of operators. The above-mentioned system has a problem in that it cannot reduce labor. Therefore, for example, when searching information stored in a database based on a recognition result by OCR and estimating data corresponding to the recognition result, it is preferable to have support for making a more accurate estimation.
[0008] An example of an objective of the present disclosure is to solve the above problems and provide assistance in achieving labor savings with higher accuracy in data entry processing using OCR. [Means for solving the problem]
[0009] In order to achieve the above object, an estimation support device according to one aspect of the present disclosure includes: a wildcard conversion unit that converts a part of a character string generated by optical character recognition, the character string including a plurality of items, into a wildcard; a search processing unit that searches a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracts a plurality of corresponding character strings; A calculation setting unit that sets some or all of the plurality of items as target items and sets a plurality of types of the target items; a similarity calculation unit that calculates, for each type of target item, a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, using the target item; Equipped with It is characterized by:
[0010] In order to achieve the above object, an estimation support method according to one aspect of the present disclosure includes: a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; Equipped with It is characterized by:
[0011] Furthermore, in order to achieve the above object, a program according to one aspect of the present disclosure includes: On the computer, a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; Execute the command. Effect of the Invention
[0012] As described above, according to the present disclosure, it is possible to provide assistance for achieving labor saving with higher accuracy in data entry processing using OCR. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 is a block diagram showing a schematic configuration of an example of an estimation support device. [Diagram 2] FIG. 2 is a block diagram specifically showing the configuration of an input support device including an estimation support device. [Diagram 3] FIG. 3 is a diagram illustrating an example of a process performed by the unreadability calculation unit. [Figure 4] FIG. 4 is a diagram illustrating an example of a conversion process performed by the wildcard conversion unit. [Diagram 5] FIG. 5 is a diagram illustrating an example of a search process performed by the search processor. [Figure 6] FIG. 6 is a diagram showing an example of a search result by the search processing unit. [Figure 7] FIG. 7 is a diagram illustrating an example of processing by the calculation setting unit and the priority setting unit. [Figure 8] FIG. 8 is a diagram illustrating an example of the processing contents by the similarity calculation unit. [Figure 9] FIG. 9 is a diagram showing an example of a result of sorting based on priority. [Figure 10] FIG. 10 is a flow diagram showing an example of the operation of the input support device including the estimation support device. [Figure 11] FIG. 11 is a block diagram showing an example of a computer that realizes an input support device (estimation support device). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] (Embodiment) Hereinafter, in the embodiment, an example of an estimation support device, an estimation support method, and a program will be described with reference to FIGS.
[0015] [Device configuration] First, a schematic configuration of an estimation support device will be described with reference to Fig. 1. Fig. 1 is a block diagram showing a schematic configuration of an example of an estimation support device.
[0016] The estimation support device 10 shown in Fig. 1 is a device for supporting an estimation process in an input process in data input using OCR. As shown in Fig. 1, the estimation support device 10 includes a wildcard conversion unit 11, a search processing unit 12, a calculation setting unit 13, and a similarity calculation unit 14.
[0017] The wildcard conversion unit 11 converts a part of a character string (hereinafter referred to as an "OCR character string") generated by optical character recognition (OCR) and including a plurality of items into a wildcard. The search processing unit 12 uses the character string partially converted into a wildcard to search a database 30 in which information consisting of character strings including a plurality of items is registered, and extracts a plurality of corresponding character strings (hereinafter, character strings from the database 30 are also referred to as "DB character strings"). The calculation setting unit 13 sets a part or all of the plurality of items as target items, and sets a plurality of types of target items. The similarity calculation unit 14 performs a process of calculating the similarity between each of the plurality of DB character strings extracted from the database 30 and the OCR character string using the target item, for each type of target item.
[0018] In this way, the estimation support device 10 does not use the OCR string as is, but instead performs a search using a string partially converted into a wildcard. This also allows multiple candidate DB strings to be obtained, and further allows various similarities between each DB string and the OCR string to be calculated. In other words, the estimation support device 10 provides various similarities, thereby providing support for achieving more accurate labor savings in data entry processing using OCR.
[0019] Next, the configuration and functions of the estimation support device will be described in more detail with reference to Figures 2 to 7. Figure 2 is a block diagram specifically showing the configuration of an input support device including the estimation support device.
[0020] 2, the estimation support device 10 is provided in an input support device 1. The input support device 1 includes the estimation support device 10 and an information estimation unit 20.
[0021] The estimation support device 10 includes, in addition to the above-mentioned wildcard conversion unit 11, search processing unit 12, calculation setting unit 13, and similarity calculation unit 14, an unreadability calculation unit 15 and a priority setting unit 16. A database 30 is connected to the input support device 1.
[0022] In this embodiment, the subject of OCR is data composed of a set of information divided into items, and a character string is generated for each item by OCR. That is, the OCR character string includes a plurality of items. The OCR character string includes a character string for each of the plurality of items, and a set of character strings for each of the plurality of items is the OCR character string. A specific example of the subject of OCR is an application form used for money exchange processing, which is written on a paper medium and divided into a plurality of items.
[0023] In this embodiment, the database 30 registers information composed of character strings for each of the above-mentioned items. Specifically, the database accumulates digital data of information written on past application forms. In the example of Figs. 1 and 2, the database 30 is provided outside the input support device 1, but this is just one example. The database 30 may be provided inside the input support device 1. The character strings registered in the database 30 are DB character strings. As with OCR character strings, DB character strings include character strings for multiple items, and a collection of character strings for multiple items is a DB character string.
[0024] The unreadability calculation unit 15 calculates the unreadability rate for each item of the OCR character string. Specifically, the unreadability calculation unit 15 calculates the unreadability rate by dividing the number of characters not recognized by OCR by the total number of characters read by OCR.
[0025] In this embodiment, the unreadability calculation unit 15 judges whether the calculated unreadability rate is equal to or greater than a threshold value. In this case, the wildcard conversion unit 11, the search processing unit 12, the calculation setting unit 13, the priority setting unit 16, the similarity calculation unit 14, and the information estimation unit 20 execute processing for items whose unreadability rate is less than the threshold value.
[0026] In this way, by excluding items with a high unread rate from the processing target, a decrease in the accuracy of the search by the search processing unit 12 is suppressed, and as a result, a decrease in the accuracy of the estimation by the information estimation unit 20 is also suppressed. Figure 3 is a diagram for explaining an example of the processing performed by the unread rate calculation unit.
[0027] In the example of FIG. 3, the OCR result is shown for each item. As shown in FIG. 3, examples of the items of the OCR character string include "phone number", "client name", "subject", "account number", "recipient name", and "other items". The items of the DB character string are the same as the examples of the items of the OCR character string. In this embodiment, any character that is not recognized as a character in the OCR result is indicated by "?". In this embodiment, any one character is indicated by "x". Also, as shown in FIG. 3, a threshold value of the unreadability rate is set for each item. The threshold value of the unreadability rate is set in advance in the unreadability rate calculation unit 15, but may be configured to be changeable. Since the threshold value of the unreadability rate is set for each item in this way, the unreadability rate calculation unit 15 compares the corresponding threshold value and the unreadability rate for each item of the OCR character string to identify a character string whose unreadability rate is equal to or greater than the threshold. Specifically, in the example of FIG. 3, the unreadability rate calculation unit 15 identifies "account number" as an item whose unreadability rate is equal to or greater than the threshold. Items of the OCR character string whose unreadability rate is equal to or greater than the threshold are excluded from the search process in the search processing unit 12. An item of the OCR character string whose unreadability rate is equal to or greater than a threshold value may be subject to a similarity calculation with a corresponding item of a DB character string, which will be described later, or may be excluded from the similarity calculation.
[0028] In this embodiment, the wildcard conversion unit 11 converts a part of the OCR character string into a wildcard for each item. Fig. 4 is a diagram showing an example of the conversion process by the wildcard conversion unit.
[0029] In the example of FIG. 4, for example, one of the characters in the OCR character string is converted to a wildcard "*" for the items "phone number" and "client name" shown in FIG. 3. As a result, a plurality of character strings (hereinafter referred to as "wildcard character strings") are generated for each item in one OCR character string. The wildcard conversion unit 11 converts unreadable characters in the OCR character string whose unreadability rate is less than a threshold value into wildcards. Note that the number of characters to be converted to the wildcard "*" in one item may be one, or may be two or more, and some or all of the characters in the item are the conversion target. The number of characters to be converted to the wildcard "*" in one item is set in advance by the designer or maintainer of the input support device 1. Also, the number of target items to be converted to the wildcard may be one, two or more, or all. The target items to be converted to the wildcard are set in advance by the designer or maintainer of the input support device 1. In the example of FIG. 4, the unreadable characters are indicated by a blank "_", but as will be described later, the unreadable characters (blanks) are also converted to wildcards.
[0030] In this embodiment, the search processing unit 12 searches the database 30 using the wildcard character string generated by the wildcard conversion unit 11, and extracts a plurality of corresponding DB character strings. Fig. 5 is a diagram showing an example of the search process by the search processing unit.
[0031] In the example of FIG. 5, the upper diagram shows an example of the same OCR character string as that shown in FIG. 3. In the upper diagram of FIG. 5, one character of the telephone number is shown as an unreadable character "?", three characters of the account number are shown as unreadable characters "?", and two characters of other items are shown as unreadable characters "?". The search processing unit 12 searches for one item of the OCR character string, or a wildcard character string for each of two or more items, using the query as a query. More specifically, the middle diagram of FIG. 5 shows a search result when a wildcard character string in the item "telephone number" of the OCR character string is used as a query. The lower diagram of FIG. 5 shows a search result when a wildcard character string in the item "client name" of the OCR character string is used as a query. In this embodiment, the database 30 manages DB character strings that group together data of each item for each application form, so the search result includes not only the data of the item that was the subject of the search, but also data of other items linked to it. In other words, the search result is a record (DB character string) that includes the corresponding character string. The search processor 12 may arbitrarily set items to be searched in the DB character string among the items specified by the OCR character string. In this case, the search processor 12 sets some or all of the items specified by the OCR character string as the items to be searched. The items to be searched are set in advance by the designer or maintainer of the input support device 1.
[0032] Fig. 6 is a diagram showing an example of a search result by the search processing unit. In the example of Fig. 6, the search results shown in the middle and lower diagrams of Fig. 5 are shown. In Fig. 6, the character string with the telephone number "85242812" is duplicated between the upper diagram and the lower diagram. For this reason, one of the duplicated characters is deleted in Fig. 6. The duplicated character string may be deleted by the search processing unit 12, the priority setting unit 16, the calculation setting unit 13, or the similarity calculation unit 14.
[0033] The calculation setting unit 13 sets some or all of the multiple items in the OCR character string (DB character string) as target items, sets multiple types of target items, and sets a character string similarity calculation formula for each type of target item.
[0034] Fig. 7 is a diagram showing an example of processing by the calculation setting unit 13 and the priority setting unit 16. In this embodiment, as shown in Fig. 7, a combination of one target item and a similarity calculation formula corresponding to this target item is called a "combination." The target items in the first combination in Fig. 7 are the items "recipient name" and "sender name" in the OCR character string (DB character string). The similarity calculation formula for character strings in the first combination in Fig. 7 is cosine similarity.
[0035] 7 further shows the target items in the second to fifth combinations and the corresponding similarity calculation formulas for character strings. The number of combinations may be one or more, or may be two or more, and is not specifically limited. The number of combinations, the target items in each combination, and the corresponding similarity calculation formulas are preset by the designer or maintainer of the input support device 1.
[0036] In each combination, the number of target items in the character string may be one, two or more, or all items may be the target items. In addition, in this embodiment, the cosine similarity and JARO Distance are exemplified as the similarity calculation formula, but the present invention is not limited to these. A known formula can be used as the formula for calculating the similarity of character strings.
[0037] The priority setting unit 16 sets priorities for the multiple combinations. In Fig. 7, the first combination is set to priority 1 (the highest priority). The priority setting unit 16 sets the second and subsequent combinations with the priorities shown in Fig. 7. The lower the priority value, the higher the priority. The priority values assigned to the multiple combinations by the priority setting unit 16 are set in advance by the designer or maintainer of the input support device 1. The items set by the designer or maintainer of the input support device 1 are configured to be appropriately changeable by the designer or maintainer.
[0038] The similarity calculation unit 14 performs a process of calculating the similarity between each DB character string extracted from the database 30 and the OCR character string using a similarity calculation formula for target items specified by the combination, for each of a plurality of combinations.
[0039] Fig. 8 is a diagram showing an example of the processing contents by the similarity calculation unit. As shown in Fig. 7 and Fig. 8, the similarity calculation unit 14 obtains the similarity of the first combination by calculating the characters of the target items identified by the first combination for the OCR string and each DB string using the similarity calculation formula specified by the first combination. That is, the similarity calculation unit 14 obtains the similarity of the nth combination by calculating the characters of the target items identified by the nth combination (n is an integer of 1 or more) for the OCR string and each DB string using the similarity calculation formula specified by the nth combination.
[0040] In this embodiment, the similarity calculation unit 14 sorts a plurality of DB character strings extracted from the database 30 in descending order of similarity obtained from the combination with the highest priority. More specifically, as shown in Fig. 9, the similarity calculation unit 14 sorts the DB character strings in descending order of similarity between the OCR character string and the DB character string calculated based on the first combination with priority 1. Fig. 9 is a diagram showing an example of a sorting result based on priority.
[0041] 9, the similarity calculation unit 14 arranges multiple DB strings in descending order of the similarity of the first combination. When there are multiple DB strings with the same similarity of the first combination, the DB string with the higher similarity of the combination with the next highest priority, priority 2 (second combination), is arranged higher. In this way, when there are multiple DB strings with the same similarity of combinations up to priority N (N is an integer equal to or greater than 1), the DB string with the higher similarity of the combination with priority N+1 is arranged higher.
[0042] Based on the calculated similarity, the information estimation unit 20 estimates that one of the multiple DB character strings extracted from the database 30 is the information that was the target of OCR. The information estimation unit 20 performs the estimation by referring to a list of the multiple DB character strings extracted from the database 30, sorted in descending order of similarity obtained from the combination with the highest priority, in descending order of rank.
[0043] The information estimation unit 20 estimates whether or not the information has been the subject of OCR, starting from the top DB character string, in accordance with a condition preset by a designer or maintainer of the input support device 1. Specifically, for example, if the preset condition is that the similarity is equal to or greater than a threshold (0.8), the information estimation unit 20 first determines whether or not all (five in this embodiment) similarities are equal to or greater than the threshold for the top DB character string (DB character string of record No. 2) shown in FIG.
[0044] When all of the multiple similarities of the top DB character string (DB character string of record No. 2) are equal to or greater than the threshold, the information estimation unit 20 estimates that the top DB character string is information that was the subject of OCR. On the other hand, when at least one of the multiple similarities of the top DB character string (DB character string of record No. 2) is less than the threshold, the information estimation unit 20 judges whether all of the multiple similarities of the next top DB character string (DB character string of record No. 3 in FIG. 9) are equal to or greater than the threshold. In this manner, until it finds a DB character string that satisfies a preset condition for the DB character string to be estimated, the information estimation unit 20 judges whether the DB character string satisfies a preset condition, starting from the top DB character string. Then, when it finds a DB character string that is the subject of estimation (determination) and all of the multiple similarities are equal to or greater than the threshold, the information estimation unit 20 estimates that the DB character string is information that was the subject of OCR. The information estimation unit 20 outputs the character string that it estimates to be information that was the subject of OCR to an external device or the like.
[0045] When none of the DB character strings extracted as records meet a predetermined standard, the information estimation unit 20 outputs information other than the DB character strings extracted from the database 30. In this embodiment, when none of the sorted DB character strings meet a predetermined standard, the information estimation unit 20 outputs the OCR result and ends the processing of the input support device 1. In this case, the information estimation unit 20 may output the result (image data) before the OCR recognition process, or may output the character string after the OCR recognition process.
[0046] The above-mentioned sorting of the multiple DB character strings may be performed by the information estimation unit 20, the similarity calculation and setting unit 13, or the priority setting unit 16.
[0047] Furthermore, the conditions for the information estimation unit 20 to estimate whether or not the information was subject to OCR are not limited to the above-mentioned conditions (all similarities are equal to or greater than the threshold). For example, the threshold of similarity may be different for each combination (priority). In this embodiment, the estimation support device 10 performs processing as a preparation for the information estimation unit 20 to perform estimation. It is not important that a specific information estimation method by the information estimation unit 20 is specifically specified.
[0048] If all DB character strings do not satisfy the above conditions, the information estimation unit 20 may relax the conditions for estimating the information that was the subject of OCR and perform the estimation again. In this case, for example, the threshold value for all similarities may be lowered to a value lower than 0.8 (for example, 0.7), or the threshold value for similarity may be different for each combination.
[0049] Furthermore, the information estimation unit 20 may estimate that the DB string having the greatest similarity among the multiple DB strings in the combination of priority 1 is the information that was the target of OCR. When there are two or more DB strings having the greatest similarity in the combination of priority 1, the information estimation unit 20 may estimate that the DB string having the greatest similarity in the combination of priority 2 among these DB strings is the information that was the target of OCR. In other words, when there are two or more DB strings having the greatest similarity in each of the combinations up to priority x (x is an integer equal to or greater than 1), the information estimation unit 20 may estimate that the DB string having the greatest similarity in the combination of priority x+1 among these DB strings is the information that was the target of OCR.
[0050] [Device operation] Next, an example of the operation of the input support device 1 including the estimation support device 10 will be described with reference to FIG. 10. FIG. 10 is a flow diagram showing an example of the operation of the input support device including the estimation support device. In the following description, FIGS. 1 to 10 will be referred to as appropriate. In addition, in this embodiment, an input support method (estimation support method) is implemented by operating the input support device 1 (estimation support device 10). Therefore, the description of the input support method (estimation support method) in this embodiment will be replaced with the following description of the operation of the input support device 1 (estimation support device 10).
[0051] As shown in Fig. 10, the unreadability calculation unit 15 selects one item of character strings from among multiple items of character strings of an OCR character string generated by OCR (step A1). The selected item is specified, for example, from the left side of the character string. Next, the unreadability calculation unit 15 calculates the unreadability rate for the character string of the selected item (step A2). Next, the unreadability calculation unit 15 determines whether the unreadability rate calculated in step A2 is equal to or greater than a threshold value (step A3).
[0052] If the result of the determination in step A3 is that the unreadability rate is equal to or greater than the threshold, the unreadability rate calculation unit 15 deletes, from the search items, the OCR items whose unreadability rate is equal to or greater than the threshold (step A4), and proceeds to step A7.
[0053] On the other hand, if the result of the determination in step A3 is that the unreadability rate is not equal to or greater than the threshold (is less than the threshold), the wildcard conversion unit 11 converts a part of the character string of the selected item into a wildcard to generate a plurality of wildcard character strings (step A5), as shown in Fig. 4. At this time, the wildcard conversion unit 11 converts all unreadable characters into wildcards.
[0054] Next, the search processing unit 12 searches the database 30 using the wildcard character string, part of which has been converted into a wildcard, generated in step A5, and extracts a plurality of records (DB character strings) including the corresponding character string, as shown in Fig. 6 (step A6). In this embodiment, the DB character strings including the corresponding character strings are the records extracted from the database 30. Also, in step A6, the search processing unit 12 holds the records (DB character strings) extracted by the search as a search list. Note that, if no records are extracted by the search, the search processing unit 12 holds an empty search list.
[0055] Next, the search processing unit 12 judges whether or not record extraction (search) from the database 30 has been completed for all items set as search target items among the items identified by the OCR character string (step A7). Then, if the result of the judgment in step A7 is that the search has not been completed for all items (NO in step A7), the search processing unit 12 instructs the unreadability calculation unit 15 to execute step A1 again. As a result, steps A1 to A7 are executed again for the newly selected item.
[0056] On the other hand, if the search has been completed for all items (YES in step A7), the search processing unit 12 judges whether a record (DB character string) has been extracted (step A8). Specifically, the search processing unit 12 judges whether a record (DB character string) is included in the search list. Then, the search processing unit 12 notifies the information estimation unit 20 of the result of the judgment.
[0057] If it is determined in step A8 that no record has been extracted by the search, the information estimation unit 20 outputs only the OCR character string to an external device or the like (step A14), and ends the process.
[0058] On the other hand, if the result of the judgment in step A8 is that a record has been extracted by the search, the calculation setting unit 13 sets some or all of the multiple items in the record (DB string) as target items, as shown in Figure 7, sets multiple types of target items, and sets a string similarity calculation formula for each type of target item (step A9).
[0059] Next, the priority setting unit 16 sets priorities for a plurality of combinations as shown in FIG. 7 (step A10).
[0060] Next, the similarity calculation unit 14 calculates the similarity between each record (DB character string) and the OCR character string (step A11). At this time, the similarity calculation unit 14 performs a process of calculating the similarity of the target items specified by the combinations shown in Fig. 7 using the corresponding similarity calculation formula for each of the multiple combinations. As a result, the result shown in Fig. 8 is obtained.
[0061] Next, the information estimation unit 20 sorts the records (DB character strings) in descending order of similarity obtained from the combination of priority 1 (step A12), as shown in Fig. 9. This results in a sorted result as shown in Fig. 9. If there are two or more records (DB character strings) with the same similarity of the combination of priority 1, the record (DB character string) with a higher similarity of the combination of priority 2 is ranked higher.
[0062] Next, the information estimation unit 20 judges whether or not the judgment of all records (DB character strings) is completed (step A13). If the judgment of all records (DB character strings) is completed without estimating the information that was the subject of OCR (YES in step A13), the information estimation unit 20 outputs information (OCR results) other than the DB character strings to an external device or the like (step A14), and ends the process.
[0063] On the other hand, if there is an undetermined record in the process of step A13 (NO in step A13), the information estimation unit 20 selects an undetermined record (DB character string) (step A15). At this time, the information estimation unit 20 selects the undetermined record (DB character string) with the highest ranking in the list shown in FIG. 9 from among the undetermined records (DB character strings).
[0064] Next, the information estimation unit 20 judges whether the similarity of the referenced record (DB character string) satisfies the above-mentioned predetermined criterion (step A16). If the similarity of the referenced record (DB character string) does not satisfy the above-mentioned predetermined criterion (NO in step A16), the information estimation unit 20 returns to step A13 and performs the process after step A13.
[0065] On the other hand, if the similarity of the referenced record (DB character string) satisfies the above-mentioned predetermined criterion (YES in step A16), the information estimation unit 20 estimates that the referenced record (DB character string) is the information that was the subject of OCR. At this time, the information estimation unit 20 outputs this record (DB character string) to an external device or the like (step A17), and ends the process.
[0066] [Effects of the embodiment] As described above, in this embodiment, a search is performed on the database 30 using a wildcard character string, so that multiple candidate records (DB character strings) are extracted. Then, for each of the multiple records (DB character strings), a process of calculating the similarity with the OCR character string using the target item is performed for each type of target item. Then, based on the similarity between the extracted record (DB character string) and the OCR character string, a record (DB character string) indicating the OCR character string is estimated. In this way, by calculating the similarity for each of multiple types of target items, more diverse similarities can be calculated. As a result, it is possible to support more accurate estimation of the target of OCR. Furthermore, for each of the multiple records (DB character strings), the similarity with the OCR character string is calculated using a similarity calculation formula corresponding to the target item specified by each combination. In this way, by combining various similarity calculation formulas, it is possible to calculate the similarity in a more multifaceted manner. As a result, it is possible to prepare information for more accurate estimation of the target of OCR, and to estimate the information that was the target of OCR with higher accuracy. Furthermore, the information estimation unit 20 performs estimation by referring to the records (DB character strings) sorted in descending order of similarity obtained from the combination with the highest priority, in descending order of rank. This allows the target item with priority 1 to be appropriately tuned based on the usage experience of the input support device 1 for a certain period of time. As a result, it is possible to make the target of OCR more accurate. Furthermore, when none of the multiple records (DB character strings) meets a predetermined standard, the information estimation unit 20 outputs information other than the records (DB character strings). This makes it possible to prevent inaccurate estimation results from being output. Therefore, according to this embodiment, accurate character string data can be obtained from the OCR-processed character string without manual correction input, thereby reducing the labor required for data input processing using OCR.
[0067] In this embodiment, the OCR result is described as a configuration in which the OCR result is a transfer request data in a currency exchange process, but this is not necessarily the case. The OCR result may be data other than currency exchange process data as long as it includes a character string for each of a plurality of items. For example, the OCR character string and the DB character string may be application data submitted by an applicant to a government office, and the specific OCR target is not limited.
[0068] Furthermore, the wildcard conversion unit 11 may convert all items of the OCR character string into wildcards, or may convert only items that are to be searched by the search processing unit 12 into wildcards.
[0069] [program] An example of the program in this embodiment is a program that causes a computer to execute steps A1 to A17 shown in Fig. 8. By installing and executing this program in a computer, an input support device (estimation support device) and an input support method (estimation support device) can be realized. In this case, the processor of the computer functions as a wildcard conversion unit 11, a search processing unit 12, a calculation setting unit 13, a similarity calculation unit 14, an unreadability calculation unit 15, a priority setting unit 16, and an information estimation unit 20 to perform processing.
[0070] In addition, in this embodiment, database 30 may be realized by storing the data files that constitute it in a storage device such as a hard disk provided in the computer, or may be realized by a storage device of another computer.
[0071] The program in this embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the wildcard conversion unit 11, the search processing unit 12, the calculation setting unit 13, the similarity calculation unit 14, the unreadability calculation unit 15, the priority setting unit 16, and the information estimation unit 20.
[0072] Here, an example of a computer that realizes the input support device 1 (estimation support device 10) by executing a program in this embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing an example of a computer that realizes the input support device (estimation support device).
[0073] 11, a computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be able to communicate data with each other.
[0074] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to or instead of the CPU 111. In this aspect, the GPU or FPGA can execute the programs in the embodiments.
[0075] The CPU 111 loads a program in the embodiment, which is composed of a group of codes and is stored in the storage device 113, into the main memory 112, and executes each code in a predetermined order to perform various calculations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).
[0076] Moreover, the program in the embodiment is provided in a state stored in a computer-readable recording medium 120. The program in the embodiment may be distributed on the Internet connected via the communication interface 117.
[0077] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.
[0078] Data reader / writer 116 mediates data transmission between CPU 111 and recording medium 120, reads programs from recording medium 120, and writes processing results in computer 110 to recording medium 120. Communication interface 117 mediates data transmission between CPU 111 and other computers.
[0079] Specific examples of recording medium 120 include general-purpose semiconductor storage devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as a flexible disk, or optical recording media such as a CD-ROM (Compact Disk Read Only Memory).
[0080] The input support device 1 (estimation support device 10) in this embodiment can be realized not by a computer with a program installed, but by hardware corresponding to each unit, for example, an electronic circuit. Furthermore, the input support device 1 (estimation support device 10) may be partially realized by a program and the remaining unit by hardware. In the embodiment, the computer is not limited to the computer shown in FIG. 11.
[0081] A part or all of the above-described embodiment can be expressed by (Additional Notes 1) to (Additional Notes 9) described below, but is not limited to the following descriptions.
[0082] (Appendix 1) a wildcard conversion unit that converts a part of a character string generated by optical character recognition, the character string including a plurality of items, into a wildcard; a search processing unit that searches a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracts a plurality of corresponding character strings; A calculation setting unit that sets some or all of the plurality of items as target items and sets a plurality of types of the target items; a similarity calculation unit that calculates, for each type of target item, a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, using the target item; Equipped with An estimation support device comprising:
[0083] (Appendix 2) the calculation setting unit sets a string similarity calculation formula for each type of the target item; the similarity calculation unit performs, for each of a plurality of combinations, a process of calculating a similarity between the character string generated by the optical character recognition and each of a plurality of character strings extracted from the database, using the similarity calculation formula for the target item identified in the combination, when a type of the target item and the corresponding similarity calculation formula are considered to be one combination.
[0084] (Appendix 3) A priority setting unit that sets priorities of the plurality of combinations, The estimation support device according to claim 2, wherein the similarity calculation unit sorts the multiple character strings extracted from the database in descending order of the similarity obtained from the combination with the highest priority.
[0085] (Appendix 4) a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; Equipped with The estimation support method according to the present invention is characterized in that
[0086] (Appendix 5) The setting step sets a formula for calculating a similarity of a character string for each type of the target item, The estimation support method according to claim 4, wherein the calculation step includes, for each of a plurality of combinations, calculating a similarity between a character string generated by the optical character recognition and each of a plurality of character strings extracted from the database, where the type of the target item and the corresponding similarity calculation formula are regarded as one combination, using the similarity calculation formula for the target item identified in the combination.
[0087] (Appendix 6) A priority setting step of setting priorities of the plurality of combinations is further provided; The estimation support method according to claim 5, wherein the calculation step includes sorting the plurality of character strings extracted from the database in descending order of the degree of similarity obtained from the combination having the highest priority.
[0088] (Appendix 7) On the computer, a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; A program to execute.
[0089] (Appendix 8) The setting step sets a formula for calculating a similarity of a character string for each type of the target item, The program according to claim 7, characterized in that the calculation step performs, for each of a plurality of combinations, a process of calculating a similarity between a character string generated by the optical character recognition and each of a plurality of character strings extracted from the database, when a type of the target item and the corresponding similarity calculation formula are considered as one combination, using the similarity calculation formula for the target item identified by the combination.
[0090] (Appendix 9) On the computer, further executing a priority setting step of setting priorities of the plurality of combinations; The program according to claim 8, wherein the calculation step sorts the plurality of character strings extracted from the database in descending order of the degree of similarity obtained from the combination having the highest priority. [Industrial Applicability]
[0091] As described above, according to the present disclosure, it is possible to provide support for achieving labor-saving with higher accuracy in data entry processing using OCR. The present disclosure is useful for systems that require processing of data obtained by OCR, such as currency exchange processing systems. [Explanation of symbols]
[0092] 1. Input support device 10 Estimation support device 11 Wildcard conversion section 12 Search processing section 13 Calculation setting section 14 Similarity calculation unit 15 Unreadability calculation part 16 Priority setting section 20 Information Estimation Department 30 Databases 110 Computer 111 CPU 112 Main memory 113 Storage device 114 Input Interface 115 Display Controller 116 Data Reader / Writer 117 Communication Interface 118 Input Devices 119 Display device 120 Recording media 121 Bus
Claims
1. a wildcard conversion unit that converts a part of a character string generated by optical character recognition, the character string including a plurality of items, into a wildcard; a search processing unit that searches a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracts a plurality of corresponding character strings; A calculation setting unit that sets some or all of the plurality of items as target items and sets a plurality of types of the target items; a similarity calculation unit that calculates, for each type of target item, a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, using the target item; Equipped with An estimation support device comprising:
2. the calculation setting unit sets a string similarity calculation formula for each type of the target item; 2. The estimation support device according to claim 1, wherein, when one type of the target item and the corresponding similarity calculation formula are considered to be one combination, the similarity calculation unit performs a process of calculating, for each of the plurality of combinations, a similarity between the character string generated by the optical character recognition and each of the plurality of character strings extracted from the database, using the similarity calculation formula for the target item identified by the combination.
3. A priority setting unit that sets priorities of the plurality of combinations, 3. The estimation support device according to claim 2, wherein the similarity calculation unit sorts the plurality of character strings extracted from the database in descending order of the similarity obtained from the combination having the highest priority.
4. a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; Equipped with The estimation support method according to the present invention is characterized in that
5. The setting step sets a formula for calculating a similarity of a character string for each type of the target item, 5. The estimation support method according to claim 4, characterized in that in the calculation step, when one type of the target item and the corresponding similarity calculation formula are considered to be one combination, a process of calculating, for each of the plurality of combinations, a similarity between the character string generated by the optical character recognition and each of the plurality of character strings extracted from the database, using the similarity calculation formula for the target item identified by the combination.
6. A priority setting step of setting priorities of the plurality of combinations is further provided; 6. The estimation support method according to claim 5, wherein the calculation step includes sorting the plurality of character strings extracted from the database in descending order of the degree of similarity obtained from the combination having the highest priority.
7. On the computer, a converting step of converting a portion of a string generated by optical character recognition, the string including multiple items, into a wildcard; a search step of searching a database in which information composed of character strings including the plurality of items is registered, using the character string partly converted into the wild card, and extracting a plurality of corresponding character strings; a setting step of setting some or all of the plurality of items as target items and setting a plurality of types of the target items; a calculation step of calculating a similarity between each of the plurality of character strings extracted from the database and the character string generated by the optical character recognition, for each type of the target items, by using the target items; A program to execute.
8. The setting step sets a formula for calculating a similarity of a character string for each type of the target item, 8. The program according to claim 7, characterized in that the calculation step, when one type of target item and the corresponding similarity calculation formula are considered to be one combination, performs a process for each of the multiple combinations of the multiple character strings extracted from the database to calculate a similarity between the character string generated by the optical character recognition and the target item identified by the combination using the similarity calculation formula.
9. On the computer, further executing a priority setting step of setting priorities of the plurality of combinations; 9. The program according to claim 8, wherein said calculating step includes sorting the plurality of character strings extracted from said database in descending order of the degree of similarity obtained from the combination having the highest priority.
Citation Information
Patent Citations
Exchange ocr system
JP2003006441A