Input support program, input support method, and input support device

The input support program enhances Japanese input error correction by associating reading candidates with semantically similar groups and evaluating character arrangement and connection strength, improving accuracy and reducing user workload.

JP2025105256APending Publication Date: 2025-07-10JUSTSYSTEMS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023223689
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing input error correction systems in Japanese character input fail to accurately correct errors and often output undesired conversion candidates due to decreased estimation accuracy at the position of the error.

Method used

An input support program that detects input errors, creates reading candidates for the error position, associates them with semantically similar groups, and refers to a storage unit to specify a reading candidate based on semantic connections with surrounding words, determining the correct reading using evaluation values for character arrangement and connection strength.

Benefits of technology

Improves the estimation accuracy of the correct reading at the error position, reducing the need for re-entry and minimizing the output of undesired conversion candidates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105256000001_ABST
    Figure 2025105256000001_ABST
Patent Text Reader

Abstract

To improve estimation precision of reaching of a place corresponding to an input mistake.SOLUTION: An input support device 100 creates one or more read candidates for a place corresponding to an input mistake from among inputted character strings by performing predetermined correction processing to the input mistake when the input mistake is detected from an inputted character string. The input support device 100 refers to a storage section 110, and specifies a read candidate corresponding to a work that belongs to a groups semantically connected to at least a word before and after a place corresponding to an input mistake in an inputted character string from among created read candidates. The storage section 110 stores a specific word that is semantically connected to a group and is inputted continuously to a word that belongs to the group in association with the group into which semantically similar words are sorted. The input support device 100 determines reading of a place corresponding to an input mistake on the basis of a specified read candidate.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an input support program, an input support method, and an input support device that assist character input performed based on a user's operation.

Background Art

[0002] Conventionally, in a Japanese input system, there is a technique for automatically correcting input errors made by a user. For example, when detecting the omission of a vowel (input error) from an input character string, a vowel is supplemented at the omitted position to estimate the reading, and the result of converting the estimated reading into a Chinese character or the like is output as a conversion candidate.

[0003] As a prior art for assisting character input, kana-kanji conversion is performed based on a basic dictionary that stores information representing the notation and category for the reading of a word, and from a rule dictionary, based on any of the reading of the word, the information representing the reading, notation, and category, a rule that matches the content of the conversion result storage area is searched, and the content of the corresponding conversion result storage means is rewritten according to the matched rule.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the prior art, there is a problem that input errors made by a user cannot be correctly corrected, the estimation accuracy of the reading at the position corresponding to the input error decreases, and a conversion candidate that the user does not desire may be output.

[0006] In one aspect, an object of the present invention is to provide an input support program, an input support method, and an input support device that improve the estimation accuracy of the reading at the position corresponding to an input error.

Means for Solving the Problem

[0007] In order to solve the above-described problems and achieve the object, when an input error is detected from an input string, the input support program according to the present invention performs a predetermined correction process on the input error, and creates one or more reading candidates for a portion corresponding to the input error in the string, associates them with a group in which words having semantic similarity are classified, and refers to a storage unit that stores a specific word that is semantically connected to the group and that is continuously input with a word belonging to the group, and among the one or more reading candidates created, specifies a reading candidate corresponding to a word belonging to a group that has a semantic connection with at least one of the words before and after the portion in the string, and based on the specified reading candidate, determines the reading of the portion, and causes a computer to execute the process.

[0008] Further, in the input support program according to the present invention, in the above invention, the storage unit further stores the strength of the semantic connection between the group and the specific word, refers to the storage unit, and for the specified reading candidate, specifies the strength of the semantic connection between the group to which the word corresponding to the reading candidate belongs and either of the words, causes the computer to execute the process, and the determining process determines the reading of the portion based on the specified strength of the connection from among the specified reading candidates.

[0009] Further, in the input support program according to the present invention, in the above invention, the determining process determines the reading of the portion based on an evaluation value regarding the arrangement of characters based on the plausibility of the transition between adjacent characters in the specified reading candidate from among the specified reading candidates.

[0010] In addition, in the input support program according to the present invention, in the above invention, the determination process determines the reading of the location based on an evaluation value regarding the arrangement of characters based on the strength of the identified connection and the likelihood of the transition between adjacent characters in the identified reading candidate among the identified reading candidates.

[0011] In addition, in the input support program according to the present invention, in the above invention, the determination process identifies the reading candidate with the strongest identified connection strength among the identified reading candidates. When a plurality of the strongest reading candidates are identified, the reading candidate with the highest evaluation of the character arrangement based on the evaluation value among the identified strongest reading candidates is determined as the reading of the location.

[0012] In addition, in the input support program according to the present invention, in the above invention, the storage unit stores the context between the group and the specific word, and the identification process refers to the storage unit and, based on the context, identifies the reading candidate corresponding to the word belonging to the group having a semantic connection with any of the one or more created reading candidates.

[0013] In addition, in the input support program according to the present invention, in the above invention, the specific word is either one or more words continuously input before or after the word belonging to the group, or two or more words continuously input with the word belonging to the group sandwiched therebetween. When the determination process includes, among the identified reading candidates, a first reading candidate corresponding to a word belonging to a group having a semantic connection with one of the words before or after the location and a second reading candidate corresponding to a word belonging to a group having a semantic connection with at least two of the words before or after the location, the second reading candidate is determined as the reading of the location.

[0014] Further, in the input support program according to the present invention, in the above invention, the computer is caused to execute a process of outputting a conversion candidate corresponding to the character string based on the reading of the determined location.

[0015] Also, in the input support method according to the present invention, when an input error is detected from the input character string, a predetermined correction process is performed on the input error, and for the location corresponding to the input error in the character string, one or more reading candidates are created, and the reading candidates are associated with a group in which words with semantic similarity are classified. Then, referring to a storage unit that stores a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group, among the one or more created reading candidates, a reading candidate corresponding to a word belonging to a group that has a semantic connection with at least one of the words before and after the location in the character string is specified, and based on the specified reading candidate, the computer executes a process of determining the reading of the location.

[0016] Further, the input support device according to the present invention, when an input error is detected from the input character string, performs a predetermined correction process on the input error, and for the location corresponding to the input error in the character string, creates one or more reading candidates, and associates the reading candidates with a group in which words with semantic similarity are classified. Then, referring to a storage unit that stores a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group, among the one or more created reading candidates, a reading candidate corresponding to a word belonging to a group that has a semantic connection with at least one of the words before and after the location in the character string is specified, and based on the specified reading candidate, the control unit executes a process of determining the reading of the location.

Effect of the Invention

[0017] According to the input support program, the input support method, and the input support device according to the present invention, there is an effect that the estimation accuracy of the reading of the location corresponding to the input error can be improved.

Brief Description of Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0019] (Embodiment) Hereinafter, with reference to the drawings, embodiments of an input support program, an input support method, and an input support apparatus according to the present invention will be described in detail.

[0020] FIG. 1 is an explanatory diagram showing an example of an input support method according to an embodiment. In FIG. 1, an input support device 100 is a computer that supports character input performed based on a user's operation. The input support device 100 is, for example, a PC (Personal Computer). The input support device 100 may be a tablet PC, a smartphone, or the like. Further, the input support device 100 may be a server that can be connected from a PC or the like used by the user.

[0021] The user's operation is performed using an input device such as a keyboard or a touch panel, for example. In the case of Japanese input, there are input modes such as Roman character input and kana input. In Roman character input, when inputting Japanese, Roman characters combining the consonants and vowels of the characters are input. Kana input is performed using kana characters written on a keyboard or the like.

[0022] In character input performed based on the user's operation, input errors may occur. Examples of input errors include omission of vowels, input of extra characters, and incorrect order of character input. If an input error occurs by the user, it is convenient if the input error can be automatically corrected, as it saves the trouble of the user having to re-enter.

[0023] As a method for correcting input errors, for example, when an omission of a vowel occurs, there is one that supplements the vowel at the omitted position. Also, based on a machine learning method, the natural order of Roman characters as Japanese can be learned, and for input errors, it is conceivable to supplement characters, delete extra characters, or swap the order of characters.

[0024] However, in the prior art, there are cases where input errors by the user cannot be corrected correctly. If the input error cannot be corrected correctly, the correct reading of the portion corresponding to the input error cannot be obtained, and a conversion result that the user does not desire is output.

[0025] For example, when inputting the string "maguro", assume that the vowel "a" after "m" is missing and the string "mguro" is input. In this case, by supplementing and correcting the missing vowel (a, i, u, e, o) at the missing position, five reading candidates, namely "maguro", "miguro", "muguro", "meguro", and "moguro", are created for the part corresponding to the input error.

[0026] In the prior art, it is difficult to determine which of these five reading candidates is the correct reading. For example, it is conceivable to determine the correct reading from the five reading candidates in consideration of the natural order of the learned Roman characters as Japanese.

[0027] For example, among the vowels (a, i, u, e, o), assume that the probability of transitioning from "m" is highest for "e", and it is determined that "meguro" is the most natural Roman character order. In this case, "mguro" is corrected to "meguro", and it is determined that "meguro" is the correct reading.

[0028] However, assume that the user is trying to input the phrase "sashimi of tuna" and has input the string "nosasimi" following "mguro". In this case, "maguro" is the correct reading, and "meguro" is an incorrect reading unintended by the user.

[0029] Therefore, in this embodiment, when estimating the reading of the part corresponding to the input error, the input support device 100 considers the semantic connection between that part and the words before and after. At this time, the input support device 100 treats words with semantic similarity as a group, and determines the semantic connection between the part corresponding to the input error and the words before and after from the semantic connection between the group and a specific word.

[0030] Here, a processing example of the input support device 100 will be described. The processing example of the input support device 100 is executed, for example, in kana-kanji conversion processing. Kana-kanji conversion processing is a process of converting an input string into a kanji or a sentence containing kanji.

[0031] (1) When an input error is detected from the input string 101, the input support device 100 creates one or more reading candidates for the portion corresponding to the input error in the input string 101 by performing a predetermined correction process on the input error. Here, the input string 101 is a string input based on a user's operation.

[0032] The input string 101 is a string before confirmation. For example, in the case of Japanese input in the Roman character input mode, the input string 101 corresponds to the input Roman character spelling or the result of converting the Roman character spelling into kana. Also, in the case of Japanese input in the kana input mode, the input string 101 corresponds to the input kana spelling (keystrokes).

[0033] Input errors are, for example, omission of vowels, input of extra characters, incorrect order of character input, etc. The predetermined correction process is a process of supplementing characters such as vowels, deleting extra characters, or swapping the order of characters for the input error. The portion corresponding to the input error is, for example, a clause including the location of the input error.

[0034] In the example of FIG. 1, it is assumed that Japanese input is in the Roman character input mode, and the input string 101 is "mぐろのさしみ (mguronosasimi)". "mguronosasimi" is the input Roman character spelling. "mぐろのさしみ" is the result of converting "mguronosasimi" into kana.

[0035] In the input string 101, the vowel after m is missing. Therefore, the missing vowel after m is detected as an input error from the input string 101. In this case, for example, a correction process of supplementing a vowel (a, i, u, e, o) for the input error is performed, and the reading candidates 102-1 to 102-5 for the location 102 corresponding to the input error are created.

[0036] The reading candidate 102-1 is the one obtained by supplementing the vowel "a" after m and converting it into kana. The reading candidate 102-2 is the one obtained by supplementing the vowel "i" after m and converting it into kana. The reading candidate 102-3 is the one obtained by supplementing the vowel "u" after m and converting it into kana. The reading candidate 102-4 is the one obtained by supplementing the vowel "e" after m and converting it into kana. The reading candidate 102-5 is the one obtained by supplementing the vowel "o" after m and converting it into kana.

[0037] (2) The input support device 100 refers to the storage unit 110 and identifies, among the created reading candidates 102-1 to 102-5, the reading candidates corresponding to the words belonging to the group that has a semantic connection with at least one of the words before and after the location 102 corresponding to the input error in the input string 101.

[0038] Here, the storage unit 110 stores, in association with a group in which words with semantic similarity are classified, a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group. For example, "maguro", "mejina", and "tai" are all words representing fish and can be said to have semantic similarity. Also, "ringo", "mikan", and "ichigo" are all words representing fruits and can be said to have semantic similarity. Such groups in which words are classified are registered in the storage unit 110.

[0039] The words that are continuously input are, for example, words that are continuously input. Also, the words that are continuously input may be words that are input with a particle or auxiliary verb such as "te ni o wa" in between. Semantic connection means a relationship in which two or more words are connected to form a new meaningful word.

[0040] For example, the semantic connection may be a relationship in which two or more words are combined to form a compound word. Further, the semantic connection may be a relationship in which two or more words are combined to form an idiomatic expression. Further, the semantic connection may be a relationship between a modifier and a modified word.

[0041] In the example of FIG. 1, it is assumed that a specific word 121 is stored in the storage unit 110 in association with a group 120 that classifies words having semantic similarity. The group 120 includes, for example, words representing fish such as "maguro", "mejina", "tai", "kanpachi", "hamachi". The specific word 121 is the word "sashimi", which has a semantic connection with the group 120 and is a word that is continuously input with the words belonging to the group 120.

[0042] Specifically, for example, the input support device 100 extracts at least one of the words before and after the portion 102 corresponding to the input error from the input character string 101. The previous word is, for example, the word immediately before the portion 102 corresponding to the input error in the input character string 101. The subsequent word is, for example, the word immediately after the portion 102 corresponding to the input error in the input character string 101. However, particles and auxiliary verbs such as "te ni o wa" may be excluded from the extraction target.

[0043] In the example of FIG. 1, the word 103 "sashimi" after the portion 102 corresponding to the input error is extracted from the input character string 101. The particle "no" immediately after the portion 102 corresponding to the input error is excluded from the extraction target. Here, the word 103 matches the specific word 121 stored in the storage unit 110 in association with the group 120.

[0044] Therefore, the word 103 has a semantic connection with the group 120. In this case, the input support device 100 refers to the storage unit 110 and specifies the reading candidate 102-1 corresponding to the word "maguro" belonging to the group 120 that has a semantic connection with the word 103 among the reading candidates 102-1 to 102-5.

[0045] Note that in the example of FIG. 1, the relationship between group 120 and specific word 121 is not defined in storage unit 110. In this case, even if the word 103 "sashimi" is before the location 102 corresponding to the input error, it is determined that there is a semantic connection with group 120.

[0046] (3) Input support device 100 determines the reading of the location 102 corresponding to the input error based on the identified reading candidate. Determining the reading corresponds to estimating the correct reading. It can be said that the identified reading candidate has a stronger connection with the surrounding words compared to other reading candidates and is likely to be the correct reading of the location 102 corresponding to the input error.

[0047] Here, as the reading candidate corresponding to the word belonging to group 120 that has a semantic connection with word 103, reading candidate 102-1 is identified. In this case, input support device 100 determines, for example, the identified reading candidate 102-1 as the reading of the location 102 corresponding to the input error.

[0048] As a result, the reading corresponding to the input character string 101 becomes "maguro no sashimi". Among "maguro no sashimi", "maguro" corresponds to the correction result 104 of the location 102 corresponding to the input error. Among "maguro no sashimi", "sashimi" corresponds to word 103. "Maguro no sashimi" is, for example, converted to "maguro no sashimi" and output as a conversion candidate.

[0049] As described above, according to the input support device 100, when estimating the reading of a location corresponding to an input error, by considering the semantic connection with the words before and after that location, the estimation accuracy of the correct reading can be improved. At this time, according to the input support device 100, words with semantic similarity are treated as a group, and from the semantic connection between the group and a specific word, the semantic connection between the location corresponding to the input error and the words before and after can be determined. For this reason, the input support device 100 can consider the semantic connection between the location corresponding to the input error and the words before and after even if, for example, the semantic connection between individual words is not accumulated as knowledge.

[0050] In the example of FIG. 1, even when a plurality of reading candidates are created for the location 102 corresponding to the input error, the input support device 100 can determine the reading candidate 102-1 corresponding to the word belonging to the group 120 having a semantic connection with the word 103 after the location 102 as the reading of the location 102 corresponding to the input error.

[0051] Thereby, even when the user makes an input error, the input support device 100 can accurately estimate the correct reading of the location corresponding to the input error and prevent the output of conversion candidates that the user does not desire. In addition, the input support device 100 can reduce the trouble of the user re-entering and reduce the work load associated with character input.

[0052] (Hardware configuration example of the input support device 100) Next, with reference to FIG. 2, a hardware configuration example of the input support device 100 will be described. Here, the case where the input support device 100 is applied to a computer such as a PC or a tablet PC will be taken as an example for explanation. However, the input support device 100 may be applied to a server that can be connected from a PC or the like used by the user.

[0053] FIG. 2 is a block diagram showing an example of the hardware configuration of the input support device 100 according to the embodiment. In FIG. 2, the input support device 100 includes a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, and a RAM (Random Access Memory) 203.

[0054] The input support device 100 also includes an HDD (Hard Disc Drive) 204, an HD 205, a CD (Compact Disc)-RW (ReWritable) drive 206, and a CD-RW 207. The input support device 100 also includes a display 208, a keyboard 209, a mouse 210, and a network I / F (Interface) 211. Each component is connected by a bus 200.

[0055] Here, the CPU 201 controls the entire input support device 100. The CPU 201 may have a plurality of cores. The ROM 202 and the HD 205 store various programs. The programs stored in the ROM 202 and the HD 205 are, for example, the present input support program. The present input support program is applied to, for example, kana-kanji conversion software. The program stored in the ROM 202 is loaded into the CPU 201 to cause the CPU 201 to execute the coded processing. The RAM 203 is used as a work area for the CPU 201.

[0056] The HDD 204 controls the reading or writing of data to / from the HD 205 according to the control of the CPU 201. The HD 205 stores the data written according to the control of the HDD 204. The CD-RW drive 206 controls the reading or writing of data to / from the CD-RW 207 according to the control of the CPU 201. The CD-RW 207 stores the data written according to the control of the CD-RW drive 206. The CD-RW 207 may be detachable from the input support device 100, for example.

[0057] The display 208 displays various data such as a cursor, icons, menus, windows, toolboxes, characters, images, or function information. The display 208 is, for example, a liquid crystal display, an organic EL (Electroluminescence) display, or the like.

[0058] The keyboard 209 has keys for inputting characters, numerical values, various instructions, etc., and inputs data. The mouse 210 selects or executes various instructions, selects a processing target, or moves the mouse pointer. Also, the display 208 may be a touch panel and may have functions corresponding to the keyboard 209 and the mouse 210. In this case, the input support device 100 may not have the keyboard 209 and the mouse 210.

[0059] The network I / F 211 is connected to the network NW through a communication line and is connected to other computers via the network NW. The network NW is, for example, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, or the like. The network I / F 211 manages the interface between the network NW and the inside of the input support device 100 and controls the input and output of data from other computers. The network I / F 211 is, for example, a modem, a LAN adapter, or the like.

[0060] In addition to the components described above, the input support device 100 may have, for example, a DVD (Digital Versatile Disc) drive, an SSD (Solid State Drive), a USB (Universal Serial Bus) port, or the like. Also, in addition to the components described above, the input support device 100 may have, for example, a printer, a scanner, a microphone, or a speaker. Also, the input support device 100 may not have, for example, the HDD 204, the HD 205, the CD-RW drive 206, the CD-RW 207, etc. among the components described above.

[0061] (Semantic connection table 300's stored content) Next, the stored content of the semantic connection table 300 used by the input support device 100 will be described. The semantic connection table 300 is realized by a storage device such as the RAM 203 or the HD 205, for example.

[0062] FIG. 3 is an explanatory diagram showing an example of the stored content of the semantic connection table 300. In FIG. 3, the semantic connection table 300 has fields for group ID, semantic attribute, member words, previous word, next word, and connection strength, and stores semantic connection information (for example, semantic connection information 300-1 to 300-3) as records by setting information in each field.

[0063] Here, the group ID is an identifier that uniquely identifies a group that classifies words with semantic similarity. The semantic attribute is for classifying words by meaning and indicates an attribute that the words belonging to the group commonly have. The member words are the words (kana) belonging to the group. The previous word is a specific word (kana) that has a semantic connection with the group and appears before the word belonging to the group.

[0064] The next word is a specific word (kana) that has a semantic connection with the group and appears after the word belonging to the group. The words belonging to the group and the specific words (previous word, next word) are input continuously, for example, during Japanese input, or are input with particles or auxiliary verbs such as "te ni o wa" in between. Note that the member words, previous word, and next word may have the conversion results of the words (kana) converted to Chinese characters associated with them.

[0065] The connection strength indicates the strength (degree) of the semantic connection between the group and the specific words (previous word, next word). The connection strength is represented by a numerical value from 0 to 1, for example, and the larger the value, the stronger the semantic connection. However, the connection strength may also be represented by levels such as "strong, medium, weak".

[0066] The connection strength is set to increase as the frequency of continuous input of words belonging to a group and a specific word increases. It can be said that the higher the connection strength, the higher the probability that words belonging to the group will be continuously input with the specific word.

[0067] For example, the semantic connection information 300-1 has a semantic connection with group G1 and indicates the subsequent word "sashimi" after being continuously input with words belonging to group G1. Group G1 is a group that classifies words with the semantic attribute "fish" such as "maguro", "mejina", and "tai". In this case, the words belonging to group G1 correspond to the previous words that appear before the subsequent word "sashimi". Also, the semantic connection information 300-1 indicates the semantic connection strength "0.75" between group G1 and the subsequent word "sashimi".

[0068] Also, the semantic connection information 300-3 has a semantic connection with group G3 and indicates the previous word "tokkyuu" before being continuously input with words belonging to group G3. Group G3 is a group that classifies words with the semantic attribute "train nickname" such as "azusa", "kajiji", and "shokaze". In this case, the words belonging to group G3 correspond to the subsequent words that appear after the previous word "tokkyuu". Also, the semantic connection information 300-3 indicates the semantic connection strength "0.51" between group G3 and the previous word "tokkyuu". "Azusa", "kajiji", and "shokaze" are registered trademarks.

[0069] In the example of FIG. 3, the case where the specific word associated with the group is one word (previous word or subsequent word) is taken as an example for explanation, but it is not limited to this. For example, in the semantic connection table 300, two or more words may be stored as specific words in association with the group. For example, in the semantic connection table 300, the front-back relationship (word order) between three words (previous word, middle word, subsequent word) with a semantic connection, with the words belonging to the group as the previous word, and the strength of their semantic connection may be stored.

[0070] The semantic connection table 300 may be created manually or automatically by a computer. For example, assume that there is a word X whose semantic connection with a specific word is known. In this case, a group may be created manually or by a computer to classify words (including the word X) that have the same semantic attribute as the word X, and the created group may be associated with the specific word to create semantic connection information.

[0071] Words to be classified into the same group may be identified from existing dictionary information (e.g., a thesaurus), or may be identified by analyzing Japanese document data and grouping together words whose semantic similarity is equal to or greater than a threshold. The semantic similarity is an index value that indicates the degree of semantic similarity between words. The semantic similarity may be calculated using any existing technology. The connection strength may be set, for example, according to the frequency with which a word belonging to a group and a specific word are input consecutively in existing document data.

[0072] The contents stored in the semantic connection table 300 may be updated at any time. For example, a semantic connection between a new group and a specific word may be registered, or the strength of the connection may be updated, depending on the number of occurrences or frequency of a combination of words having a semantic connection in Japanese document data.

[0073] (Example of Functional Configuration of Input Support Device 100) Next, an example of a functional configuration of the input support device 100 will be described with reference to FIG.

[0074] FIG. 4 is a block diagram showing a functional configuration example of the input support device 100 according to the embodiment. In FIG. 4, the input support device 100 includes a reception unit 401, an input error detection unit 402, a candidate creation unit 403, an extraction unit 404, a candidate identification unit 405, a correct reading determination unit 406, an output unit 407, and a storage unit 410. The reception unit 401 to the output unit 407 have functions as a control unit. For example, by causing the CPU 201 to execute a program stored in a storage device such as the ROM 202, RAM 203, or HD 205 shown in FIG. 2, or by the network I / F 211, these functions are realized. The processing results of each functional unit are stored in a storage device such as the RAM 203 or HD 205, for example. The storage unit 410 is realized by a storage device such as the RAM 203 or HD 205, for example. Specifically, for example, the storage unit 410 stores the semantic connection table 300 shown in FIG. 3. The storage unit 110 shown in FIG. 1 corresponds to the storage unit 410, for example.

[0075] The reception unit 401 receives an input of a character string. The character string to be input is, for example, a character string to be subject to kana-kanji conversion. Specifically, for example, the reception unit 401 receives an input of a character string by a user operation using the keyboard 209 or the mouse 210 shown in FIG. 2.

[0076] The input error detection unit 402 detects an input error from the input character string. Input errors include, for example, omission of vowels, input of extra characters, or incorrect order of character input. Specifically, for example, the input error detection unit 402 learns features of Japanese from information on character strings input in the past during Japanese input by a method based on machine learning such as deep learning. The features of Japanese are, for example, the natural arrangement of Roman characters as Japanese.

[0077] Then, the input error detection unit 402 detects an input error from the input character string based on the learned features of Japanese. For example, the input error detection unit 402 detects as an input error a place where the omission of a vowel or the input of an extra character deviates from the natural arrangement of Roman characters as learned Japanese.

[0078] More specifically, for example, input errors include insertion errors, deletion errors, substitution errors, and transposition errors. An insertion error is a mistake where an extra key character is entered (e.g., nyuuryoku (にゅうりょく) → nyuuuryoku (にゅううりょく)). A deletion error is a mistake where any key character is missing (e.g., nyuuryoku (にゅうりょく) → nyuuyoku (にゅうよく)). A substitution error is a mistake where any key character is mistaken for another character (such as a character on a neighboring key) (e.g., nyuuryoku (にゅうりょく) → nuuuryoku (ぬううりょく)). A transposition error is a mistake where the keying order of any key characters is reversed (e.g., nyuuryoku (にゅうりょく) → nyuuyroku (にゅうyろく)). Note that any existing technology may be used as the technology for detecting input errors.

[0079] When an input error is detected from the input string, the candidate creation unit 403 creates one or more reading candidates for the portion corresponding to the input error in the input string by performing a predetermined correction process on the input error. The portion corresponding to the input error is, for example, a clause including the location of the input error.

[0080] Specifically, for example, the candidate creation unit 403 performs a correction process on the input error, such as supplementing vowels, deleting extra characters, or swapping the order of characters, so that the learned Roman characters in Japanese are in a natural order. As a result, the candidate creation unit 403 creates one or more reading candidates for the portion corresponding to the input error.

[0081] The extraction unit 404 extracts at least one of the words before and / or after the portion corresponding to the input error from the input string. Specifically, for example, the extraction unit 404 extracts the word immediately before the portion corresponding to the input error from the input string. Also, the extraction unit 404 extracts the word immediately after the portion corresponding to the input error from the input string.

[0082] However, particles and auxiliary verbs such as "tenioha" may be excluded from the extraction target. For example, if the word immediately before the location corresponding to the input error is a particle, the extraction unit 404 extracts the word immediately before that particle. Also, if the word immediately after the location corresponding to the input error is a particle, the extraction unit 404 extracts the word immediately after that particle. Further, the extraction unit 404 may extract two or more words before the location corresponding to the input error. Also, the extraction unit 404 may extract two or more words after the location corresponding to the input error. In this case, the maximum number of words to be extracted may be set in advance.

[0083] The candidate identification unit 405 refers to the storage unit 410 and identifies, among the one or more created reading candidates, the reading candidates corresponding to the words belonging to the group that has a semantic connection with at least one of the words before and after the location corresponding to the input error in the input character string.

[0084] Here, the storage unit 410 stores, in association with a group in which words with semantic similarity are classified, a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group. In the storage unit 410, for example, a specific word that has a semantic connection with each of a plurality of groups and is continuously input with the words belonging to each of the groups is stored in association with each group.

[0085] Further, the storage unit 410 may further store the strength of the semantic connection between the group and the specific word. In this case, the candidate identification unit 405 may refer to the storage unit 410 and identify the strength of the semantic connection between the group to which the word corresponding to the identified reading candidate belongs and at least one of the words before and after the location corresponding to the input error.

[0086] Further, the storage unit 410 may store the context between a group and a specific word. In this case, the candidate specifying unit 405 may refer to the storage unit 410 and specify, among one or more reading candidates, a reading candidate corresponding to a word belonging to a group having a semantic connection with at least one of the words before and after the portion corresponding to the input error based on the context between the group and the specific word.

[0087] For example, assume that a word before the portion corresponding to the input error is extracted from the input character string. In this case, the candidate specifying unit 405 refers to, for example, the semantic connection table 300 shown in FIG. 3, and among the created reading candidates, specifies a reading candidate corresponding to a word belonging to a group that appears after the extracted word (corresponding to the "previous word") and has a semantic connection with the extracted word.

[0088] At this time, the candidate specifying unit 405 may refer to the semantic connection table 300 and specify the strength of the semantic connection between the group to which the word corresponding to the specified reading candidate belongs and the extracted word (previous word). The strength of the semantic connection is represented by, for example, a numerical value from 0 to 1, and the larger the value, the stronger the semantic connection (see FIG. 3). Further, the strength of the semantic connection may be represented by levels such as "strong, medium, weak", for example.

[0089] However, the specification of the strength of the semantic connection may be performed, for example, when a plurality of reading candidates are specified. For example, when there are a plurality of groups having a semantic connection with the extracted word and there are words corresponding to the reading candidates in each of the groups, a plurality of reading candidates are specified.

[0090] Further, assume that a word after the portion corresponding to the input error is extracted from the input character string. In this case, the candidate specifying unit 405 refers to, for example, the semantic connection table 300 and among the created reading candidates, specifies a reading candidate corresponding to a word belonging to a group that appears before the extracted word (corresponding to the "subsequent word") and has a semantic connection with the extracted word.

[0091] At this time, the candidate identification unit 405 may refer to the semantic connection table 300 and identify the strength of the semantic connection between the group to which the word corresponding to the identified reading candidate belongs and the extracted word (the subsequent word) for the identified reading candidate. However, the identification of the strength of the semantic connection may be performed, for example, when a plurality of reading candidates are identified.

[0092] In the following description, the strength of the semantic connection between the group to which the word corresponding to the identified reading candidate belongs and at least one of the words before and / or after the location corresponding to the input error may be simply referred to as the "strength of the semantic connection with the surrounding words".

[0093] The correct reading determination unit 406 determines the reading of the location corresponding to the input error based on the identified reading candidate. For example, assume that only one reading candidate is identified. In this case, the correct reading determination unit 406 may determine the identified reading candidate as the reading of the location corresponding to the input error.

[0094] Further, the correct reading determination unit 406 may also determine the reading of the location corresponding to the input error based on the strength of the semantic connection with the surrounding words among the identified reading candidates. For example, assume that a plurality of reading candidates are identified, and for each of the identified plurality of reading candidates, the strength of the semantic connection between the group to which the word corresponding to the reading candidate belongs and the surrounding words is identified.

[0095] In this case, the correct reading determination unit 406 may determine, among the identified plurality of reading candidates, the reading candidate with the strongest strength of the semantic connection with the identified surrounding words as the reading of the location corresponding to the input error. The reading candidate with the strongest semantic connection is, for example, the reading candidate with the largest value indicating the strength of the semantic connection or the reading candidate with the highest level (strong, medium, weak) indicating the strength of the semantic connection.

[0096] Further, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, a reading candidate among the identified plurality of reading candidates whose semantic connection strength with the words before and after the identification is equal to or greater than the threshold α. Further, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the top N reading candidates among the identified plurality of reading candidates whose semantic connection strength with the words before and after the identification is strong.

[0097] The threshold α can be arbitrarily set. For example, when the semantic connection strength is represented by a numerical value from 0 to 1, the threshold α may be set to a value of about 0.5. Also, when the semantic connection strength is represented by levels such as "strong, medium, weak", the threshold α may be set to "medium". Also, the number N is a natural number of 2 or more and can be arbitrarily set.

[0098] Also, when there are two or more reading candidates whose semantic connection strength with the words before and after is equal to or greater than the threshold α, the two or more reading candidates are determined as the reading of the location corresponding to the input error. The same applies when there are the top N reading candidates with a strong semantic connection with the words before and after. In this case, the correct reading determination unit 406 may set a priority corresponding to the semantic connection strength with the words before and after for each of the reading candidates determined as the reading of the location corresponding to the input error.

[0099] Here, the priority indicates the degree of priority as the correct reading of the location corresponding to the input error. The priority is set, for example, to increase as the semantic connection with the words before and after becomes stronger.

[0100] Further, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine all of the identified plurality of reading candidates as the reading of the location corresponding to the input error. In this case, the correct reading determination unit 406 may set a priority corresponding to the semantic connection strength with the words before and after for each of the plurality of reading candidates.

[0101] In addition, the correct reading determination unit 406 may determine the reading of the portion corresponding to the input error from among the specified reading candidates based on the evaluation value regarding the character sequence of the specified reading candidates. Here, the evaluation value regarding the character sequence indicates the unnaturalness (or plausibility) of the character sequence. The evaluation value is calculated based on, for example, the plausibility of the transition between adjacent characters in the reading candidate.

[0102] The plausibility of the transition between adjacent characters is represented by, for example, the transition probability between adjacent characters. The transition probability between adjacent characters is derived based on, for example, the correct Roman character sequence input in the past during Japanese input. Note that the information indicating the transition probability between adjacent characters is stored in advance in a storage device such as the RAM 203 or the HD 205.

[0103] In addition, the evaluation value may be calculated based on, for example, the plausibility that the leading Roman character of the Roman character spelling corresponding to the reading candidate is at the beginning in the case of Japanese input in the Roman character input mode. Also, the evaluation value may be calculated based on the plausibility that the trailing Roman character of the Roman character spelling corresponding to the reading candidate is at the end. Further, the evaluation value may be calculated based on the transition probability from the Roman character immediately before the portion corresponding to the input error in the input character string to the leading Roman character of the Roman character spelling corresponding to the reading candidate. Additionally, the evaluation value may be calculated based on the transition probability from the trailing Roman character of the Roman character spelling corresponding to the reading candidate to the Roman character immediately after the portion corresponding to the input error in the input character string.

[0104] In the following description, the evaluation value regarding the character sequence of the reading candidate may be referred to as the "unnaturalness score". It is assumed that the greater the value of the unnaturalness score, the more unnatural the character sequence.

[0105] Note that the unnaturalness score of the reading candidate may be calculated, for example, by the candidate creation unit 403 when the reading candidate is created. Further, the unnaturalness score of the reading candidate may be calculated, for example, by the candidate identification unit 405 when the reading candidate is identified. Further, the unnaturalness score of the reading candidate may also be calculated when a plurality of reading candidates are identified by the candidate identification unit 405.

[0106] For example, assume that a plurality of reading candidates are identified and an unnaturalness score is calculated for each of the identified reading candidates. In this case, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the reading candidate with the lowest calculated unnaturalness score among the plurality of identified reading candidates. The reading candidate with the lowest unnaturalness score is the reading candidate with the highest evaluation of the character arrangement.

[0107] Further, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the reading candidate among the plurality of identified reading candidates whose calculated unnaturalness score is equal to or less than the threshold β. The threshold β can be arbitrarily set. Further, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the top N reading candidates from the ones with the lower calculated unnaturalness scores among the plurality of identified reading candidates.

[0108] Further, when there are two or more reading candidates whose unnaturalness scores are equal to or less than the threshold β, the two or more reading candidates are determined as the reading of the location corresponding to the input error. The same applies when there are the top N reading candidates from the ones with the lower unnaturalness scores. In this case, the correct reading determination unit 406 may set a priority corresponding to the calculated unnaturalness score for each of the reading candidates determined as the reading of the location corresponding to the input error. The priority is set, for example, to be higher as the calculated unnaturalness score is lower.

[0109] In addition, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine that all of the identified plurality of reading candidates are the readings of the portions corresponding to input errors. In this case, the correct reading determination unit 406 may set a priority according to the unnaturalness score for each of the plurality of reading candidates.

[0110] Alternatively, the correct reading determination unit 406 may determine the reading of the portion corresponding to the input error based on the strength of the semantic connection with the words before and after the identified reading candidate and the unnaturalness score. Specifically, for example, the correct reading determination unit 406 identifies the reading candidate with the strongest semantic connection with the words before and after among the identified plurality of reading candidates.

[0111] At this time, there may be a case where a plurality of the strongest reading candidates are identified. For example, there may be a case where there are a plurality of reading candidates with the stage representing the strength of the semantic connection with the words before and after being "strong". Also, there may be a case where there are a plurality of reading candidates with the value indicating the strength of the semantic connection with the words before and after being the maximum.

[0112] In addition, there may be a case where there are a plurality of reading candidates for which it can be said that the value indicating the strength of the semantic connection with the words before and after is the maximum (for example, the difference from the maximum value is within 0.1). In this case, the correct reading determination unit 406 may determine that the reading candidate with the lowest unnaturalness score among the identified strongest reading candidates is the reading of the portion corresponding to the input error.

[0113] Alternatively, the correct reading determination unit 406 may extract reading candidates with the strength of the semantic connection with the words before and after being equal to or greater than the threshold α from among the identified plurality of reading candidates. Then, the correct reading determination unit 406 may determine that the reading candidate with the lowest unnaturalness score among the extracted reading candidates is the reading of the portion corresponding to the input error.

[0114] Further, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine that all of the identified plurality of reading candidates are the readings of the portions corresponding to input errors. In this case, the correct reading determination unit 406 may set a priority for each of the plurality of reading candidates according to the strength of the semantic connection with the preceding and succeeding words and the unnaturalness score. The priority is set, for example, so that the higher the strength of the semantic connection with the preceding and succeeding words, the higher the priority, and the lower the unnaturalness score, the higher the priority.

[0115] The output unit 407 outputs conversion candidates corresponding to the input character string based on the reading of the portion corresponding to the determined input error. Here, the conversion candidates are, for example, those obtained by converting the reading corresponding to the input character string into Chinese characters or a sentence containing Chinese characters.

[0116] Specifically, for example, first, the output unit 407 creates a reading corresponding to the input character string based on the reading of the portion corresponding to the determined input error. The reading corresponding to the input character string is created, for example, by using the reading of the portion corresponding to the input error in the input character string as the reading candidate determined for that portion.

[0117] Next, the output unit 407 refers to the mapping information associating readings with Chinese characters and creates conversion candidates corresponding to the input character string by performing kana-Chinese character conversion on the created reading. When there are a plurality of reading candidates determined as the readings of the portion corresponding to the input error, conversion candidates corresponding to each of the plurality of reading candidates are created.

[0118] Then, the output unit 407 outputs the created conversion candidates. More specifically, for example, the output unit 407 displays the created conversion candidates on the display 208 so that they can be selected for the input character string. At this time, the output unit 407 may display the created conversion candidates together with a message indicating that the input error has been corrected.

[0119] In addition, there may be cases where there are multiple readings for the part corresponding to the input error and multiple reading candidates are determined. In such a case, the output unit 407 may display, on the display 208, the conversion candidates corresponding to each of the multiple reading candidates in descending order of the priority set for each of the multiple reading candidates. An output example of the conversion candidates corresponding to the input character string will be described later with reference to FIG. 6.

[0120] Also, the specific word stored in the storage unit 410 in association with the group may be one word or two or more words. For example, the specific word may be one or two or more words that are continuously input before the word belonging to the group. Also, the specific word may be one or two or more words that are continuously input after the word belonging to the group. Also, the specific word may be two or more words that are continuously input sandwiching the word belonging to the group.

[0121] Specifically, for example, in the semantic connection table 300, semantic connection information may be stored in which words belonging to a group classifying words with the semantic attribute "fish" (tuna, sea bream, snapper,...) are used as the previous word, and the specific words are the two words "sashimi" (raw fish) as the middle word and "teishoku" (set meal) as the following word.

[0122] Here, assume that two words after the part corresponding to the input error are extracted from the input character string. In this case, the candidate specifying unit 405 may specify, for example, with reference to the semantic connection table 300, a reading candidate corresponding to a word (corresponding to the "previous word") belonging to a group that appears before the two extracted words (corresponding to the "middle word" and the "following word") and has a semantic connection with the two extracted words, from among the created reading candidates.

[0123] Also, assume that one word each before and after the location corresponding to the input error is extracted from the input string. In this case, the candidate identification unit 405 may, for example, refer to the semantic connection table 300, and from the created reading candidates, identify the reading candidates corresponding to the words (corresponding to "middle words") that appear after the extracted word (corresponding to "previous word"), appear before the extracted word (corresponding to "next word"), and belong to a group having a semantic connection with the two extracted words.

[0124] Also, assume that two words before the location corresponding to the input error are extracted from the input string. In this case, the candidate identification unit 405 may, for example, refer to the semantic connection table 300, and from the created reading candidates, identify the reading candidates corresponding to the words (corresponding to "next word") that appear after the two extracted words (corresponding to "previous word" and "middle word") and belong to a group having a semantic connection with the two extracted words.

[0125] Also, there may be cases where a first reading candidate and a second reading candidate are included in the reading candidates identified by the candidate identification unit 405. Here, the first reading candidate is a reading candidate corresponding to a word belonging to a group having a semantic connection with one of the words before and after the location corresponding to the input error.

[0126] The second reading candidate is a reading candidate corresponding to a word belonging to a group having a semantic connection with at least two or more of the words before and after the location corresponding to the input error. In this case, the correct reading determination unit 406 may determine the second reading candidate as the reading of the location corresponding to the input error (giving priority to the connection of longer words).

[0127] Also, there may be cases where there are multiple second reading candidates among the reading candidates identified by the candidate identification unit 405. In this case, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the second reading candidate having the strongest semantic connection strength with the identified previous and next words among the multiple second reading candidates.

[0128] Note that the functional units (reception unit 401 to output unit 407) of the input support device 100 may be realized by, for example, a server that can be connected via a network NW (see FIG. 2) from a PC or the like used by the user. In this case, the input support device 100 receives an input of a character string from the PC used by the user by the reception unit 401, and outputs the created conversion candidates to the PC used by the user by the output unit 407.

[0129] Also, the functional units (reception unit 401 to output unit 407) of the input support device 100 may be realized by a plurality of computers (for example, a server and a PC). For example, the reception unit 401 to the correct reading determination unit 406 may be realized by a server, and the output unit 407 may be realized by a PC. In this case, the communication between the functional units of different computers is performed, for example, by transmission and reception between the functional units via the network NW. Thereby, on the server side, it is possible to determine the reading of the portion corresponding to the input error, and on the PC side, based on the reading estimated by the server, create and output conversion candidates corresponding to the input character string.

[0130] (Example of determination of reading of portion corresponding to input error) Next, with reference to FIG. 5, an example of determination of the reading of the portion corresponding to the input error detected from the input character string will be described. The process described in FIG. 5 is executed, for example, in kana-kanji conversion processing.

[0131] FIG. 5 is an explanatory diagram showing an example of determination of the reading of the portion corresponding to the input error. In FIG. 5, it is assumed that the user tries to input the phrase "Medina's sashimi" and inputs the character string 501 "mjinanosasimi". "mjinanosasimi" is the input Romanization. "mjinanosasimi" is the result of converting "mjinanosasimi" into kana.

[0132] In the input string 501, the vowel after m is missing. The input error detection unit 402 detects the missing vowel after m in the input string 501 as an input error. In this case, the candidate creation unit 403 creates reading candidates 502-1 to 502-5 for the location 502 corresponding to the input error by performing a correction process of supplementing a vowel (a, i, u, e, o) for the detected input error. Each of the reading candidates 502-1 to 502-5 is obtained by supplementing the vowel "a, i, u, e, o" after m and converting it into kana.

[0133] The extraction unit 404 extracts at least one word before and / or after the location 502 corresponding to the input error from the input string 501. At this time, the extraction unit 404 excludes particles and auxiliary verbs such as "てにをは" from the extraction target. Here, the input string 501 is divided into "mじな (mjina)", "の (no)", and "さしみ (sasimi)", and the word 503 "さしみ (sasimi)" after the location 502 corresponding to the input error is extracted. The "の (no)" immediately after the location 502 corresponding to the input error is excluded from the extraction target.

[0134] The candidate identification unit 405 refers to the semantic connection table 300 and identifies, from among the reading candidates 502-1 to 502-5, the reading candidates corresponding to the words belonging to the group that appears before the extracted word 503 (corresponding to the "post-word") and has a semantic connection with the extracted word 503 by referring to the semantic connection table 300.

[0135] Here, the extracted word 503 matches the post-word "さしみ" of the semantic connection information 300-1, 300-2 in the semantic connection table 300. The reading candidate 502-3 matches the word "むじな" belonging to the group G2 of the semantic connection information 300-2. Therefore, the candidate identification unit 405 identifies, from among the reading candidates 502-1 to 502-5, the reading candidate 502-3 corresponding to the word "むじな" belonging to the group G2 that appears before the extracted word 503 and has a semantic connection with the extracted word 503.

[0136] Further, the reading candidate 502-4 matches the word "mejina" belonging to the group G1 of the semantic connection information 300-1. Therefore, the candidate specifying unit 405 specifies, from among the reading candidates 502-1 to 502-5, the reading candidate 502-4 corresponding to the word "mejina" belonging to the group G1 that appears before the extracted word 503 and has a semantic connection with the extracted word 503.

[0137] Also, the candidate specifying unit 405 refers to the semantic connection table 300 and specifies, for each of the specified reading candidates 502-3 and 502-4, the strength of the semantic connection between each of the groups G1 and G2 to which the words corresponding to the reading candidates 502-3 and 502-4 belong and the extracted word 503.

[0138] Here, the strength of the semantic connection between the group G2 to which the word corresponding to the reading candidate 502-3 belongs and the word 503 is "0.25" from the semantic connection information 300-2. Also, the strength of the semantic connection between the group G1 to which the word corresponding to the reading candidate 502-4 belongs and the word 503 is "0.75" from the semantic connection information 300-1.

[0139] The correct reading determination unit 406 determines the reading of the location 502 corresponding to the input error based on the specified reading candidates 502-3 and 502-4. Specifically, for example, the correct reading determination unit 406 determines, from among the reading candidates 502-3 and 502-4, the reading candidate with the strongest specified connection strength as the reading of the location 502 corresponding to the input error.

[0140] Here, among the reading candidates 502-3 and 502-4, the strength of the semantic connection between the group G1 to which the word corresponding to the reading candidate 502-4 belongs and the extracted word 503 is the strongest. Therefore, the correct reading determination unit 406 determines the reading candidate 502-4 as the reading (correct reading) of the location 502 corresponding to the input error.

[0141] As a result, the reading corresponding to the input string 501 is "めじなのさしみ". Among "めじなのさしみ", "めじな" corresponds to the correction result 504 of the portion 502 corresponding to the input error. Among "めじなのさしみ", "さしみ" corresponds to the word 503.

[0142] In this way, when estimating the reading of the portion 502 corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the semantic connection with the word "さしみ" after the portion 502.

[0143] In the example of FIG. 5, since the appearance frequency of the string "めじなのさしみ" is not high, even if the semantic connection between "めじな" and "さしみ" is not accumulated as knowledge, the input support device 100 can determine the semantic connection between the portion 502 corresponding to the input error and the word 503 from the semantic connection between the group G1 of the semantic attribute "fish" and "さしみ".

[0144] Also, since the appearance frequency of the string "むじなのさしみ" is not high, even if the semantic connection between "むじな" and "さしみ" is not accumulated as knowledge, the input support device 100 can determine the semantic connection between the portion 502 corresponding to the input error and the word 503 from the semantic connection between the group G2 of the semantic attribute "game animal" and "さしみ".

[0145] Also, even when a plurality of reading candidates 502-3, 502-4 are specified from the semantic connection with the word 503, the input support device 100 can estimate the correct reading by considering the strength of the semantic connection with the word 503. Here, since the reading candidate 502-4 has a stronger semantic connection with the word 503 and it can be said that there is a high possibility of being continuously input with the word 503, the reading candidate 502-4 is estimated to be the correct reading.

[0146] As a result, the input assistance device 100 can improve the estimation accuracy of the reading of the portion 502 corresponding to the input error, and can output the conversion candidate "Medina sashimi" desired by the user.

[0147] (Output example of conversion candidate) Next, with reference to FIG. 6, an output example of a conversion candidate corresponding to the input character string will be described. Here, taking the input character string 501 shown in FIG. 5 as an example, the output example of the conversion candidate will be described.

[0148] FIG. 6 is an explanatory diagram showing an output example of a conversion candidate. In (6-1) of FIG. 6, due to the user's operation, the character string 501 is input into the document 600. The user's operation is performed using, for example, the keyboard 209 or the mouse 210 shown in FIG. 2. The document 600 is displayed on, for example, the display 208 shown in FIG. 2.

[0149] In the input character string 501, an input error occurs where the vowel after k is missing. Therefore, as described in FIG. 5, in the kana-kanji conversion process, the reading of the portion 502 (see FIG. 5) corresponding to the input error is estimated, and the reading "mejina no sashimi" corresponding to the input character string 501 is created.

[0150] Note that the kana-kanji conversion process is executed, for example, in response to a conversion instruction from the user. The conversion instruction is performed, for example, by pressing a specific key on the keyboard 209.

[0151] In (6-2) of FIG. 6, based on the created reading "mejina no sashimi", the conversion candidate 601 corresponding to the input character string 501 is displayed selectably. The conversion candidate 601 is "Medina sashimi". Also, among the conversion candidates 601, the portion 602 (completed) where correction processing has been performed on the input error is displayed in a different manner from the other portions.

[0152] Here, the background color of part 602 is different from that of other parts. As another aspect, for example, part 602 may be made bold or underlined. Also, a message 603 is displayed in association with the conversion candidate 601. The message 603 is a message indicating that the input error included in the input string 501 has been corrected.

[0153] As a result, even when the user makes an input error during Japanese input, the user can obtain the desired conversion candidate 601 without re-entering. Also, the user can grasp that the input error has been automatically corrected by referring to the message 603. Further, the user can grasp that the part 602, whose background color is different from that of other parts among the conversion candidates 601, is the corrected part.

[0154] Note that after the conversion candidate 601 is displayed, for example, when a confirmation instruction (selection operation) is performed by the user's operation, the conversion candidate 601 is selected and the input is confirmed.

[0155] (Input support processing procedure of the input support device 100) Next, with reference to FIG. 7, the input support processing procedure of the input support device 100 will be described.

[0156] FIG. 7 is a flowchart showing an example of the input support processing procedure of the input support device 100 according to the embodiment. In the flowchart of FIG. 7, first, the input support device 100 determines whether an input error is detected from the input string (step S701). Here, the input support device 100 waits for an input error to be detected (step S701: No).

[0157] Then, when the input support device 100 detects an input error (step S701: Yes), it creates a reading candidate for the part corresponding to the input error in the input string by performing a predetermined correction process on the input error (step S702). At this time, the input support device 100 may calculate an unnaturalness score for the created reading candidate based on the plausibility of the transition between adjacent characters.

[0158] Next, the input support device 100 determines whether a plurality of reading candidates have been created (step S703). Here, when a plurality of reading candidates have been created (step S703: Yes), the input support device 100 executes a correct reading determination process (step S704) and proceeds to step S706.

[0159] The correct reading determination process is a process of determining the reading of the portion corresponding to the input error from among a plurality of reading candidates. The specific processing procedure of the correct reading determination process will be described later with reference to FIGS. 8 and 9.

[0160] Also, in step S703, when one reading candidate has been created (step S703: No), the input support device 100 determines the created reading candidate as the reading of the portion corresponding to the input error (step S705).

[0161] Next, the input support device 100 creates a conversion candidate corresponding to the input character string based on the reading of the portion corresponding to the determined input error (step S706). Then, the input support device 100 outputs the created conversion candidate for the input character string (step S707), and ends the series of processes according to this flowchart.

[0162] Thereby, when an input error is detected from the input character string, the input support device 100 can estimate the correct reading of the portion corresponding to the input error and output a conversion candidate corresponding to the input character string.

[0163] Next, the specific processing procedure of the correct reading determination process in step S704 will be described with reference to FIGS. 8 and 9.

[0164] FIGS. 8 and 9 are flowcharts showing an example of the specific processing procedure of the correct reading determination process. In the flowchart of FIG. 8, first, the input support device 100 extracts at least one word before and / or after the portion corresponding to the input error from the input character string (step S801).

[0165] Then, the input support device 100 determines whether the word before the location corresponding to the input error has been extracted (step S802). Here, if the previous word has not been extracted (step S802: No), the input support device 100 proceeds to step S901 shown in FIG. 9.

[0166] On the other hand, if the previous word has been extracted (step S802: Yes), the input support device 100 refers to the semantic connection table 300 to identify a group that appears after the extracted word and has a semantic connection with the extracted word (step S803). Then, the input support device 100 determines whether the group has been identified (step S804).

[0167] Here, if the group has not been identified (step S804: No), the input support device 100 proceeds to step S901 shown in FIG. 9. On the other hand, if the group has been identified (step S804: Yes), the input support device 100 refers to the semantic connection table 300 to identify a reading candidate corresponding to the word belonging to the identified group from among the plurality of created reading candidates (step S805).

[0168] Then, the input support device 100 determines whether the reading candidate has been identified (step S806). Here, if the reading candidate has not been identified (step S806: No), the input support device 100 proceeds to step S901 shown in FIG. 9.

[0169] On the other hand, if the reading candidate has been identified (step S806: Yes), the input support device 100 refers to the semantic connection table 300 to identify the strength of the semantic connection between the identified reading candidate, the identified group, and the extracted word (step S807), and proceeds to step S901 shown in FIG. 9.

[0170] In the flowchart of FIG. 9, first, the input support device 100 determines whether a word after the portion corresponding to the input error has been extracted (step S901). Here, if the subsequent word has not been extracted (step S901: No), the input support device 100 proceeds to step S907.

[0171] On the other hand, if the subsequent word has been extracted (step S901: Yes), the input support device 100 refers to the semantic connection table 300 to identify a group that appears before the extracted word and has a semantic connection with the extracted word (step S902). Then, the input support device 100 determines whether a group has been identified (step S903).

[0172] Here, if a group has not been identified (step S903: No), the input support device 100 proceeds to step S907. On the other hand, if a group has been identified (step S903: Yes), the input support device 100 refers to the semantic connection table 300 to identify a reading candidate corresponding to a word belonging to the identified group from among the plurality of created reading candidates (step S904).

[0173] Then, the input support device 100 determines whether a reading candidate has been identified (step S905). Here, if a reading candidate has not been identified (step S905: No), the input support device 100 proceeds to step S907.

[0174] On the other hand, if a reading candidate has been identified (step S905: Yes), the input support device 100 refers to the semantic connection table 300 to identify the strength of the semantic connection between the identified reading candidate, the identified group, and the extracted word (step S906).

[0175] Next, the input support device 100 determines whether a plurality of reading candidates have been identified from among the plurality of created reading candidates (step S907).

[0176] Here, when a plurality of reading candidates are specified (step S907: Yes), the input support device 100 specifies, from among the plurality of specified reading candidates, the reading candidate with the strongest semantic connection strength (step S908). Then, the input support device 100 determines the reading candidate specified in step S908 as the reading at the location corresponding to the input error (step S909), and returns to the step where the correct reading determination process was called.

[0177] Also, in step S907, when one reading candidate is specified (step S907: No), the input support device 100 determines the specified reading candidate as the reading at the location corresponding to the input error (step S910), and returns to the step where the correct reading determination process was called.

[0178] Thereby, even when multiple reading candidates are assumed for the location corresponding to the input error, the input support device 100 can estimate the correct reading considering the strength of the semantic connection with the words before and after the location corresponding to the input error.

[0179] Note that in step S907, when no reading candidate is specified, the input support device 100 may, for example, determine, as the reading at the location corresponding to the input error, the reading candidate with the lowest unnaturalness score among the plurality of created reading candidates.

[0180] Also, in step S908, when a plurality of the strongest reading candidates are specified, the input support device 100 may, for example, determine, as the reading at the location corresponding to the input error, the reading candidate with the lowest unnaturalness score among the specified strongest reading candidates.

[0181] Also, when words before and after the location corresponding to the input error are extracted, the input support device 100 may identify reading candidates corresponding to words that appear after the extracted previous word and before the extracted subsequent word and belong to a group having a semantic connection with the extracted previous word and subsequent word. Then, the input support device 100 may determine the identified reading candidates as the reading of the location corresponding to the input error. Thereby, when there is a reading candidate having a semantic connection with the surrounding words, the input support device 100 can preferentially determine the reading candidate as the reading of the location corresponding to the input error.

[0182] As described above, according to the input support device 100 according to the embodiment, when an input error is detected from the input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the input character string, one or more reading candidates can be created. Also, according to the input support device 100, with reference to the storage unit 410, among the one or more created reading candidates, a reading candidate corresponding to a word belonging to a group having a semantic connection with at least one of the words before and after the location corresponding to the input error in the input character string can be identified. The storage unit 410 stores, in association with a group in which words with semantic similarity are classified, a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group. Then, according to the input support device 100, based on the identified reading candidates, the reading of the location corresponding to the input error can be determined.

[0183] Accordingly, when estimating the reading of the portion corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the semantic connection with the words before and after that portion. At this time, the input support device 100 treats words with semantic similarity as a group, and can determine the semantic connection between the portion corresponding to the input error and the words before and after from the semantic connection between the group and a specific word. For example, even when there are multiple assumed reading candidates for the portion corresponding to the input error, the input support device 100 can determine as the reading of the portion corresponding to the input error a reading candidate that has a semantic connection with the words before and after. Also, by treating words with semantic similarity as a group, the input support device 100 can determine the semantic connection with a specific word even for a word that has a semantic connection with a specific word but is not stored as knowledge because its appearance frequency is not high.

[0184] Further, according to the input support device 100, with reference to the storage unit 410 in which the strength of the semantic connection between the group and a specific word is stored, for the identified reading candidate, the strength of the semantic connection between the group to which the word corresponding to the reading candidate belongs and at least one of the words before and after the portion corresponding to the input error can be identified. Then, according to the input support device 100, the reading of the portion corresponding to the input error can be determined from among the identified reading candidates based on the identified connection strength.

[0185] Accordingly, when estimating the reading of the portion corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the strength of the semantic connection with the words before and after that portion. For example, even when there are multiple assumed reading candidates for the portion corresponding to the input error, the input support device 100 can estimate as the reading of the portion corresponding to the input error a reading candidate that has a stronger semantic connection with the words before and after.

[0186] Further, according to the input support device 100, it is possible to determine the reading of the portion corresponding to the input error from among the identified reading candidates based on the unnaturalness score of the identified reading candidates.

[0187] Thereby, when estimating the reading of the portion corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the semantic connection with the words before and after that portion and the likelihood of the character arrangement.

[0188] Further, according to the input support device 100, it is possible to determine the reading of the portion corresponding to the input error from among the identified reading candidates based on the identified connection strength and the identified unnaturalness score.

[0189] Thereby, when estimating the reading of the portion corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the semantic connection with the words before and after that portion, the strength of that connection, and the likelihood of the character arrangement.

[0190] Further, according to the input support device 100, among the identified reading candidates, the reading candidate with the strongest identified connection strength is identified. When a plurality of the strongest reading candidates are identified, the reading candidate with the lowest unnaturalness score among the identified strongest reading candidates can be determined as the reading of the portion corresponding to the input error.

[0191] Thereby, even when there are a plurality of reading candidates that can be said to have the strongest semantic connection with the words before and after, the input support device 100 can determine the reading candidate with a more likely character arrangement as the reading of the portion corresponding to the input error.

[0192] Further, according to the input support device 100, by referring to the storage unit 410 in which the context between a group and a specific word is stored, based on the context, among the one or more reading candidates created, a reading candidate corresponding to a word belonging to a group having a semantic connection with at least one of the words before and after the location corresponding to the input error can be specified.

[0193] Thereby, the input support device 100 can improve the estimation accuracy of the correct reading by determining the semantic connection in consideration of whether the word is before or after the location corresponding to the input error. For example, the semantic connection information 300-1 shown in FIG. 3 shows the semantic connection between the group G1 and the word "sashimi" that is continuously input after the word belonging to the group G1. Therefore, when there is a word "sashimi" after the location 502 corresponding to the input error as in the input character string 501 shown in FIG. 5, the semantic connection information 300-1 is applied. On the other hand, even if there is a word "sashimi" before the location 502 corresponding to the input error, the semantic connection information 300-1 is not applied. Note that the semantic connection table 300 may not define the context between a group and a specific word (previous word or next word). In this case, the input support device 100 may refer to the semantic connection table 300, for example, regardless of whether the extracted word is the word before or after the location corresponding to the input error, and specify a reading candidate corresponding to a word belonging to a group having a semantic connection with the extracted word (corresponding to the "specific word") from among the created reading candidates.

[0194] In addition, according to the input support device 100, when the first reading candidate and the second reading candidate are included in the specified reading candidates, the second reading candidate can be determined as the reading of the location corresponding to the input error. The first reading candidate is a reading candidate corresponding to a word belonging to a group that has a semantic connection with one of the words before and after the location corresponding to the input error. The second reading candidate is a reading candidate corresponding to a word belonging to a group that has a semantic connection with at least two or more of the words before and after the location corresponding to the input error. However, a specific word is either one or two or more words continuously input before or after the words belonging to the group, or two or more words continuously input with the words belonging to each group sandwiched therebetween.

[0195] Thereby, the input support device 100 can preferentially determine, as the reading of the location corresponding to the input error, a reading candidate with a longer connection as the semantic connection between the words before and after, and can improve the estimation accuracy of the correct reading.

[0196] In addition, according to the input support device 100, a conversion candidate corresponding to the input character string can be output based on the reading of the location corresponding to the determined input error.

[0197] Thereby, the input support device 100 can suppress the output of conversion candidates that the user does not want.

[0198] From these facts, according to the input support program, the input support method, and the input support device according to the embodiment, even when the user makes an input error, the correct reading of the location corresponding to the input error can be accurately estimated, and the conversion accuracy of Japanese input can be improved. For example, when creating a document using a document creation application, the time and effort for the user to re-enter can be reduced, the workload related to Japanese input can be reduced, and the efficiency of document creation can be improved.

[0199] In addition, when, for example, the semantic connection between a certain word and a specific word is obtained as knowledge, the app developer can group together those words that are semantically similar to that word and register the semantic connection with the specific word, thereby taking into account the semantic connection between words. For this reason, in order to consider the semantic connection between words, compared with the case where the relationship between words is registered one by one, the workload and working time required for app development can be reduced, and the app development period can be shortened.

[0200] Note that the input support method described in this embodiment can be realized by executing a pre-prepared program on a computer such as a personal computer or a workstation. This input support program is recorded on a computer-readable recording medium such as a hard disk, a flexible disk, a CD-ROM, a DVD, or a USB memory, and is executed by being read from the recording medium by a computer. Further, this input support program may be distributed via a network such as the Internet.

Industrial Applicability

[0201] The input support program, input support method, and input support device according to this invention are useful for a computer system that supports character input performed based on a user's operation, and in particular, are suitable for a computer system that inputs Japanese based on a user's operation.

Explanation of Signs

[0202] 100 Input support device 101, 501 Character string 102, 502 Location 102-1 to 102-5, 502-1 to 502-5 Reading candidates 103, 121, 503 Word 104, 504 Correction result 110, 410 Storage unit 120, G1, G2, G3 Group 200 Bus 201 CPU 202 ROM 203 RAM 204 HDD 205 HD 206 CD-RW drive 207 CD-RW 208 Display 209 Keyboard 210 Mouse 211 Network I / F 300 Semantic connection table 401 Reception section 402 Input error detection section 403 Candidate creation section 404 Extraction section 405 Candidate identification section 406 Correct reading determination section 407 Output section 600 Document 601 Conversion candidate 602 Part 603 Message NW Network

Claims

Claim 1 When an input error is detected in the input string, one or more reading candidates are created for the part of the string corresponding to the input error by performing a predetermined correction process on the input error, in association with a group in which words with semantic similarity are classified, and referring to a storage unit that stores a specific word that has a semantic connection with the group and is continuously input with the words belonging to the group, among the one or more reading candidates created, identify a reading candidate corresponding to a word belonging to a group that has a semantic connection with at least one of the words before and after the part in the string, Based on the identified reading candidate, determine the reading of the part, An input support program characterized by causing a computer to execute the process. Claim 2 The storage unit further stores the strength of the semantic connection between the group and the specific word, Referring to the storage unit, cause the computer to execute a process of identifying the strength of the semantic connection between the group to which the word corresponding to the identified reading candidate belongs and either of the words for the identified reading candidate, The process of determining is Based on the identified strength of the connection, determine the reading of the part from among the identified reading candidates. The input support program according to claim 1, characterized in that. Claim 3 The process of determining is Based on an evaluation value regarding the arrangement of characters based on the plausibility of the transition between adjacent characters in the identified reading candidate, determine the reading of the part from among the identified reading candidates. The input support program according to claim 1, characterized in that. Claim 4 The process of determining is Based on the identified strength of the connection and an evaluation value regarding the arrangement of characters based on the plausibility of the transition between adjacent characters in the identified reading candidate, determine the reading of the part from among the identified reading candidates. The input support program according to claim 2, characterized in that. Claim 5 The process of determining is Among the identified reading candidates, identify the reading candidate with the strongest identified strength of the connection. When a plurality of the strongest reading candidates are identified, determine the reading of the part as the reading candidate with the highest evaluation of the arrangement of characters based on the evaluation value among the identified strongest reading candidates. The input support program according to claim 4, characterized in that. Claim 6 The memory unit stores the context between the group and the specific word. The specifying process refers to the memory unit, and based on the context, among the one or more reading candidates created, specifies a reading candidate corresponding to a word belonging to a group having a semantic connection with any of the words, the input support program according to claim 1, characterized in that.

7. The specific word is either one or more words continuously input before or after a word belonging to the group, or two or more words continuously input with a word belonging to the group sandwiched therebetween. The determining process If among the specified reading candidates, there is a first reading candidate corresponding to a word belonging to a group having a semantic connection with one word before or after the location, and a second reading candidate corresponding to a word belonging to a group having a semantic connection with at least two or more words before or after the location, determines the second reading candidate as the reading of the location, the input support program according to claim 1, characterized in that.

8. Outputs a conversion candidate corresponding to the character string based on the determined reading of the location. The input support program according to any one of claims 1 to 7, characterized in that the computer is caused to execute the process.

9. When an input error is detected from the input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, create one or more reading candidates. Refer to a memory unit that stores a specific word that is associated with a group formed by classifying words with semantic similarity and has a semantic connection with the group and is continuously input with a word belonging to the group, and among the one or more reading candidates created, specify a reading candidate corresponding to a word belonging to a group having a semantic connection with at least one of the words before and after the location in the character string. Based on the specified reading candidate, determine the reading of the location. An input support method, characterized in that a computer executes the process.

10. When an input error is detected from the input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, create one or more reading candidates. In association with a group obtained by classifying words having semantic similarity, referring to a storage unit that stores a specific word that has a semantic connection with the group and that is continuously input with the words belonging to the group, among the one or more reading candidates created, a reading candidate corresponding to a word belonging to a group having a semantic connection with at least one of the words before and after the position in the character string is specified. Based on the specified reading candidate, the reading of the position is determined. An input support device characterized by having a control unit that executes processing.

Citation Information

Patent Citations

  • 'KANA' / 'kanji' converter

    JP1993081238A