Input support program, input support method, and input support apparatus
By performing a correction process and utilizing semantic connections to select the most appropriate reading candidate, the system effectively addresses the challenge of accurately correcting user input mistakes in Japanese character input, enhancing estimation accuracy and user experience.
Patent Information
- Application Number
- JP2023198461
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-06-03
AI Technical Summary
Existing input support systems struggle to accurately correct user input mistakes in Japanese character input, leading to decreased estimation accuracy of readings and the output of undesired conversion candidates.
The system performs a predetermined correction process on detected input errors, creates reading candidates for the affected portion, and refers to a storage unit that stores semantic connections between continuously input words to identify the most appropriate reading candidate based on semantic connections and evaluation values.
This approach significantly improves the estimation accuracy of the correct reading at the position corresponding to the input error, reducing the output of unwanted conversion candidates and enhancing user experience by minimizing the need for re-entering characters.
Smart Images

Figure 2025084505000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an input support program, an input support method, and an input support device that assist character input performed based on a user's operation.
Background Art
[0002] Conventionally, in a Japanese input system, there is a technique for automatically correcting input mistakes made by a user. For example, when a vowel omission (input mistake) is detected from an input character string, a vowel is supplemented at the omitted position to estimate the reading, and a result obtained by converting the estimated reading into kanji or the like is output as a conversion candidate.
[0003] As a prior art for assisting character input, a character string being input is divided into clauses, and kana-kanji conversion is performed for each clause using a dictionary that holds information indicating at least the reading, notation, part of speech, and conjugation form, and a group of candidate character strings indicating the kana-kanji conversion result and their respective occurrence probabilities are output (see, for example, Patent Document 1 below).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the prior art, there is a problem that input mistakes made by a user cannot be correctly corrected, the estimation accuracy of the reading at the position corresponding to the input mistake decreases, and conversion candidates not desired by the user may be output.
[0006] In one aspect, an object of the present invention is to provide an input support program, an input support method, and an input support device that improve the estimation accuracy of the reading at the position corresponding to an input mistake.
Means for Solving the Problem
[0007] In order to solve the above-described problems and achieve the object, when an input error is detected from an input string, the input support program according to the present invention performs a predetermined correction process on the input error, and creates one or more reading candidates for a portion corresponding to the input error in the string, and refers to a storage unit that stores the context between words having a semantic connection that are continuously input, and among the one or more reading candidates created, identifies a reading candidate having a semantic connection with at least one of the words before and after the portion in the string, and based on the identified reading candidate, determines the reading of the portion, and causes a computer to execute the process.
[0008] Further, in the input support program according to the present invention, in the above invention, the storage unit further stores the strength of the semantic connection between the words, and refers to the storage unit to identify the strength of the semantic connection between the identified reading candidate and any of the words, and causes the computer to execute the process of identifying, and the determining process determines the reading of the portion based on the identified strength of the connection among the identified reading candidates.
[0009] Further, in the input support program according to the present invention, in the above invention, the determining process determines, as the reading of the portion, the reading candidate having the strongest semantic connection with any of the words among the identified reading candidates based on the identified strength of the connection.
[0010] Further, in the input support program according to the present invention, in the above invention, the determining process determines the reading of the portion based on the identified strength of the connection and an evaluation value regarding the arrangement of characters based on the plausibility of the transition between adjacent characters in the identified reading candidate among the identified reading candidates.
[0011] Further, in the input support program according to the present invention, in the above invention, the determination process is based on the identified strength of the connection, among the identified reading candidates, to identify the reading candidate with the strongest semantic connection with any of the words, and when a plurality of the strongest reading candidates are identified, among the identified strongest reading candidates, the reading candidate with the highest evaluation of the character arrangement based on the evaluation value is determined as the reading of the location.
[0012] Further, in the input support program according to the present invention, in the above invention, the computer is caused to execute a process of outputting a conversion candidate corresponding to the character string based on the determined reading of the location.
[0013] Further, in the input support method according to the present invention, when an input error is detected from an input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, one or more reading candidates are created, and referring to a storage unit that stores the context between words having a semantic connection that are continuously input, among the one or more created reading candidates, a reading candidate having a semantic connection with at least any of the words before and after the location in the character string is identified, and based on the identified reading candidate, the computer executes a process of determining the reading of the location.
[0014] Further, the input support device according to the present invention has a control unit that, when an input error is detected from an input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, one or more reading candidates are created, and referring to a storage unit that stores the context between words having a semantic connection that are continuously input, among the one or more created reading candidates, a reading candidate having a semantic connection with at least any of the words before and after the location in the character string is identified, and based on the identified reading candidate, executes a process of determining the reading of the location.
Advantages of the Invention
[0015] According to the input support program, input support method, and input support device of the present invention, it is possible to achieve the effect of improving the estimation accuracy of the reading of the part corresponding to the input error.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0017] (Embodiment) Hereinafter, with reference to the drawings, embodiments of the input support program, input support method, and input support device according to the present invention will be described in detail.
[0018] FIG. 1 is an explanatory diagram showing an example of an input support method according to an embodiment. In FIG. 1, an input support device 100 is a computer that supports character input performed based on a user's operation. The input support device 100 is, for example, a PC (Personal Computer). The input support device 100 may be a tablet PC, a smartphone, or the like. Further, the input support device 100 may be a server connectable from a PC or the like used by the user.
[0019] The user's operation is performed using an input device such as a keyboard or a touch panel, for example. In the case of Japanese input, there are input modes such as Roman character input and kana input. In Roman character input, when inputting Japanese, Roman characters combining the consonants and vowels of the characters are input. Kana input is performed using kana characters written on a keyboard or the like.
[0020] In character input performed based on the user's operation, input errors may occur. Examples of input errors include omission of vowels, input of extra characters, and incorrect order of character input. If an input error occurs by the user, it is convenient if the input error can be automatically corrected, as it saves the trouble of the user having to re-enter.
[0021] As a method for correcting input errors, for example, when an omission of a vowel occurs, there is one that supplements the vowel at the omitted position. Also, by a method based on machine learning, the natural order of Roman characters as Japanese is learned in advance, and for input errors, it is conceivable to supplement characters, delete extra characters, or swap the order of characters.
[0022] However, in the prior art, there are cases where input errors by the user cannot be correctly corrected. If the input error cannot be correctly corrected, the correct reading of the portion corresponding to the input error cannot be obtained, and a conversion result that the user does not desire is output.
[0023] For example, when inputting the character string "kansei", the vowel "a" after the k is omitted, resulting in the input character string "knsei". In this case, if you correct the missing part by adding a vowel (a, i, u, e, o), five possible readings for the part corresponding to the input mistake will be created: "kansei, kinsei, kunsei, kensei, konsei".
[0024] In the conventional technology, it is difficult to determine which of the five candidate reading patterns is the correct reading. For example, it is possible to determine the correct reading from the five candidate reading patterns by taking into account the natural order of the romaji characters in the learned Japanese language. However, in the above example, all of the five candidate reading patterns exist as actual words, and it is difficult to determine the correct reading from the order of the romaji characters.
[0025] For example, among the vowels (a, i, u, e, o), "e" has the highest probability of transitioning from "k," and it is determined that "kensei" is the most natural romanization. In this case, "knsei" is corrected to "kensei," and "kensei" is determined to be the correct reading.
[0026] However, suppose that the user is trying to input the word "kanseisho" (completed presentation) and inputs the character string "hirou" after "knsei." In this case, "kansei" is the correct reading, and "kensei" is an incorrect reading that the user did not intend.
[0027] Therefore, in this embodiment, when estimating the reading of a portion corresponding to an input error, an input support method that improves the estimation accuracy of the correct reading by considering the semantic connection with the surrounding words will be described. Here, a processing example of the input support device 100 will be described. The processing example of the input support device 100 is executed, for example, in kana-kanji conversion processing. Kana-kanji conversion processing is a process of converting an input string into a string of kanji or a sentence containing kanji.
[0028] (1) When an input error is detected from the input string 101, the input support device 100 creates one or more reading candidates for the portion corresponding to the input error in the input string 101 by performing a predetermined correction process on the input error. Here, the input string 101 is a string input based on a user's operation.
[0029] The input string 101 is a string before confirmation. For example, in the case of Japanese input in the Roman character input mode, the input string 101 corresponds to the input Roman character spelling or the result of converting the Roman character spelling into kana. Also, in the case of Japanese input in the kana input mode, the input string 101 corresponds to the input kana spelling (keystrokes).
[0030] Input errors include, for example, omission of vowels, input of extra characters, and incorrect order of character input. The predetermined correction process is a process of supplementing characters such as vowels, deleting extra characters, or swapping the order of characters for the input error. The portion corresponding to the input error is, for example, a clause including the location of the input error.
[0031] In the example of FIG. 1, assuming Japanese input in the Roman character input mode, let the input string 101 be "kんせいひろう (knseihirou)". "knseihirou" is the input Roman character spelling. "kんせいひろう" is the result of converting "knseihirou" into kana.
[0032] In the input string 101, the vowel after "k" is missing. Therefore, the missing vowel after "k" is detected as an input error in the input string 101. In this case, for example, a correction process of supplementing a vowel (a, i, u, e, o) for the input error is performed, and reading candidates 102-1 to 102-5 for the portion 102 corresponding to the input error are created.
[0033] The reading candidate 102-1 is the one obtained by supplementing the vowel "a" after "k" and converting it into kana. The reading candidate 102-2 is the one obtained by supplementing the vowel "i" after "k" and converting it into kana. The reading candidate 102-3 is the one obtained by supplementing the vowel "u" after "k" and converting it into kana. The reading candidate 102-4 is the one obtained by supplementing the vowel "e" after "k" and converting it into kana. The reading candidate 102-5 is the one obtained by supplementing the vowel "o" after "k" and converting it into kana.
[0034] (2) The input support device 100 refers to the storage unit 110 and specifies, from among the created reading candidates 102-1 to 102-5, a reading candidate that has a semantic connection with at least one of the words before and after the portion 102 corresponding to the input error in the input string 101. Here, the storage unit 110 stores the sequential relationship between words with a semantic connection that are continuously input.
[0035] Words that are continuously input are, for example, words that are continuously input. Also, words that are continuously input may be words that are input with a particle or auxiliary verb such as "てにをは" in between. A semantic connection is a relationship in which two or more words are combined to form a new meaningful word.
[0036] For example, a semantic connection may be a relationship in which two or more words are combined to form a compound word. Also, a semantic connection may be a relationship in which two or more words are combined to form an idiomatic expression. Also, a semantic connection may be a relationship between a modifier and a modified word.
[0037] In the example of Fig. 1, the context (word order) of the word "kansei" and the word "hiro" is stored in the storage unit 110 as a context between words that are semantically related. The words "kansei" and "hiro" are in a relationship that forms a compound word "kanketsuhanan." The context between the words is that the word "kansei" comes first and the word "hiro" comes after.
[0038] Specifically, for example, the input support device 100 extracts at least one of the words before or after the portion 102 corresponding to the input error from the input character string 101. The previous word is, for example, the word immediately before the portion 102 corresponding to the input error in the input character string 101. The following word is, for example, the word immediately after the portion 102 corresponding to the input error in the input character string 101. However, particles and auxiliary verbs such as "te, ni, o, ha" may be excluded from the extraction targets.
[0039] In the example of FIG. 1, the word 103 "hirou" immediately following the portion 102 corresponding to the input error is extracted from the input character string 101. The word 103 matches the word "hirou" stored in the storage unit 110. The word "hirou" has a semantic connection with the word "kansei" and appears after the word "kansei". In this case, the input support device 100 refers to the storage unit 110 and specifies the reading candidate 102-1, which appears before the extracted word 103 and has a semantic connection with the extracted word 103, from among the reading candidates 102-1 to 102-5.
[0040] (3) The input support device 100 determines the reading of the portion 102 corresponding to the input error based on the identified reading candidate. Determining the reading corresponds to estimating the correct reading. The identified reading candidate has a stronger connection with the words before and after it compared to other reading candidates, and is therefore more likely to be the correct reading of the portion 102 corresponding to the input error.
[0041] Here, one reading candidate 102-1 is identified as a reading candidate having a semantic connection with the word 103. In this case, the input support device 100, for example, determines the identified reading candidate 102-1 as the reading of the portion 102 corresponding to the input error. As a result, the reading corresponding to the input character string 101 becomes "kansei hirou." "Kansei" in "kansei hirou" corresponds to the correction result 104 of the portion 102 corresponding to the input error. "Hirou" in "kansei hirou" corresponds to the word 103. "Kansei hirou" is converted to, for example, "kanketsuhanan" and output as a conversion candidate.
[0042] In this way, according to the input support device 100, when estimating the reading of a portion corresponding to an input error, the accuracy of estimating the correct reading can be improved by considering the semantic connection with the words before and after the portion. In the example of Fig. 1, even if multiple reading candidates for the portion 102 corresponding to the input error are created, the input support device 100 can determine the reading candidate 102-1 that has a semantic connection with the word 103 following the portion 102 as the reading of the portion 102 corresponding to the input error.
[0043] As a result, even if the user makes an input error, the input support device 100 can accurately estimate the correct reading of the part corresponding to the input error and prevent the output of conversion candidates that the user does not want. Also, the input support device 100 can reduce the user's effort to re-enter characters and reduce the workload involved in character input.
[0044] (Example of Hardware Configuration of Input Support Device 100) Next, a hardware configuration example of the input support device 100 will be described with reference to Fig. 2. Here, the input support device 100 will be described as being applied to a computer such as a PC or a tablet PC. However, the input support device 100 may also be applied to a server that can be connected from a PC or the like used by a user.
[0045] FIG. 2 is a block diagram showing a hardware configuration example of the input support device 100 according to the embodiment. In FIG. 2, the input support device 100 includes a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, and a RAM (Random Access Memory) 203.
[0046] Further, the input support device 100 includes an HDD (Hard Disc Drive) 204, an HD 205, a CD (Compact Disc)-RW (ReWritable) drive 206, and a CD-RW 207. The input support device 100 also includes a display 208, a keyboard 209, a mouse 210, and a network I / F (Interface) 211. Each component is connected by a bus 200.
[0047] Here, the CPU 201 controls the entire input support device 100. The CPU 201 may have a plurality of cores. The ROM 202 and the HD 205 store various programs. The programs stored in the ROM 202 and the HD 205 are, for example, the present input support program. The present input support program is applied to, for example, kana-kanji conversion software. The program stored in the ROM 202 is loaded into the CPU 201 to cause the CPU 201 to execute the coded processing. The RAM 203 is used as a work area for the CPU 201.
[0048] The HDD 204 controls the read or write of data to the HD 205 according to the control of the CPU 201. The HD 205 stores the data written according to the control of the HDD 204. The CD-RW drive 206 controls the read or write of data to the CD-RW 207 according to the control of the CPU 201. The CD-RW 207 stores the data written according to the control of the CD-RW drive 206. The CD-RW 207 may be, for example, detachable from the input support device 100.
[0049] The display 208 displays various data such as a cursor, an icon, a menu, a window, a toolbox, characters, an image, or function information. The display 208 is, for example, a liquid crystal display, an organic EL (Electroluminescence) display, or the like.
[0050] The keyboard 209 has keys for inputting characters, numerical values, various instructions, etc., and inputs data. The mouse 210 selects or executes various instructions, selects a processing target, or moves the mouse pointer. Further, the display 208 may be a touch panel and may have functions corresponding to the keyboard 209 and the mouse 210. In this case, the input support device 100 may not have the keyboard 209 and the mouse 210.
[0051] The network I / F 211 is connected to the network NW through a communication line and is connected to other computers via the network NW. The network NW is, for example, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, or the like. The network I / F 211 manages the interface between the network NW and the inside of the input support device 100 and controls the input / output of data from other computers. The network I / F 211 is, for example, a modem, a LAN adapter, or the like.
[0052] In addition to the components described above, the input support device 100 may have, for example, a DVD (Digital Versatile Disc) drive, or an SSD (Solid State Drive). Further, in addition to the components described above, the input support device 100 may have, for example, a USB (Universal Serial Bus) port. Further, in addition to the components described above, the input support device 100 may have, for example, a printer, a scanner, a microphone, or a speaker. Further, the input support device 100 may not have, for example, an HDD 204, an HD 205, a CD-RW drive 206, a CD-RW 207, etc. among the components described above.
[0053] (Stored content of the semantic connection table 300) Next, the stored content of the semantic connection table 300 used by the input support device 100 will be described. The semantic connection table 300 is realized by a storage device such as a RAM 203 or an HD 205, for example.
[0054] FIG. 3 is an explanatory diagram showing an example of the stored content of the semantic connection table 300. In FIG. 3, the semantic connection table 300 has fields for connection strength, previous word, and next word, and stores semantic connection information (for example, semantic connection information 300-1 to 300-6) as records by setting information in each field.
[0055] Here, the connection strength indicates the strength (degree) of the semantic connection between the previous word and the next word. The connection strength is represented by a numerical value from 0 to 1, for example, and the larger the value, the stronger the semantic connection. However, the connection strength may be represented by levels such as "strong, medium, weak".
[0056] The previous word is a word (in kana) that has a semantic connection with the subsequent word that appears after the previous word. The subsequent word is a word (in kana) that has a semantic connection with the previous word that appears before the subsequent word. The previous word and the subsequent word are, for example, continuously input when inputting Japanese, or are input with a particle or auxiliary verb such as "te ni o wa" in between. A conversion result obtained by converting the word (in kana) into Chinese characters may be associated with the previous word and the subsequent word.
[0057] For example, the semantic connection information 300-1 indicates the strength of the semantic connection "0.87" between the previous word "kansei" and the subsequent word "hiroう" that are continuously input.
[0058] Note that in the example of FIG. 3, the front-back relationship between two words with a semantic connection was used as an example for explanation, but it is not limited to this. For example, the semantic connection table 300 may store the front-back relationship (word order) between three or more words with a semantic connection. Specifically, for example, the semantic connection table 300 may store the front-back relationship between three words (previous word, middle word, subsequent word) with a semantic connection, and the strength of the semantic connection between the three words.
[0059] The semantic connection table 300 may be automatically created, for example, by analyzing Japanese document data, or may be created manually. Also, the stored content of the semantic connection table 300 may be updated at any time. For example, according to the number of occurrences or the frequency of occurrence of combinations of words with a semantic connection in Japanese document data, new semantic connections between words may be registered or the connection strength may be updated.
[0060] (Functional configuration example of the input support device 100) Next, a functional configuration example of the input support device 100 will be described with reference to FIG. 4.
[0061] FIG. 4 is a block diagram showing a functional configuration example of the input support device 100 according to the embodiment. In FIG. 4, the input support device 100 includes a reception unit 401, an input error detection unit 402, a candidate creation unit 403, an extraction unit 404, a candidate identification unit 405, a correct reading determination unit 406, an output unit 407, and a storage unit 410. The reception unit 401 to the output unit 407 are functions that serve as a control unit. For example, by causing the CPU 201 to execute a program stored in a storage device such as the ROM 202, RAM 203, or HD 205 shown in FIG. 2, or by means of the network I / F 211, their functions are realized. The processing results of each functional unit are stored in a storage device such as the RAM 203 or HD 205, for example. The storage unit 410 is realized by a storage device such as the RAM 203 or HD 205, for example. Specifically, for example, the storage unit 410 stores the semantic connection table 300 shown in FIG. 3. The storage unit 110 shown in FIG. 1 corresponds to the storage unit 410, for example.
[0062] The reception unit 401 receives the input of a character string. The character string to be input is, for example, a character string that is the target of kana-kanji conversion. Specifically, for example, the reception unit 401 receives the input of a character string by a user operation using the keyboard 209 or mouse 210 shown in FIG. 2.
[0063] The input error detection unit 402 detects input errors from the input character string. Input errors are, for example, omission of vowels, input of extra characters, incorrect order of character input, etc. Specifically, for example, the input error detection unit 402 learns the characteristics of Japanese from information on character strings input in the past during Japanese input by a method based on machine learning such as deep learning. The characteristics of Japanese are, for example, the natural arrangement of Roman letters as Japanese.
[0064] Then, the input error detection unit 402 detects input errors from the input character string based on the learned characteristics of Japanese. For example, the input error detection unit 402 detects as input errors the places where vowels are omitted or extra characters are input and deviate from the natural arrangement of Roman letters as learned Japanese.
[0065] More specifically, for example, input errors include insertion errors, deletion errors, substitution errors, and transposition errors. An insertion error is a mistake where an arbitrary keyed character is entered additionally (e.g., nyuuryoku (にゅうりょく) → nyuuuryoku (にゅううりょく)). A deletion error is a mistake where an arbitrary keyed character is missing (e.g., nyuuryoku (にゅうりょく) → nyuuyoku (にゅうよく)). A substitution error is a mistake where an arbitrary keyed character is mistaken for another character (such as a character on an adjacent key) (e.g., nyuuryoku (にゅうりょく) → nuuuryoku (ぬううりょく)). A transposition error is a mistake where the keying order of an arbitrary keyed character is reversed (e.g., nyuuryoku (にゅうりょく) → nyuuyroku (にゅうyろく)). Note that any existing technology may be used as the technology for detecting input errors.
[0066] When an input error is detected from the input string, the candidate creation unit 403 creates one or more reading candidates for the portion corresponding to the input error in the input string by performing a predetermined correction process on the input error. The portion corresponding to the input error is, for example, a clause including the location of the input error.
[0067] Specifically, for example, the candidate creation unit 403 performs a correction process on the input error, such as supplementing vowels, deleting extra characters, or swapping the order of characters, so that the learned Roman characters in Japanese are in a natural order. Thereby, the candidate creation unit 403 creates one or more reading candidates for the portion corresponding to the input error.
[0068] The extraction unit 404 extracts at least one of the words before and after the portion corresponding to the input error from the input string. Specifically, for example, the extraction unit 404 extracts the word immediately before the portion corresponding to the input error from the input string. Also, the extraction unit 404 extracts the word immediately after the portion corresponding to the input error from the input string.
[0069] However, particles and auxiliary verbs such as "te ni o wa" may be excluded from the extraction target. For example, if the word immediately before the location corresponding to the input error is a particle, the extraction unit 404 extracts the word immediately before that particle. Also, if the word immediately after the location corresponding to the input error is a particle, the extraction unit 404 extracts the word immediately after that particle. Further, the extraction unit 404 may extract two or more words before the location corresponding to the input error. Also, the extraction unit 404 may extract two or more words after the location corresponding to the input error. In this case, the maximum number of words to be extracted may be set in advance.
[0070] The candidate identification unit 405 refers to the storage unit 410 and identifies, among the one or more created reading candidates, a reading candidate that has a semantic connection with at least one of the words before and after the location corresponding to the input error in the input character string. Here, the storage unit 410 stores the sequential relationship between words with a semantic connection that are continuously input.
[0071] Further, the storage unit 410 may further store the strength of the semantic connection between continuously input words. In this case, the candidate identification unit 405 may refer to the storage unit 410 and identify the strength of the semantic connection between the identified reading candidate and at least one of the words before and after the location corresponding to the input error.
[0072] For example, it is assumed that the word before the location corresponding to the input error is extracted from the input character string. In this case, the candidate identification unit 405 refers to, for example, the semantic connection table 300 shown in FIG. 3, and among the created reading candidates, identifies a reading candidate (corresponding to the "subsequent word") that appears after the extracted word (corresponding to the "previous word") and has a semantic connection with the extracted word.
[0073] At this time, the candidate identification unit 405 may refer to the semantic connection table 300 to identify the strength of the semantic connection between the identified reading candidates and the extracted words. The strength of the semantic connection is represented by a numerical value from 0 to 1, for example, and the larger the value, the stronger the semantic connection (see FIG. 3). Further, the strength of the semantic connection may be represented by levels such as "strong, medium, weak", for example.
[0074] Note that the identification of the strength of the semantic connection may be performed when a plurality of reading candidates having a semantic connection with the extracted word are identified. Specifically, for example, when the candidate identification unit 405 identifies a plurality of reading candidates, for each of the identified plurality of reading candidates, the candidate identification unit 405 may identify the strength of the semantic connection with at least one of the words before and after the location corresponding to the input error.
[0075] Further, assume that a word after the location corresponding to the input error is extracted from the input string. In this case, the candidate identification unit 405 refers to the semantic connection table 300, for example, and identifies, from among the created reading candidates, a reading candidate (corresponding to the "previous word") that appears before the extracted word (corresponding to the "subsequent word") and has a semantic connection with the extracted word. At this time, the candidate identification unit 405 may refer to the semantic connection table 300 to identify the strength of the semantic connection between the identified reading candidate and the extracted word.
[0076] In the following description, the strength of the semantic connection between the identified reading candidate and at least one of the words before and after the location corresponding to the input error may be simply referred to as the "strength of the semantic connection with the surrounding words".
[0077] The correct reading determination unit 406 determines the reading of the location corresponding to the input error based on the identified reading candidate. For example, assume that only one reading candidate is identified. In this case, the correct reading determination unit 406 may determine the identified reading candidate as the reading of the location corresponding to the input error.
[0078] Further, the correct reading determination unit 406 may determine the reading of the portion corresponding to the input error from among the specified reading candidates based on the strength of the semantic connection with the preceding and succeeding words. For example, assume that a plurality of reading candidates are specified, and for each of the specified reading candidates, the strength of the semantic connection with the preceding and succeeding words is specified.
[0079] In this case, the correct reading determination unit 406 may determine, based on the strength of the semantic connection with the specified preceding and succeeding words, that among the plurality of specified reading candidates, the reading candidate with the strongest semantic connection with the preceding and succeeding words is the reading of the portion corresponding to the input error. The reading candidate with the strongest semantic connection is, for example, the reading candidate with the largest value indicating the strength of the semantic connection or the reading candidate with the highest stage (strong, medium, weak) indicating the strength of the semantic connection.
[0080] Further, when a plurality of reading candidates are specified, the correct reading determination unit 406 may determine that among the plurality of specified reading candidates, the reading candidates with the strength of the semantic connection with the preceding and succeeding words being equal to or greater than the threshold α are the readings of the portions corresponding to the input errors. Also, the correct reading determination unit 406 may determine that among the plurality of specified reading candidates, the top N reading candidates with a strong semantic connection with the preceding and succeeding words are the readings of the portions corresponding to the input errors.
[0081] The threshold α can be arbitrarily set. For example, when the strength of the semantic connection is represented by a numerical value from 0 to 1, the threshold α may be set to a value of about 0.5. Also, when the strength of the semantic connection is represented by the stages of "strong, medium, weak", the threshold α may be set to "medium". Also, the number N is a natural number of 2 or more and can be arbitrarily set.
[0082] Also, when there are two or more reading candidates whose semantic connection strength with the preceding and succeeding words is equal to or greater than the threshold α, the two or more reading candidates are determined to be the readings of the portions corresponding to input errors. The same applies when there are the top N reading candidates with a strong semantic connection with the preceding and succeeding words. In this case, the correct reading determination unit 406 may set a priority corresponding to the strength of the semantic connection with the preceding and succeeding words for each of the reading candidates determined to be the readings of the portions corresponding to input errors.
[0083] Here, the priority indicates the degree of priority as the correct reading of the portion corresponding to the input error. The priority is set, for example, to increase as the semantic connection with the preceding and succeeding words becomes stronger.
[0084] Also, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine that all of the identified plurality of reading candidates are the readings of the portions corresponding to input errors. In this case, the correct reading determination unit 406 may set a priority corresponding to the strength of the semantic connection with the preceding and succeeding words for each of the plurality of reading candidates.
[0085] Also, the correct reading determination unit 406 may determine the reading of the portion corresponding to the input error based on the evaluation value regarding the character sequence of the identified reading candidate. Here, the evaluation value regarding the character sequence indicates the unnaturalness (or plausibility) of the character sequence. The evaluation value is calculated, for example, based on the plausibility of the transition between adjacent characters of the reading candidate.
[0086] The plausibility of the transition between adjacent characters is represented, for example, by the transition probability between adjacent characters. The transition probability between adjacent characters is derived, for example, based on the correct Roman character sequence input in the past during Japanese input. Note that information indicating the transition probability between adjacent characters is stored in advance in a storage device such as the RAM 203 or the HD 205.
[0087] In the following description, the evaluation value regarding the character sequence of the reading candidate may be referred to as an "unnaturalness score". The higher the value of the unnaturalness score, the more unnatural the character sequence is considered to be.
[0088] Note that the unnaturalness score of the reading candidate may be calculated, for example, by the candidate creation unit 403 when the reading candidate is created. Further, the unnaturalness score of the reading candidate may be calculated, for example, by the candidate identification unit 405 when the reading candidate is identified. Further, the unnaturalness score of the reading candidate may also be calculated when a plurality of reading candidates having a semantic connection with the word extracted by the candidate identification unit 405 are identified.
[0089] For example, it is assumed that a plurality of reading candidates are identified and an unnaturalness score is calculated for each of the identified reading candidates. In this case, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the reading candidate having the lowest calculated unnaturalness score among the plurality of identified reading candidates. The reading candidate with the lowest unnaturalness score is the reading candidate with the highest evaluation of the character arrangement.
[0090] Further, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the reading candidate among the plurality of identified reading candidates whose calculated unnaturalness score is equal to or less than the threshold β. The threshold β can be arbitrarily set. Further, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the top N reading candidates from the plurality of identified reading candidates whose calculated unnaturalness scores are lower.
[0091] Further, when there are two or more reading candidates whose unnaturalness scores are equal to or less than the threshold β, the two or more reading candidates are determined as the reading of the location corresponding to the input error. The same applies when there are the top N reading candidates from the ones with lower unnaturalness scores. In this case, the correct reading determination unit 406 may set a priority corresponding to the calculated unnaturalness score for each of the reading candidates determined as the reading of the location corresponding to the input error. The priority is set, for example, to be higher as the calculated unnaturalness score is lower.
[0092] In addition, when multiple reading candidates are identified, the correct reading determination unit 406 may determine that all of the identified multiple reading candidates are the readings of the portions corresponding to input errors. In this case, the correct reading determination unit 406 may set a priority according to the unnaturalness score for each of the multiple reading candidates.
[0093] In addition, the correct reading determination unit 406 may determine the reading of the portion corresponding to the input error based on the strength of the semantic connection with the words before and after the identified reading candidate and the unnaturalness score. Specifically, for example, the correct reading determination unit 406 identifies the reading candidate with the strongest semantic connection with the words before and after among the identified reading candidates based on the strength of the semantic connection with the words before and after.
[0094] At this time, there may be a case where multiple reading candidates with the strongest semantic connection with the words before and after are identified. For example, there may be a case where there are multiple reading candidates with a stage indicating the strength of the semantic connection with the words before and after being "strong". Also, there may be a case where there are multiple reading candidates with the value indicating the strength of the semantic connection with the words before and after being the maximum.
[0095] In addition, there may be a case where there are multiple reading candidates (for example, the difference from the maximum value is within 0.1) for which it can be said that the value indicating the strength of the semantic connection with the words before and after is the maximum. In this case, the correct reading determination unit 406 may determine that the reading candidate with the lowest unnaturalness score among the identified strongest reading candidates is the reading of the portion corresponding to the input error.
[0096] In addition, the correct reading determination unit 406 may extract reading candidates with a strength of semantic connection with the words before and after of a threshold α or more from among the identified multiple reading candidates. Then, the correct reading determination unit 406 may determine that the reading candidate with the lowest unnaturalness score among the extracted reading candidates is the reading of the portion corresponding to the input error.
[0097] Also, when a plurality of reading candidates are identified, the correct reading determination unit 406 may determine that all of the identified plurality of reading candidates are the readings of the portions corresponding to input errors. In this case, the correct reading determination unit 406 may set a priority for each of the plurality of reading candidates according to the strength of the semantic connection with the preceding and succeeding words and the unnaturalness score. The priority is set, for example, so that it becomes higher as the semantic connection with the preceding and succeeding words is stronger and lower as the unnaturalness score is lower.
[0098] The output unit 407 outputs conversion candidates corresponding to the input character string based on the reading of the portion corresponding to the determined input error. Here, the conversion candidates are, for example, those obtained by converting the reading corresponding to the input character string into Chinese characters or a sentence containing Chinese characters.
[0099] Specifically, for example, first, the output unit 407 creates a reading corresponding to the input character string based on the reading of the portion corresponding to the determined input error. The reading corresponding to the input character string is created, for example, by setting the reading of the portion corresponding to the input error among the input character string as the reading candidate determined for that portion.
[0100] Next, the output unit 407 creates conversion candidates corresponding to the input character string by performing kana-to-Chinese-character conversion on the created reading with reference to mapping information associating readings with Chinese characters. When there are a plurality of reading candidates determined as the readings of the portion corresponding to the input error, conversion candidates corresponding to each of the plurality of reading candidates are created.
[0101] Then, the output unit 407 outputs the created conversion candidates. More specifically, for example, the output unit 407 displays the created conversion candidates on the display 208 so that they can be selected for the input character string. At this time, the output unit 407 may display the created conversion candidates together with a message indicating that the input error has been corrected.
[0102] In addition, there may be cases where there are multiple readings for the part corresponding to the input error and the determined reading candidates. In this case, the output unit 407 may display the conversion candidates corresponding to each of the multiple reading candidates on the display 208 in descending order of the priority set for each of the multiple reading candidates. An output example of the conversion candidates corresponding to the input character string will be described later with reference to FIG. 7.
[0103] Further, the storage unit 410 may store the context (word order) between words of three or more words that are continuously input and have a semantic connection. Furthermore, the storage unit 410 may store the strength of the semantic connection between words of three or more words that are continuously input.
[0104] For example, it is assumed that the semantic connection table 300 stores the context (word order) between words of three words (previous word, middle word, and subsequent word) having a semantic connection and the strength of the semantic connection of the three words. Also, it is assumed that the words before and after the part corresponding to the input error are extracted from the input character string.
[0105] In this case, the candidate specifying unit 405 may refer to, for example, the semantic connection table 300 and specify a reading candidate (corresponding to the "middle word") that appears after the extracted previous word (corresponding to the "previous word") and before the extracted subsequent word (corresponding to the "subsequent word") among one or more reading candidates and has a semantic connection with the extracted previous word and subsequent word.
[0106] At this time, the candidate specifying unit 405 may refer to the semantic connection table 300 and specify the strength of the semantic connection between the three words of the specified reading candidate, the extracted previous word, and the subsequent word. In this case, the correct reading determination unit 406 may determine the reading of the part corresponding to the input error from among the specified reading candidates based on the strength of the semantic connection between the three specified words.
[0107] Also, assume that two words after the location corresponding to the input error are extracted from the input string. In this case, for example, the candidate identification unit 405 may refer to the semantic connection table 300 to identify, among one or more reading candidates, a reading candidate (corresponding to the "preceding word") that appears before the two extracted words (corresponding to the "middle word" and the "following word") and has a semantic connection with the two extracted words.
[0108] Also, among the reading candidates identified by the candidate identification unit 405, there may be a reading candidate included in two words with a semantic connection and a reading candidate included in three or more words with a semantic connection. In this case, the correct reading determination unit 406 may determine the reading candidate included in three or more words as the reading at the location corresponding to the input error (prioritizing the connection of longer words).
[0109] For example, assume that the string "kんせいきねんかんない(knseikinenkannai)" is input. In this case, the omission of the vowel after k is detected as an input error, and by performing a correction process to supplement the vowel (a, i, u, e, o) for the input error, reading candidates for the location corresponding to the input error are created. The location corresponding to the input error is "kんせい(knsei)". The created reading candidates are "かんせい(kansei), きんせい(kinsei), くんせい(kunsei), けんせい(kensei), こんせい(konsei)".
[0110] Here, assume that the word "きねん(kinen)" immediately after the location corresponding to the input error is extracted. In this case, for example, the candidate identification unit 405 may refer to the semantic connection table 300 to identify, from among the created reading candidates, a reading candidate (corresponding to the "preceding word") that appears before the extracted word (corresponding to the "following word") and has a semantic connection with the extracted word.
[0111] As an example, assume that a semantic connection between the previous word "かんせい (kansei, completion)" and the subsequent word "きねん (kinen, commemoration)" is registered in the semantic connection table 300. In this case, among the generated reading candidates, a reading candidate "かんせい" that appears before the extracted word "きねん" and has a semantic connection with the extracted word "きねん" is identified.
[0112] Also, assume that the two subsequent words "きねん (kinen)" and "かんない (kannai)" after the location corresponding to the input error are extracted. In this case, the candidate identification unit 405 refers to, for example, the semantic connection table 300 and identifies, among the generated reading candidates, a reading candidate (corresponding to the "previous word") that appears before the two extracted words (corresponding to the "middle word" and the "subsequent word") and has a semantic connection with the two extracted words.
[0113] As an example, assume that a semantic connection between the previous word "けんせい (kensei, constitutional government)", the middle word "きねん (kinen, commemoration)", and the subsequent word "かんない (kannai, inside the museum)" is registered in the semantic connection table 300. In this case, among the generated reading candidates, a reading candidate "けんせい" that appears before the two extracted words "きねん" and "かんない" and has a semantic connection with the two extracted words "きねん" and "かんない" is identified.
[0114] As a result, the reading candidate "かんせい" included in the two words with a semantic connection (かんせい / きねん) and the reading candidate "けんせい" included in the three words with a semantic connection (けんせい / きねん / かんない) are identified. In this case, the correct reading determination unit 406 may determine the reading candidate "けんせい" included in the three words with a semantic connection (けんせい / きねん / かんない) as the reading at the location corresponding to the input error.
[0115] As a result, the correct reading determination unit 406 can prioritize a reading candidate with a longer (more words) connection as the semantic connection between the preceding and following words. For example, in the input "kんせいきねんかんない", it is determined that the semantic connection consisting of three words "けんせい(Constitution) / きねん(commemoration) / かんない(inside the museum)" is stronger than the semantic connection consisting of two words "かんせい(completion) / きねん(commemoration)", and it is repaired to "けんせいきねんかんない".
[0116] Note that there may be cases where there are two or more reading candidates among the reading candidates specified by the candidate specification unit 405 that are included in three or more words with a semantic connection. In this case, the correct reading determination unit 406 may determine, as the reading of the location corresponding to the input error, the reading candidate with the strongest semantic connection between the words of three or more words among the two or more reading candidates.
[0117] Note that the functional units (reception unit 401 to output unit 407) of the input support device 100 may be realized by, for example, a server that can be connected via a network NW (see FIG. 2) from a PC or the like used by the user. In this case, the input support device 100 receives an input of a character string from the PC used by the user by the reception unit 401, and outputs the created conversion candidate to the PC used by the user by the output unit 407.
[0118] Also, the functional units (reception unit 401 to output unit 407) of the input support device 100 may be realized by a plurality of computers (for example, a server and a PC). For example, the reception unit 401 to the correct reading determination unit 406 may be realized by a server, and the output unit 407 may be realized by a PC. In this case, the communication between the functional units of different computers is performed, for example, by transmission and reception between the functional units via the network NW. As a result, on the server side, the reading of the location corresponding to the input error can be determined, and on the PC side, based on the reading estimated by the server, a conversion candidate corresponding to the input character string can be created and output.
[0119] (Example of determination of reading of location corresponding to input error) Next, an example of determining the reading of a part corresponding to an input error detected from an input character string will be described with reference to Figures 5 and 6. The process described with reference to Figures 5 and 6 is executed, for example, in a kana-kanji conversion process.
[0120] Fig. 5 is an explanatory diagram showing a first judgment example of the reading of a portion corresponding to an input error. In Fig. 5, assume that a user is trying to input the word "completed presentation" and inputs the character string 501 "knseihirou". "knseihirou" is the inputted romanized spelling. "knseihirou" is "knseihirou" converted into kana.
[0121] In the input character string 501, a vowel after k is missing. The input error detection unit 402 detects the missing vowel after k as an input error from the input character string 501. In this case, the candidate creation unit 403 creates reading candidates 500-1 to 500-5 for the part 502 corresponding to the input error by performing a correction process to add vowels (a, i, u, e, o) to the detected input error. Each of the reading candidates 500-1 to 500-5 is converted into kana by adding the vowels "a, i, u, e, o" after k.
[0122] The extraction unit 404 extracts at least one of the words before or after the portion 502 corresponding to the input error from the input character string 501. Here, the input character string 501 is divided into "knsei" and "hirou", and the word "hirou" 503 immediately after the portion 502 corresponding to the input error is extracted.
[0123] The candidate specification unit 405 refers to the semantic connection table 300 and specifies, from among the reading candidates 500-1 to 500-5, a reading candidate that has a semantic connection with the extracted word 503. Here, the word 503 matches the subsequent word "hirou" of the semantic connection information 300-1 to 300-3 in the semantic connection table 300.
[0124] Therefore, the candidate identification unit 405 refers to the semantic connection table 300 and identifies the reading candidates 500-1, 500-2, and 500-5 that appear before the extracted word 503 and have a semantic connection with the extracted word 503 from among the reading candidates 500-1 to 500-5. Further, the candidate identification unit 405 refers to the semantic connection table 300 and identifies the strength of the semantic connection between the identified reading candidates 500-1, 500-2, and 500-5 and the extracted word 503, respectively.
[0125] Here, the strength of the semantic connection between the reading candidate 500-1 and the word 503 is "0.87" from the semantic connection information 300-1. Also, the strength of the semantic connection between the reading candidate 500-2 and the word 503 is "0.10" from the semantic connection information 300-2. Further, the strength of the semantic connection between the reading candidate 500-5 and the word 503 is "0.03" from the semantic connection information 300-3.
[0126] The correct reading determination unit 406 determines the reading of the location 502 corresponding to the input error based on the identified reading candidates 500-1, 500-2, and 500-5. Specifically, for example, the correct reading determination unit 406 determines the reading candidate with the strongest semantic connection with the extracted word 503 among the reading candidates 500-1, 500-2, and 500-5 as the reading of the location 502 corresponding to the input error.
[0127] Here, among the reading candidates 500-1, 500-2, and 500-5, the reading candidate 500-1 has the strongest semantic connection with the word 503. Therefore, the reading candidate 500-1 is determined as the reading (correct reading) of the location 502 corresponding to the input error. As a result, the reading corresponding to the input character string 501 becomes "かんせいひろう". Among "かんせいひろう", "かんせい" corresponds to the correction result 504 of the location 502 corresponding to the input error. Among "かんせいひろう", "ひろう" corresponds to the word 503.
[0128] In this way, even when multiple reading candidates are assumed for the portion 502 corresponding to the input error, the input support device 100 can estimate the correct reading in consideration of the strength of the semantic connection with the word 503 (hirou) after the portion 502 corresponding to the input error. Thereby, the input support device 100 can improve the estimation accuracy of the reading of the portion 502 corresponding to the input error, and can output the conversion candidate "completed disclosure" desired by the user.
[0129] FIG. 6 is an explanatory diagram showing a second determination example of the reading of the portion corresponding to the input error. In FIG. 6, it is assumed that the user attempts to input the phrase "Venus exploration" and inputs the character string 601 "kんせいたんさ (knseitansa)". "knseitansa" is the input Romanized spelling. "kんせいたんさ" is the result of converting "knseitansa" into kana.
[0130] In the input character string 601, the vowel after k is missing. The input error detection unit 402 detects the missing vowel after k as an input error from the input character string 601. In this case, the candidate creation unit 403 creates reading candidates 600-1 to 600-5 for the portion 602 corresponding to the input error by performing a correction process of supplementing the detected input error with a vowel (a, i, u, e, o). Each of the reading candidates 600-1 to 600-5 is the result of supplementing the vowel "a, i, u, e, o" after k and converting it into kana.
[0131] The extraction unit 404 extracts at least one of the words before and after the portion 602 corresponding to the input error from the input character string 601. Here, the input character string 601 is divided into "kんせい (knsei)" and "たんさ (tansa)", and the word 603 "たんさ (tansa)" immediately after the portion 602 corresponding to the input error is extracted.
[0132] The candidate specifying unit 405 refers to the semantic connection table 300 and specifies, from among the reading candidates 600-1 to 600-5, the reading candidates that have a semantic connection with the extracted word 603. Here, the word 603 matches the post-word "tan sa" of the semantic connection information 300-4 to 300-6 in the semantic connection table 300.
[0133] Therefore, the candidate specifying unit 405 refers to the semantic connection table 300 and specifies, from among the reading candidates 600-1 to 600-5, the reading candidates 600-1, 600-2, and 600-5 that appear before the extracted word 603 and have a semantic connection with the extracted word 603. Further, the candidate specifying unit 405 refers to the semantic connection table 300 and respectively specifies the strength of the semantic connection between the specified reading candidates 600-1, 600-2, and 600-5 and the extracted word 603.
[0134] Here, the strength of the semantic connection between the reading candidate 600-1 and the word 603 is "0.25" from the semantic connection information 300-4. Also, the strength of the semantic connection between the reading candidate 600-2 and the word 603 is "0.77" from the semantic connection information 300-5. Further, the strength of the semantic connection between the reading candidate 600-5 and the word 603 is "0.08" from the semantic connection information 300-6.
[0135] The correct reading determination unit 406 determines the reading of the location 602 corresponding to the input error based on the specified reading candidates 600-1, 600-2, and 600-5. Specifically, for example, the correct reading determination unit 406 determines, as the reading of the location 602 corresponding to the input error, the reading candidate that has the strongest semantic connection with the extracted word 603 among the reading candidates 600-1, 600-2, and 600-5.
[0136] Here, among the reading candidates 600-1, 600-2, and 600-5, the reading candidate 600-2 has the strongest semantic connection with the word 603. Therefore, the reading candidate 600-2 is determined to be the reading (correct reading) of the location 602 corresponding to the input error. As a result, the reading corresponding to the input character string 601 is "kinseitansa". Among "kinseitansa", "kinsei" corresponds to the correction result 604 of the location 602 corresponding to the input error. Among "kinseitansa", "tansa" corresponds to the word 603.
[0137] In this way, even when multiple reading candidates for the location 602 corresponding to the input error are assumed, the input support device 100 can estimate the correct reading by considering the strength of the semantic connection with the word 603 (tansa) after the location 602 corresponding to the input error. Thereby, the input support device 100 can improve the estimation accuracy of the reading of the location 602 corresponding to the input error and can output the conversion candidate "Venus exploration" desired by the user.
[0138] (Output example of conversion candidate) Next, with reference to FIG. 7, an output example of the conversion candidate corresponding to the input character string will be described. Here, taking the input character string 501 shown in FIG. 5 as an example, the output example of the conversion candidate will be described.
[0139] FIG. 7 is an explanatory diagram showing an output example of the conversion candidate. In (7-1) of FIG. 7, the character string 501 is input into the document 700 by the user's operation. The user's operation is performed using, for example, the keyboard 209 or the mouse 210 shown in FIG. 2. The document 700 is displayed on, for example, the display 208 shown in FIG. 2.
[0140] In the input character string 501, an input error occurs where the vowel after k is missing. Therefore, as described in FIG. 5, in the kana-kanji conversion process, the reading of the location 502 corresponding to the input error (see FIG. 5) is estimated, and the reading "kansei hiro" corresponding to the input character string 501 is created.
[0141] Note that the kana-kanji conversion process is executed, for example, in response to a conversion instruction from the user. The conversion instruction is given, for example, by pressing a specific key on the keyboard 209.
[0142] In FIG. 7(7-2), based on the created reading "kansai hirou", the conversion candidate 701 corresponding to the input character string 501 is displayed so as to be selectable. The conversion candidate 701 is "complete disclosure". Also, among the conversion candidates 701, the portion 702 (completed) where correction processing has been performed for input errors is displayed in a different manner from the other portions.
[0143] Here, the background color of the portion 702 is different from that of the other portions. As other modes, for example, the portion 702 may be made bold or underlined. Also, a message 703 is displayed in association with the conversion candidate 701. The message 703 is a message indicating that the input errors included in the input character string 501 have been repaired.
[0144] As a result, even when the user makes an input error during Japanese input, the user can obtain the desired conversion candidate 701 without re-entering. Also, the user can grasp that the input errors have been automatically repaired by referring to the message 703. Also, the user can grasp that the portion 702, whose background color is different from that of the other portions among the conversion candidates 701, is the repaired portion.
[0145] Note that after the conversion candidate 701 is displayed, for example, when a confirmation instruction (selection operation) is performed by the user's operation, the conversion candidate 701 is selected and the input is confirmed.
[0146] (Input support processing procedure of the input support device 100) Next, with reference to FIG. 8, the input support processing procedure of the input support device 100 will be described.
[0147] FIG. 8 is a flowchart showing an example of the input support processing procedure of the input support device 100 according to the embodiment. In the flowchart of FIG. 8, first, the input support device 100 determines whether an input error is detected from the input character string (step S801). Here, the input support device 100 waits for an input error to be detected (step S801: No).
[0148] Then, when an input error is detected (step S801: Yes), the input support device 100 creates a reading candidate for the portion corresponding to the input error in the input character string by performing a predetermined correction process on the input error (step S802). At this time, the input support device 100 may calculate an unnaturalness score for the created reading candidate based on the plausibility of the transition between adjacent characters.
[0149] Next, the input support device 100 determines whether a plurality of reading candidates have been created (step S803). Here, when a plurality of reading candidates have been created (step S803: Yes), the input support device 100 executes a correct reading determination process (step S804) and proceeds to step S806.
[0150] The correct reading determination process is a process of determining the reading of the portion corresponding to the input error from among a plurality of reading candidates. The specific processing procedure of the correct reading determination process will be described later with reference to FIG. 9.
[0151] Also, in step S803, when one reading candidate is created (step S803: No), the input support device 100 determines the created reading candidate as the reading of the portion corresponding to the input error (step S805).
[0152] Next, the input support device 100 creates a conversion candidate corresponding to the input character string based on the reading of the portion corresponding to the determined input error (step S806). Then, the input support device 100 outputs the created conversion candidate for the input character string (step S807) and ends a series of processes according to this flowchart.
[0153] As a result, when an input error is detected from the input string, the input support device 100 can estimate the correct reading of the portion corresponding to the input error and output a conversion candidate corresponding to the input string.
[0154] Next, with reference to FIG. 9, the specific processing procedure of the correct reading determination process in step S804 will be described.
[0155] FIG. 9 is a flowchart showing an example of the specific processing procedure of the correct reading determination process. In the flowchart of FIG. 9, first, the input support device 100 extracts at least one of the words before and after the portion corresponding to the input error from the input string (step S901).
[0156] Then, the input support device 100 determines whether the word before the portion corresponding to the input error has been extracted (step S902). Here, when the previous word has not been extracted (step S902: No), the input support device 100 proceeds to step S905.
[0157] On the other hand, when the previous word has been extracted (step S902: Yes), the input support device 100 refers to the semantic connection table 300 and identifies, from among the plurality of created reading candidates, the reading candidate that appears after the extracted word and has a semantic connection with the extracted word (step S903). Then, the input support device 100 refers to the semantic connection table 300 and identifies the strength of the semantic connection between the identified reading candidate and the extracted word (step S904).
[0158] Next, the input support device 100 determines whether the word after the portion corresponding to the input error has been extracted (step S905). Here, when the subsequent word has not been extracted (step S905: No), the input support device 100 proceeds to step S908.
[0159] On the other hand, when a subsequent word is extracted (step S905: Yes), the input support device 100 refers to the semantic connection table 300 and identifies, from among the plurality of created reading candidates, a reading candidate that appears before the extracted word and has a semantic connection with the extracted word (step S906). Then, the input support device 100 refers to the semantic connection table 300 and identifies the strength of the semantic connection between the identified reading candidate and the extracted word (step S907).
[0160] Next, the input support device 100 determines whether or not a plurality of reading candidates having a semantic connection with the extracted word have been identified (step S908).
[0161] Here, when a plurality of reading candidates are identified (step S908: Yes), the input support device 100 identifies, from among the plurality of identified reading candidates, the reading candidate having the strongest semantic connection with the extracted word (step S909). Then, the input support device 100 determines the reading candidate identified in step S909 as the reading at the location corresponding to the input error (step S910), and returns to the step where the correct reading determination process was called.
[0162] Also, in step S908, when one reading candidate is identified (step S908: No), the input support device 100 determines the identified reading candidate as the reading at the location corresponding to the input error (step S911), and returns to the step where the correct reading determination process was called.
[0163] Thereby, even when a plurality of reading candidates corresponding to the input error are assumed, the input support device 100 can estimate the correct reading in consideration of the strength of the semantic connection with the words before and after the location corresponding to the input error.
[0164] Note that in step S901, when neither of the words before and after the location corresponding to the input error is extracted, the input support device 100 may determine, for example, the reading candidate having the lowest unnaturalness score among the plurality of created reading candidates as the reading at the location corresponding to the input error.
[0165] Also, in step S909, when a plurality of the strongest reading candidates are identified, the input support device 100 may determine, for example, the reading candidate with the lowest unnaturalness score among the identified strongest reading candidates as the reading of the location corresponding to the input error.
[0166] Also, in step S901, when the words before and after the location corresponding to the input error are extracted, the input support device 100 may identify a reading candidate that appears after the extracted previous word and before the extracted next word and has a semantic connection with the extracted previous word and the next word. Then, the input support device 100 may determine the identified reading candidate as the reading of the location corresponding to the input error. Thereby, when there is a reading candidate having a semantic connection with the previous and next words, the input support device 100 can preferentially determine the reading candidate as the reading of the location corresponding to the input error.
[0167] As described above, according to the input support device 100 according to the embodiment, when an input error is detected from the input character string, by performing a predetermined correction process on the input error, one or more reading candidates can be created for the location corresponding to the input error in the input character string. Also, according to the input support device 100, by referring to the storage unit 410, among the one or more created reading candidates, a reading candidate having a semantic connection with at least one of the words before and after the location corresponding to the input error in the input character string can be identified. The storage unit 410 stores the sequential relationship between semantically connected words that are continuously input. Then, according to the input support device 100, based on the identified reading candidate, the reading of the location corresponding to the input error can be determined.
[0168] As a result, when estimating the reading of the part corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the semantic connection with the words before and after that part. For example, even when multiple reading candidates are assumed for the part corresponding to the input error, the input support device 100 can determine a reading candidate with a semantic connection to the words before and after as the reading of the part corresponding to the input error.
[0169] Further, according to the input support device 100, by referring to the storage unit 410 in which the strength of the semantic connection between words is stored, it is possible to specify the strength of the semantic connection (the strength of the semantic connection with the words before and after) between the specified reading candidate and at least one of the words before and after the part corresponding to the input error. Then, according to the input support device 100, it is possible to determine the reading of the part corresponding to the input error based on the strength of the semantic connection with the specified words before and after among the specified reading candidates.
[0170] As a result, when estimating the reading of the part corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the strength of the semantic connection with the words before and after that part. For example, even when multiple reading candidates are assumed for the part corresponding to the input error, the input support device 100 can estimate a reading candidate with a stronger semantic connection to the words before and after as the reading of the part corresponding to the input error.
[0171] Further, according to the input support device 100, based on the strength of the semantic connection with the specified words before and after, it is possible to determine the reading candidate with the strongest semantic connection to the words before and after among the specified reading candidates as the reading of the part corresponding to the input error.
[0172] As a result, the input support device 100 can estimate the reading candidate with the strongest semantic connection to the words before and after as the reading of the part corresponding to the input error.
[0173] In addition, according to the input support device 100, it is possible to determine the reading of the location corresponding to the input error from among the specified reading candidates based on the strength of the semantic connection with the specified words before and after, and the unnaturalness score.
[0174] Thereby, when estimating the reading of the location corresponding to the input error, the input support device 100 can improve the estimation accuracy of the correct reading by considering the strength of the semantic connection with the words before and after that location and the likelihood of the character sequence.
[0175] In addition, according to the input support device 100, based on the strength of the semantic connection with the specified words before and after, among the specified reading candidates, the reading candidate with the strongest semantic connection with the words before and after is specified. When multiple reading candidates with the strongest connection are specified, the reading candidate with the lowest unnaturalness score among the specified reading candidates with the strongest connection can be determined as the reading of the location corresponding to the input error.
[0176] Thereby, even when there are multiple reading candidates that can be said to have the strongest semantic connection with the words before and after, the input support device 100 can determine the reading candidate with a more likely character sequence as the reading of the location corresponding to the input error.
[0177] In addition, according to the input support device 100, among the specified reading candidates, the reading candidates with a strength of semantic connection with the words before and after that is equal to or greater than the threshold α can be used to determine the reading of the location corresponding to the input error. Also, according to the input support device 100, among the specified reading candidates, the top N reading candidates with a strong semantic connection with the words before and after may be determined as the reading of the location corresponding to the input error.
[0178] As a result, the input support device 100 can estimate, as the reading of the part corresponding to the input error, a reading candidate whose semantic connection with the preceding and succeeding words is stronger than a predetermined standard, or a reading candidate whose semantic connection with the preceding and succeeding words is relatively strong. For this reason, the input support device 100 can prevent, for example, a reading candidate with a weak semantic connection with the preceding and succeeding words from being determined as the correct reading. This determination method is useful, for example, when the stored content of the semantic connection table 300 (for example, the strength of the connection between words) is updated at any time.
[0179] Further, according to the input support device 100, when among the specified reading candidates, there are reading candidates included in two words with a semantic connection and reading candidates included in three or more words with a semantic connection, the reading candidates included in three or more words can be determined as the reading of the part corresponding to the input error.
[0180] As a result, the input support device 100 can preferentially determine, as the reading of the part corresponding to the input error, a reading candidate with a longer semantic connection as the semantic connection between the preceding and succeeding words.
[0181] Further, according to the input support device 100, based on the reading of the part corresponding to the determined input error, conversion candidates corresponding to the input character string can be output.
[0182] As a result, the input support device 100 can suppress the output of conversion candidates that the user does not want.
[0183] From these facts, according to the input support program, the input support method, and the input support device according to the embodiment, even when the user makes an input error, the correct reading of the part for the input error can be accurately estimated, and the conversion accuracy of Japanese input can be improved. For example, when creating a document using a document creation application, the labor of the user re-entering can be reduced, the work load related to Japanese input can be reduced, and the efficiency of document creation can be improved.
[0184] Incidentally, the input support method described in this embodiment can be realized by executing a pre-prepared program on a computer such as a personal computer or a workstation. This input support program is recorded on a computer-readable recording medium such as a hard disk, a flexible disk, a CD-ROM, a DVD, or a USB memory, and is executed by being read from the recording medium by the computer. Further, this input support program may be distributed via a network such as the Internet.
Industrial Applicability
[0185] The input support program, the input support method, and the input support device according to this invention are useful for a computer system that supports character input performed based on a user's operation, and in particular, are suitable for a computer system that inputs Japanese based on a user's operation.
Explanation of Signs
[0186] 100 Input support device 101, 501, 601 Character string 102, 502, 602 Location 103, 503, 603 Word 104, 504, 604 Correction result 110, 410 Storage unit 200 Bus 201 CPU 202 ROM 203 RAM 204 HDD 205 HD 206 CD-RW drive 207 CD-RW 208 Display 209 Keyboard 210 Mouse 211 Network I / F 300 Semantic connection table 401 Reception unit 402 Input error detection unit 403 Candidate creation unit 404 Extraction Unit 405 Candidate Identification Unit 406 Correct Reading Judgment Unit 407 Output Unit 700 Document 701 Conversion Candidate 702 Portion 703 Message NW Network
Claims
1. When an input error is detected from the input string, by performing a predetermined correction process on the input error, one or more reading candidates are created for the portion corresponding to the input error in the string, referring to a storage unit that stores the context between words with a semantic connection that are continuously input, among the one or more created reading candidates, identifying a reading candidate that has a semantic connection with at least one of the words before and after the portion in the string, determining the reading of the portion based on the identified reading candidate, An input assistance program characterized by causing a computer to execute the process.
2. The storage unit further stores the strength of the semantic connection between the words, causing the computer to execute a process of referring to the storage unit and identifying the strength of the semantic connection between the identified reading candidate and the one word, The process of determining is as follows: Among the identified reading candidates, determining the reading of the portion based on the identified strength of the connection. The input assistance program according to claim 1, characterized in that.
3. The process of determining is as follows: Based on the identified strength of the connection, among the identified reading candidates, determining the reading candidate with the strongest semantic connection with the one word as the reading of the portion. The input assistance program according to claim 2, characterized in that.
4. The process of determining is as follows: Among the identified reading candidates, determining the reading of the portion based on the identified strength of the connection and an evaluation value regarding the arrangement of characters based on the plausibility of the transition between adjacent characters in the identified reading candidate. The input assistance program according to claim 2, characterized in that.
5. The process of determining is as follows: Based on the identified strength of the connection, among the identified reading candidates, identifying the reading candidate with the strongest semantic connection with the one word, When a plurality of the strongest reading candidates are identified, among the identified strongest reading candidates, determining the reading candidate with the highest evaluation of the arrangement of characters based on the evaluation value as the reading of the portion. The input assistance program according to claim 4, characterized in that.
6. The storage unit further stores the context between words of three or more words with a semantic connection that are continuously input, The process of determining is as follows: If, among the specified reading candidates, there are reading candidates included in two words with a semantic connection and reading candidates included in three or more words with a semantic connection, the input support program according to claim 1, wherein the reading candidate included in the three or more words is determined as the reading of the location.
7. Output a conversion candidate corresponding to the character string based on the determined reading of the location. The input support program according to any one of claims 1 to 6, characterized in that the computer is caused to execute the process.
8. When an input error is detected from the input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, create one or more reading candidates. With reference to a storage unit that stores the context between continuously input words with a semantic connection, among the one or more created reading candidates, identify a reading candidate that has a semantic connection with at least one of the words before and after the location in the character string. Based on the identified reading candidate, determine the reading of the location. An input support method, characterized in that the computer executes the process.
9. When an input error is detected from the input character string, by performing a predetermined correction process on the input error, for the location corresponding to the input error in the character string, create one or more reading candidates. With reference to a storage unit that stores the context between continuously input words with a semantic connection, among the one or more created reading candidates, identify a reading candidate that has a semantic connection with at least one of the words before and after the location in the character string. Based on the identified reading candidate, determine the reading of the location. An input support device, characterized by having a control unit that executes the process.
Citation Information
Patent Citations
Input assist device, input assist system, and program
JP2016033780A
Cited By
Transformer having a tertiary winding
US12456573B2