Text input method and text input device

The sentence input method improves efficiency by sequentially presenting words in a specific order, reducing eye movement and optimizing word presentation, thus enhancing the speed and accuracy of sentence input using brain wave signals.

JP7675526B2Active Publication Date: 2025-05-13NISSAN MOTOR CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021018978
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-09
Publication Date
2025-05-13
Estimated Expiration
2041-02-09

AI Technical Summary

Technical Problem

Existing text input methods using event-related potentials in brain wave signals are inefficient for inputting sentences, as they require users to recognize and respond to multiple word options displayed simultaneously, leading to increased response time.

Method used

A sentence input method that estimates intended sentences by sequentially presenting constituent words in a specific order, allowing users to respond to each word individually, thereby reducing the need for eye movement between options and optimizing the presentation of words with high similarity.

Benefits of technology

This approach significantly reduces the time required for users to input sentences by minimizing eye movement and optimizing the presentation of words, enhancing the accuracy and speed of sentence input using brain wave signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675526000001
    Figure 0007675526000001
  • Figure 0007675526000002
    Figure 0007675526000002
  • Figure 0007675526000003
    Figure 0007675526000003
Patent Text Reader

Abstract

To shorten time required for inputting words constituting a text intended by a user using an event-related potential contained in a user electroencephalogram signal.SOLUTION: In a text input method, a selected text proposing unit 101 estimates a text that a user intends to input and proposes a plurality of text options, a word presentation control unit 102 classifies constituent words that constitute each of the plurality of text options in the order for each text option and sequentially presents the plurality of constituent words classified in the same order one by one to the user on a presentation unit 120, a desired word recognition signal analysis unit 103 and a signal determination unit 106 analyze the event-related potentials starting from the presentation of constituent word for each constituent word sequentially presented one by one, wherein the event-related potentials are contained in electroencephalogram signals measured from the user, and a desired word determining unit 104 repeatedly determines the constituent words that the user intends to input in the same order to determine the text that the user intends to input based on the analysis result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a text input method and a text input device. [Background technology]

[0002] A technology has been proposed that estimates a user's intention by utilizing an event-related potential included in the user's electroencephalogram signal (Patent Document 1). In this technology, a menu screen displays multiple options for a question, and each option is sequentially blinked on and off. The event-related potential corresponding to each option, which is included in the electroencephalogram signal after a certain time has passed since each blink, is used to estimate the option that the user intended to select. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 4856791 Summary of the Invention [Problem to be solved by the invention]

[0004] When the technology of Patent Document 1 is used to input a sentence intended by a user, the user can recognize which word in a choice is blinking on a screen displaying all the choices of words that make up the sentence, and then react to the recognized word. The user's reaction to each word in a choice appears in the blinking event-related potential included in the electroencephalogram signal after a certain time has passed since the recognition of the blinking word.

[0005] The present invention has been made in consideration of the above circumstances, and an object of the present invention is to reduce the time required for inputting words that constitute a sentence intended by a user, by utilizing event-related potentials contained in the user's electroencephalogram signal. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, a sentence input method according to one embodiment of the present invention estimates a sentence that a user intends to input and proposes multiple sentence options. The constituent words constituting each of the multiple sentence options are classified into the order in each sentence option, and the multiple constituent words classified into the same order are sequentially presented to the user one by one. An event-related potential starting from the presentation of the constituent words, which is included in an electroencephalogram signal measured from the user, is analyzed for each of the sequentially presented constituent words one by one, and the constituent words in the same order that the user intends to input are determined. The sequential presentation of the multiple constituent words and the determination of the constituent words that the user intends to input are repeated for all orders, and the sentence that the user intends to input is determined. Effect of the Invention

[0007] According to the present invention, it is possible to reduce the time required for inputting words that constitute a sentence intended by the user by utilizing an event-related potential contained in the user's electroencephalogram signal. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram of a text input device according to an embodiment of the present invention. [Diagram 2] FIG. 2 is an explanatory diagram for comparing the relationship between the presentation pattern of constituent words and the user's reaction between the text input device of FIG. 1 and the conventional device. [Figure 3A] FIG. 3A is an explanatory diagram of an example of a sentence candidate proposed by the selected sentence proposing unit. [Figure 3B] FIG. 3B is an explanatory diagram of an example of a group of word options generated by the word presentation control unit. [Figure 4] FIG. 4 is an explanatory diagram of an example of a presentation pattern of constituent words by the word presentation control unit. [Diagram 5] FIG. 5 is a flowchart of an example of a procedure for determining the presentation order of constituent words. [Figure 6A] FIG. 6A is a graph showing an example of electroencephalographic potential changes in response to words presented. [Figure 6B]FIG. 6B is a graph showing an example of electroencephalographic potential changes in response to words presented. [Figure 7A] FIG. 7A is a graph of an electroencephalogram analyzed by the desired word recognition signal analysis unit. [Figure 7B] FIG. 7B is a graph of an electroencephalogram analyzed by the desired word recognition signal analysis unit. [Figure 8] FIG. 8 is an explanatory diagram showing the analysis result of the desired word recognition signal analysis unit in FIG. [Figure 9A] FIG. 9A is a graph showing the relationship between the presentation interval of constituent words and the accuracy rate of judgment. [Figure 9B] FIG. 9B is a graph showing an example of the relationship between the number of sequential presentations of constituent words, the number of repetitions, and the accuracy rate of judgments. [Figure 10A] FIG. 10A is a graph showing another example of the relationship between the presentation interval of constituent words and the correct answer rate of the judgment. [Figure 10B] FIG. 10B is a graph showing an example of the relationship between the number of sequential presentations of constituent words, the number of repetitions, and the accuracy rate of judgments. [Figure 11] FIG. 11 is a flowchart of an example of a procedure for the selection process. [Figure 12] FIG. 12 is a flowchart of an example of a learning process procedure. [Figure 13A] FIG. 13A is an explanatory diagram showing a classification boundary between targets and non-targets using a retrained classifier. [Figure 13B] FIG. 13B is an explanatory diagram showing the analysis result of the desired word recognition signal analysis unit in which both a determination error due to misclassification by the classifier and a determination error due to misrecognition by the user exist. [Figure 14A] FIG. 14A is an explanatory diagram showing a boundary line for distinguishing between targets and non-targets by a classifier when event-related potential data erroneously recognized by a user is excluded from data used for re-learning of the classifier. [Figure 14B] FIG. 14B is an explanatory diagram showing the analysis result of the desired word recognition signal analysis unit in which only a determination error due to misclassification by the classifier exists. [Figure 15]FIG. 15 is an explanatory diagram showing an example of a schedule for optimizing the control parameters of the word presentation control unit and re-learning the classifier of the signal determination unit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same parts are given the same reference numerals and the description will be omitted.

[0010] (Embodiment) The configuration of a text input device 100 will be described with reference to Fig. 1. As shown in Fig. 1, the text input device 100 according to the present embodiment executes a text input method according to an embodiment of the present invention.

[0011] The sentence input device 100 presents options for constituent words of a sentence that the user intends to input, and uses an event-related potential contained in the user's electroencephalogram signal to determine which constituent word the user intends to input from among the presented options.

[0012] For example, when the technology of Patent Document 1 described in the Prior Art section is used to present candidate constituent words, the user visually watches the constituent words of each option blinking in sequence on a screen on which all options of constituent words of a sentence are presented all at once. The user's reaction to the constituent words of each option that he / she has visually seen appears as an event-related potential in the electroencephalogram signal after the "eye movement time" and "presentation time (blinking time)" have elapsed from the start of the display of all options on the screen, as shown by the bar graph of all at once presentation in FIG. 2. The "eye movement time" is the time it takes for the user to move the viewpoint on the screen to find the blinking part. The "presentation time (blinking time)" is the time it takes for the user's reaction to the constituent words of the blinking option to appear in the electroencephalogram signal.

[0013] In the sentence input device 100 of this embodiment, options for constituent words of a sentence are presented to the user one by one in sequence. By presenting the options in sequence, as shown by the bar graph of sequential presentation in FIG. 2, the user recognizes the constituent words presented one by one without moving the viewpoint within the screen of the presentation unit 120 to search for them. Since there is no need to search for the presented constituent words within the screen of the presentation unit 120, in the sentence input device 100 of this embodiment, there is no need to spend "time moving the eyes" until the user's reaction to the constituent words of the presented options appears as an event-related potential in the electroencephalogram signal.

[0014] The text input device 100 of this embodiment includes a controller 110, a presentation unit 120, and a biological signal detection unit .

[0015] The text input device 100 can be used as a man-machine interface for a device to which a user needs to input text. Examples of the device to which a user needs to input text include a personal computer, a tablet terminal, and a smartphone. When the text input device 100 is mounted in a vehicle, an example of the device to which a user needs to input text includes an in-vehicle multimedia player. These devices may be devices that incorporate the text input device 100, or may be external devices connected to the text input device 100.

[0016] In the following embodiment and its modified examples, the text input device 100 is described as being mounted on a vehicle, but it goes without saying that the text input device 100 does not have to be mounted on a vehicle. The user of the text input device 100 mounted on a vehicle is a vehicle occupant. The vehicle occupant may be either the driver or a passenger.

[0017] The presentation unit 120 sequentially displays options for words that make up a sentence one by one. The contents of the words displayed by the presentation unit 120 are determined by executing a process described below by the controller 110. The biological signal detection unit 130 detects the user's brain waves measured by the brain wave sensor unit 200 worn by the user as an brain wave signal.

[0018] The user generates brain waves from the head. The brain waves are generated from specific brain activation areas of the user's head. The brain activation areas change according to the user's mental state, emotions, and behaviors such as vision, smell, taste, and motor control. For example, when the user responds to a change in visual information, the occipital lobe of the brain, which is related to visual information, becomes active. When the visual information is a word or sentence containing letters, the parietal lobe of the brain, which is related to language comprehension, becomes active.

[0019] In this embodiment, the EEG sensor unit 200 can measure EEG in the parietal lobe and occipital lobe of the user's head. For example, when using the 10-20 electrode placement method (10-20 method), which is the international standard, the EEG sensor unit 200 can have at least an electrode placed at Cz (central midline). When using the 10-10 electrode placement method (10% electrode placement method), which is suitable for recording high-resolution EEG, the EEG sensor unit 200 can have at least an electrode placed at Cpz.

[0020] In this embodiment, since the text input device 100 is mounted on a vehicle, the electroencephalogram sensor unit 200 can be disposed, for example, in a headrest portion of the seat where the user sits.

[0021] The biosignal detection unit 130 detects the user's electroencephalogram signal measured by the electroencephalogram sensor unit 200. The user's electroencephalogram signal contains an event-related potential that occurs due to the user's reaction to an event. The user's reaction appears in the electroencephalogram signal a certain time after the occurrence of the event. The biosignal detection unit 130 can extract a feature amount of the event-related potential.

[0022] The feature amount of the event-related potential may be, for example, a component called the P300 component (also called the P3 component) of the event-related potential. The P300 component of the event-related potential is the positive maximum value of the event-related potential at a point in time when a certain time has elapsed since the occurrence of an event. The certain time may be, for example, around 300 ms. The certain time may increase or decrease depending on the content of the event (the active area of ​​the user's brain reacting to the event).

[0023] The controller 110 can be realized by using a microcomputer including a CPU (Central Processing Unit), a memory, and an input / output unit. The controller 110 installs and executes a computer program in the microcomputer for causing the microcomputer to function as the controller 110. This allows the microcomputer to function as a plurality of information processing units (101 to 106) included in the controller 110.

[0024] Here, an example is shown in which the information notification device is realized by software, but it is of course possible to prepare dedicated hardware for executing each information process and configure the controller 110. Dedicated hardware includes devices such as application specific integrated circuits (ASICs) and conventional circuit components arranged to execute the functions described in the embodiments.

[0025] Although the controller 110 has been described as an independent component, it is of course possible to configure the controller 110 together with other devices as one device. Alternatively, the multiple information processing units (101-106) may be divided and configured using two or more different devices. Furthermore, all or part of the multiple information processing units (101-106) may be configured using an ECU (Electronic Control Unit) mounted on a vehicle.

[0026] The controller 110 includes a selected sentence proposing unit 101, a word presentation control unit 102, a desired word recognition signal analyzing unit 103, a desired word determining unit 104, a selected word output unit 105, and a signal determining unit 106 as a plurality of information processing units (101 to 106).

[0027] The selected sentence suggestion unit 101 can estimate a sentence that the user intends to input and propose (generate) constituent words of a plurality of sentence options. For example, as shown in FIG. 3A, the selected sentence suggestion unit 101 generates sentence candidates 1 to 3 as a plurality of sentence options. Sentence candidate 1 is a sentence in which three constituent words A to C are arranged in the order of A, B, C. Sentence candidate 2 is a sentence in which three constituent words D to F are arranged in the order of D, E, F. Sentence candidate 3 is a sentence in which three constituent words G to I are arranged in the order of G, H, I.

[0028] The selected sentence suggestion unit 101 (sentence suggestion unit) may estimate sentence options to be proposed from the environment surrounding the user, the communication status of the device connected to the sentence input device 100, and the like. The environment surrounding the user can be obtained, for example, from traffic congestion information of a car navigation system. For example, if the traffic congestion information indicates that a traffic congestion has occurred at the current location of the vehicle, the environment surrounding the user can be determined as "traffic congestion." The communication status can be, for example, the operating status of a device that requests an operation by inputting a sentence, or the status of sending and receiving an e-mail or the like using a sentence input to the sentence input device 100 by a device that incorporates the sentence input device 100 or a connected device.

[0029] For example, in a state where an email is being sent or received in a vehicle traveling through a traffic jam, the selected sentence suggestion unit 101 can estimate, as a sentence option, a sentence that the user is expected to intend to input as a sentence for an email notifying the user of the traffic jam. The selected sentence suggestion unit 101 can use, for example, AI (Artificial Intelligence) to estimate the sentence option.

[0030] The word presentation control unit 102 can classify the constituent words constituting each of the multiple sentence options proposed by the selected sentence suggestion unit 101 according to the order in which they are arranged in each sentence option.

[0031] 3B, the word presentation control unit 102 classifies the constituent words A to C, D to F, and G to I of each of the sentence candidates 1 to 3 into three word option groups 1 to 3. The constituent words A, D, and G in the first order are classified into word option group 1. The constituent words B, E, and H in the second order are classified into word option group 2. The constituent words C, F, and I in the third order are classified into word option group 3.

[0032] The word presentation control unit 102 can present the plurality of constituent words, each classified in the same order, one by one to the user on the presentation unit 120.

[0033] For example, the word presentation control unit 102 sequentially presents to the user the first constituent words A, D, and G of each of the sentence candidates 1 to 3 classified into the word option group 1. The word presentation control unit 102 can sequentially present to the user the constituent words A, D, and G by repeating them multiple times.

[0034] The word presentation control unit 102 presents the constituent words A, D, and G of the word option group 1 to the user in a random order in each word presentation. In the example shown in Fig. 4, in the first word presentation, the constituent words A, D, and G are presented to the user in this order, and in the second word presentation, the constituent word D is presented to the user first. The presentation interval of the constituent words A, D, and G can be, for example, 200 msec.

[0035] The word presentation control unit 102 can also sequentially present the second constituent words B, E, and H of each of the sentence candidates 1 to 3 classified into the word choice group 2 to the user one by one in the presentation unit 120 in the same manner as the constituent words A, D, and G of the word choice group 1. The word presentation control unit 102 can also sequentially present the third constituent words C, F, and I of each of the sentence candidates 1 to 3 classified into the word choice group 3 to the user one by one in the presentation unit 120 in the same manner as the constituent words A, D, and G of the word choice group 1.

[0036] When the constituent words of each of the word choice groups 1 to 3 are presented one by one in sequence, the user's reaction to the presentation of the same constituent word differs between when two constituent words with high similarity are presented in succession and when two constituent words with low similarity are presented in succession.

[0037] For example, when a target word that the user intends to input is presented next to a non-target word that has low similarity to the target word, the preceding and following words are easily distinguished, and the user's reaction to the target word is prominent. However, when the target word is presented next to a non-target word that has high similarity to the target word, the preceding and following words are not easily distinguished, and the user's reaction to the target word is dull.

[0038] If the user's response to the target constituent words is slow, it becomes difficult to determine the constituent words that the user intends to input from the event-related potentials of the user's electroencephalogram signals.

[0039] The word presentation control unit 102 presents the constituent words of each of the word option groups 1 to 3 to the user in a random order, in principle. However, when a combination of constituent words that is evaluated to have high similarity based on a distance value between two words calculated by a distance function is present in each of the word option groups 1 to 3, the word presentation control unit 102 makes the presentation interval between the constituent words included in the combination longer than usual.

[0040] In order to determine the order of sequential presentation of constituent words taking similarity into consideration, the word presentation control unit 102 executes the process shown in the flowchart of Fig. 5. For constituent words classified into each of the word option groups 1 to 3, the word presentation control unit 102 calculates the edit distance of each word to other words (step S101), and compares the calculated edit distance with a threshold value that is used as a criterion for determining similarity (step S103).

[0041] The word presentation control unit 102 determines the similarity of each word to other words based on the comparison result between the edit distance and a threshold value (step S105), and records the similarity determination result in the memory of the controller 110 (step S107). The word presentation control unit 102 generates a presentation order of constituent words in which highly similar words are not consecutive, based on the record in the memory (step S109), and ends the series of processes.

[0042] The presentation order of the constituent words generated by the word presentation control unit 102 in step S109 can be, for example, an order in which two constituent words with high similarity are sandwiched between other constituent words with low similarity to the two constituent words. Note that, instead of calculating the edit distance in step S101, the distance value between two words may be calculated using a distance function such as the Hamming distance.

[0043] In the process performed by the word presentation control unit 102, the parameters of the number of sequentially presented constituent words in one word presentation, the interval between sequentially presented constituent words, and the number of repetitions of word presentation are each set to an appropriate value.

[0044] When a combination of highly similar constituent words is included in each of the word option groups 1 to 3, increasing the number of constituent words presented and the presentation interval increases the probability that the presentation interval between highly similar constituent words will be longer. When a combination of highly similar constituent words is included in each of the word option groups 1 to 3, increasing the number of repetitions of word presentation increases the number of times the target constituent word is presented, and the user's reaction to the target constituent word gradually becomes stronger. By increasing the number of constituent words presented, the presentation interval, and the number of repetitions of word presentation, the accuracy of judging the target constituent word can be improved.

[0045] However, regardless of whether the number of constituent words presented, the presentation interval, or the number of times the words are presented is increased, the time required to present all of the constituent words classified into one word option group 1 to 3 will be longer, and the time at which the target constituent word can be identified will be delayed.

[0046] For this reason, the parameters of the number of constituent words to be presented, the presentation interval, and the number of times the word presentation is repeated may be set by taking into consideration the balance between the accuracy of determining the constituent words of the target and the time at which the constituent words of the target can be determined.

[0047] The desired word recognition signal analysis unit 103 (signal analysis unit) extracts a portion of the user's electroencephalogram signal detected by the biological signal detection unit 130, from the 0 point on the horizontal axis when the constituent word is presented to the presentation unit 120 to the time when the P300 component of the event-related potential appears. When the constituent word of the target is presented to the user, the P300 component appears about 500 msec after the 0 point on the horizontal axis when the constituent word is presented to the presentation unit 120.

[0048] In the example of the user's electroencephalogram signal when the constituent words of the non-target are presented to the user shown in Fig. 6A, only frequency fluctuations corresponding to the period (100 to 200 msec) during which the constituent words are presented one by one sequentially by the presentation unit 120 are observed, and no event-related potential is observed. In the example of the user's electroencephalogram signal when the constituent words of the target are presented to the user shown in Fig. 6B, in addition to frequency fluctuations corresponding to the presentation period of the constituent words, an event-related potential due to the user's reaction to the presentation of the constituent words of the target is observed.

[0049] The desired word recognition signal analysis unit 103 inputs the extracted electroencephalogram signal portion to the signal determination unit 106. The signal determination unit 106 (signal analysis unit) has, for example, a classifier that classifies the extracted portion of the electroencephalogram signal input from the desired word recognition signal analysis unit 103 according to whether or not it is a constituent word of the target that the user intends to input.

[0050] The classifier can be configured, for example, using a neural network that receives an extracted portion of the electroencephalogram signal from the desired word recognition signal analysis unit 103 as input and outputs a judgment value indicating whether or not the word is a constituent word of the target.

[0051] 7A and 7B, from the presentation of the constituent words by the presentation unit 120 to the time when the event-related potential appears, and inputs the feature value of the extracted part of the electroencephalogram signal. The feature value input to the classifier may be, for example, the slope of the waveform part approaching the maximum value of the P300 component of the event-related potential or the waveform part passing the maximum value, or the maximum value (peak value), etc.

[0052] The classifier can output a judgment value indicating whether or not the constituent word presented to the presentation unit 120 is a constituent word of the target based on the input feature value. Fig. 7A shows a case where a constituent word of a non-target is presented to the user, and Fig. 7B shows a case where a constituent word of a target is presented to the user.

[0053] For example, when a non-target word is presented to a user, the user's reaction to the non-target word appears in the portion extracted from the EEG signal. A classifier that receives the feature value of the EEG signal of the extracted portion including this reaction outputs a judgment value of "0" indicating that the word is a non-target.

[0054] When the constituent words of the target are presented to the user, the user's reaction to the constituent words of the target (the P300 component of the event-related potential) appears in the portion extracted from the EEG signal. The classifier, to which the feature value of the EEG signal of the extracted portion including this reaction is input, outputs a judgment value of "1" indicating that it is the target.

[0055] The data of the feature values ​​input to the classifier and the judgment values ​​output by the classifier are associated with each other and stored in a learning database (DB) of the controller 110. The stored data in the learning DB can be used when re-learning the classifier.

[0056] The desired word recognition signal analysis unit 103 outputs the judgment values ​​output by the classifier of the signal judgment unit 106 for all classified constituent words for each of the word option groups 1 to 3 as the analysis results of the P300 component of the event-related potential contained in the user's electroencephalogram signal corresponding to each constituent word.

[0057] The desired word determination unit 104 (word determination unit, sentence determination unit) compares the analysis results of the P300 components of the event-related potential corresponding to each constituent word output by the desired word recognition signal analysis unit 103 for all constituent words classified into each of the word option groups 1 to 3. The desired word determination unit 104 compares the analysis results for each of the word option groups 1 to 3. The desired word determination unit 104 executes a process of determining the constituent words of the target (desired words) according to the analysis results of the desired word recognition signal analysis unit 103 shown in FIG. 8.

[0058] In Fig. 8, the judgment values ​​(judgment results) of the P300 component of the event-related potential corresponding to all constituent words (option words) of eight constituent words A to H classified into a certain word option group are arranged by the number of times the words were presented. In the example shown in Fig. 8, constituent word A is the constituent word of the target, and constituent words B to H are the constituent words of the non-target. The part of the judgment value in Fig. 8 surrounded by a thick line frame is where the constituent word of the target has been classified as a non-target, or the constituent word of the non-target has been classified as a target, due to misclassification by the classifier.

[0059] The rightmost column of the table in Fig. 8 shows the average value of the judgment value obtained in five repetitions (presentations). The desired word judgment unit 104 judges the constituent word with the highest average judgment value to be the constituent word of the target. In the example shown in Fig. 8, the desired word judgment unit 104 judges the constituent word A with the highest average judgment value to be the constituent word of the target.

[0060] The selected word output unit 105 determines that a sentence in which the constituent words determined by the desired word determination unit 104 as targets for each of the word option groups 1 to 3 are connected in order is a sentence that the user intends to input. The selected word output unit 105 presents the determined sentence to the user at the presentation unit 120. In addition, the selected word output unit 105 can output text data of the determined sentence as a sentence input by the user to a device where the user requires input of a sentence.

[0061] In this embodiment, the options of the constituent words of a sentence are classified in order, and the constituent words of the options (word option groups) in each order are presented one by one to the presentation unit 120. The user recognizes the constituent words presented one by one without moving the viewpoint within the screen of the presentation unit 120 to search for them. Therefore, in the user's electroencephalogram signal, a reaction according to whether the presented constituent word is a target or a non-target appears without spending "eye movement time" shown in the bar graph of the collective presentation in FIG. 2. Therefore, by utilizing the P300 component of the event-related potential included in the user's electroencephalogram signal, it is possible to shorten the time required for inputting the words that constitute the intended sentence when the user inputs them.

[0062] In this embodiment, when a combination of constituent words that is evaluated to have high similarity based on the distance value between two words calculated by the distance function exists in each of the word option groups 1 to 3, the presentation interval between the constituent words included in the combination is made longer than usual. Therefore, the user's reaction to the constituent words of the next target appears after the user's reaction to the highly similar non-target constituent words has subsided, so that the user can react significantly to the constituent words of the target.

[0063] Furthermore, in this embodiment, the constituent words are presented sequentially in the order of two highly similar constituent words sandwiched between other constituent words with low similarity, so that the user can easily distinguish between the preceding and following constituent words, and the user can react significantly to the target constituent word.

[0064] In this embodiment, the word presentation control unit 102 repeats each of the constituent words A, D, and G classified into the word option group in the same order multiple times to sequentially present them to the user in random order. This allows the constituent words of the target to be intermittently and repeatedly presented to the user, so that the user's reaction to the constituent words of the target increases as the number of presentations increases.

[0065] (First Modification) The accuracy rate when determining constituent words of a sentence intended by a user using the P300 component of the event-related potential varies depending on the presentation time for which the constituent words are presented to the user by the presentation unit 120. The presentation time refers to the interval at which the constituent words are presented sequentially when the word presentation control unit 102 presents words for each of the word option groups 1 to 3.

[0066] When the presentation interval (presentation time) of constituent words becomes longer, the accuracy rate of the judgment by the desired word judgment unit 104 increases at a certain time length, for example, as shown in Fig. 9A. The longer the presentation interval (presentation time) of constituent words, the higher the probability that the presentation interval between highly similar constituent words will be longer, and the easier it is for the user to recognize whether the constituent word is the one he or she intended while it is being presented.

[0067] Also, the accuracy rate of the judgment by the desired word judgment unit 104 decreases when the number of sequentially presented constituent words in one word presentation performed by the word presentation control unit 102 for each of the word option groups 1 to 3 is small, and increases when the number of presentations is large. The accuracy rate of the judgment by the desired word judgment unit 104 decreases when the number of repetitions of word presentation performed by the word presentation control unit 102 for each of the word option groups 1 to 3 is small, and increases when the number of repetitions is large. The graph in FIG. 9B shows the relationship between the accuracy rate of the judgment by the desired word judgment unit 104, the number of constituent words presented by the word presentation control unit 102 (number of presented words), and the number of repetitions of word presentation (number of word presentations).

[0068] The sensitivity of the reaction to the presentation of constituent words varies from user to user. Therefore, the relationship between the accuracy rate of the judgment by the desired word judgment unit 104 and the presentation interval (presentation time) of the constituent words by the word presentation control unit 102 shown in Fig. 9A may differ from user to user. For the same reason, the relationship between the accuracy rate of the judgment by the desired word judgment unit 104, the number of constituent words presented by the word presentation control unit 102 (number of presented words), and the number of times the word presentation is repeated (number of times the word is presented) shown in Fig. 9B may also differ from user to user.

[0069] As shown in Fig. 10A, for a user other than the user shown in Fig. 9A, the accuracy rate of the judgment by the desired word judgment unit 104 increases when the interval (presentation time) of the constituent words is longer than that of the user shown in Fig. 9A. Also, as shown in Fig. 10B, for this user, the accuracy rate of the judgment by the desired word judgment unit 104 is lower than that of the user shown in Fig. 9A for the same number of constituent words presented (number of presented words) and the same number of repetitions of word presentation (number of word presentations) as the user shown in Fig. 9A.

[0070] Thus, when the word presentation control unit 102 sequentially presents constituent words, the number of constituent words presented in one word presentation, the sequential presentation interval of the constituent words, and the number of repetitions of word presentation may be set as control parameters, and may be optimized using the classification results by the classifier of the signal determination unit 106. To optimize each parameter, the word presentation control unit 102 selects an optimal combination of the number of constituent words presented (number of presented words) and the number of repetitions of word presentation (number of presentations).

[0071] The optimal combination is the combination of the number of constituent words presented (number of presented words) and the number of repetitions of word presentation (number of presentations) that results in the correct answer rate of the constituent words that the user intends to input exceeding a predetermined threshold and the shortest word selection time (word judgment time). The optimal combination is premised on a constant presentation interval (presentation time). The word selection time is the product of three parameters: the number of repetitions of word presentation, the number of constituent words presented in one word presentation, and the sequential presentation interval of the constituent words (number of word presentations x number of word presentations x presentation time = word selection time).

[0072] If there are multiple optimal combinations of the number of constituent words presented (number of presented words) and the number of times the word presentation is repeated (number of presentations), the optimal combination is, for example, the combination with the fewest number of times the word presentation is repeated (number of presentations).

[0073] In order to select an optimal combination of the number of constituent words to be presented (number of presented words) and the number of times that word presentation is repeated (number of presentations), the word presentation control unit 102 executes a selection process shown in the flowchart of FIG. 11. The word presentation control unit 102 causes the controller 110 to determine constituent words of a sentence intended by the user using the P300 component of the event-related potential while sequentially changing the value of each parameter. The word presentation control unit 102 acquires the numerical values ​​of the number of constituent words to be presented (number of presented words) at the time of judgment, the number of times that word presentation is repeated (number of presentations), and the accuracy rate of the determined constituent words. Then, the word presentation control unit 102 accumulates data indicating each acquired numerical value in an optimization database (DB) of the controller 110 (the above is step S201).

[0074] The word presentation control unit 102 analyzes the data stored in the optimization DB, and selects an optimal combination of the number of constituent words to be presented (number of presented words) and the number of times the word presentation is repeated (number of presentations) that will maximize the correct answer rate (step S203). The word presentation control unit 102 also analyzes the data stored in the optimization DB, and selects an optimal combination of the number of constituent words to be presented (number of presented words) and the number of times the word presentation is repeated (number of presentations) that will minimize the word selection time (step S205). Furthermore, the word presentation control unit 102 analyzes the data stored in the optimization DB, and selects an optimal combination of the number of constituent words to be presented (number of presented words) and the number of times the word presentation is repeated (number of presentations) that will minimize the number of times the word presentation is repeated (number of presentations) (step S207). This ends the selection process.

[0075] The optimal interval (shortest time) for presenting the constituent words is determined separately.

[0076] If there are multiple optimal combinations selected by the selection process of Figure 11, the word presentation control unit 102 selects, as described above, for example, the combination with the fewest number of repetitions of word presentation (number of presentations) as the optimal combination.

[0077] By performing the above selection process for each user, the parameters of the number of constituent words presented in one word presentation, the interval between successive presentations of the constituent words, and the number of repetitions of word presentation when the constituent words are presented successively in the presentation unit 120 can be optimized for each user. By optimizing the parameters for each user, the accuracy rate of the judgment by the desired word judgment unit 104 can be increased for each user. In addition, by optimizing each parameter, the time required for the user to judge the constituent words of the sentence that the user intends to input can be further shortened.

[0078] (Second Modification) The classifier in the signal determination unit 106 can improve the accuracy of the determination value indicating whether or not a word is a constituent word of the target by, for example, adjusting the bias and gain in the neural network through re-learning. By improving the accuracy of the determination value of the classifier, it is possible to further reduce the time required to determine the constituent words of the sentence that the user intends to input.

[0079] In order to re-learn the classifier, the signal determination unit 106 executes a learning process shown in the flowchart of Fig. 12. Every time the classifier classifies an input feature value and outputs a judgment value, the signal determination unit 106 accumulates a combination of the feature value and the judgment value as event-related potential data in the learning DB of the controller 110 (step S301).

[0080] The signal determination unit 106 selects target data with a judgment value of "1" from the past event-related potential data of the user stored in the learning DB (step S303). Also, the signal determination unit 106 selects non-target data with a judgment value of "0" from the accumulated data in the learning DB (step S305).

[0081] The signal determination unit 106 performs machine learning of the classifier using the utilization data of the selected target data and non-target data (step S307), and ends the learning process. The machine learning of the classifier can be, for example, a content that optimizes (adjusts) the bias and gain in the neural network of the classifier to suit the user by re-learning using the backpropagation method.

[0082] When selecting the target data and the non-target data to be utilized, the signal determination unit 106 excludes the event-related potential data erroneously recognized by the user from the data to be utilized for re-learning the classifier. If the event-related potential data erroneously recognized by the user is utilized for re-learning the classifier, as shown in FIG. 13A, the classification boundary line of the data classified (discriminated) by the classifier may be biased toward either the target side or the non-target side.

[0083] Fig. 13B shows the analysis result of the desired word recognition signal analysis unit 103, which includes a mixture of judgment errors due to misclassification by the classifier and judgment errors due to misrecognition by the user. The bold judgment value part in Fig. 13B is where the constituent words of the target are classified as non-targets, or the constituent words of the non-targets are classified as targets, due to misclassification by the classifier. The bold-lined judgment value part in Fig. 13B is where the constituent words of the target are classified as non-targets, or the constituent words of the non-targets are classified as targets, due to misrecognition by the user.

[0084] In the learning process performed by the signal determination unit 106, if only data misclassified by the classifier is excluded as the utilization data of the target data and non-target data, and data misrecognized by the user is utilized, the bias and gain of the classifier are inappropriately adjusted. As a result, the classifier after re-learning performs classification with a biased boundary line between the target data and non-target data as shown in FIG. 13A.

[0085] Therefore, in the example of Figure 13A, the four constituent words A, B, E, and H have the highest average classifier judgment values ​​obtained after five iterations, and the constituent words of the target cannot be determined.

[0086] For this reason, when selecting utilization data of target data and non-target data in the learning process, the signal determination unit 106 excludes data that has been erroneously recognized by the user in addition to data that has been erroneously classified by the classifier.

[0087] This allows the bias and gain of the classifier after re-learning to be appropriately adjusted, thereby reliably improving the accuracy of the judgment value through re-learning of the classifier and reliably shortening the time required to determine the constituent words of the sentence that the user intends to input.

[0088] The utilization data of the target data and non-target data can be selected, for example, by comparing the peak area of ​​the portion marked with diagonal lines in Figures 7A and 7B from the presentation of the constituent words in the presentation unit 120 to the time when the event-related potential appears, with a threshold value.

[0089] For example, in the case of target data, when the peak area is equal to or greater than the threshold value corresponding to the minimum area, the data can be selected as target data, whereas in the case of non-target data, when the peak area is equal to or less than the threshold value corresponding to the maximum area, the data can be selected as non-target data.

[0090] By determining the data to be used for re-learning the classifier by comparing the peak area with a threshold value, even if the peak shape of the event-related potential varies from user to user, the data to be used for re-learning for each user can be selected with high accuracy. In addition, by determining the data to be used for re-learning the classifier by comparing the peak area with a threshold value, it is possible to prevent data that has been determined as an error due to a user's misrecognition from being used for re-learning the classifier.

[0091] When both the optimization of the control parameters of the word presentation control unit 102 in the first modified example and the re-learning of the classifier in the second modified example are performed, the execution schedule for each can be, for example, as shown in Fig. 15. First, at the initial stage after purchase (start of operation) of the sentence input device 100, a general-purpose classifier with bias and gain set to initial values ​​is used as the classifier of the signal determination unit 106.

[0092] While increasing or decreasing the control parameters of the word presentation control unit 102, the user is fixed and the controller 110 is made to determine the constituent words of the sentence intended by the user by utilizing the P300 component of the event-related potential. This allows data for optimizing the control parameters according to the user to be accumulated in the optimization DB of the controller 110.

[0093] When sufficient data has been accumulated in the optimization DB, the data is analyzed, and the optimal combination of the number of constituent words (number of presented words) and the number of repetitions of word presentation (number of presentations) that minimizes the number of repetitions of word presentation (number of presentations) is selected, and the control parameters are optimized. At this point, the classifier remains a general-purpose classifier.

[0094] After optimizing the control parameters of the word presentation control unit 102, event-related potential data that combines the feature values ​​input to the classifier and the judgment values ​​output by the classifier is stored in the learning DB of the controller 110 as data for re-learning the classifier.

[0095] When sufficient data is accumulated in the learning DB, data to be used for re-learning is selected from the accumulated data, and the selected data is used to perform machine learning of the classifier, thereby updating the classifier to an optimized classifier. At this point, the control parameters of the word presentation control unit 102 are optimized parameters, and the classifier becomes an optimized classifier. After that, every time sufficient data is newly accumulated in the learning DB, or periodically, machine learning of the classifier is performed to update the classifier.

[0096] In this way, by first optimizing the control parameters of the word presentation control unit 102 and then re-learning the classifier, the data in the learning DB used for re-learning becomes data with contents that improve the classification accuracy of the classifier by re-learning. Therefore, it is possible to further reduce the time required for the user to input constituent words of a sentence that he or she intends to input.

[0097] In this embodiment, the text input device 100 is mounted on a vehicle, but the text input device 100 can be used without being mounted on a vehicle. In this embodiment, the brain wave sensor unit 200 is used to detect brain activity in the biosignal detection unit 130, but instead of this, various known means for detecting electromagnetic waves generated due to brain nerve activity, such as magnetic field detection using a non-contact biomagnetic sensor, can be applied.

[0098] The above-described embodiment and modified examples are merely examples of the present invention. Therefore, the present invention is not limited to the above-described embodiment and modified examples, and various modifications can be made according to the design and the like, even if the modifications are in other forms than the above-described embodiment, as long as they do not deviate from the technical concept of the present invention. [Explanation of symbols]

[0099] 100 Text input device (text suggestion section) 101 Selection Suggestion Section 102 Word presentation control unit 103 Desired word recognition signal analysis unit 104 Desired word determination unit (word determination unit, sentence determination unit) 106 Signal judgment unit (classifier) 120 Presentation section 200 Brain wave sensor unit

Claims

1. The system predicts the sentence the user intends to input and suggests multiple options for the sentence. classifying constituent words constituting each of the plurality of sentence options into an order in which each of the sentence options is arranged; presenting the plurality of constituent words, each classified into the same sequence, one by one to the user; analyzing an event-related potential, which is included in an electroencephalogram signal measured from the user and is generated based on the presentation of the constituent words, for each of the constituent words presented one by one in sequence to determine the constituent words that the user intends to input in the same order; repeating the sequential presentation of the plurality of constituent words and the determination of the constituent words that the user intends to input for all of the permutations to determine a sentence that the user intends to input; the sequential presentation of the plurality of constituent words in the same order is repeated, and the event-related potential is analyzed for each of the constituent words at each repetition; at least one of the parameters of the number of constituent words presented and the presentation interval in one sequential presentation and the number of repetitions of the sequential presentation is optimized based on the result of the judgment of the sentence that the user intends to input; and each of the parameters of the number of presentations and the number of repetitions is optimized to a combination of the number of presentations and the number of repetitions such that, when the presentation interval is constant, the accuracy rate in the judgment of the constituent words that the user intends to input exceeds a predetermined threshold value and the word judgment time from the start of the sequential presentation of the constituent words in the same order to the judgment of the constituent words that the user intends to input is the shortest. How to enter text.

2. The system predicts the sentence the user intends to input and suggests multiple options for the sentence. classifying constituent words constituting each of the plurality of sentence options into an order in which each of the sentence options is arranged; presenting the plurality of constituent words, each classified into the same sequence, one by one to the user; analyzing an event-related potential, which is included in an electroencephalogram signal measured from the user and is generated based on the presentation of the constituent words, for each of the constituent words presented one by one in sequence to determine the constituent words that the user intends to input in the same order; repeating the sequential presentation of the plurality of constituent words and the determination of the constituent words that the user intends to input for all of the permutations to determine a sentence that the user intends to input; A sentence input method in which, when a combination of constituent words that is evaluated to have a high similarity based on a distance value between the two words calculated by a distance function is present among the plurality of constituent words classified in the same order, the two constituent words included in the combination are presented at an interval longer than the sequential presentation interval of the plurality of constituent words.

3. The sentence input method according to claim 2 , wherein the constituent words not included in the combination are presented in an order sandwiched between two of the constituent words included in the combination.

4. 4. The sentence input method according to claim 2 or 3, wherein the sequential presentation of the plurality of constituent words is repeated in the same order, the event-related potential is analyzed for each of the constituent words at each repetition, and at least one of the parameters of the number of presentations and the presentation interval of the constituent words in one session of the sequential presentation and the number of repetitions of the sequential presentation is optimized based on a result of a judgment of a sentence that the user intends to input, and each of the parameters of the number of presentations and the number of repetitions is optimized to a combination of the number of presentations and the number of repetitions such that, when the presentation interval is constant, a correct answer rate in the judgment of the constituent words that the user intends to input exceeds a predetermined threshold value and the word judgment time from the start of the sequential presentation of the constituent words in the same order to the judgment of the constituent words that the user intends to input is the shortest.

5. The method for inputting text according to claim 4 , wherein, when there are a plurality of said combinations, the parameters of the number of presentations and the number of repetitions are optimized for the combination with the smallest number of repetitions.

6. 6. The text input method according to claim 1, wherein in the analysis of the event-related potential for each of the constituent words, a classifier that classifies the event-related potential according to whether or not the constituent word is one that the user intends to input is optimized for the user by machine learning based on past analysis results of the event-related potential for the user.

7. 7. The text input method according to claim 6, further comprising: separating past analysis results of the event-related potential based on a peak area of ​​a waveform portion of the event-related potential in the electroencephalogram signal; and using a portion of the separated past analysis results of the event-related potential for the machine learning.

8. Inferring a plurality of sentence options to be proposed as sentences that the user intends to input, based on the environment surrounding the user and a communication state of a device connected to a sentence input device used as a man-machine interface of a device to which the user needs to input sentences; The text input method according to claim 1 .

Citation Information

Patent Citations

  • JP1973056791A

  • Translation device

    JP2004295578A

  • Method for determining human psychological state, or the like, using event-related potential and system thereof

    JP2005034620A

  • Communication device, communication method and program

    JP2011013871A

  • Character input device and character input system

    JP2015219762A