Speech recognition device, speech recognition method, and speech recognition program
The speech recognition device addresses noise-induced command misrecognition by allowing operation-based correction and database association, enhancing accuracy and reducing user intervention.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-04
AI Technical Summary
Existing voice recognition systems fail to accurately identify commands due to differences in noise sounds between the time of command registration and utterance, particularly in mobile objects like vehicles, leading to incorrect recognition and the need for user intervention to select synonyms.
A speech recognition device that acquires input voice data, compares it with registered voice data, presents an estimation result, and allows users to correct errors through operation-based determination of the correct command, associating it with the input data in a database.
Ensures accurate recognition of commands by associating the correct command with input voice data, reducing user burden and improving recognition accuracy even in noisy environments.
Smart Images

Figure 2026035881000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technology for voice recognition of commands for a mobile object. [Background technology]
[0002] Patent Document 1 discloses a voice recognition device that, if it is unable to recognize speech input by a vehicle user, stores the speech as an unrecognized word in association with the vehicle's driving conditions, selects multiple synonyms for the unrecognized word from a voice recognition dictionary based on the vehicle's driving conditions, presents the selected multiple synonyms to the user, and stores the synonym selected by the user from the presented multiple synonyms in association with the unrecognized word.
[0003] However, Patent Document 1 does not take into consideration the possibility that correct speech recognition may not be possible due to differences in noise sounds when registering a registration command and when speaking an input command, and therefore further improvement is necessary. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-233542 Summary of the Invention
[0005] The present disclosure has been made to solve such problems, and aims to provide a technology that can correctly identify the correct command for an input command even if the noise sound differs between when the registered command is registered and when the input command is spoken.
[0006] A voice input device according to one aspect of the present disclosure is a voice recognition device that performs voice recognition for commands for a moving body, and includes: a first acquisition unit that acquires input voice data for an input command spoken by a speaker riding on the moving body; a database that stores registered voice data for a plurality of registered commands previously spoken by the speaker; an estimation unit that estimates a registered command corresponding to the input command by comparing the plurality of registered voice data with the input voice data; a presentation unit that presents the estimation result; a second acquisition unit that acquires an error indication indicating that the estimation result is incorrect; a determination unit that, when the error indication is acquired, determines a correct command corresponding to the input command based on the speaker's operation; and a database management unit that associates the correct command with the input voice data and stores them in a database.
[0007] According to the present disclosure, even if the noise sound differs between when the registered command is registered and when the input command is uttered, the correct command for the input command can be correctly identified. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram illustrating an example of a configuration of a voice recognition device according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating an example of a data configuration of a database. [Figure 3] 4 is a flowchart showing an example of processing performed by the voice recognition device according to the first embodiment. [Figure 4] FIG. 10 is a diagram showing an example of a scene in which an input command is uttered. [Figure 5] FIG. 10 is a diagram showing an example of a confirmation screen displaying a confirmation message. [Figure 6] FIG. 10 is a diagram showing an example of a scene in which an error button is operated. [Figure 7] FIG. 10 is a diagram showing an example of a list screen of correct answer candidate commands. [Figure 8] FIG. 10 is a diagram showing an example of a scene in which a correct command is selected. [Figure 9]FIG. 10 is a diagram showing an example of an end screen that displays an end message. [Figure 10] FIG. 10 is a block diagram showing an example of the configuration of a voice recognition device according to a second embodiment. [Figure 11] 10 is a flowchart showing an example of processing performed by the voice recognition device according to the second embodiment. [Figure 12] FIG. 10 is a diagram showing a cancellation screen presenting a cancellation message. [Figure 13] FIG. 1 is a diagram illustrating an example of a monitoring scene. [Figure 14] FIG. 11 is a block diagram showing an example of the configuration of a voice recognition device according to a third embodiment. [Figure 15] 11 is a flowchart showing an example of processing performed by a voice recognition device 1B according to the third embodiment. [Figure 16] 16 is a continuation of the flowchart in FIG. 15. [Figure 17] FIG. 10 is a diagram showing an example of a screen presenting an estimated correct command. [Figure 18] FIG. 10 is a diagram showing an example of a list screen of correct answer candidate commands. [Figure 19] FIG. 10 is a diagram showing an example of a presentation screen according to a modified example. [Figure 20] 10 is a flowchart according to a modification of the first embodiment. [Figure 21] FIG. 10 is a diagram showing an example of a confirmation screen displaying a confirmation message in a modification of the first embodiment. [Figure 22] 10 is a flowchart according to a modification of the second embodiment. [Figure 23] 13 is a flowchart according to a modification of the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] (Findings underlying this disclosure) A speaker identification technique is known that acquires speech data uttered by a target speaker to be identified, compares the feature values of the acquired speech data with the feature values of the speech data of multiple registered speakers, and identifies which of multiple registered speakers the target speaker corresponds to. It has been found that with such speaker identification techniques, the recognition rate decreases if the speech content is different, even for the same speaker. It has also been found that the decrease in recognition rate is particularly noticeable for short utterances such as device commands.
[0010] On the other hand, by utilizing this decrease in recognition rate, it is possible to recognize commands to devices uttered by the same speaker. For example, by registering the features of multiple registered commands uttered by a target speaker in advance, comparing the feature of an input command uttered by the target speaker with each of the features of the multiple registered commands, and determining the registered command with the greatest similarity as the input command, the input command can be recognized.
[0011] However, in a device such as a mobile object, the ambient noise differs depending on the driving situation, such as when the vehicle is moving or when the vehicle is stopped. Therefore, if the driving situation differs between the time of registration and the time of utterance, such as when the feature amounts of multiple registered commands are registered when the vehicle is stopped and the input command is uttered while the vehicle is moving, only a low similarity value can be obtained, and the input command cannot be recognized accurately.
[0012] Patent Document 1 does not mention the difference in noise sound between registration and speech as a cause of voice recognition failure, and therefore differs in concept from the present disclosure. Patent Document 1 also presents multiple synonyms for an unrecognized word to the user, but these multiple synonyms are simply selected based on the driving situation and are not necessarily synonymous with the unrecognized word. Furthermore, Patent Document 1 requires the user to select a synonym from the multiple synonyms presented, which requires the user to memorize a synonym different from the word originally intended. Therefore, Patent Document 1 needs further improvement in terms of associating the correct synonym corresponding to the unrecognized word with the unrecognized word and memorizing it.
[0013] The present disclosure has been made to solve such problems. Hereinafter, aspects of the present disclosure will be described.
[0014] A speech recognition device according to one aspect of the present disclosure is a speech recognition device that performs speech recognition on commands for a moving body, and includes: a first acquisition unit that acquires input speech data for an input command spoken by a speaker riding on the moving body; a database that stores registered speech data for a plurality of registered commands previously spoken by the speaker; an estimation unit that estimates a registered command corresponding to the input command by comparing the plurality of registered speech data with the input speech data; a presentation unit that presents the estimation result; a second acquisition unit that acquires an error indication indicating that the estimation result is incorrect; a determination unit that, when the error indication is acquired, determines a correct command corresponding to the input command based on the speaker's operation; and a database management unit that associates the correct command with the input speech data and stores them in a database.
[0015] According to this configuration, a registered command corresponding to an input command is estimated, and if an error indication is obtained for the estimation result, a correct command for the input command is determined based on the speaker's operation, and the correct command and the input voice data are associated and stored in a database.
[0016] Therefore, even if an input command cannot be correctly recognized due to a difference in noise between when the registered command is registered and when the input command is uttered, the correct command for the input command can be correctly identified and the identified correct command can be registered in the database in association with the input voice data. As a result, the input voice data including the noise sound when the input command cannot be recognized is registered in the database in association with the correct command. Therefore, if an incorrectly estimated input command is uttered in the same scene as when an erroneous instruction was input, the correct command for that input command can be correctly recognized. Furthermore, with this configuration, the input utterance data of the words (input command) that the speaker originally intended to say is registered as is, and the act of registering this input utterance data leaves an impression of the words that the speaker originally intended to say in the speaker's mind, making the input of commands by voice extremely smooth thereafter.
[0017] In the above-described speech recognition device, the determination unit may present a plurality of correct candidate commands, acquire a selection instruction to select one correct candidate command from the plurality of correct candidate commands, and determine the one correct candidate command as the correct command.
[0018] According to this configuration, one correct candidate command selected from a plurality of correct candidate commands is determined as the correct command, so that the correct command corresponding to the input command can be correctly recognized.
[0019] In the above voice recognition device, the determination unit may present, as the correct candidate commands, the plurality of registered commands sorted in descending order of similarity between the input voice data and the plurality of registered voice data.
[0020] According to this configuration, multiple registered commands sorted in descending order of similarity between the input voice data and multiple registered voice data are presented as multiple correct candidate commands, making it easier to select the correct command.
[0021] In the above voice recognition device, the determination unit may monitor an operation input to the moving object after the input of the erroneous instruction, and determine the correct command based on a monitoring result.
[0022] According to this configuration, after an incorrect instruction is input, the correct command is determined based on the results of monitoring the operation input to the moving object, so the correct command can be determined without explicitly confirming the correct command with the speaker, thereby reducing the processing burden and the burden on the speaker.
[0023] In the above-described voice recognition device, the determination unit may store in a memory the input voice data corresponding to the input command in which the erroneous instruction was input, estimate a correct command based on a monitoring result of an operation on the moving object after the erroneous instruction is input, and when it detects that the moving object has stopped, play back the input voice data stored in the memory and present the estimated correct command, obtain a confirmation instruction for the estimated correct command, and determine the correct command based on the confirmation instruction.
[0024] According to this configuration, when a stop of the mobile body is detected, input voice data corresponding to the erroneous instruction stored in memory is played back, and an estimated correct command is presented based on the monitoring results of operations input to the mobile body after the erroneous instruction is input. The correct command is determined based on a confirmation request from the speaker for the estimated correct command. This ensures the speaker's safety while driving, since the correct command is confirmed after the vehicle has stopped, rather than while the vehicle is moving. Furthermore, because the estimated correct command is presented along with the playback of the input voice data, even if multiple input commands are erroneously estimated while driving, the speaker can easily confirm which input command the played back input voice data corresponds to. Furthermore, because the correct command presented together with the input voice data is the correct command estimated based on the monitoring results of operations input to the mobile body, the correct command presented together with the input voice data is more likely to be the correct command, facilitating the speaker's confirmation work.
[0025] In the above-described speech recognition device, when the determination unit receives the confirmation instruction indicating that the estimated correct command is incorrect, the determination unit may present the plurality of registered commands as correct candidate commands, receive a selection instruction to select one correct candidate command from the plurality of registered commands, and determine the one correct candidate command as the correct command.
[0026] According to this configuration, even if the estimated correct command is incorrect, a plurality of registered commands are presented as correct candidate commands, and the speaker decides the correct command from among the presented correct candidate commands, so that the correct command can ultimately be decided.
[0027] In the above voice recognition device, the determination unit may present, as the correct candidate command, registered commands sorted in descending order of similarity between the plurality of registered voice data and the input voice data.
[0028] According to this configuration, multiple registered commands sorted in descending order of similarity between the input voice data and multiple registered voice data are presented as multiple correct candidate commands, making it easier to select the correct command.
[0029] In the above voice recognition device, the estimation unit may estimate a registered command corresponding to the input command by comparing features of the plurality of registered voice data with the input voice data.
[0030] According to this configuration, since the feature amounts of the plurality of registered voice data and the input voice data are compared, it is possible to accurately estimate the registered command corresponding to the input command.
[0031] In the above-described speech recognition device, the estimation unit may identify a registered speaker corresponding to the speaker by comparing features of the input speech data with features of the speech of a plurality of registered speakers, and estimate a registered command corresponding to the input command by comparing registered speech data of the plurality of registered commands with the input speech data for the identified registered speaker.
[0032] According to this configuration, the registered command corresponding to the input command is estimated by comparing the input voice data and the registered voice data of the same speaker, so that the registered command can be estimated with high accuracy.
[0033] In the above speech recognition device, the presentation unit may present a message prompting the user to input the error indication only when the estimation result is incorrect.
[0034] According to this configuration, it is only necessary to input an error indication when the estimation result is incorrect, thereby reducing the input burden.
[0035] In the above-described speech recognition device, if the second acquisition unit acquires the erroneous instruction within a predetermined timeout period, the determination unit may determine the correct command based on an operation of the speaker, and if the second acquisition unit does not acquire the erroneous instruction within the timeout period, the determination unit may determine that the estimation result is correct.
[0036] According to this configuration, if an erroneous instruction is obtained within the timeout period, a process for determining the correct command is executed, and if an erroneous instruction is not obtained within the timeout period, the estimation result is determined to be correct, thereby avoiding a situation where the accuracy of the estimation result is never determined.
[0037] Another aspect of the present disclosure is a speech recognition method in a speech recognition device that speech-recognizes commands for a moving body, which method acquires input speech data of an input command spoken by a speaker riding on the moving body, acquires from a database a plurality of registered commands previously spoken by the speaker, estimates a registered command corresponding to the input command by comparing each registered command with the input speech data, presents the estimated result, acquires an error indication indicating that the estimated result is incorrect, and if the error indication is acquired, determines a correct command corresponding to the input command based on an operation of the speaker, and stores the correct command and the input speech data in association with each other in the database.
[0038] According to this configuration, it is possible to provide a speech recognition method that can achieve the same effects as those of the above-described speech recognition device.
[0039] In yet another aspect of the present disclosure, a voice recognition program causes a computer to function as a voice recognition device that voice recognizes commands for a moving body, and causes the computer to perform the following processes: acquire input voice data of an input command spoken by a speaker riding on the moving body; acquire registered voice data of a plurality of registered commands previously spoken by the speaker from a database; estimate a registered command corresponding to the input command by comparing each registered voice data with the input voice data; present the estimated result; acquire an error indication indicating that the estimated result is incorrect; and, if the error indication is acquired, determine a correct command corresponding to the input command based on the speaker's operation; and store the correct command in a database in association with the input voice data.
[0040] According to this configuration, it is possible to provide a voice recognition program that can achieve the same effects as the voice recognition device described above.
[0041] It goes without saying that the present disclosure allows such a speech recognition program to be distributed via a computer-readable non-transitory recording medium such as a CD-ROM or a communication network such as the Internet.
[0042] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.
[0043] (Embodiment 1) 1 is a block diagram showing an example of the configuration of a speech recognition device 1 according to a first embodiment of the present disclosure. The speech recognition device 1 is mounted on, for example, a mobile object. Examples of the mobile object include a gasoline-powered automobile, an electric automobile, an electric bicycle, an electric kick scooter, and an electric motorcycle. However, this is just one example, and the speech recognition device 1 may also be a mobile information terminal carried by a user riding in the mobile object.
[0044] The speech recognition device 1 includes a microphone 2 , a processor 3 , a database 4 , a speaker 5 , a display 6 , an operation unit 7 , and a memory 8 .
[0045] The microphone 2 picks up sounds such as voices spoken by a user riding on the vehicle, converts the picked up sounds into sound signals, A / D converts the converted sound signals, and inputs the A / D converted sound signals to the first acquisition unit 31.
[0046] The speaker 5 converts sound signals such as messages generated by the processor 3 into sounds, and outputs the converted sounds to the outside.
[0047] The display 6 is a display device such as a liquid crystal display panel, and displays images including various messages generated by the presentation unit 33.
[0048] The operation unit 7 is configured with, for example, a touch panel and physical buttons, and accepts user operations. The operation unit 7 includes, for example, an operation unit that accepts user operations on the voice recognition device 1 and an operation unit that accepts operations on the mobile object. The operation unit that accepts operations on the mobile object is, for example, an operation unit for operating devices provided on the mobile object. The devices provided on the mobile object include, for example, audio equipment, air conditioning equipment, windshield wipers, a car navigation system, and lighting equipment.
[0049] The memory 8 is configured by, for example, a rewritable nonvolatile storage device, and stores a program that causes the processor 3 to function as the voice recognition device 1, etc.
[0050] The processor 3 is configured by, for example, a central processing unit. The processor 3 includes a first acquisition unit 31, an estimation unit 32, a presentation unit 33, a second acquisition unit 34, a determination unit 35, and a database management unit 36. The first acquisition unit 31 to the database management unit 36 are realized by the processor 3 executing a voice recognition program. However, this is just one example, and the first acquisition unit 31 to the database management unit 36 may also be configured by dedicated semiconductor circuits such as ASICs.
[0051] The first acquisition unit 31 detects a voice section from the sound signal input from the microphone 2 and acquires the detected voice section as input voice data. The input voice data includes, for example, voice data of an input command uttered by a user riding in a mobile object. The input command is a command for controlling the mobile object. For example, the input command may be a command for setting a destination in a car navigation system, a command for switching the display screen of the car navigation system to heading up or north up, a command for operating the drive system of the mobile object such as the engine and accelerator, or a command for operating various equipment of the mobile object such as the air conditioner, wipers, windows, and doors.
[0052] The estimation unit 32 compares the input voice data with a plurality of registered voice data to estimate a registered command corresponding to the input command. For example, the estimation unit 32 calculates the similarity between each of the registered voice data, which is voice data of a plurality of registered commands pre-registered in the database 4, and the input voice data, and estimates the registered command with the largest similarity equal to or greater than a threshold as the registered command corresponding to the input command. For example, the estimation unit 32 may calculate the similarity between the feature quantities of the registered voice data and the feature quantities of the input voice data. The feature quantities are, for example, feature quantities suitable for voice recognition, such as i-vector, x-vector, and d-vector. The similarity can be, for example, the reciprocal of the distance between the feature quantities of the registered voice data and the feature quantities of the input voice data. The distance is, for example, the Euclidean distance.
[0053] The presentation unit 33 presents the estimation result by the estimation unit 32. For example, the presentation unit 33 presents a confirmation message as the estimation result to confirm whether the registered command estimated by the estimation unit 32 is correct or not. The estimation unit 32 may present the confirmation message by outputting the confirmation message to at least one of the display 6 and the speaker 5.
[0054] The second acquisition unit 34 acquires a confirmation instruction from the user, indicating whether the estimation result is correct or incorrect, via the operation unit 7. The second acquisition unit 34 may acquire the confirmation instruction by voice recognition. In this case, the second acquisition unit 34 acquires the confirmation instruction by performing voice recognition on input voice data picked up by the microphone 2 within a predetermined period after the confirmation message is presented. The confirmation instruction includes, for example, an error instruction indicating that the estimation result is incorrect, or a correct instruction indicating that the estimation result is correct. When a correct instruction is acquired, the second acquisition unit 34 outputs the registration command estimated by the estimation unit 32 to the in-vehicle controller 100.
[0055] When an incorrect instruction is acquired, the determination unit 35 determines a correct command corresponding to the input command based on the speaker's operation. For example, the determination unit 35 presents a plurality of correct candidate commands, acquires a selection instruction to select one correct candidate command from the plurality of correct candidate commands, and determines the acquired one correct candidate command as the correct command. The determination unit 35 may present, as the plurality of correct candidate commands, a plurality of registered commands sorted in descending order of similarity between the input voice data and the plurality of registered voice data. For example, the determination unit 35 may display a list of the plurality of registered commands sorted in descending order on the display 6.
[0056] The database management unit 36 stores the correct command and the input voice data in association with each other in the database 4.
[0057] The database 4 is configured, for example, with a hard disk drive or a solid state drive. The database 4 stores features of registered voice data for each of a plurality of registered commands for each of one or more speakers scheduled to ride the moving object. FIG. 2 is a diagram showing an example of the data configuration of the database 4. The database 4 stores speaker IDs, registered command IDs, and features in association with each other. The speaker ID is an identifier that uniquely identifies one or more speakers scheduled to ride the moving object. The registered command ID is an identifier that uniquely identifies a registered command. The features are features of the registered voice data. In the example of FIG. 2, the database 4 stores features for n registered commands with registered command IDs "C1" to "Cn" for a user with user ID "U1".
[0058] The registration voice data is voice data acquired through a pre-registration process in which a speaker utters registration commands one by one. This pre-registration process is performed, for example, immediately after purchasing a mobile object while the mobile object is stopped. "Stopped" refers to a state in which power is supplied from the battery of the mobile object to at least the voice recognition device 1, but the mobile object is not moving.
[0059] In the pre-registration process, the database management unit 36 outputs messages prompting the speaker to speak each of the multiple registration commands that are recognition candidates to the display 6 and the speaker 5 in a predetermined order. The speaker speaks the registration commands one by one in accordance with these messages. Prior to starting the pre-registration process, the database management unit 36 prompts the speaker to input a speaker ID.
[0060] This allows the database management unit 36 to know which speaker spoke the currently uttered registered command and which command it was. The database management unit 36 then acquires registered voice data using the microphone 2, calculates the feature amounts of the acquired registered voice data, and stores the calculated feature amounts in the database 4 in association with the user ID and registered command ID.
[0061] FIG. 3 is a flowchart showing an example of the processing of the speech recognition device 1 in the first embodiment. In step S11, a speaker (user) utters an input command. In step S12, the first acquisition unit 31 acquires input voice data of the input command from a sound signal collected by the microphone 2. FIG. 4 is a diagram showing an example of a scene in which an input command is uttered. In the example of FIG. 4, a speaker operating the steering wheel 400 utters "turn up the temperature" as an input command. The steering wheel 400 includes an error button 601 for inputting an incorrect instruction and a decision button 602 for inputting a correct instruction. The error button 601 and the decision button 602 are examples of the operation unit 7.
[0062] In step S22, the first acquisition unit 31 records the input voice data by storing the input voice data in the memory 8. In the example of Fig. 4, the input voice data of "raise the temperature" is recorded.
[0063] In step S23, the estimation unit 32 identifies the speaker who uttered the input command from the features of the input voice data, calculates the similarity between each of the features of multiple registered commands stored in the database 4 and the features of the input voice data for the identified speaker, and estimates the registered command with the largest similarity equal to or greater than a threshold as the registered command corresponding to the input command. When identifying the speaker, the estimation unit 32 calculates the similarity between the features of the voice of each of one or more registered speakers who are scheduled to ride the vehicle and are registered in advance in the memory 8 of the vehicle, and the features of the input voice data, and identifies the registered speaker with the largest calculated similarity as the speaker who uttered the input command.
[0064] In step S24, the presentation unit 33 presents a confirmation message to the speaker. Fig. 5 is a diagram showing an example of a confirmation screen G1 showing a confirmation message. In the example of Fig. 5, the estimation result for the input command is "Turn up the volume," so the confirmation screen G1 displays the estimation result "Turn up the volume." The confirmation screen G1 also displays a message "Command estimation result" indicating that "Turn up the volume" is the estimation result for the input command, and a message "Press OK or NG" prompting the speaker to input a confirmation instruction as to whether the estimation result is correct or incorrect.
[0065] In step S12, the speaker inputs the confirmation result using the operation unit 7. Fig. 6 is a diagram showing an example of a scene in which the error button 601 is operated. In the example of Fig. 6, although the speaker uttered "turn up the temperature," it was estimated to be "turn up the volume," so the speaker presses the error button 601. If the estimation result of the estimation unit 32 is correct, the speaker presses the decision button 602.
[0066] In step S25, the second acquisition unit 34 determines whether the input confirmation instruction is an incorrect instruction or a correct instruction. Here, since the speaker has pressed the error button 601, the second acquisition unit 34 determines that the confirmation instruction is an incorrect instruction (YES in step S25). On the other hand, if the speaker has pressed the enter button 602, the second acquisition unit 34 determines that the confirmation instruction is a correct instruction (NO in step S25), and outputs the registration command estimated by the estimation unit 32 to the in-vehicle controller 100 (step S32). As a result, the registration command is accepted by the in-vehicle controller 100.
[0067] In step S26, the determination unit 35 displays a plurality of registered commands stored in the database 4 as correct candidate commands on the display 6. In this case, the determination unit 35 sorts the registered commands in descending order of similarity of the registered voice data to the input voice data and displays them on the display 6. FIG. 7 is a diagram showing an example of a list screen G2 of correct candidate commands. In this example, the similarity of the registered commands to the input voice data is highest in the order of "increase temperature," "decrease volume," and "increase airflow," so the list screen G2 displays the registered commands in this order. Note that, when displaying the correct candidate commands, the determination unit 35 may exclude registered commands that have been erroneously estimated by the estimation unit 32.
[0068] In step S13, the speaker selects a correct command from the registered commands displayed on the list screen G2. In the example of Fig. 7, the speaker selects the correct command by inputting an operation to position the rectangular cursor 702 on the correct command and pressing the decision button 602.
[0069] 8 is a diagram showing an example of a scene in which a correct command is selected. As shown in FIG. 8, the operation of positioning cursor 702 can be performed by operating up button 603 or down button 604 provided on handle 400. Up button 603 is a button that moves cursor 702 upward, and down button 604 is a button that moves cursor 702 downward. Up button 603 and down button 604 are examples of operation unit 7.
[0070] 7, when the down button 604 is pressed while the cursor 702 is positioned on the bottom or second-to-bottom registered command, the list screen G2 scrolls through the registered commands displayed, and displays the registered command with the lowest similarity to the input voice data. The list screen G2 may also scroll through the registered commands displayed when an operation is input on the scroll bar 701.
[0071] The speaker may select the correct command by touching the registered command that corresponds to the correct command from the registered commands displayed on the list screen G2. In this example, the registered command "Raise the temperature" is selected.
[0072] In step S27, the determination unit 35 acquires a selection instruction indicating the selected registered command via the operation unit 7. In step S28, the determination unit 35 determines the registered command indicated by the correct instruction as the correct command.
[0073] In step S29, the determination unit 35 outputs the correct command to the in-vehicle controller 100. As a result, the in-vehicle controller 100 acquires the correct command and executes the acquired correct command. In this case, the correct command "increase temperature" is executed. As a result, the in-vehicle controller 100 performs control to increase the temperature of the air conditioner equipped in the moving object.
[0074] In step S30, the database management unit 36 stores the correct command in the database 4 in association with the speaker ID of the speaker identified in step S23 and the feature amounts of the input voice data acquired in step S21. As a result, a record is added to the database 4 in which the feature amounts of the input voice data, the registered command ID of the registered command corresponding to the correct command, and the speaker ID are associated with each other. Here, the feature amounts of the added input voice data are feature amounts that include noise sounds that are environmental sounds when the erroneous instruction is input. Therefore, if an incorrectly estimated input command is uttered in a scene similar to the scene in which the erroneous instruction was input, the estimation unit 32 can compare the feature amounts of the input voice data and the registered voice data, which have similar noise sounds, thereby improving the estimation accuracy of the registered command.
[0075] In step S31, the determination unit 35 uses at least one of the microphone 2 and the speaker 5 to present to the user an end message indicating that the correct command has been accepted. FIG. 9 is a diagram showing an example of an end screen G3 that displays an end message. The end screen G3 displays a message saying "Command accepted," which indicates that the input command has been accepted. The end screen G3 also displays a message saying that the accepted command, "Raise temperature," will be "executed."
[0076] In this way, according to the speech recognition device 1, a registered command corresponding to an input command is estimated, and if an error indication is obtained for the estimation result, a correct command for the input command is determined based on the speaker's operation, and the correct command and the input speech data are associated with each other and stored in the database 4.
[0077] Therefore, even if an input command uttered by the same speaker cannot be correctly recognized due to a difference in noise sound between when the registered command is registered and when the input command is uttered, it is possible to correctly identify the correct command for the input command and to associate the identified correct command with the input voice data and the speaker ID and register it in the database 4. As a result, input voice data including the noise sound when the input command cannot be recognized is registered in the database 4 in association with the correct command and the speaker ID. Therefore, if an incorrectly estimated input command is uttered by the same speaker in the same scene as when the erroneous instruction was input, the correct command for that input command can be correctly recognized.
[0078] (Embodiment 2) In the second embodiment, a correct command is determined based on the results of monitoring operations input to a moving object. Fig. 10 is a block diagram showing an example of the configuration of a speech recognition device 1A in the second embodiment. In the second embodiment, the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof will be omitted.
[0079] The processor 3A of the speech recognition device 1A includes a first acquisition unit 31, an estimation unit 32, a presentation unit 33, a second acquisition unit 34, a determination unit 35A, and a database management unit 36. After the second acquisition unit 34 acquires an erroneous instruction, the determination unit 35A monitors the operation input to the operation unit 7 and determines the correct command based on the monitoring result. Here, the determination unit 35A may determine, as the correct command, the first registered command input after the second acquisition unit 34 acquires the erroneous instruction. For example, the determination unit 35A may set a certain period of time after the erroneous instruction is acquired as a monitoring period, and may determine, as the correct command, the first registered command input during the monitoring period. The monitoring period is an estimated period from the input of the erroneous instruction to the input of an operation corresponding to the spoken input command, and may be an appropriate period of time, such as one minute or two minutes.
[0080] Fig. 11 is a flowchart showing an example of processing by the speech recognition device 1A in the second embodiment. In Fig. 11, steps S101 and S102 are the same as S11 and S12 in Fig. 3. In Fig. 11, steps S201, S202, S203, S204, S205, and S212 are the same as steps S21, S22, S23, S24, S25, and S32 in Fig. 3. Also, Fig. 11 illustrates a case in which "Turn up the volume" is estimated in response to an utterance of "Turn up the temperature," as in the first embodiment. Therefore, in step S204, the confirmation screen G1 shown in Fig. 5 is displayed.
[0081] In step S206, the second acquisition unit 34, which has acquired the error instruction, presents a cancellation message indicating that the input command has been canceled to the speaker using at least one of the microphone 2 and the speaker 5. Fig. 12 is a diagram showing a cancellation screen G4 presenting the cancellation message. The cancellation screen G4 displays the message "Cancellation accepted," indicating that the input command has been canceled.
[0082] In step S103, the speaker inputs an operation for the device to the operation unit 7.
[0083] In step S207, the determination unit 35A monitors the device operation input to the operation unit 7. In step S208, the determination unit 35A determines the correct command from the monitoring result.
[0084] FIG. 13 is a diagram illustrating an example of a monitoring scene. In the example of FIG. 13, the dial 500 for adjusting the temperature of the air conditioner is adjusted clockwise, and an operation to increase the temperature of the air conditioner is input. This operation was the first operation input during the monitoring period. Therefore, the determination unit 35A determines that the input command is to "increase the temperature" of the air conditioner. Note that the determination unit 35A may determine the registered command as the correct command when the registered command is first input during the monitoring period, or may wait until the end of the monitoring period, identify the registered command that was first input during the monitoring period, and determine the identified registered command as the correct command. Note that if no registered command is input during the monitoring period, the determination unit 35A may terminate the process without determining a correct command.
[0085] Steps S209, S210, and S211 are the same as steps S29, S30, and S31 in FIG. 3, and the correct command is output to the in-vehicle controller 100, the correct command is associated with the features of the input voice data and the speaker ID and stored in the database 4, and an end message is displayed.
[0086] In this way, according to the speech recognition device 1A, after an erroneous instruction is input, the correct command is determined based on the monitoring results of the operation input to the moving object, so that the correct command can be determined without explicitly confirming the correct command with the speaker, thereby reducing the processing load and the burden on the speaker.
[0087] (Embodiment 3) In the third embodiment, the speaker is made to confirm the correct command after the moving object has stopped. Fig. 14 is a block diagram showing an example of the configuration of a speech recognition device 1B in the third embodiment. In the third embodiment, the same components as those in the first and second embodiments are denoted by the same reference numerals, and the description thereof will be omitted.
[0088] The processor 3B of the speech recognition device 1B includes a first acquisition unit 31, an estimation unit 32, a presentation unit 33, a second acquisition unit , a determination unit 35B, and a database management unit .
[0089] The determination unit 35B stores input voice data corresponding to the input command in which the erroneous instruction was input in the memory 8. After the erroneous instruction is input, the determination unit 35B estimates the correct command based on the monitoring results of the operation on the moving object. Details of the process of estimating the correct command based on the monitoring results are the same as the process of determining the correct command based on the monitoring results in the second embodiment.
[0090] When the determination unit 35B detects that the mobile object has stopped, it plays back the input voice data stored in the memory 8, presents the estimated correct command, receives a confirmation instruction for the estimated correct command, and determines the correct command based on the confirmation instruction. For example, the determination unit 35B may detect that the mobile object has stopped when it receives a stop notification from the in-vehicle controller 100 informing the user that the mobile object has stopped. The stop notification may be sent, for example, when the ignition key of the mobile object is turned off, or when the shift lever of the mobile object is placed in park or neutral. The determination unit 35B may also recognize that the mobile object has stopped when it has remained in a predetermined location for a certain period of time using location information such as GPS. When the determination unit 35B detects that the mobile object has stopped, it may play back the input voice data stored in the memory 8 and present the correct command to a mobile information terminal (e.g., a smartphone) of a user riding in the mobile object.
[0091] FIG. 15 is a flowchart showing an example of the processing of the speech recognition device 1B in the third embodiment. In FIG. 15, steps S111 and S112 are the same as steps S101 and S103 in FIG. 11. In FIG. 15, steps S221, S222, S223, S224, S225, S226, and S236 are the same as steps S201, S202, S203, S204, S205, S206, and S212 in FIG. 11. Also, FIG. 15 illustrates a case in which the utterance "Turn up the temperature" is estimated to be "Turn up the volume" in response to the utterance "Turn up the temperature," as in the second embodiment. Therefore, in step S224, the confirmation screen G1 shown in FIG. 5 is displayed, and in step S226, the cancel screen G4 shown in FIG. 12 is displayed.
[0092] Fig. 16 is a flowchart continuing from Fig. 15. In Fig. 16, step S113 is the same as step S102 in Fig. 11. In Fig. 16, step S227 is the same as step S207 in Fig. 11.
[0093] In step S228, the determination unit 35B estimates a correct command from the monitoring result of the device operation input to the operation unit 7. In step S301, the in-vehicle controller 100 outputs a vehicle stop notification. Hereinafter, the estimated correct command will be referred to as an estimated correct command.
[0094] In step S229, the determination unit 35B detects that the moving object has stopped by receiving the stop notification. In step S230, the determination unit 35B plays back the input voice data recorded in step S222 and presents the estimated correct command.
[0095] FIG. 17 is a diagram showing an example of a display screen G5 for displaying an estimated correct command. The display screen G5 displays the estimated correct command "Raise the temperature" and a message inquiring whether the displayed estimated correct command is correct, "Is this correct?". With the display screen G5 displayed, the determination unit 35B outputs recorded input voice data from the speaker 5. The speaker determines whether the estimated correct command is correct by comparing the output input voice data "Raise the temperature" with the estimated correct command displayed on the display screen G5.
[0096] In step S114, the speaker who has checked the presentation screen G5 inputs a confirmation instruction indicating whether the estimated correct command is correct. If the speaker determines that the command is correct, he or she presses the decision button 602. As a result, the decision unit 35B acquires the correct instruction as a confirmation instruction (NO in step S231), decides that the estimated correct command is the correct command, and stores the decided correct command in the database 4 in association with the feature amount of the input voice data and the speaker ID (step S236).
[0097] On the other hand, a speaker who determines that the answer is incorrect presses the error button 601. As a result, the decision unit 35B acquires an error instruction as a confirmation instruction (YES in step S231), and the process proceeds to step S232.
[0098] In step S232, the determination unit 35B displays a plurality of registered commands stored in the database 4 as correct candidate commands on the display 6. FIG. 18 is a diagram showing an example of a list screen G6 of correct candidate commands. In this example, the similarity of the registered commands to the input voice data is highest in the order of "Turn down the volume," "Increase the airflow," and "Decrease the airflow," so the list screen G6 displays the registered commands in this order. Note that, when displaying the correct candidate commands, the determination unit 35B may exclude registered commands and estimated correct commands that have been erroneously estimated by the estimation unit 32. The basic configuration of the list screen G6 is the same as that of the list screen G2 shown in FIG. 7, and therefore a detailed description thereof will be omitted.
[0099] In step S115, the speaker selects a correct command from the registered commands displayed on the list screen G6. In the example of Fig. 18, the speaker operates the up button 603 or the down button 604 to position the cursor 702 on the corresponding registered command, and selects the correct command by pressing the enter button 602. The enter unit 35B may output the input voice data "Raise the temperature" from the speaker 5 again while the list screen G6 is displayed.
[0100] In step S233, the determination unit 35B acquires a selection instruction indicating the selected registered command via the operation unit 7. In step S234, the determination unit 35B determines the registered command indicated by the correct instruction as the correct command.
[0101] In step S235, the database management unit 36 stores the correct command in the database 4 in association with the feature amount of the input voice data and the speaker ID.
[0102] In this way, according to the speech recognition device 1B, when the stopping of the moving body is detected, the input speech data corresponding to the erroneous instruction stored in the memory 8 is reproduced, and an estimated correct command estimated based on the monitoring result of the operation input to the moving body after the input of the erroneous instruction is presented, and the correct command is determined based on a confirmation instruction from the speaker for the estimated correct command. As a result, the correct command is confirmed after the moving body has stopped, rather than while the moving body is moving, and the safety of the speaker while the moving body is moving can be ensured.
[0103] Note that multiple erroneous commands may be input while the vehicle is traveling. In this case, the first acquisition unit 31 may store multiple pieces of input voice data containing erroneous commands in the memory 8 while the vehicle is traveling. The determination unit 35B may estimate an estimated correct command for each piece of input voice data while the vehicle is traveling based on the results of monitoring device operations for each piece of input voice data. After the vehicle has stopped, the determination unit 35B may present the estimated correct command along with playback for each piece of input voice data, and have the speaker sequentially confirm the correct command, thereby determining the correct command corresponding to the multiple pieces of input voice data. In this case, the database management unit 36 may store the feature values for each piece of input voice data in the database 4 in association with the determined correct command and speaker ID.
[0104] In this way, by presenting an estimated correct command along with the playback of input voice data, even if multiple erroneous instructions are input while driving, the speaker can easily confirm which input command the input voice data being played back corresponds to.
[0105] Furthermore, since the estimated correct command presented along with the input voice data is a correct command estimated from the monitoring results of the operation input to the mobile object, the correct command presented along with the input voice data is more likely to be the correct command, making it easier to confirm the speaker.
[0106] The present disclosure can employ the following modifications.
[0107] (1) In step S230 of Fig. 16, the determination unit 35B may display a presentation screen G7 shown in Fig. 19 on the display 6. Fig. 19 is a diagram showing an example of the presentation screen G7 according to a modified example. The presentation screen G7 displays the estimated correct command at the top, and displays registered commands sorted in descending order of similarity to the input voice data from second to last, excluding the registered command erroneously estimated by the estimation unit 32.
[0108] The determination unit 35B outputs the input voice data "Raise the temperature" from the speaker 5 in conjunction with the display of the presentation screen G5. The speaker who hears this voice operates the up button 603 or the down button 604 to position the cursor 702 on the target registration command, and presses the determination button 602. The registration command selected in this manner is determined as the correct command for the input voice data, and is stored in the database 4 in association with the feature amount of the input voice data. In this modified example, the speaker does not need to input the confirmation instruction shown in step S114, and the display of the list screen G6 shown in step S232 is also unnecessary, thereby simplifying the processing.
[0109] (2) In the first embodiment, some of the components of the processor 3 and the database 4 may be held by a cloud server.
[0110] (3) The database stores the feature quantities of the registered voice data, but may store the registered voice data.
[0111] (4) In the first embodiment, the presentation unit 33 may present the estimation result by outputting a confirmation message to at least one of the display 6 and the speaker 5, prompting the user to input an error instruction only when the estimation result is incorrect. Fig. 20 is a flowchart according to a modification of the first embodiment. The same processes in Fig. 20 and Fig. 3 are denoted by the same numbers, and descriptions thereof will be omitted.
[0112] In step S1001 following step S23, the presentation unit 33 presents a confirmation message to the speaker. FIG. 21 is a diagram showing an example of a confirmation screen G1' showing a confirmation message in a variation of the first embodiment. The difference between the confirmation screen G1' and the confirmation screen G1 is that the confirmation screen G1 displays a message saying "Press OK or NG," whereas the confirmation screen G1' displays a message saying "Press NG if incorrect." In other words, the message on the confirmation screen G1' prompts the speaker to input an error indication only if the estimation result is incorrect.
[0113] In step S1002, if the speaker determines that the estimated result "Turn up the volume" displayed on the confirmation screen G1' is incorrect, the speaker inputs an error instruction using the operation unit 7. In this case, the speaker presses the error button 601 as described in FIG.
[0114] On the other hand, in step S1002, if the speaker determines that the estimation result "increase the volume" displayed on the confirmation screen G1' is accurate, he or she does not input anything to the operation unit .
[0115] In step S1003, the second acquiring unit 34 determines whether or not an error instruction has been acquired within a timeout period. The timeout period is a period of time during which the second acquiring unit 34 waits for the speaker to input an error instruction, and is a preset period of time.
[0116] If the second acquisition unit 34 acquires an error indication within the timeout period (YES in step S1003), the process proceeds to step S26. On the other hand, if the second acquisition unit 34 does not acquire an error indication within the timeout period (NO in step S1003), the process proceeds to step S1004.
[0117] In step S1004, the determining unit 35 determines that the estimation result obtained by the estimating unit 32 is correct, and the process proceeds to step S32. The subsequent processes are the same as those in the first embodiment.
[0118] According to this modification, an error indication only needs to be input when the estimation result is incorrect, thereby reducing the input burden on the speaker.
[0119] (5) In the second embodiment, the presentation unit 33 may present the estimation result by outputting a confirmation message to at least one of the display 6 and the speaker 5, prompting the user to input an error indication only when the estimation result is incorrect. Fig. 22 is a flowchart of a modified example of the second embodiment. The same processes in Fig. 22 and Fig. 11 are denoted by the same numbers, and their explanations will be omitted.
[0120] In step S2001, the presentation unit 33 presents a confirmation message to the speaker. In this case, the presentation unit 33 presents the confirmation screen G1' described in the modification of the first embodiment to the speaker.
[0121] In step S2002, if the speaker determines that the estimated result "Turn up the volume" displayed on the confirmation screen G1' is incorrect, the speaker inputs an error instruction using the operation unit 7. In this case, the speaker presses the error button 601 as described in FIG.
[0122] On the other hand, in step S2002, if the speaker determines that the estimation result "increase the volume" displayed on the confirmation screen G1' is accurate, he or she does not input anything to the operation unit 7.
[0123] In step S2003, the second acquisition unit 34 determines whether or not an error indication was acquired within the timeout period. If the second acquisition unit 34 acquired an error indication within the timeout period (YES in step S2003), the process proceeds to step S206. On the other hand, if the second acquisition unit 34 did not acquire an error indication within the timeout period (NO in step S2003), the process proceeds to step S2004.
[0124] In step S2004, the determining unit 35A determines that the estimation result obtained by the estimating unit 32 is correct, and the process proceeds to step S212. After that, the same processes as those in the second embodiment are executed.
[0125] According to this modification, an error indication only needs to be input when the estimation result is incorrect, thereby reducing the input burden on the speaker.
[0126] (6) In the third embodiment, the presentation unit 33 may present the estimation result by outputting to at least one of the display 6 and the speaker 5 a confirmation message prompting the user to input an error instruction only when the estimation result is incorrect. FIG. 23 is a flowchart according to a modification of the third embodiment. The same processes in FIG. 23 and FIG. 15 are assigned the same numbers, and their explanations will be omitted. The flowchart following FIG. 23 is the same as FIG. 16.
[0127] In step S3001, the presentation unit 33 presents a confirmation message to the speaker. In this case, the presentation unit 33 presents the confirmation screen G1' described in the modification of the first embodiment to the speaker.
[0128] In step S3002, if the speaker determines that the estimated result "Turn up the volume" displayed on the confirmation screen G1' is incorrect, the speaker inputs an error instruction using the operation unit 7. In this case, the speaker presses the error button 601 as described in FIG. On the other hand, in step S3002, if the speaker determines that the estimation result "increase the volume" displayed on the confirmation screen G1' is accurate, he or she does not input anything to the operation unit .
[0129] In step S3003, the second acquisition unit 34 determines whether or not an error indication was acquired within the timeout period. If the second acquisition unit 34 acquired an error indication within the timeout period (YES in step S3003), the process proceeds to step S226. On the other hand, if the second acquisition unit 34 did not acquire an error indication within the timeout period (NO in step S3003), the process proceeds to step S3004.
[0130] In step S3004, the determining unit 35B determines that the estimation result obtained by the estimating unit 32 is correct, and the process proceeds to step S236. After that, the same processes as those in the third embodiment are executed.
[0131] According to this modification, an error indication only needs to be input when the estimation result is incorrect, thereby reducing the input burden on the speaker. [Industrial Applicability]
[0132] The present disclosure is useful in the technical field of voice input of input commands for devices such as mobile objects.
Claims
1. A voice recognition device that recognizes voice commands for a mobile object, a first acquisition unit that acquires input voice data of an input command uttered by a speaker riding on the vehicle; a database that stores registered voice data of a plurality of registered commands; an estimation unit that estimates a registered command corresponding to the input command based on the plurality of registered voice data and the input voice data; a presentation unit that presents the estimation result; a second acquisition unit that acquires an error indication indicating that the estimation result is an error; a determination unit that, when the error indication is acquired, determines a correct command corresponding to the input command; the determination unit estimates a correct command based on a monitoring result of an operation on the moving object after the input of the erroneous instruction; presenting the correct command estimated by the determination unit; presenting at least one of the plurality of registered commands in a presentation order following the correct command estimated by the determination unit; Voice recognition device.
2. the determination unit presents the correct command estimated by the determination unit in the highest presentation order; presenting at least one of the plurality of registered commands sorted in descending order of similarity between the plurality of registered voice data and the input voice data in a presentation order following the correct command estimated by the determination unit; The speech recognition device according to claim 1 .
3. At least one of the plurality of registered commands presented excludes a registered command that has been erroneously estimated by the estimation unit.
2. The speech recognition device according to claim 1.
4. It also has a display, the determination unit highlights the correct command estimated by the determination unit on the display.
2. The speech recognition device according to claim 1.
5. the estimation unit estimates a registered command corresponding to the input command by comparing features of the plurality of registered voice data with the input voice data.
2. The speech recognition device according to claim 1.
6. the estimation unit identifies a registered speaker corresponding to the speaker by comparing a feature of the input voice data with feature of voices of a plurality of registered speakers, and estimates a registered command corresponding to the input command by comparing registered voice data of the plurality of registered commands with the input voice data for the identified registered speaker.
2. The speech recognition device according to claim 1.
7. A speech recognition method for a speech recognition device that recognizes a command for a mobile object by speech, comprising: acquiring input voice data of an input command uttered by a speaker riding on the vehicle; Retrieves multiple registered commands from the database, estimating a registered command corresponding to the input command based on each registered command and the input voice data; Present the estimation results, obtaining an error indication indicating that the estimation result is incorrect; If the error indication is obtained, determining a correct command corresponding to the input command; In determining the correct command, After the input of the erroneous instruction, a correct command is estimated based on a result of monitoring the operation of the moving object; Present the estimated correct command, presenting at least one of the plurality of registered commands in a presentation order following the estimated correct command; Speech recognition methods.
8. A speech recognition program that causes a computer to function as a speech recognition device that recognizes speech of a command for a moving object, acquiring input voice data of an input command uttered by a speaker riding on the vehicle; Acquires registered voice data for multiple registered commands from the database, estimating a registered command corresponding to the input command based on each registered voice data and the input voice data; Present the estimation results, obtaining an error indication indicating that the estimation result is incorrect; If the error indication is obtained, determining a correct command corresponding to the input command; In determining the correct command, After the input of the erroneous instruction, a correct command is estimated based on a result of monitoring the operation of the moving object; Present the estimated correct command, causing a computer to execute a process of presenting at least one of the plurality of registered commands in a presentation order following the estimated correct command; Speech recognition program.
Citation Information
Patent Citations
Speech recognition equipment
JP2004233542A