Speech recognition device, speech recognition method, and program
The speech recognition device addresses user operability issues by assessing proficiency and tailoring command output to the user's skill level, enhancing ease of use and reducing confusion in voice command inputs.
Patent Information
- Application Number
- JP2024120641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Existing voice recognition devices face challenges in user operability, particularly in determining appropriate commands based on user proficiency, leading to confusion and difficulty in inputting commands.
A speech recognition device equipped with a proficiency determination unit that assesses user proficiency through successful voice input attempts, and a selection processing unit that sets and outputs commands tailored to the user's skill level, allowing easier operation by displaying or vocalizing commands appropriate for the user's proficiency.
Enhances user experience by providing commands suitable for the user's skill level, reducing confusion and improving the ease of operation, especially in environments like vehicle navigation systems.
Smart Images

Figure 2026019226000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech recognition device, a speech recognition method, and a program. [Background technology]
[0002] In recent years, voice input to information processing devices has become common. For example, voice input is suitable for vehicles, since drivers often have both hands full while driving. For example, Patent Document 1 describes a voice recognition device equipped with a command restriction means for restricting the actual use of certain operation commands from among a large number of operation commands that can be used, and a command increase means for increasing the number of operation commands that a user is permitted to use depending on the user's usage of the voice recognition device. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-282284 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in a voice recognition device that allows free speech, there is room for improvement in terms of operation, such as the difficulty in knowing what to say when inputting. One example of the problem that the present invention aims to solve is to provide a voice recognition device that is easy for users to operate. [Means for solving the problem]
[0005] The invention described in claim 1 is A speech recognition device, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a selection processing unit that allows a user to select a command; Equipped with The selection processing unit, in the selection processing, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; The speech recognition device acquires the speech input by the user, processes the speech, and uses the results to identify the command selected by the user.
[0006] The invention described in claim 14 is A speech recognition method performed by a computer, comprising: determining a user's proficiency with the speech recognition method; Let the user select a command, In the selection process, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; The speech recognition method acquires speech input by the user, processes the speech, and uses the results to identify the command selected by the user.
[0007] The invention described in claim 15 is A program that causes a computer to function as a speech recognition device, Computer, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a selection processing unit that allows a user to select a command; It functions as The selection processing unit, in the selection processing, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; The program acquires the voice input by the user, processes the voice, and uses the results to identify the command selected by the user. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram showing an overview of a voice recognition device according to a first embodiment together with a usage environment. [Figure 2] 4 is a diagram showing an example of data stored in a storage unit and used in a first example of a method for determining a level of proficiency. FIG. [Figure 3] FIG. 10 is a diagram showing an example of a command that the selection processing unit causes the screen display device to display. [Figure 4] FIG. 10 is a diagram showing an example of a screen including a command for performing a selection process multiple times. [Figure 5] 10 is an example of commands and difficulty levels stored in a storage unit. [Figure 6] 10 is an example of rules, predetermined values, and priorities stored in a storage unit. [Figure 7] FIG. 1 is a first diagram showing output targets set for each level of proficiency. [Figure 8] FIG. 2 is a second diagram showing output targets set for each level of proficiency. [Figure 9] FIG. 3 is a third diagram showing output targets set for each level of proficiency. [Figure 10] FIG. 4 is a fourth diagram showing output targets set for each proficiency level. [Figure 11] FIG. 2 is a block diagram illustrating an example of a hardware configuration of a voice recognition device. [Figure 12] FIG. 3 is a flowchart illustrating an example of the operation of the voice recognition device according to the first embodiment. [Figure 13] FIG. 10 is a diagram showing an example of a screen displayed while waiting for a predetermined operation from a user. [Figure 14] FIG. 10 is a tree diagram showing screens that are displayed after a predetermined operation is accepted. [Figure 15] FIG. 10 is a block diagram showing an overview of a voice recognition device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are denoted by like reference numerals, and the description thereof will be omitted as appropriate.
[0010] [First embodiment] 1 is a diagram showing an overview of a speech recognition device 10 according to this embodiment together with a usage environment. The speech recognition device 10 is used together with, for example, a speech input device 20, a speech output device 30, and a screen display device 40. The speech input device 20, the speech output device 30, and the screen display device 40 may all be part of the speech recognition device 10, or may be devices separate from the speech recognition device 10. In the latter case, the speech input device 20, the speech output device 30, and the screen display device 40 communicate with the speech recognition device 10 at least when the speech recognition device 10 is used.
[0011] The speech recognition device 10 can be used for various purposes, and as an example, it is used as an input interface for a car navigation device that supports vehicle driving. However, the speech recognition device 10 may also be used for other purposes. In the following description, the speech recognition device 10 will be described as being used as at least a part of a car navigation device.
[0012] The vehicle may be, for example, an automobile, a motorcycle, a moped, or a bicycle, but is not limited to these. When the vehicle is an automobile, for example, the voice recognition device 10 is integrated with the voice input device 20, the voice output device 30, and the screen display device 40 and attached to the vehicle. When the vehicle is a motorcycle, a moped, or a bicycle, for example, the voice input device 20 and the voice output device 30 are separate devices from the voice recognition device 10. The voice input device 20 and the voice output device 30 are attached to, for example, the driver's helmet. However, these devices may also be attached to wearable items other than a helmet, such as eyeglasses or sunglasses.
[0013] The voice recognition device 10 also has a wireless communication function. As an example, the voice recognition device 10 is a smartphone on which a predetermined application is installed. In this case, the microphone of the smartphone functions as the voice input device 20, the speaker of the smartphone functions as the voice output device 30, and the screen of the smartphone functions as the screen display device 40.
[0014] 1, the speech recognition device 10 includes a proficiency determination unit 100, a selection processing unit 200, and a storage unit 400. However, the storage unit 400 may be included inside the speech recognition device 10 or may be provided externally. Each component will be described in detail below.
[0015] [Proficiency Assessment Unit 100] The proficiency determination unit 100 determines the user's proficiency with the speech recognition device 10. The proficiency with the speech recognition device 10 is an index that indicates how familiar the user is with operating the speech recognition device 10. The speech recognition device 10 uses this proficiency to make it easier for the user to use the speech recognition device 10. A specific method by which the proficiency determination unit 100 determines the user's proficiency with the speech recognition device 10 will be described below.
[0016] <Example 1> The proficiency determination unit 100 determines the proficiency of the voice recognition device 10 based on the number of times the user successfully inputs voice. The proficiency determination unit 100 determines that the user's proficiency with the voice recognition device 10 increases as the number of times the user successfully inputs voice increases. For example, the proficiency determination unit 100 determines whether the voice input is successful or not using the recognition result for the input voice. For example, if the recognition result is consistent with the content prompting the user to input, for example, if the result is included in pre-prepared options, the proficiency determination unit 100 determines that the voice input is successful. For example, if the recognition result is consistent with the content prompting the user to input, for example, if the result is included in pre-prepared options, the proficiency determination unit 100 determines that the voice input is successful. For example, if the recognition result is recognized when the user is prompted to input a destination, the proficiency determination unit 100 determines that the voice input is successful. On the other hand, if the recognition result is recognized when the user is prompted to input a destination, the proficiency determination unit 100 determines that the voice input is unsuccessful.
[0017] As another example, the proficiency assessment unit 100 may display the recognition result for the voice input to the user. The user then inputs whether the displayed recognition result matches what was intended. If the user inputs that the recognition result matches what was intended, the proficiency assessment unit 100 determines that the voice input was successful.
[0018] As yet another example, the proficiency level determination unit 100 determines that the voice input was successful if the user does not attempt voice input again. In this case, if the user attempts voice input again, the proficiency level determination unit 100 determines that the initial voice input was unsuccessful.
[0019] FIG. 2 shows an example of data stored in the storage unit 400 and used in the proficiency determination method of the first example. As shown in FIG. 2(a), the storage unit 400 stores the number of times a user has successfully performed voice input in association with user identification information. The user identification information is information for identifying a user, and is, for example, a character string including alphanumeric characters. As shown in FIG. 2(b), the storage unit 400 may also store a proficiency level corresponding to the number of times the user has successfully performed voice input. The proficiency determination unit 100 then uses this information to determine the proficiency level. Note that in the example shown in FIG. 2(b), the proficiency level is shown as a plurality of levels, for example, three levels of "1," "2," and "3," but the proficiency level may also be a value calculated using the number of times the user has successfully performed voice input.
[0020] <Example 2> The proficiency determination unit 100 determines the proficiency of the speech recognition device 10 using information that can identify the probability (success rate) that the user has successfully input speech. The success rate is, for example, the ratio of the "number of times the user's speech input was successful" to the "number of times the user's speech input was attempted" or the "number of times the user's speech input was unsuccessful." As an example, the information that can identify the success rate includes at least one of the "number of times the user's speech input was attempted," the "number of times the user's speech input was unsuccessful," and the "number of times the user's speech input was successful." Note that the "number of times the user's speech input was attempted" is the sum of the "number of times the user's speech input was successful" and the "number of times the user's speech input was unsuccessful." The proficiency determination unit 100 determines that the higher the user's success rate of speech input, the higher the user's proficiency with the speech recognition device 10. Note that the determination of whether the user's speech input was successful is as described in the first example. Also, in the second example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0021] Although not shown, in the second example, the storage unit 400 stores the success rate of the voice input of the user in association with the user identification information.
[0022] <Example 3> The proficiency determination unit 100 determines the user's proficiency using the time elapsed since the user started using the voice recognition device 10. The proficiency determination unit 100 determines that the longer the time elapsed since the user started using the voice recognition device 10, the higher the user's proficiency with the voice recognition device 10. This time is, for example, the cumulative usage time of the voice recognition device 10 by the user. Note that the time when the user started using the voice recognition device 10 may be, for example, when the user's user identification information is registered in the voice recognition device 10 or when the user makes a voice input for the first time. Also, in the third example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0023] Although not shown, in the third example, the storage unit 400 stores the elapsed time since the user started using the voice recognition device 10 and the accumulated usage time in association with the user identification information.
[0024] <Example 4> The proficiency determination unit 100 determines the user's proficiency using the time from the user's predetermined operation to the start of voice input. The user's predetermined operation is, for example, an operation for starting the voice recognition device 10, such as pressing a predetermined button. The proficiency determination unit 100 determines that the shorter the time from the user's predetermined operation to the start of voice input, the higher the user's proficiency with the voice recognition device 10. Also, in the fourth example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0025] Although not shown, in the fourth example, the storage unit 400 stores the time from a predetermined operation by the user to the start of voice input in association with the user identification information. This time may not be the value for the latest voice input, but may be the average value of that time for voice inputs up to that point, or the average value of that time for a predetermined number of recent inputs (for example, 5 to 10 times).
[0026] The proficiency determination unit 100 may determine the user's proficiency by combining the methods of the above-mentioned first to fourth examples. For example, scores may be associated with the results of the conditions given in the first to fourth examples, and the user's proficiency may be determined using these scores.
[0027] In addition, the proficiency level determination unit 100 may manage the user's proficiency level in multiple stages, and when the specified conditions after the user reaches each stage satisfy the criteria for that stage, change the user's stage to the next higher stage.
[0028] For example, in the first example described above, the initial proficiency level is set to "1," and when the user successfully inputs by voice a predetermined number of times or more, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the user successfully inputs by voice a predetermined number of times or more, the proficiency level is set to "3."
[0029] In the second example described above, for example, the initial proficiency level is set to "1," and when the user's success rate of voice input reaches a predetermined level or higher, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the user's success rate of voice input reaches a predetermined level or higher, the proficiency level is set to "3."
[0030] In the third example described above, for example, the initial proficiency level is set to "1," and when a predetermined time has elapsed since the user started using the voice recognition device 10, the proficiency level is changed to "2." Furthermore, after the proficiency level reaches "2," when a predetermined time has elapsed since the user started using the voice recognition device 10, the proficiency level is changed to "3."
[0031] In the fourth example described above, for example, the initial proficiency level is set to "1," and when the average value of the last five times from the user's predetermined operation to the start of voice input reaches a predetermined value or less, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the average value of the last five times from the user's predetermined operation to the start of voice input reaches a predetermined value or less, the proficiency level is set to "3."
[0032] Furthermore, the management of proficiency levels according to the above multiple stages may be performed by combining multiple conditions. For example, the proficiency determination unit 100 manages the user's proficiency level into four stages: "1," "2," "3," and "4." The proficiency determination unit 100 then sets the user's initial proficiency level to "1," and when a predetermined number of days have passed since the user started using the voice recognition device 10, changes the proficiency level to "2." Next, when the number of times the user successfully inputs voice reaches a predetermined number or more, the proficiency level is changed to "3." Finally, when the user's success rate of voice input reaches a predetermined number or more, the proficiency level is changed to "4."
[0033] [Selection processing unit 200] The selection processing unit 200 allows the user to select a command. A command is an instruction to the car navigation system to perform an operation, such as an instruction to input a destination. Alternatively, a command may be an instruction to set music or an instruction to set air conditioning. When an instruction to input a destination is made, the command may be, for example, the name of the destination, or the genre of the destination or facility such as "restaurant" or "tourist spot," or the purpose of visiting the destination such as "dining" or "lodging." When an instruction to set music is made, the command may be, for example, the name or genre of music to be played, and when an instruction to set air conditioning is made, the command may be an air conditioning operation such as "heating" or "cooling."
[0034] The command may also be related to guidance to a destination. For example, the command may be "zoom in" or "zoom out," which instructs the user to zoom in or out on a map when guiding the user to a destination. Alternatively, the command may be in the form of a question that prompts the speech recognition device 10 to respond, such as "Where should I turn next?" or "How long until we arrive?"
[0035] The selection processing unit 200 sets an output object including at least one command according to the user's level of proficiency. The output object is output to the user in a process described below. As an example, for a user with a high level of proficiency, the selection processing unit 200 sets an output object including a command with a high level of difficulty, which will be described later. In other words, for a user with a low level of proficiency, the selection processing unit 200 sets an output object that does not include a command with a high level of difficulty, which will be described later.
[0036] Then, the selection processing unit 200 outputs at least one set command in a manner that can be recognized by the user. For example, outputting in a manner that can be recognized by the user means causing the screen display device 40 to display the command on the screen, or causing the audio output device 30 to output the command by voice.
[0037] 3 shows an example of commands that the selection processing unit 200 causes to be displayed on the screen display device 40. In the example shown in FIG. 3, the left side of the screen is a speedometer screen 1, and the right side of the screen is a map screen 2. The selection processing unit 200 then displays four commands 3 superimposed on the map screen 2.
[0038] The selection processing unit 200 then acquires the voice input by the user and identifies the command selected by the user by using the results of processing the voice. Specifically, the selection processing unit 200 acquires the user's voice acquired by the voice input device 20 and processes the voice to identify the command selected by the user.
[0039] The selection processing unit 200 will be described in more detail below.
[0040] First, the difficulty of commands will be explained. The difficulty of a command refers to the difficulty of inputting the command by voice. Therefore, it is estimated that for users with low proficiency, it is difficult to input a command with a high level of difficulty by voice, and the possibility of successful voice input is low. Therefore, if a command with a high level of difficulty is set as the output target, it may confuse users with low proficiency. On the other hand, from the perspective of supporting users with low proficiency so that they can input commands with a high level of difficulty, it may be better to display commands with a high level of difficulty to users with low proficiency (set commands with a high level of difficulty as the output target). Below, commands with a high level of difficulty will be explained based on specific examples.
[0041] An example of a difficult command is a command that is linked to at least one alternative expression. Specifically, for example, in a command specifying a "convenience store" as a destination, a command that includes the expression "convenience store" is a difficult command.
[0042] Other examples of difficult commands are commands that require a long time for voice input, such as commands that require input of a long string of characters (e.g., a name) or commands that require multiple selection processes.
[0043] FIG. 4 shows an example of a screen including a command that causes the selection process to be performed multiple times. The screen shown in FIG. 4 is a screen that is displayed when the "convenience store" command is specified in FIG. 3. The screen shown in FIG. 4 displays multiple commands indicating the company name or store name of the "convenience store," and the user further specifies the store name. In this embodiment, the command that causes the selection process to be performed multiple times may be a command with a high level of difficulty. In other words, the "convenience store" command may be a command with a higher level of difficulty than the "home" command in FIG. 3.
[0044] Another example of a command with a high level of difficulty is a command that requires the user to input additional information. Specifically, this is a command displayed as a string of characters such as "Take me to the convenience store in XX-cho XX-chome," "Play XX music," or "Set the air conditioning temperature to XX degrees Celsius," and the user inputs the additional information by entering characters (values) selected by the user in place of the XX. A command that requires the user to input additional information may include a command that is selected in the second or subsequent selection process among commands that cause the above-mentioned selection process to be performed multiple times. For example, the command "Convenience store in XX-cho XX-chome" may be more difficult than the command "convenience store."
[0045] Next, the identification of a command selected by a user will be described. As described above, the selection processing unit 200 identifies a command selected by a user. In this case, the selection processing unit 200 may set only commands set as output targets as the identified commands. In other words, the selection processing unit 200 may identify a command input by the user from only commands set as output targets. More specifically, for a user with low proficiency, the selection processing unit 200 may set only commands set as output targets as the identified commands. In this case, when a user with low proficiency inputs an incorrect voice, it is possible to prevent a command unintended by the user from being identified from commands that are not set as output targets (not displayed to the user). This prevents confusion in the operation of a user with low proficiency. On the other hand, the selection processing unit 200 may set commands other than those set as output targets as the identified commands. For example, since a user with high proficiency can input a large number of commands, it is preferable to allow such users to input more commands than can be displayed on the screen. 3, a user with a high level of proficiency may be able to input commands such as "restaurant" and "meal" (not shown), or commands related to music operation. In other words, the commands set as output targets or the commands output in a manner recognizable to the user may be only a part of the commands that the user can input, rather than all of them.
[0046] Furthermore, for commands that require multiple selection processes, such as those shown in FIG. 4, the command used in the second selection process (e.g., the process of selecting "convenience store A" in the example of FIG. 4) that is performed after the first selection process (e.g., the process of selecting "convenience store A" in the example of FIG. 3) may be set as a candidate command to be identified, regardless of whether it is output in a manner that is recognizable to the user. In particular, the above configuration is preferable because it is thought that highly skilled users may want to be guided to "convenience store A" from the start. An example of this will be described later using FIG. 14.
[0047] In the above explanation, the selection processing unit 200 has been described as setting output targets including difficult commands for users with high proficiency (setting output targets not including difficult commands for users with low proficiency), but the operation of the selection processing unit 200 is not limited to this.
[0048] The selection processing unit 200 may set an output target including a command with a high level of difficulty for a user with a low level of proficiency. In this case, the command with a high level of difficulty is output to the user with a low level of proficiency in a manner that the user can recognize, thereby supporting the user with a low level of proficiency in inputting a command with a high level of difficulty. For example, an output target including the command "Take me to a convenience store in XX town, XX chome" may be set for a user with a low level of proficiency. In this case, for example, a user who has set a convenience store as a destination by entering multiple steps, starting with the command "Take me to a convenience store," can be encouraged to set a convenience store as a destination by entering fewer steps.
[0049] Furthermore, the selection processing unit 200 may set output targets that do not include commands with low difficulty levels for highly proficient users. In this case, even if the commands are not output in a recognizable manner, it is possible to prevent users who are proficient enough to input highly proficient commands from feeling annoyed by the display of the commands.
[0050] [Storage section 400] The storage unit 400 stores the user's proficiency in association with the user's user identification information. The user's proficiency is updated, for example, every time the proficiency determination unit 100 determines the user's proficiency. The proficiency determination unit 100 determines the user's proficiency, for example, at predetermined time intervals. Alternatively, the proficiency determination unit 100 determines the user's proficiency at the time of a predetermined operation (for example, when the voice recognition device 10 is started).
[0051] The storage unit 400 may also store output targets to be set according to the user's proficiency, or a method for setting output targets according to the user's proficiency. Specifically, the storage unit 400 stores rules such as "for a user with a proficiency level equal to or higher than a first predetermined value, include commands with a difficulty level equal to or higher than a second predetermined value in the output targets," "for a user with a proficiency level equal to or higher than a third predetermined value, do not include commands with a difficulty level equal to or lower than a fourth predetermined value in the output targets," "for a user with a proficiency level equal to or lower than a fifth predetermined value, do not include commands with a difficulty level equal to or higher than a sixth predetermined value in the output targets," and "for a user with a proficiency level equal to or lower than a seventh predetermined value, include commands with a difficulty level equal to or higher than an eighth predetermined value in the output targets," as well as the respective predetermined values for the rules. The storage unit 400 may also store rules regarding the number of commands to be output targets and the priority of including each command in the output targets. Furthermore, when the storage unit 400 stores multiple rules, the storage unit 400 preferably stores a priority associated with each rule. In this case, when setting an output target, if a plurality of rules conflict with each other, the selection processing unit 200 sets the output target according to the rule with the highest priority.
[0052] Furthermore, the storage unit 400 stores information for determining the level of proficiency, as described in the description of the proficiency determination unit 100. Furthermore, the storage unit 400 stores a plurality of commands in association with the difficulty level of the commands.
[0053] Furthermore, when there are multiple users, it is preferable that the storage unit 400 stores user voice features for identifying users from input voices in association with user identification information.
[0054] 5 and 6 show an example of information stored in the storage unit 400 in this embodiment. FIG. 5 shows an example of commands stored in the storage unit 400 and the difficulty levels corresponding to the commands. In the example shown in FIG. 5, the storage unit 400 stores the commands "Turn on the air conditioner," "Display a message," "Send a message," "Take me home," and "Take me to a convenience store" in association with the difficulty levels.
[0055] Fig. 6 shows rules, predetermined values, and priorities for setting output targets stored in the storage unit 400. In the example shown in Fig. 6, the storage unit 400 stores four rules associated with priorities: "The number of commands to be output targets shall be 3 or less," "Commands with a difficulty level of 5 shall be set as output targets only for users with a proficiency level of 4 or higher," "Commands with a difficulty level of 4 shall be set as output targets only for users with a proficiency level of 3 or higher," and "Commands with a difficulty level of 1 shall not be set as output targets for users with a proficiency level of 2 or higher."
[0056] Next, Figs. 7-10 show examples of screens that display output targets set for each user's proficiency level when the storage unit 400 stores the information shown in Figs. 5 and 6. Fig. 7 shows a screen that displays output targets set for a user with a proficiency level of 1. For a user with a proficiency level of 1, commands with difficulty levels of 4 and 5 are not set as output targets due to the rules that "commands with difficulty level 5 are set as output targets only for users with a proficiency level of 4 or higher" and "commands with difficulty level 4 are set as output targets only for users with a proficiency level of 3 or higher." Therefore, three commands with difficulty levels ranging from 1 to 3, "Turn on the air conditioner," "Display a message," and "Send a message," are set as output targets.
[0057] FIG. 8 is a screen displaying output targets set for a user with a proficiency level of 2. Even for a user with a proficiency level of 2, commands with difficulty levels of 4 and 5 are not set as output targets due to the rules "commands with difficulty level 5 are set as output targets only for users with a proficiency level of 4 or higher" and "commands with difficulty level 4 are set as output targets only for users with a proficiency level of 3 or higher." Furthermore, due to the rule "commands with difficulty level 1 are not set as output targets for users with a proficiency level of 2 or higher," the command "Turn on the air conditioner," which is a command with difficulty level 1, is also excluded from the output targets. In this way, the voice recognition device 10 according to this embodiment does not display commands with particularly low difficulty for users who are accustomed to operating the voice recognition device 10, thereby reducing annoyance.
[0058] 9 is a screen displaying output targets set for a user with a proficiency level of 3. For a user with a proficiency level of 3, the command "Show me home" is included in the output targets due to the rule that "commands with difficulty level 4 are set as output targets only for users with a proficiency level of 3 or higher." In this way, the speech recognition device 10 according to this embodiment displays new commands with a high level of difficulty to a user who is accustomed to operating the speech recognition device 10.
[0059] FIG. 10 is a screen displaying output targets set for a user with a proficiency level of 4. For a user with a proficiency level of 4, the command "Take me to a convenience store" is included in the output targets according to the rule that "commands with a difficulty level of 5 are set as output targets only for users with a proficiency level of 4 or higher." However, according to the rule that "the number of commands to be output targets shall be 3 or less," one command is excluded from the output targets. In the example shown in FIG. 10, the command with the lowest difficulty level, "Display a message," is excluded from the output targets. In this way, the speech recognition device 10 according to this embodiment displays commands appropriate for each user.
[0060] [Example of hardware configuration of speech recognition device 10] FIG. 11 is a block diagram showing an example of the hardware configuration of the voice recognition device 10. As shown in FIG. Each functional component of the speech recognition device 10 may be realized by hardware (e.g., a hardwired electronic circuit) that realizes each functional component, or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). Below, a case where each functional component of the speech recognition device 10 is realized by a combination of hardware and software will be further described.
[0061] The speech recognition device 10 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method for connecting the processor 1040 and the like to each other is not limited to bus connection.
[0062] The processor 1040 is one of various processors such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, or a read only memory (ROM) or the like.
[0063] The input / output interface 1100 is an interface for connecting the voice recognition device 10 to the above-mentioned voice input device 20, voice output device 30, screen display device 40, and other input / output devices.
[0064] The network interface 1120 is an interface for connecting the speech recognition device 10 to a communication network. This communication network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1120 may be connected to the communication network by wireless connection or by wired connection.
[0065] The storage device 1080 stores program modules that realize each functional component of the speech recognition device 10. The processor 1040 reads each of these program modules into the memory 1060 and executes them to realize the function corresponding to each program module.
[0066] Next, a description will be given of the operation of the voice recognition device 10 according to this embodiment. Fig. 12 is a flow chart showing an example of the operation of the voice recognition device 10 according to this embodiment.
[0067] First, the voice recognition device 10 is started (step S10). The voice recognition device 10 is started, for example, when an engine key is inserted into a vehicle. Alternatively, the voice recognition device 10 is started when a user presses a start button on the voice recognition device 10.
[0068] Next, the skill determination unit 100 determines the user's skill (step S20). The skill determination unit 100 reads information about the user from the storage unit 400, for example, and determines the user's skill.
[0069] Alternatively, if the user's proficiency level has been determined in advance, such as during a previous operation, and the determined proficiency level has been stored in the memory unit 400, the proficiency determination unit 100 may read out the proficiency level stored in the memory unit 400.
[0070] Furthermore, if there are multiple users of the voice recognition device 10, a process of identifying the user may be performed at this timing.
[0071] Next, the selection processing unit 200 sets an output target according to the user's proficiency level (step S30). The selection processing unit 200 may, for example, read out an output target according to the user's proficiency level from the storage unit 400 and set the output target. After that, the speech recognition device 10 waits until it receives a predetermined action when the user performs speech input (step S40). The predetermined action is, for example, pressing a button attached to a steering wheel or the like.
[0072] When a predetermined action is received (step S40: Yes), the selection processing unit 200 displays at least some of the commands to be output in a manner that can be recognized by the user (S50).
[0073] Thereafter, the voice input by the user is processed to identify the command instructed by the user (S60).
[0074] Next, the operation of the speech recognition device 10 according to this embodiment in the above-mentioned steps S50 to S60 will be described using an example of an actually displayed screen.
[0075] Fig. 13 shows an example of a screen that is displayed when waiting for a predetermined operation from the user after the processing of steps S10 to S30. As shown in Fig. 13, at this point, only a speedometer screen 1 on the left side of the screen and a map screen 2 on the right side of the screen are displayed.
[0076] Next, Fig. 14 is a tree diagram showing screens that are displayed after a predetermined action is accepted, but Fig. 14 does not show screens for all selection patterns.
[0077] When a predetermined action is accepted on the screen shown in Fig. 13 (step S40: Yes), a command is displayed superimposed on the map screen 2, as shown on screen T1 in Fig. 14 (step S50). Here, when the command "Command 2" is entered on screen T1, the process proceeds to step S60, where the command "Command 2" is identified.
[0078] On the other hand, in the example shown in Fig. 14, "Command 1" is a command that includes multiple selection processes. For example, like the above-mentioned "convenience store" command, it is a command that, after being input, prompts the user to further select the name of the convenience store company. When "Command 1" is selected, step S50 continues (step S50-2) to display screen T2 that includes the commands "Command 1-1," "Command 1-2," "Command 1-3," and "Command 1-4."
[0079] Then, for example, when the command "Convenience store 1-2" is input on screen T2, step S60 is reached and the command "Command 1-2" is identified. On the other hand, "Command 1-1" is a command that further includes a selection process. When "Command 1-1" is selected, step S50 continues (step S50-3) to display screen T3, which includes the commands "Command 1-1-1," "Command 1-1-2," "Command 1-1-3," and "Command 1-1-4." Finally, the input command is identified on screen T3 (step S60).
[0080] As described above, when a selection process is performed multiple times, in the first selection process, commands used in selection processes performed after the first selection process may be set as inputtable commands in the first selection process. For example, even when screen T1 is displayed, commands 1-1 to 1-4 displayed on screen T2 and commands 1-1-1 to 1-1-4 displayed on screen T3 may also be set as inputtable commands in the first selection process. Alternatively, when a selection process is performed multiple times, after the first selection process, commands used in the second selection process performed after the first selection process may be set as candidate commands to be identified regardless of whether they are output in a manner recognizable to the user. For example, when "Command 1" is selected on screen T1, commands including "Command 1-1," "Command 1-2," "Command 1-3," and "Command 1-4" on screen T2 may be set as inputtable commands before screen T2 is displayed.
[0081] As described above, the voice recognition device 10 according to this embodiment can display commands according to the user's level of proficiency, making it easier for the user to operate.
[0082] [Second embodiment] Next, a speech recognition device 10 according to a second embodiment will be described. This embodiment is the same as the other embodiments except for the following points.
[0083] 15 is a block diagram showing an outline of a speech recognition device 10 according to this embodiment. The speech recognition device 10 according to this embodiment is similar to the information processing device 10 according to the first embodiment, except that it further includes an execution unit 300.
[0084] [Executive Division 300] The execution unit 300 executes a process corresponding to the command identified by the selection processing unit 200. For example, when a command specifying a destination is identified, the execution unit 300 executes "guidance to destination." Specifically, the execution unit 300 activates a navigation function or the like to execute guidance to the destination. Furthermore, when a command specifying music to play is identified, the execution unit 300 executes "play music." Specifically, the execution unit 300 reads music information from the storage unit 400 and activates a speaker or the like to play the music. Furthermore, when a command in the form of a question, such as "Where do I turn next?" or "How long until we arrive?" is identified, the execution unit 300 answers the question. Specifically, the execution unit 300 reads map information and traffic information, etc., to the destination from the storage unit 400 and calculates the next turn, the estimated arrival time, etc. Then, the execution unit 300 creates text data including the calculated information and outputs the created text data via the audio output device 30.
[0085] In this embodiment, the storage unit 400 stores, for example, a predetermined process associated with a command and information used for the process. For example, a command specifying a destination may store a route to the destination.
[0086] As described above, the speech recognition device 10 according to this embodiment can also display commands according to the user's level of proficiency, making it easier for the user to operate. In addition, it can execute processing for the identified commands.
[0087] [Third embodiment] Next, a speech recognition device 10 according to a third embodiment will be described. This embodiment is the same as the other embodiments except for the following points.
[0088] In this embodiment, the proficiency determination unit 100 determines the user's proficiency for each command. For example, some users may be accustomed to inputting commands that specify destinations, but may not be accustomed to inputting commands that specify music settings. In this embodiment, by determining the proficiency for each command and setting an output target, it is possible to set an output target for the user that takes into account the proficiency for each command.
[0089] The method for determining the proficiency level for each command is the same as the method for determining the proficiency level of the voice recognition device 10 described in relation to the first embodiment. That is, the proficiency level for each command may be determined using "the number of times the user has successfully input by voice" or "the success rate of the user's voice input" for each command. Alternatively, similar to the proficiency level of the voice recognition device 10, the proficiency level for each command may be determined using "the elapsed time since the user started using the voice recognition device 10" or "the time from the user's predetermined operation until the start of voice input."
[0090] Furthermore, the level of proficiency for each command may be managed in multiple stages, similar to the level of proficiency for the voice recognition device 10 .
[0091] In this embodiment, the selection processing unit 200 sets output targets using the proficiency level for each command. For example, the selection processing unit 200 sets commands whose proficiency level for each command is equal to or lower than a predetermined level as output targets. In this case, only commands that the user is unfamiliar with can be displayed, thereby efficiently supporting the user in becoming accustomed to inputting new commands.
[0092] Alternatively, the selection processing unit 200 sets commands for which the proficiency level for each command is equal to or higher than a predetermined level as output targets. In this case, commands that the user is unfamiliar with are not displayed, which can prevent the user from becoming confused about the operation. In addition, if the familiar commands are frequently used commands, it is possible to display the frequently used commands.
[0093] In this embodiment, as in the other embodiments, the proficiency level of the speech recognition device 10 may be further used to set the output target.
[0094] In this embodiment, the storage unit 400 stores the proficiency level for each command. The storage unit 400 also stores a method for setting an output target using the proficiency level for each command. The selection processing unit 200 then uses the above information to set an output target.
[0095] As described above, the speech recognition device 10 according to this embodiment can also display commands according to the user's level of proficiency, making it possible to facilitate operation for the user. Furthermore, commands can be displayed taking into account the level of proficiency for each command. [Explanation of symbols]
[0096] 10 Voice recognition device 20 Voice input device 30 Audio output device 40 Screen display device 100 Proficiency Assessment Section 200 Selection processing section 300 Executive Department 400 Storage section 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface
Claims
1. A speech recognition device, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a selection processing unit that allows a user to select a command; Equipped with The selection processing unit, in the selection processing, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; A speech recognition device that acquires speech input by the user and identifies a command selected by the user by using a result of processing the speech.
2. the proficiency level determination unit determines the proficiency level of the user for each of the commands; the selection processing unit sets the output target including the command whose proficiency level is equal to or lower than a predetermined level in the selection process. The speech recognition device according to claim 1 .
3. an execution unit that executes a process corresponding to the command specified by the selection processing unit; 3. The speech recognition device according to claim 1.
4. the selection processing unit sets the command with increasing difficulty as the proficiency level of the user increases.
3. The speech recognition device according to claim 1.
5. A command that takes a long time to input the voice is defined as a command with a high level of difficulty.
5. The speech recognition device according to claim 4.
6. A command associated with at least one paraphrase expression is defined as the command with a high level of difficulty.
5. The speech recognition device according to claim 4.
7. the selection processing unit causes the user to perform the selection process a plurality of times; a command to be used in a second selection process that is performed after the first selection process is set to a command with a high level of difficulty in the first selection process; 5. The speech recognition device according to claim 4.
8. the proficiency level determination unit determines the proficiency level of the user using the number of times the user has successfully input by voice.
3. The speech recognition device according to claim 1.
9. the proficiency level determination unit determines the proficiency level of the user using information capable of identifying a probability that the user has successfully input voice.
3. The speech recognition device according to claim 1.
10. the proficiency level determination unit determines the proficiency level of the user using an elapsed time since the user started using the speech recognition device.
3. The speech recognition device according to claim 1.
11. The proficiency determination unit managing the user's proficiency level in a plurality of stages; changing the level of the user's proficiency to the next higher level when a predetermined condition after the user's proficiency level reaches each level satisfies a criterion for the level; 3. The speech recognition device according to claim 1.
12. the proficiency level determination unit determines the proficiency level of the user using a time period from a predetermined operation to a start of voice input.
3. The speech recognition device according to claim 1.
13. The selection processing unit having the user perform the selection process multiple times; If the user's proficiency level is equal to or higher than a predetermined level, the command used in a second selection process that is performed after the first selection process is included in the candidates for the command to be identified, regardless of whether the command is output in a manner that is recognizable to the user.
3. The speech recognition device according to claim 1.
14. A speech recognition method performed by a computer, comprising: determining a user's proficiency with the speech recognition method; Let the user select a command, In the selection process, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; A speech recognition method for identifying a command selected by the user by acquiring speech input by the user and using a result of processing the speech.
15. A program that causes a computer to function as a speech recognition device, Computer, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a selection processing unit that allows a user to select a command; It functions as The selection processing unit, in the selection processing, setting an output target including at least one of the commands according to the level of proficiency; outputting the set at least one command in a manner recognizable by the user; A program that acquires a voice input by the user and identifies a command selected by the user by using a result of processing the voice.
Citation Information
Patent Citations
Voice recognition device
JP2001282284A