Speech recognition device, speech recognition method, and program
The speech recognition device adapts input and output modes based on user proficiency, addressing redundancy and complexity issues by optimizing operations for varying skill levels.
Patent Information
- Application Number
- JP2024120642
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Existing voice recognition devices are not optimized for user proficiency, leading to redundant operations for experienced users and difficulty for novice users.
A speech recognition device that determines user proficiency through various methods, including successful voice input counts, success rates, and elapsed time, and adjusts input and output modes accordingly to enhance user experience.
The device provides a seamless user experience by tailoring operations to individual proficiency levels, reducing redundancy for experienced users and simplifying interactions for novice users.
Smart Images

Figure 2026019227000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech recognition device, a speech recognition method, and a program. [Background technology]
[0002] In recent years, voice input to information processing devices has become common. For example, voice input is suitable for vehicles, since drivers often have both hands full while driving. For example, Patent Document 1 describes a voice recognition device equipped with a command restriction means for restricting the actual use of certain operation commands from among a large number of operation commands that can be used, and a command increase means for increasing the number of operation commands that a user is permitted to use depending on the user's usage of the voice recognition device. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-282284 Summary of the Invention [Problem to be solved by the invention]
[0004] However, for users who are accustomed to voice input, there is room for improvement in terms of operation, such as the fact that voice input may feel redundant. One example of the problem that the present invention aims to solve is to provide a voice recognition device that is easy for users to operate. [Means for solving the problem]
[0005] The invention described in claim 1 is A speech recognition device, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a setting unit that sets a voice input mode according to the proficiency level; an execution unit that executes voice input in the mode set by the setting unit; The speech recognition device has the following.
[0006] The invention described in claim 12 is A speech recognition method performed by a computer, comprising: determining a user's proficiency with the speech recognition method; setting a voice input mode according to the proficiency level; The speech recognition method includes executing the input of speech in the set manner.
[0007] The invention described in claim 13 is A program that causes a computer to function as a speech recognition device, Computer, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a setting unit that sets a voice input mode according to the proficiency level; an execution unit that executes voice input in the mode set by the setting unit; It is a program that functions as a [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram showing an overview of a voice recognition device according to a first embodiment together with a usage environment. [Figure 2] 4 is a diagram showing an example of data stored in a storage unit and used in a first example of a method for determining a level of proficiency. FIG. [Figure 3] FIG. 10 is a tree diagram showing a second example of an input mode. [Figure 4] FIG. 2 is a block diagram illustrating an example of a hardware configuration of a voice recognition device. [Figure 5] FIG. 3 is a flowchart illustrating an example of the operation of the voice recognition device according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0010] [First embodiment] 1 is a diagram showing an overview of a speech recognition device 10 according to this embodiment together with a usage environment. The speech recognition device 10 is used together with, for example, a speech input device 20, a speech output device 30, and a screen display device 40. The speech input device 20, the speech output device 30, and the screen display device 40 may all be part of the speech recognition device 10, or may be devices separate from the speech recognition device 10. In the latter case, the speech input device 20, the speech output device 30, and the screen display device 40 communicate with the speech recognition device 10 at least when the speech recognition device 10 is used.
[0011] The speech recognition device 10 can be used for various purposes, and as an example, it is used as an input interface for a car navigation device that supports vehicle driving. However, the speech recognition device 10 may also be used for other purposes. In the following description, the speech recognition device 10 will be described as being used as at least a part of a car navigation device.
[0012] The vehicle may be, for example, an automobile, a motorcycle, a moped, or a bicycle, but is not limited to these. When the vehicle is an automobile, for example, the voice recognition device 10 is integrated with the voice input device 20, the voice output device 30, and the screen display device 40 and attached to the vehicle. When the vehicle is a motorcycle, a moped, or a bicycle, for example, the voice input device 20 and the voice output device 30 are separate devices from the voice recognition device 10. The voice input device 20 and the voice output device 30 are attached to, for example, the driver's helmet. However, these devices may also be attached to wearable items other than a helmet, such as eyeglasses or sunglasses.
[0013] The voice recognition device 10 also has a wireless communication function. As an example, the voice recognition device 10 is a smartphone on which a predetermined application is installed. In this case, the microphone of the smartphone functions as the voice input device 20, the speaker of the smartphone functions as the voice output device 30, and the screen of the smartphone functions as the screen display device 40.
[0014] 1, the speech recognition device 10 includes a proficiency determination unit 100, a setting unit 200, an execution unit 300, and a storage unit 400. However, the storage unit 400 may be included inside the speech recognition device 10 or may be provided externally. Each component will be described in detail below.
[0015] [Proficiency Assessment Unit 100] The proficiency determination unit 100 determines the user's proficiency with the speech recognition device 10. The proficiency with the speech recognition device 10 is an index that indicates how familiar the user is with operating the speech recognition device 10. The speech recognition device 10 uses this proficiency to make it easier for the user to use the speech recognition device 10. A specific method by which the proficiency determination unit 100 determines the user's proficiency with the speech recognition device 10 will be described below.
[0016] <Example 1> The proficiency determination unit 100 determines the proficiency of the voice recognition device 10 based on the number of times the user successfully inputs voice. The proficiency determination unit 100 determines that the user's proficiency with the voice recognition device 10 increases as the number of times the user successfully inputs voice increases. For example, the proficiency determination unit 100 determines whether the voice input is successful or not using the recognition result for the input voice. For example, if the recognition result is consistent with the content prompting the user to input, for example, if the result is included in pre-prepared options, the proficiency determination unit 100 determines that the voice input is successful. For example, if the recognition result is consistent with the content prompting the user to input, for example, if the result is included in pre-prepared options, the proficiency determination unit 100 determines that the voice input is successful. For example, if the recognition result is recognized when the user is prompted to input a destination, the proficiency determination unit 100 determines that the voice input is successful. On the other hand, if the recognition result is recognized when the user is prompted to input a destination, the proficiency determination unit 100 determines that the voice input is unsuccessful.
[0017] As another example, the proficiency assessment unit 100 may display the recognition result for the voice input to the user. The user then inputs whether the displayed recognition result matches what was intended. If the user inputs that the recognition result matches what was intended, the proficiency assessment unit 100 determines that the voice input was successful.
[0018] As yet another example, the proficiency level determination unit 100 determines that the voice input was successful if the user does not attempt voice input again. In this case, if the user attempts voice input again, the proficiency level determination unit 100 determines that the initial voice input was unsuccessful.
[0019] FIG. 2 shows an example of data stored in the storage unit 400 and used in the proficiency determination method of the first example. As shown in FIG. 2(a), the storage unit 400 stores the number of times a user has successfully performed voice input in association with user identification information. The user identification information is information for identifying a user, and is, for example, a character string including alphanumeric characters. As shown in FIG. 2(b), the storage unit 400 may also store a proficiency level corresponding to the number of times the user has successfully performed voice input. The proficiency determination unit 100 then uses this information to determine the proficiency level. Note that in the example shown in FIG. 2(b), the proficiency level is shown as a plurality of levels, for example, three levels of "1," "2," and "3," but the proficiency level may also be a value calculated using the number of times the user has successfully performed voice input.
[0020] <Example 2> The proficiency determination unit 100 determines the proficiency of the speech recognition device 10 using information that can identify the probability (success rate) that the user has successfully input speech. The success rate is, for example, the ratio of the "number of times the user's speech input was successful" to the "number of times the user's speech input was attempted" or the "number of times the user's speech input was unsuccessful." As an example, the information that can identify the success rate includes at least one of the "number of times the user's speech input was attempted," the "number of times the user's speech input was unsuccessful," and the "number of times the user's speech input was successful." Note that the "number of times the user's speech input was attempted" is the sum of the "number of times the user's speech input was successful" and the "number of times the user's speech input was unsuccessful." The proficiency determination unit 100 determines that the higher the user's success rate of speech input, the higher the user's proficiency with the speech recognition device 10. Note that the determination of whether the user's speech input was successful is as described in the first example. Also, in the second example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0021] Although not shown, in the second example, the storage unit 400 stores the success rate of the voice input of the user in association with the user identification information.
[0022] <Example 3> The proficiency determination unit 100 determines the user's proficiency using the time elapsed since the user started using the voice recognition device 10. The proficiency determination unit 100 determines that the longer the time elapsed since the user started using the voice recognition device 10, the higher the user's proficiency with the voice recognition device 10. This time is, for example, the cumulative usage time of the voice recognition device 10 by the user. Note that the time when the user started using the voice recognition device 10 may be, for example, when the user's user identification information is registered in the voice recognition device 10 or when the user makes a voice input for the first time. Also, in the third example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0023] Although not shown, in the third example, the storage unit 400 stores the elapsed time since the user started using the voice recognition device 10 and the accumulated usage time in association with the user identification information.
[0024] <Example 4> The proficiency determination unit 100 determines the user's proficiency using the time from the user's predetermined operation to the start of voice input. The user's predetermined operation is, for example, an operation for starting the voice recognition device 10, such as pressing a predetermined button. The proficiency determination unit 100 determines that the shorter the time from the user's predetermined operation to the start of voice input, the higher the user's proficiency with the voice recognition device 10. Also, in the fourth example, as in the first example, the proficiency may be expressed in multiple stages or by a value.
[0025] Although not shown, in the fourth example, the storage unit 400 stores the time from a predetermined operation by the user to the start of voice input in association with the user identification information. This time may not be the value for the latest voice input, but may be the average value of that time for voice inputs up to that point, or the average value of that time for a predetermined number of recent inputs (for example, 5 to 10 times).
[0026] The proficiency determination unit 100 may determine the user's proficiency by combining the methods of the above-mentioned first to fourth examples. For example, scores may be associated with the results of the conditions given in the first to fourth examples, and the user's proficiency may be determined using these scores.
[0027] In addition, the proficiency level determination unit 100 may manage the user's proficiency level in multiple stages, and when the specified conditions after the user reaches each stage satisfy the criteria for that stage, change the user's stage to the next higher stage.
[0028] For example, in the first example described above, the initial proficiency level is set to "1," and when the user successfully inputs by voice a predetermined number of times or more, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the user successfully inputs by voice a predetermined number of times or more, the proficiency level is set to "3."
[0029] In the second example described above, for example, the initial proficiency level is set to "1," and when the user's success rate of voice input reaches a predetermined level or higher, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the user's success rate of voice input reaches a predetermined level or higher, the proficiency level is set to "3."
[0030] In the third example described above, for example, the initial proficiency level is set to "1," and when a predetermined time has elapsed since the user started using the voice recognition device 10, the proficiency level is changed to "2." Furthermore, after the proficiency level reaches "2," when a predetermined time has elapsed since the user started using the voice recognition device 10, the proficiency level is changed to "3."
[0031] In the fourth example described above, for example, the initial proficiency level is set to "1," and when the average value of the last five times from the user's predetermined operation to the start of voice input reaches a predetermined value or less, the proficiency level is set to "2." Furthermore, after the proficiency level reaches "2," when the average value of the last five times from the user's predetermined operation to the start of voice input reaches a predetermined value or less, the proficiency level is set to "3."
[0032] Furthermore, the management of proficiency levels according to the above multiple stages may be performed by combining multiple conditions. For example, the proficiency determination unit 100 manages the user's proficiency level into four stages: "1," "2," "3," and "4." The proficiency determination unit 100 then sets the user's initial proficiency level to "1," and when a predetermined number of days have passed since the user started using the voice recognition device 10, changes the proficiency level to "2." Next, when the number of times the user successfully inputs voice reaches a predetermined number or more, the proficiency level is changed to "3." Finally, when the user's success rate of voice input reaches a predetermined number or more, the proficiency level is changed to "4."
[0033] [Settings section 200] The setting unit 200 sets the mode of voice input to the voice recognition device 10 according to the proficiency determined by the proficiency determination unit 100. This allows the user to use the voice recognition device 10 according to this embodiment comfortably. Setting of the mode of voice input will be described below with a specific example.
[0034] <Example 1> The setting unit 200 sets an input mode for accepting voice input including at least one of the user's abbreviation or paraphrase according to the user's proficiency. For example, for a user with a high level of proficiency, the setting unit 200 sets an input mode for accepting voice input including at least one of the user's abbreviation or paraphrase, and for a user with a low level of proficiency, sets an input mode for not accepting voice input including at least one of the user's abbreviation or paraphrase. Specifically, for a user with a high level of proficiency, the setting unit 200 sets an input mode for accepting voice input including an abbreviation such as "Take me to a convenience store" in addition to the voice input "Take me to a convenience store" for a voice input instructing guidance to a convenience store. Alternatively, for a user with a high level of proficiency, the setting unit 200 sets an input mode for accepting voice input including paraphrases such as "Take me to a restaurant" in addition to the voice input "Take me to a restaurant."
[0035] <Example 2> The setting unit 200 sets an input mode that accepts a voice input that skips a portion of the voice input when multiple voice inputs are required, depending on the user's proficiency. For example, for a highly proficient user, the setting unit 200 sets an input mode that accepts a voice input that skips a portion of the voice input when multiple voice inputs are required, and for a less proficient user, the setting unit 200 sets an input mode that does not accept a voice input that skips a portion of the voice input. Alternatively, for a highly proficient user, the setting unit 200 may accept a voice input in an order different from the predetermined order of the voice inputs when multiple voice inputs are required. Specifically, for example, the voice recognition device 10 may ask the user questions such as "Do you want to use a toll road?" or "Do you want to prioritize main roads?" when receiving a voice input specifying a destination. For example, for a highly proficient user, the setting unit 200 may set an input mode that accepts an input specifying a destination without asking the above questions. Alternatively, a mode may be set in which an input specifying a destination is accepted even when questions such as "Do you want to use toll roads?" or "Do you want to take main roads first?" have been asked.
[0036] The second example will be described in detail using the tree diagram in FIG. 3. First, a specific example of processing for a user with low proficiency will be described. First, upon receiving a predetermined action from the user to start voice input of a destination, such as pressing a predetermined button (step S1), the voice recognition device 10 outputs a voice prompting the user to input whether or not to use a toll road, such as "Do you want to use the toll road?", and the user responds to the voice prompt to input whether or not to use the toll road (step S2). Next, the voice recognition device 10 outputs a voice prompting the user to input whether or not to give priority to main roads, such as "Do you want to give priority to main roads?", and the user responds to the voice prompt to input whether or not to give priority to main roads (step S3). After that, the voice recognition device 10 prompts the user to input a destination, and the user inputs the destination (step S4). The above flow is easy to use for users with low proficiency because it prompts the user to input each piece of information.
[0037] Next, a specific example of processing for a highly proficient user will be described. When a predetermined action to start voice input of a destination is received from the user, such as pressing a predetermined button (step S1), the speech recognition device 10 then receives voice input indicating the destination. That is, the inputs in steps S2 and S3 are omitted. This makes it possible to omit processing that may be perceived as redundant by a highly proficient user.
[0038] However, contrary to the above example, the setting unit 200 may set an input mode for a user with low proficiency to accept a voice input that skips part of the voice input when multiple voice inputs are required. A user with low proficiency may enter a voice input different from the input prompted by the voice recognition device 10. For example, when a voice prompting the user to enter whether or not to use a toll road, such as "Do you want to use a toll road?", is output, the user may enter a destination by voice. In anticipation of such a situation, the setting unit 200 may set an input mode for a user with low proficiency to accept a voice input that skips part of the voice input when multiple voice inputs are required. Alternatively, the setting unit 200 may accept a user with low proficiency to accept a voice input in an order different from the predetermined order of the voice inputs when multiple voice inputs are required.
[0039] In the first and second examples, a user with a high level of proficiency is, for example, a user whose proficiency is equal to or greater than a predetermined reference value, and a user with a low level of proficiency is a user whose proficiency is less than the predetermined reference value. The reference value of proficiency is a predetermined level when proficiency is managed in multiple levels, and a predetermined value when proficiency is managed as a value. The reference value of proficiency is, for example, preset and stored in the storage unit 400.
[0040] In the first and second examples, if the level of proficiency is managed in multiple stages, the input mode may be set according to the multiple stages of proficiency. For example, in the first example, an input mode may be set in which the higher the level of the user's proficiency, the more voice inputs that include at least one of the abbreviated name and paraphrases are accepted.
[0041] The execution unit 300 accepts voice input from the user in accordance with the input mode set by the setting unit 200. For example, the execution unit 300 processes the voice input from the user in accordance with the input mode set by the setting unit 200.
[0042] [Storage section 400] The storage unit 400 stores the user's proficiency in association with the user's user identification information. The user's proficiency is updated, for example, every time the proficiency determination unit 100 determines the user's proficiency. The proficiency determination unit 100 determines the user's proficiency, for example, at predetermined time intervals. Alternatively, the proficiency determination unit 100 determines the user's proficiency at the time of a predetermined operation (for example, when the voice recognition device 10 is started).
[0043] Furthermore, the storage unit 400 may store, for example, depending on the user's level of proficiency, an input mode that should be set for that level of proficiency.
[0044] As described in the description of the skill level determining unit 100, the storage unit 400 also stores information for determining the skill level.
[0045] Furthermore, when there are multiple users, it is preferable that the storage unit 400 stores user voice features for identifying users from input voices in association with user identification information.
[0046] [Example of hardware configuration of speech recognition device 10] FIG. 4 is a block diagram showing an example of the hardware configuration of the voice recognition device 10. As shown in FIG. Each functional component of the speech recognition device 10 may be realized by hardware (e.g., a hardwired electronic circuit) that realizes each functional component, or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). Below, a case where each functional component of the speech recognition device 10 is realized by a combination of hardware and software will be further described.
[0047] The speech recognition device 10 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method for connecting the processor 1040 and the like to each other is not limited to bus connection.
[0048] The processor 1040 is one of various processors such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, or a read only memory (ROM) or the like.
[0049] The input / output interface 1100 is an interface for connecting the voice recognition device 10 to the above-mentioned voice input device 20, voice output device 30, screen display device 40, and other input / output devices.
[0050] The network interface 1120 is an interface for connecting the speech recognition device 10 to a communication network. This communication network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1120 may be connected to the communication network by wireless connection or by wired connection.
[0051] The storage device 1080 stores program modules that realize each functional component of the speech recognition device 10. The processor 1040 reads each of these program modules into the memory 1060 and executes them to realize the function corresponding to each program module.
[0052] Next, a description will be given of the operation of the voice recognition device 10 according to this embodiment. Fig. 5 is a flow chart showing an example of the operation of the voice recognition device 10 according to this embodiment.
[0053] First, the voice recognition device 10 is started (step S10). The voice recognition device 10 is started, for example, when an engine key is inserted into a vehicle. Alternatively, the voice recognition device 10 is started when a user presses a start button on the voice recognition device 10.
[0054] Next, the skill determination unit 100 determines the user's skill (step S20). The skill determination unit 100 reads information about the user from the storage unit 400, for example, and determines the user's skill.
[0055] Alternatively, if the user's proficiency level has been determined in advance, such as during a previous operation, and the determined proficiency level has been stored in the memory unit 400, the proficiency determination unit 100 may read out the proficiency level stored in the memory unit 400.
[0056] Furthermore, if there are multiple users of the voice recognition device 10, a process of identifying the user may be performed at this timing.
[0057] Next, the setting unit 200 sets an input mode according to the user's proficiency level (step S30). The setting unit 200 may, for example, read out an input mode according to the proficiency level from the storage unit 400 and set the input mode. Then, the execution unit 300 processes the voice input by the user according to the set input mode (step S40).
[0058] As described above, the speech recognition device 10 according to this embodiment sets the input mode according to the user's proficiency, so that it is possible to prevent a user with a high level of proficiency from feeling that the operation is redundant, and to prevent a user with a low level of proficiency from feeling that the operation is difficult, thereby providing a speech recognition device 10 that is easy for users to use.
[0059] [Second embodiment] Next, a speech recognition device 10 according to a second embodiment will be described. This embodiment is the same as the other embodiments except for the following points.
[0060] In this embodiment, the setting unit 200 further sets the mode of speech output from the speech recognition device 10 in accordance with the user's proficiency. This allows the user to more comfortably use the speech recognition device 10. Hereinafter, the setting of the speech output mode will be described with a specific example.
[0061] <Example 1> The setting unit 200 sets an output mode in which audio including abbreviations and paraphrases is output according to the user's level of proficiency. For example, the setting unit 200 sets an output mode in which audio including abbreviations and paraphrases is output for a user with a high level of proficiency, and audio including abbreviations and paraphrases is not output for a user with a low level of proficiency. Specifically, for example, the setting unit 200 sets an output mode in which, when starting audio guidance to a convenience store, audio including paraphrases such as "We will guide you to the convenience store" is output for a user with a high level of proficiency, instead of audio output of "We will guide you to the convenience store."
[0062] <Example 2> The setting unit 200 sets an output mode in which the audio playback speed varies depending on the user's proficiency. For example, the setting unit 200 sets an output mode in which the audio playback speed is increased for users with high proficiency and the audio playback speed is normal for users with low proficiency. However, even in this example, it is preferable to slow down the playback speed for important words and words that are likely to be misheard. An important word is, for example, a specific name such as the name of a destination in the audio output when guidance to a destination begins.
[0063] <Example 3> The setting unit 200 sets an output mode that outputs voice with different amounts of information depending on the user's proficiency. For example, the setting unit 200 sets the output mode so that voice with less information is output to a user with high proficiency and voice with a normal amount of information is output to a user with low proficiency. Specifically, for example, the setting unit 200 sets the output mode so that voice output of navigation such as "We will guide you to convenience store C in town A, block B" is output to a user with low proficiency, whereas the setting unit 200 sets the output mode so that voice output of navigation only includes important information such as "We will guide you to convenience store C" is output to a user with high proficiency. In this case, it is possible to prevent a user with high proficiency from feeling that a series of operations is redundant due to the output of unnecessary information. Alternatively, the setting unit 200 may set the output mode so that voice with less information is output to a user with low proficiency and voice with a normal amount of information is output to a user with high proficiency. In this case, it is possible to prevent an excessive amount of information from being output to a user with low proficiency, and to prevent a user with low proficiency from becoming confused during a series of operations.
[0064] In the above first to third examples, when proficiency is managed in multiple stages, the output mode may be set according to the multiple stages of proficiency. For example, in the above first example, an output mode may be set in which the audio output containing at least one of abbreviations or paraphrases increases as the user's proficiency level increases. Also, for example, in the above second example, an output mode in which the audio playback speed increases as the user's proficiency level increases, or an output mode in which fewer terms that slow down the playback speed may be set as the user's proficiency level increases. Also, for example, in the above third example, an output mode in which the amount of information to be output increases or decreases may be set for each stage of the user's proficiency.
[0065] In this embodiment, the storage unit 400 stores, for example, in accordance with the user's level of proficiency, the output mode that should be set for that level of proficiency.
[0066] In this embodiment, the setting unit 200 reads out the output mode from the storage unit 400 and sets the read-out output mode. The execution unit 300 then outputs a voice to the user in accordance with the output mode set by the setting unit 200.
[0067] As described above, the speech recognition device 10 according to this embodiment can also provide a speech recognition device 10 that is easy for users to use. Furthermore, in this embodiment, the speech output mode is further set according to the user's proficiency, so that a speech recognition device 10 that is even easier for users to use can be provided. [Explanation of symbols]
[0068] 10 Voice recognition device 20 Voice input device 30 Audio output device 40 Screen display device 100 Proficiency Assessment Section 200 Settings 300 Executive Department 400 Storage section 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface
Claims
1. A speech recognition device, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a setting unit that sets a voice input mode according to the proficiency level; an execution unit that executes voice input in the mode set by the setting unit; A speech recognition device having:
2. the setting unit uses an abbreviated name in the voice input by the user when the user has a high level of proficiency; 2. The speech recognition device according to claim 1.
3. the setting unit sets a manner of audio output in accordance with the proficiency level.
3. The speech recognition device according to claim 1.
4. the setting unit increases the playback speed of the audio to be output when the user has a high level of proficiency.
4. The speech recognition device according to claim 3.
5. When the setting unit increases the playback speed of the audio to be output, the setting unit reduces the playback speed of a portion of the audio corresponding to a specific content.
5. The speech recognition device according to claim 4.
6. the setting unit uses an abbreviation as the name to be included in the output voice when the user has a high level of proficiency.
4. The speech recognition device according to claim 3.
7. the proficiency level determination unit determines the proficiency level of the user using the number of times the user has successfully input by voice.
3. The speech recognition device according to claim 1.
8. the proficiency level determination unit determines the proficiency level of the user using information capable of identifying a probability that the user has successfully input voice.
3. The speech recognition device according to claim 1.
9. the proficiency level determination unit determines the proficiency level of the user using an elapsed time since the user started using the speech recognition device.
3. The speech recognition device according to claim 1.
10. The proficiency determination unit managing the user's proficiency level in a plurality of stages; changing the level of the user's proficiency to the next higher level when a predetermined condition after the user's proficiency level reaches each level satisfies a criterion for the level; 3. The speech recognition device according to claim 1.
11. the proficiency level determination unit determines the proficiency level of the user using a time period from a predetermined operation to a start of voice input.
3. The speech recognition device according to claim 1.
12. A speech recognition method performed by a computer, comprising: determining a user's proficiency with the speech recognition method; setting a voice input mode according to the proficiency level; A speech recognition method comprising: executing speech input in the set manner.
13. A program that causes a computer to function as a speech recognition device, Computer, a proficiency determination unit that determines a user's proficiency with the speech recognition device; a setting unit that sets a voice input mode according to the proficiency level; an execution unit that executes voice input in the mode set by the setting unit; A program that makes it work.
Citation Information
Patent Citations
Voice recognition device
JP2001282284A