Voice recognition device, voice recognition system, and voice recognition method
The speech recognition system addresses the challenge of recognizing slightly differing utterances by using a two-tiered recognition approach, allowing for accurate recognition without the need for extensive word registration.
Patent Information
- Application Number
- JP2021106166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-25
- Publication Date
- 2025-06-16
- Estimated Expiration
- 2041-06-25
AI Technical Summary
Conventional speech recognition technologies struggle to recognize utterances that slightly differ from registered words, leading to misrecognition and the need to register multiple variations or synonyms, which is impractical.
A speech recognition system comprising a first voice recognition unit and a second voice recognition unit with a higher recognition level, where the first unit outputs recognition results based on a reference value, and a determination unit transmits unrecognizable utterances to the second unit for further recognition.
Enables accurate speech recognition even when utterances do not exactly match registered words, reducing the need for extensive word registration and improving recognition rates without increasing response delays.
Smart Images

Figure 0007692667000001 
Figure 0007692667000002 
Figure 0007692667000003
Abstract
Description
Technical Field
[0001] The present invention relates to a voice recognition technology for recognizing a user's speech.
Background Art
[0002] As a voice recognition technology for recognizing a user's speech, for each word registered in a voice recognition dictionary in advance, a likelihood indicating the plausibility that the word is the word represented by the uttered voice is calculated, and the word with the maximum likelihood is recognized as the word uttered by the user when the likelihood exceeds a predetermined threshold. This technology is known.
[0003] Also, in the voice recognition technology as described above, in order to improve the recognition rate, there is a technology for performing voice recognition by using another voice recognition unit with a high voice recognition level in combination. For example, when a voice recognition terminal cannot recognize the uttered voice of a user, the voice data of the uttered voice is transmitted to a server, the server performs voice recognition using this voice data, and the recognition result is transmitted to the voice recognition terminal (Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Problems in conventional speech recognition technology will be described. For example, assume that "じたくにかえる" is registered as a word in a speech recognition dictionary. Here, when a user utters "じたくへかえる" to utilize the function of the said word, by mispronouncing the "に" part of the registered word as "へ", the likelihood does not exceed the reference value, and the problem occurs that the uttered word cannot be recognized. Also, other than the above, for a registered word "ちかくのこうえん", when uttered as "ちかくのこーえん" with the "こうえん" part pronounced with a long vowel in the same way, the likelihood does not exceed the reference value, so it could not be recognized.
[0006] According to the invention described in Patent Document 1, by performing speech recognition on the speech that could not be recognized using a speech recognition unit with a higher recognition level, even when the utterance is somewhat different from the registered word as described above, it becomes possible to recognize and execute it. However, since all the unrecognizable uttered speech is sent to the server for analysis, there arises a problem that even utterances not related to speech recognition are sent. Also, for example, it is possible to take measures by registering both "じたくへかえる" and "じたくにかえる" in the speech recognition dictionary, but if it is necessary to register synonyms and words corresponding to long vowel pronunciation for all words, a huge number of words need to be registered.
[0007] The present invention has been made to solve such problems, and an object thereof is to provide a speech recognition apparatus, a speech recognition system, and a speech recognition method capable of performing speech recognition even when the input uttered speech does not match the word registered in the speech recognition dictionary.
Means for Solving the Problems
[0008] The voice recognition device according to the present invention includes a microphone, a voice recognition dictionary in which a plurality of words each having a reference value representing the degree of matching are registered, and words registered in the voice recognition dictionary, wherein the microphone picks up the uttered voice, and a first voice recognition unit that outputs, as a recognition result, a word that matches the uttered voice with a degree higher than the degree represented by the reference value set for the word, a second voice recognition unit having a higher recognition level than the first voice recognition unit, a determination unit that determines whether or not a transition of a score indicating the degree of matching with the uttered voice satisfies a predetermined condition in the recognition result by the first voice recognition unit, and a transmission unit that transmits the uttered voice to the second voice recognition unit when it is determined by the determination unit that the transition of the score satisfies the predetermined condition.
[0009] In one embodiment, as the predetermined condition, when at least one transition of a score indicating a temporary deterioration in the degree of matching is detected, the transmission unit transmits the uttered voice to the second voice recognition unit. In one embodiment, as the predetermined condition, when the recognition result by the first voice recognition unit does not satisfy the reference value, the transmission unit transmits the uttered voice to the second voice recognition unit. In one embodiment, the voice recognition dictionary further has a preliminary reference value set to represent a degree of matching lower than the reference value, and as the predetermined condition, when the recognition result by the first voice recognition unit matches with a degree higher than the preliminary reference value set for the word, the transmission unit transmits the uttered voice to the second voice recognition unit. In one embodiment, the second voice recognition unit performs voice recognition based on the uttered voice transmitted from the transmission unit and outputs a recognition result. In one embodiment, the transmission unit transmits text data of the word that matches with a degree higher than the preliminary reference value together with the uttered voice to the second voice recognition unit. In one embodiment, the second voice recognition unit performs voice recognition based on the uttered voice and the text data transmitted from the transmission unit and outputs a recognition result.
[0010] The voice recognition system according to the present invention comprises the voice recognition device described above, and a voice recognition server that receives the uttered voice from the voice recognition device, performs voice recognition, and transmits the recognition result to the voice recognition device. In one embodiment, the second voice recognition unit described above is the voice recognition server. In one embodiment, the voice recognition device outputs the voice recognition result obtained from the voice recognition server.
[0011] The voice recognition method according to the present invention is a voice recognition method comprising a microphone, a first voice recognition unit, and a second voice recognition unit having a higher recognition level than the first voice recognition unit, wherein the first voice recognition unit outputs, as a recognition result, a word that matches the uttered voice picked up by the microphone at a degree equal to or higher than the degree represented by a reference value set for each of a plurality of words registered in a first voice recognition dictionary in which reference values representing the degree of matching are set; determines whether or not a transition of a score indicating the degree of matching with the uttered voice satisfies a predetermined condition in the recognition result by the first voice recognition unit; and, when it is determined in the determining step that the score transition satisfies the predetermined condition, transmits the uttered voice to the second voice recognition unit.
[0012] In one embodiment, as the predetermined condition, when at least one score transition indicating a temporary deterioration in the degree of matching is detected, the transmitting step transmits the uttered voice to the second voice recognition unit. As the predetermined condition, when the recognition result by the first voice recognition unit does not satisfy the reference value, the transmitting step transmits the uttered voice to the second voice recognition unit. In one embodiment, the voice recognition dictionary further has a preliminary reference value set to represent a degree of matching lower than the reference value, and as the predetermined condition, when the recognition result by the first voice recognition unit matches at a degree higher than the preliminary reference value set for the word, the transmitting step transmits the uttered voice to the second voice recognition unit.
Advantages of the Invention
[0013] According to the present invention, when an utterance for which the score at the time of speech recognition satisfies a predetermined condition is input, speech recognition is performed by a speech recognition unit with a high recognition level. Therefore, speech recognition is possible even if the uttered speech does not match the words registered in the speech recognition dictionary. can be done.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0015] Next, embodiments of the present invention will be described in detail with reference to the drawings. Hereinafter, the details of the present invention will be described by taking an example of an in-vehicle device equipped with a speech recognition function according to the present invention.
Examples
[0016] First, the system configuration of the speech recognition system according to the embodiment of the present invention will be described. FIG. 1 is a block diagram showing the system configuration of the speech recognition system according to the embodiment of the present invention. The speech recognition system 10 includes an in-vehicle device 100 and a speech recognition server 300. The in-vehicle device 100 and the speech recognition server 300 transmit and receive information via a network 900 such as the Internet, for example.
[0017] The in-vehicle device 100 is a so-called car navigation mounted on a vehicle, and has functions such as detecting the current position of the vehicle and searching for and guiding a route to a destination set by a user boarding the vehicle. Details of these functions will be omitted. The in-vehicle device 100 does not have to be a device fixedly attached to the vehicle. For example, the in-vehicle device 100 may be a mobile terminal brought into the vehicle by the user.
[0018] The speech recognition server 300 is a cloud server that provides services related to speech recognition of speech collected by a client terminal. A mobile terminal equipped with a virtual assistant, the in-vehicle device 100 according to the present embodiment, etc. function as client terminals for the speech recognition server 300. The speech recognition server 300 performs speech recognition on the speech data of the uttered speech transmitted from the client terminal, and transmits the recognition result to the client terminal. The speech recognition server 300 in this embodiment has abundant hardware resources and high CPU processing power, so its recognition performance is higher than that of the speech recognition function mounted on the stand-alone in-vehicle device 100. Also, the number of words that can be registered is larger, and the freedom of the uttered speech that can be received is higher. Therefore, it is possible to perform recognition on various speech inputs such as combinations of vocabulary and turns of phrase. For example, when executing an operation to turn on the TV, even if it is input with different turns of phrase such as "Turn on the TV" or "TV ON", it is recognized as a word related to "Turn on the TV", and a predetermined operation can be executed.
[0019] Next, the configuration of the in-vehicle device 100 will be described. The in-vehicle device 100 includes a storage unit 110, an input unit 120, an output unit 130, a control unit 140, and a communication unit 150. The storage unit 110 can store application software, programs, road map data, content data such as music and videos, etc. that the in-vehicle device 100 executes. In addition, the storage unit 110 stores a speech recognition dictionary necessary for performing speech recognition. Here, FIG. 2 shows the storage mode of the speech recognition dictionary. As shown in FIG. 2, in the speech recognition dictionary, a plurality of words to be the target of speech recognition, an identification number (No.) of the word, and a reference value Th are registered. Note that the storage mode of the speech recognition dictionary shown in FIG. 2 is an example, and the content of the registered words and the value of the reference value can be changed according to the embodiment.
[0020] The input unit 120 includes voice input, a touch panel, and an input key device. The voice input picks up the voice spoken by the user via a sound collection device such as a microphone. Speech recognition processing is performed based on the picked-up spoken voice. The output unit 130 performs output by display on a screen and by voice.
[0021] The control unit 140 includes a microcontroller or a central processor, etc., and executes a predetermined process based on a control program stored in the storage unit. The communication unit 150 enables wired or wireless data transmission and reception with external devices, external networks, etc. The communication unit 150 can transmit the voice data of the spoken voice to the voice recognition server 300. Further, the communication unit 150 can receive the voice recognition result by the voice recognition server 300.
[0022] In the present embodiment, the control unit 140 executes a speech recognition program 500. FIG. 3 is a diagram showing the functional configuration of the speech recognition program 500 according to the embodiment of the present invention. The speech recognition program 500 includes a speech recognition unit 510, an execution unit 520, a condition determination unit 530, and a data transmission unit 540.
[0023] The voice recognition unit 510 calculates the score of each word stored in the voice recognition dictionary for the uttered voice in parallel with the input of the uttered voice input from the input unit 120.
[0024] Here, the score of each word registered in the voice recognition dictionary for the uttered voice represents a predicted value of the magnitude of the difference between the word and the phrase represented by the uttered voice. The larger the predicted difference, the larger the score.
[0025] More specifically, the score calculation is performed by setting a predetermined initial value as the score, and each time the sound of each voice section (for example, the voice section for each phoneme) of the uttered voice is input, the sound of the voice section and the pronunciation of the corresponding part of each word registered in the voice recognition dictionary for the voice section are calculated for the presence or absence of matching. If they match, the score is decreased by a predetermined value, and if they do not match, the score is increased by a predetermined value. Note that the increase value / decrease value of the score for each voice section of the uttered voice is, for example, the magnitude obtained by multiplying the ratio of the voice section of the word to all voice sections by the initial value of the score.
[0026] Here, a specific example shown in FIG. 4 will be referred to for explaining the score transition in the voice recognition according to the embodiment of the present invention. For example, as shown in FIG. 4(a), the score between the uttered voice "じたくにかえる" and the registered word "じたくにかえる" decreases sequentially during the period when the utterance of "じたくにかえる" of the uttered voice is input. When it exceeds the reference value Th of "じたくにかえる" registered in the voice recognition dictionary shown in FIG. 2, it is recognized that the registered word "じたくにかえる" has been uttered. That is, the reference value Th indicates the similarity. When the similarity exceeds the reference value, it is recognized that the input uttered voice corresponds to any of the registered words in the voice recognition dictionary.
[0027] Also, in the case of the word "じたくにかえる" shown in FIG. 4(a), the voice recognition unit 510 sets the score for the word "じたくにかえる" to be below the reference value Th when the "え" of "じたくにかえる" of the uttered voice is input. At this point, "じたくにかえる" is recognized.
[0028] On the other hand, when the uttered voice "return to the home" shown in Fig. 4(b) is input, during the period when the voice in the part of "home" is input, the score decreases sequentially as in Fig. 4(a), but the score increases at the timing when the voice of "to" is input. This is because the part of "to" in the registered word "return to the home" is different, so the similarity deteriorates and the value of the score transitions in the direction away from the reference value Th. Thereafter, during the period when the voice of "return" is input, it decreases sequentially, but as the final recognition result, it does not exceed the reference value Th. Therefore, the uttered voice "return to the home" is not recognized because it does not fall below the reference value Th of the registered word "return to the home".
[0029] The voice recognition unit 510 performs voice recognition on the input uttered voice by the above-described processing. Note that the voice recognition unit 510 in this embodiment has fewer hardware resources and lower CPU processing capabilities compared to the voice recognition server 300. Also, since there is a limit to the number of words that can be registered in the voice recognition function, the recognition performance and the freedom of the uttered speech that can be accepted are inferior to those of the voice recognition server 300.
[0030] Returning to Fig. 3, the execution unit 520 executes the function corresponding to the registered word of the recognition result of the voice recognition unit 510. When executing, the execution result is output by display or voice via the output unit 130. For example, when the uttered voice is "return to the home", as the recognition result, the route with the home set as the destination is displayed on the map and the route guidance is started. Also, the recognition result of the word obtained from the voice recognition server 300 described later is executed.
[0031] The condition determination unit 530 determines whether it is necessary to transmit the uttered voice to the voice recognition server 300. In this embodiment, when the conditions described later are satisfied, it is determined that it is necessary to transmit the uttered voice to the voice recognition server 300.
[0032] When the data transmission unit 540 is determined by the condition determination unit 530 to satisfy a predetermined condition, it transmits data to the voice recognition server 300 via the communication unit 150 for voice recognition processing at the voice recognition server 300. Here, the data to be transmitted preferably transmits voice data of the uttered voice, or data including text data of registered words having a high similarity in addition to the voice data.
[0033] Then, the voice recognition server 300 performs voice recognition based on various data transmitted by the data transmission unit 540, and transmits the recognition result to the in-vehicle device 100. In one embodiment, the voice recognition server 300 performs voice recognition processing using a voice recognition dictionary held on the voice recognition server 300 side for the transmitted voice data. Then, the recognition result is fed back to the in-vehicle device 100. Also, in one embodiment, the voice recognition server 300 analyzes whether words having a high similarity to the transmitted voice data are synonymous. Then, the recognition result is fed back to the in-vehicle device 100. Regarding this voice recognition processing, it can be changed depending on the form of the voice recognition server 300.
[0034] The in-vehicle device 100 executes a function corresponding to the word by the execution unit 520 based on the recognition result received from the voice recognition server 300.
[0035] Next, the control operation of voice recognition according to the embodiment of the present invention will be described with reference to the flowcharts of FIGS. 4 and 5. In the following description, examples will be described in which "return to the garage" which is a registered word and "return to the garage" which is different from the registered word are each input as the uttered voice. First, the utterance by the user is detected via the input unit 120 (S101). If it is detected that the uttered voice has been input, proceed to the next step.
[0036] When speech is input, the speech recognition unit 510 determines, based on a score, which word registered in the speech recognition dictionary the input speech corresponds to (S103). For example, as shown in Fig. 4(a), when the user speaks "jitaku ni kaeru", the speech recognition unit 510 calculates the score for the registered word "jitaku ni kaeru" in parallel with the speech input by the speech, and when the score exceeds the reference value Th, it recognizes the speech "jitaku ni kaeru".
[0037] As described above, when the score exceeds the reference value Th, the execution unit 520 executes a predetermined process related to the registered word recognized (S105). In this embodiment, since the recognized registered word is "jitaku ni kaeru", processes such as setting the destination of the route guidance to the home are executed.
[0038] Returning to S103, when the score does not exceed the reference value Th, it proceeds to S107. For example, as shown in Fig. 4(b), when the user speaks "jitaku e kaeru", the speech recognition unit 510 similarly calculates the score, and since it does not exceed the reference value Th at the end of the speech, it determines that there is no target registered word and the recognition fails.
[0039] As described above, when the speech recognition by the speech recognition unit 510 fails, the condition determination unit 530 determines whether it is necessary to transmit the input speech to the speech recognition server 300. In this embodiment, it is determined that it is necessary to transmit the input speech to the speech recognition server 300 when two conditions are satisfied.
[0040] Here, the first condition in the embodiment is whether or not the degree of matching set to a preliminary reference value PTh lower than the reference value Th is exceeded as the speech recognition result (S107). Here, as shown in Fig. 4(b), when "jitaku e kaeru" is spoken, the final recognition result does not exceed the reference value Th but exceeds the preliminary reference value PTh. In this case, it can be determined that although recognition is not achieved, speech with a high similarity to the registered word is spoken. Then, upon determining that the preliminary reference value PTh is exceeded, it proceeds to the next step.
[0041] For the uttered speech that does not exceed the preliminary reference value PTh, it can be presumed that speech unrelated to the registered word is being input. Therefore, without determining the second condition described later, the recognition process is terminated or an error display is performed (S118).
[0042] Next, the second condition determines whether there is a timing when the value of the score deteriorates temporarily (S109). As shown in FIG. 4(b), when "じたくへかえる" (return to the desk) is uttered, the score deteriorates and increases temporarily only at the timing when "へ" (to) is input in the uttered speech. When such a score transition is detected, it is determined that the second condition is satisfied, and the process proceeds to the next step. On the other hand, when such a score transition is not detected, the recognition process is terminated or an error display is performed (S118).
[0043] Next, when it is determined by the condition determination unit 530 that the two conditions are satisfied, the data transmission unit 540 transmits the data necessary for performing the speech recognition process to the speech recognition server 300 via the communication unit 150 (S111). Here, the data transmitted by the data transmission unit 540 includes, in this embodiment, the speech data of the uttered speech and the text data of the registered word that exceeds the preliminary reference value PTh. For example, taking FIG. 4(b) as an example, the speech data when "じたくへかえる" is uttered and the text data of the word "じたくにかえる" (return to the desk) that exceeds the preliminary reference value PTh are transmitted to the speech recognition server 300.
[0044] The speech recognition server 300 performs speech recognition based on the speech data and text data transmitted by the data transmission unit 540. In this embodiment, the speech recognition server 300 determines whether the speech data and the text data are synonymous, and when it is determined that the speech data is synonymous with the text data, the result is fed back to the in-vehicle device 100. For example, when the speech data "じたくへかえる" and the text data "じたくにかえる" are transmitted, the speech recognition server 300 analyzes whether the speech data "じたくへかえる" is synonymous with the text data "じたくにかえる".
[0045] Return to S111 and receive, via the communication unit 150, the determination result as to whether or not the uttered voice transmitted from the voice recognition server 300 is synonymous (S113). When the voice recognition server 300 receives a recognition result indicating that the voice data and the text data are synonymous, the registered word determined to be synonymous is transmitted to the execution unit 540 (S115). Then, the execution unit 540 executes the function corresponding to the word (S105).
[0046] In S115, when the voice recognition server 300 receives a recognition result indicating that the voice data and the text data are not synonymous, the recognition process is terminated or an error display is performed (S118).
[0047] As described above, even when uttering different paraphrases for the words registered in the voice recognition dictionary, it is possible to detect a mispronunciation from the transition of the score and transmit only the utterances that failed to be recognized due to the mispronunciation to the voice recognition server 300 with a high recognition level for voice recognition. Therefore, even paraphrases that could not be recognized by conventional recognition can now be recognized. As a result, it is no longer necessary to register all synonyms in the dictionary. Also, a stand-alone type voice recognition device has the drawback of having a high response speed but low recognition performance, while a server type voice recognition device has the drawback that the response speed decreases due to data transmission and reception. However, as in this embodiment, since only specific utterances are recognized by the server type voice recognition device, it is possible to achieve both an improvement in the recognition rate and suppression of the delay in response.
[0048] Also, in this embodiment, the condition determination unit 530 determines two conditions in order. However, depending on the embodiment, the determination order may be reversed, or the conditions may be determined simultaneously. Further, in a certain embodiment, the condition determination unit 530 may determine only by the determination based on the score transition, which is the second condition, whether it is necessary to transmit the uttered voice to the speech recognition server. Thereby, for example, it becomes possible to make a determination even when ambient noise is mixed and the preliminary reference value PTh is not exceeded, and it is also possible to suppress all uttered voices that do not exceed the reference value Th from being transmitted to the speech recognition server.
[0049] Next, in this embodiment, synonymous words such as "じたくへかえる" (return to the apartment) were used as examples for explanation, but it is also possible to handle prolonged sounds. Here, FIG. 6 shows a method of uttered voice recognition processing when a prolonged sound word is uttered. FIG. 6(a) shows the speech recognition result by the speech recognition unit 510 when the uttered voice "ちかくのえいがかん" (nearby movie theater) is uttered when the word "ちかくのえいがかん" is registered in the speech recognition dictionary. Since the same uttered voice as the registered word is input, the score decreases each time it is uttered, and speech recognition is successful by exceeding the reference value Th.
[0050] Next, FIG. 6(b) shows the speech recognition result by the speech recognition unit 510 when the uttered voice "ちかくのえーがかん" (with the prolonged sound of "い") is uttered when the word "ちかくのえいがかん" is registered in the speech recognition dictionary. Since the "い" part of the registered word is prolonged for the voice input, only that part has a deteriorated score. And due to the deterioration of the score, the reference value Th is not exceeded, so speech recognition fails.
[0051] Even when the uttered speech sounds as described above are prolonged, similar to the speech recognition method for synonyms shown in the above embodiment, the condition determination unit 530 detects whether the score exceeds the preliminary reference value PTh and whether there is no timing at which the score temporarily deteriorates, transmits the voice data and text data to the speech recognition server to perform high-precision speech recognition processing, and receives the determination result as to whether the uttered speech is synonymous with the registered word, thereby suppressing recognition failures.
[0052] Also, in this embodiment, an example where there is only one timing at which the score deteriorates as shown in FIGS. 4(b) and 6(b) has been described, but in a certain embodiment, there may be a plurality of locations where the score temporarily deteriorates. Here, FIG. 7 shows a method of uttered speech recognition processing when there are a plurality of locations where the score temporarily deteriorates. FIG. 7(a) shows the speech recognition result by the speech recognition unit 510 when the word "えいがかんにいく" is registered in the speech recognition dictionary and "えいがかんにいく" is uttered. Since the same speech as the registered word is being uttered, the score decreases with each utterance and exceeds the reference value Th, so the speech recognition is successful.
[0053] Next, FIG. 7(b) shows the speech recognition result by the speech recognition unit 510 when the word "えいがかんにいく" is registered in the speech recognition dictionary and the uttered speech is misspoken and input as "えーがかんへいく" due to prolongation. Since the "い" part of the registered word is prolonged and the "に" part is misspoken as "へ" for the speech input, the score deteriorates in each part. And due to this deterioration of the score, the reference value Th is not exceeded, so the speech recognition fails.
[0054] Even in a case where there are a plurality of locations where the score temporarily deteriorates as described above, by performing the same speech recognition processing as in the above embodiment, it is possible to recognize and execute the uttered speech even with some differences in the way of speaking.
[0055] In this embodiment, the data transmission unit 540 transmitted voice data and text data to the voice recognition server 300. However, in certain embodiments, only the voice data may be transmitted, and the voice recognition server 300 may perform voice recognition based on the received voice data and feedback the recognition result to the in-vehicle device 100.
[0056] In this embodiment, a voice recognition system using an external voice recognition server has been described as an example. However, in certain embodiments, the present invention is applicable even if the in-vehicle device 100 is equipped with two voice recognition devices with different voice recognition levels. For example, the in-vehicle device 100 performs voice recognition using the voice recognition device with a lower recognition level. When the recognition fails due to differences in expressions as in the above embodiments, based on the determination result by the condition determination unit 530, the other voice recognition device with a higher recognition level is made to perform voice recognition, thereby reducing unnecessary voice recognition processing and the waiting time caused thereby.
[0057] In this embodiment, communication was directly performed with the voice recognition server 300. However, in certain embodiments, for example, the in-vehicle device 100 is communicatively connected to a mobile terminal such as a mobile phone equipped with a virtual assistant. When the recognition fails due to differences in expressions as in the above embodiments in the in-vehicle device 100, the uttered voice may be transmitted to the mobile terminal, and the recognition result may be received using the voice recognition function on the mobile terminal side.
[0058] In this embodiment, the in-vehicle device 100 has been described as an example. However, in certain embodiments, for example, it may be a portable terminal equipped with a voice recognition function capable of cooperating with an external voice recognition server. In this case, the voice recognition function provided on the portable terminal side is usually used, and when the recognition fails due to differences in expressions as in the above embodiments, the uttered voice may be transmitted to the external voice recognition server to receive the recognition result.
[0059] Although the preferred embodiments of the present invention have been described in detail, the present invention is not limited to specific embodiments, and various modifications and changes are possible within the scope of the gist of the invention described in the claims.
Explanation of Signs
[0060] 10 Voice recognition system 100 Vehicle-mounted device 300 Voice recognition server 500 Voice recognition program 900 NW
Claims
1. A microphone, A voice recognition dictionary in which a plurality of words each set with a reference value representing the degree of matching for recognizing the input spoken voice are registered, A first voice recognition unit that calculates a score for evaluating the degree of matching between a word registered in the voice recognition dictionary and a word registered for each input voice of the spoken voice picked up by the microphone, and outputs, as a recognition result, a word for which the score satisfies the reference value, A second voice recognition unit having a higher recognition level than the first voice recognition unit, A determination unit that determines whether or not the transition of the score satisfies a predetermined condition in the recognition result by the first voice recognition unit, A transmission unit that transmits the spoken voice to the second voice recognition unit when it is determined by the determination unit that the transition of the score satisfies the predetermined condition, As the predetermined condition, when at least one transition of the score indicating a temporary deterioration in the degree of matching is detected in the score for each input voice of the spoken voice by the first voice recognition unit, the transmission unit transmits the spoken voice to the second voice recognition unit. A voice recognition device.
2. As the predetermined condition, when at least one transition of the score indicating a temporary deterioration in the degree of matching is detected in the score for each input voice of the spoken voice by the first voice recognition unit, and the recognition result by the first voice recognition unit does not satisfy the reference value, the transmission unit transmits the spoken voice to the second voice recognition unit. The voice recognition device according to claim 1.
3. The voice recognition dictionary further has a preliminary reference value set to represent a degree of matching lower than the reference value, and as the predetermined condition, when the recognition result by the first voice recognition unit matches at a degree higher than the preliminary reference value set for the word, the transmission unit transmits the spoken voice to the second voice recognition unit. The voice recognition device according to claims 1 and 2.
4. The second voice recognition unit executes voice recognition based on the uttered voice transmitted from the transmission unit and outputs a recognition result. The voice recognition device according to any one of claims 1 to 3.
5. The transmission unit transmits text data of the word that matches to a higher degree than the preliminary reference value to the second voice recognition unit together with the uttered voice. The voice recognition device according to claim 3.
6. The second voice recognition unit executes voice recognition based on the uttered voice and the text data transmitted from the transmission unit and outputs a recognition result. The voice recognition device according to claim 5.
7. A voice recognition system comprising: the voice recognition device according to any one of claims 1 to 6; and a voice recognition server that receives the uttered voice from the voice recognition device, performs voice recognition, and transmits the recognition result to the voice recognition device.
8. The second voice recognition unit is the voice recognition server. The voice recognition system according to claim 7.
9. The voice recognition device outputs a voice recognition result acquired from the voice recognition server. The voice recognition system according to claims 7 and 8.
10. A voice recognition method including a microphone, a first voice recognition unit, and a second voice recognition unit having a higher recognition level than the first voice recognition unit, A step of calculating a score for evaluating the degree of matching of the uttered voice picked up by the microphone with the words registered for each input voice from a voice recognition dictionary in which a plurality of words are registered, each having a reference value representing the degree of matching for recognizing the input uttered voice by the first voice recognition unit, and outputting, as a recognition result, the words for which the score satisfies the reference value; A step of determining whether or not a transition of the score indicating the degree of matching with the uttered voice satisfies a predetermined condition in the recognition result by the first voice recognition unit; When it is determined in the determining step that the transition of the score satisfies the predetermined condition, a step of transmitting the uttered voice to the second voice recognition unit; As the predetermined condition, when at least one transition of the score indicating a deterioration in the degree of temporary matching is detected in the score for each input sound of the uttered voice by the first voice recognition unit, the transmitting step is to transmit the uttered voice to the second voice recognition unit, a voice recognition method.
11. As the predetermined condition, when at least one transition of the score indicating a deterioration in the degree of temporary matching is detected in the score for each input sound of the uttered voice by the first voice recognition unit, and the recognition result by the first voice recognition unit does not satisfy the reference value, the transmitting step is to transmit the uttered voice to the second voice recognition unit, the voice recognition method according to claim 10.
12. The voice recognition dictionary further has a preliminary reference value set to represent a degree of matching lower than the reference value. As the predetermined condition, when the recognition result by the first voice recognition unit matches at a degree higher than the preliminary reference value set for the word, the transmitting step is to transmit the uttered voice to the second voice recognition unit, the voice recognition method according to claims 10 and 11.
Citation Information
Patent Citations
Voice recognition device
JP2000099082A
Speech recognition service mediation system, speech recognition master reference method used for same, program, and storage medium storing same program
JP2004184858A
Voice recognition apparatus
JP2011170087A
Voice recognition system, electronic device, server, voice recognition method and voice recognition program
JP2015102805A
Voice recognition client device and server-type voice recognition device
JP2016095383A