Speech recognition device and program

The speech recognition device addresses the challenge of recognizing unknown words by using an unknown word dictionary and prompt input system, allowing for accurate recognition without retraining the neural networks, thus enhancing the precision of speech recognition systems.

JP2025071488APending Publication Date: 2025-05-08NIPPON HOSO KYOKAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023181692
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

End-to-end speech recognition systems struggle with recognizing unknown words, such as names or place names, that were not encountered during training, due to the need for retraining the neural networks, which is computationally expensive and may result in reduced cognitive ability.

Method used

A speech recognition device and program that utilize an unknown word dictionary to map speech expression symbols for unknown words to recognition result symbols, with a prompt input unit generating prompts for the decoder, allowing for accurate recognition without extensive relearning.

Benefits of technology

Enables high-precision recognition of unknown words without the need for retraining the neural networks, improving the accuracy of speech recognition systems by directly mapping unknown words to their corresponding recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071488000001_ABST
    Figure 2025071488000001_ABST
Patent Text Reader

Abstract

To provide a speech recognition device and program to enable recognition of unknown words without need for relearning.SOLUTION: An unknown word dictionary storage unit stores dictionary data representing correspondence between phonetic expression symbol strings and symbol strings for recognition result output with respect to unknown words. An unknown word prompt input unit generates and outputs a prompt containing the phonetic expression symbol strings with respect to the unknown word by referring to dictionary data stored in the unknown word dictionary storage unit. A speech recognition decoder unit obtains and outputs recognition result text based on encoding result information output from a speech recognition encoder unit and the prompt. When the recognition result text contains a special token indicating that it is an unknown word unit, a dictionary utilization unit replaces parts of the special token in the recognition result text with symbol strings for recognition result output obtained from the dictionary data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a voice recognition device and a program. [Background technology]

[0002] In end-to-end speech recognition using a neural network, a network with a very simple structure is constructed, in which input and output are connected by a single deep learning model. By learning this single model with a simple structure, tasks based on the relationship between input and output can be realized. In other words, the device can be constructed more simply than conventional speech recognition methods that consist of multiple models such as acoustic models and language models. In end-to-end speech recognition, learning is performed by defining the correct vocabulary in advance.

[0003] Non-Patent Document 1 discloses a technique for providing a prompt indicating the relationship between an expression in a pre-translation language (source language) and an expression in a post-translation language (target language) (e.g., "Japan means Japan") in a neural machine translation process using a large-scale language model. It is said that the use of the technique disclosed in Non-Patent Document 1 improves the accuracy of language translation.

[0004] Non-Patent Document 2 proposes a domain adaptation method that improves the accuracy of speech recognition in a specific domain by using domain-specific text prompts.

[0005] Non-Patent Document 3 describes an end-to-end voice recognition technology. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Hongyuan Lu, Haoyang Huang, Dongdong Zhang, Haoran Yang, Wai Lam, Furu Wei, “Chain-of-Dictionary Prompting Elicits Translation in Large Language Models”, arXiv:2305.06575v3 [cs.CL], 24 May 2023. [Non-Patent Document 2] Yuang Li,Yu Wu,Jinyu Li,Shujie Liu,“PROMPTING LARGE LANGUAGE MODELS FOR ZERO-SHOT DOMAIN ADAPTATION IN SPEECH RECOGNITION”,arXiv:2306.16007v1 [cs.CL],28 Jun 2023. [Non-Patent Document 3] Article "End-to-End Speech Recognition", URL https: / / olaris.jp / poststag / Ds00Wfma, November 19, 2021. Summary of the Invention [Problem to be solved by the invention]

[0007] Due to the characteristic of end-to-end speech recognition processing, which performs learning by predefining the correct vocabulary, there are many cases where unknown words that did not appear during learning (such as people's names or place names) cannot be recognized correctly when they appear during inference. However, depending on the situation in which speech recognition processing is performed, people's names and place names may be important recognition targets in which errors are particularly unacceptable. Because it is end-to-end processing, if an attempt is made to train the neural network to add new vocabulary, a huge amount of computational cost is required to retrain the neural network's enormous internal parameters. In addition, there is a risk that the recognition ability obtained during previous training will be significantly lost.

[0008] The present invention has been made based on the above-mentioned problem recognition, and aims to provide a speech recognition device and a program that enable recognition of unknown words without requiring extensive re-learning. [Means for solving the problem]

[0009] [1] In order to solve the above problem, a speech recognition device according to one aspect of the present invention includes an unknown word dictionary storage unit that stores dictionary data indicating a correspondence between a phonetic expression symbol string for an unknown word and a recognition result output symbol string, an unknown word prompt input unit that generates and outputs a prompt including a phonetic expression symbol string for an unknown word by referring to the dictionary data stored in the unknown word dictionary storage unit, a speech recognition encoder unit that inputs a voice and outputs encoding result information corresponding to the voice, a speech recognition decoder unit that obtains and outputs a recognition result text based on the encoding result information output from the speech recognition encoder unit and the prompt passed from the unknown word prompt input unit, and a dictionary utilization unit that, when the recognition result text output from the speech recognition decoder unit includes a special token indicating an unknown word portion, obtains a recognition result output symbol string corresponding to the special token by referring to the dictionary data stored in the unknown word dictionary storage unit and replaces the portion of the special token in the recognition result text with the recognition result output symbol string.

[0010] [2] Moreover, one aspect of the present invention is a speech recognition device according to [1] above, wherein the speech recognition encoder unit and the speech recognition decoder unit are each configured using a neural network, and are configured so that values ​​of internal parameters of the speech recognition encoder unit and the speech recognition decoder unit can be updated based on learning data.

[0011] [3] Also, one aspect of the present invention is a speech recognition device according to [2] above, in which in a first learning mode, learning is performed using a set of pairs of a speech expression symbol string input to the speech recognition decoder unit and a recognition result output symbol string that is a correct answer output from the speech recognition decoder unit as learning data; if there is no prompt passed from an unknown word prompt input unit to the speech recognition decoder unit, the correct recognition result output symbol string does not include a special token indicating that it is an unknown word portion; if there is a prompt passed from the unknown word prompt input unit to the speech recognition decoder unit and the speech expression symbol string input to the speech recognition decoder unit includes a word corresponding to the prompt, the correct recognition result output symbol string includes a special token indicating that it is an unknown word portion; and if there is a prompt passed from the unknown word prompt input unit to the speech recognition decoder unit and the speech expression symbol string input to the speech recognition decoder unit does not include a word corresponding to the prompt, the correct recognition result output symbol string does not include a special token indicating that it is an unknown word portion.

[0012] [4] Furthermore, in one aspect of the present invention, in the speech recognition device of [3] above, in a second learning mode, learning is performed using a set of pairs of speech input to the speech recognition encoder unit and the recognition result text which is the correct output from the speech recognition decoder unit as learning data, and the recognition result text is a pair of the speech expression symbol string and the symbol string for outputting the recognition result.

[0013] [5] Also, one aspect of the present invention is a speech recognition device according to any one of [1] to [4] above, wherein the dictionary data stored in the unknown word dictionary storage unit represents a correspondence relationship between a speech expression symbol string and a recognition result output symbol string for a plurality of unknown words, the unknown word prompt input unit generates and outputs the prompt associated with the special token specific to each of the plurality of unknown words, the speech recognition decoder unit obtains and outputs a recognition result text based on the prompt passed from the unknown word prompt input unit, and when the recognition result text output from the speech recognition decoder unit includes a special token indicating an unknown word portion, the dictionary utilization unit identifies which unknown word among the plurality of unknown words the special token corresponds to and refers to the dictionary data to obtain a recognition result output symbol string corresponding to the special token, and replaces the portion of the special token in the recognition result text with the recognition result output symbol string.

[0014] [6] Also, one aspect of the present invention is that in the speech recognition device of [5] above, the special token specific to each of the plurality of unknown words includes index value information indicating the position of the specific unknown word in the dictionary data.

[0015] [7] Also, one aspect of the present invention is a program that causes a computer to function as a speech recognition device including: an unknown-word dictionary storage unit that stores dictionary data indicating a correspondence between a phonetic expression symbol string for an unknown word and a recognition-result output symbol string; an unknown-word prompt input unit that generates and outputs a prompt including a phonetic expression symbol string for an unknown word by referring to the dictionary data stored in the unknown-word dictionary storage unit; a speech recognition encoder unit that inputs speech and outputs encoding result information corresponding to the speech; a speech recognition decoder unit that obtains and outputs a recognition result text based on the encoding result information output from the speech recognition encoder unit and the prompt passed from the unknown-word prompt input unit; and a dictionary utilization unit that, when the recognition result text output from the speech recognition decoder unit includes a special token indicating an unknown word portion, obtains a recognition result output symbol string corresponding to the special token by referring to the dictionary data stored in the unknown-word dictionary storage unit and replaces the portion of the special token in the recognition result text with the recognition result output symbol string. Effect of the Invention

[0016] According to the present invention, the speech recognition decoder performs a decoding process based on the prompt passed from the unknown word prompt input unit to obtain the recognition result text. When the recognition result text output from the speech recognition decoder contains a special token indicating an unknown word portion, the dictionary utilization unit refers to the dictionary data and replaces the special token with a word (symbol string for outputting the recognition result) registered in the dictionary data. This makes it possible to output the recognition result text with high accuracy without the need to re-train the speech recognition encoder and the speech recognition decoder regarding unknown words. [Brief description of the drawings]

[0017] [Figure 1] 1 is a block diagram showing a schematic functional configuration of a voice recognition device according to an embodiment of the present invention. [Diagram 2]2 is a schematic diagram showing an example of a more detailed functional configuration for realizing a voice recognition encoder unit and a voice recognition decoder unit in the voice recognition device according to the embodiment. FIG. [Diagram 3] 2 is a schematic diagram showing an example of the configuration of dictionary data stored in an unknown word dictionary storage unit in the embodiment. FIG. [Figure 4] FIG. 2 is a schematic diagram for explaining a pre-learning pattern (first learning mode of the voice recognition device) of a voice recognition decoder unit (decoder) in the embodiment. [Diagram 5] 3 is a schematic diagram showing the flow of data when the encoder-decoder type end-to-end speech recognition model (FIG. 2) in the embodiment is operated in the first learning mode (learning using only text). FIG. [Figure 6] 13 is a schematic diagram showing an example of a pair of input data and correct answer data in a second learning mode (learning based on voice) in the embodiment. FIG. [Figure 7] 10 is a schematic diagram showing an example of input and output when inference is performed using a trained speech recognition encoder unit and a trained speech recognition decoder unit according to the embodiment. FIG. [Figure 8] 13 is a flowchart showing a procedure of processing during inference performed by the voice recognition device according to the embodiment. [Figure 9] 2 is a block diagram showing an example of the internal configuration of a device for realizing the voice recognition device according to the embodiment. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] Next, an embodiment of the present invention will be described with reference to the drawings.

[0019] The speech recognition device 1 of this embodiment has a means for storing information of an unknown word dictionary as a means for correctly recognizing unknown words. The unknown word dictionary in this embodiment is information that indicates the correspondence between phonetic notation and ideographic notation. For example, in the case of Japanese, the phonetic notation is expressed, for example, in katakana and corresponds to the reading (way of speaking). The text in phonetic notation may be referred to as a "phonetic expression symbol string" below. The ideographic notation is a notation when displaying the recognition result of the speech. The ideographic notation is a notation that mixes kana and kanji. However, the kana and kanji mixed notation may be a notation that uses only kanji, may be a notation that uses only kana, or may be a notation that uses a mixture of kanji and kana. The text in ideographic notation may be referred to as a "symbol string for outputting recognition results" below. In the case of targeting languages ​​other than Japanese, the speech recognition device 1 holds and uses the correspondence between the "phonetic expression symbol string" and the "symbol string for outputting recognition results" as dictionary data in the same manner as above. For example, in the case of English, the same pronunciation may be written differently for outputting the recognition result (for example, Stephen and Steven).

[0020] FIG. 1 is a block diagram showing a schematic functional configuration of a voice recognition device according to this embodiment. As shown in the figure, the voice recognition device 1 includes a voice input unit 11, a voice recognition encoder unit 12, a voice recognition decoder unit 13, an unknown word dictionary storage unit 14, an unknown word prompt input unit 15, a dictionary utilization unit 16, and a text output unit 17. These functions constituting the voice recognition device 1 can be realized, for example, by a computer and a program. Furthermore, each function stores information using a storage means as necessary. The storage means is, for example, a variable in a program or a memory allocated by the execution of a program. Furthermore, non-volatile storage means such as a magnetic hard disk device or a solid state drive (SSD) may be used as necessary. Furthermore, at least a part of the functions of each functional unit may be realized as a dedicated electronic circuit rather than a program.

[0021] The voice input unit 11 acquires data of a voice to be recognized from outside, and passes it to the voice recognition encoder unit 12. The voice input unit 11 passes time-series data of an acoustic feature, such as a log-mel spectrogram, to the voice recognition encoder unit 12. The voice input unit 11 may calculate the above acoustic feature based on an acquired voice waveform, or may acquire information on the above acoustic feature from outside.

[0022] The voice recognition encoder unit 12 calculates a predetermined numerical vector based on the acoustic feature quantity passed from the voice input unit 11, and passes the numerical vector to the voice recognition decoder unit 13. This numerical vector may be called a "state vector." The voice recognition encoder unit 12 is realized, for example, by using a neural network, and is configured to be able to adjust the values ​​of parameters inside the network by performing machine learning.

[0023] The voice recognition decoder unit 13 outputs text regarding the voice recognition result, based on the numerical vector passed from the voice recognition encoder unit 12 and the prompt (instruction information) passed from the unknown language prompt input unit 15. The voice recognition decoder unit 13 is realized, for example, using a neural network, and is configured to make it possible to adjust the parameter values ​​inside the network by performing machine learning.

[0024] The voice recognition process is realized by cooperation between the voice recognition encoder unit 12 and the voice recognition decoder unit 13. A more detailed internal configuration example of the voice recognition encoder unit 12 and the voice recognition decoder unit 13 will be described later with reference to another figure.

[0025] The unknown word dictionary storage unit 14 stores information on the correspondence between kana notation (sound expression symbol string) and kanji-kana mixed notation (symbol string for outputting recognition result) as dictionary information on unknown words. The configuration of the dictionary stored in the unknown word dictionary storage unit 14 will be described later with reference to another figure.

[0026] The unknown word prompt input unit 15 generates an unknown word prompt by referring to the unknown word dictionary storage unit 14, and passes the unknown word prompt to the speech recognition decoder unit 13. When the speech recognition result includes a predetermined phonetic expression symbol string, the unknown word prompt has the effect of instructing the speech recognition decoder unit 13 to output the part corresponding to the unknown word as a special token ([OOV] described later) corresponding to the unknown word. Specifically, the unknown word prompt input unit 15 creates a prompt using the phonetic expression symbol string (kana notation) of the unknown word registered in the unknown word dictionary storage unit 14, and inputs it to the speech recognition decoder unit 13. The format of the input data to the speech recognition decoder unit 13 will be explained later.

[0027] When the text (speech recognition result) output by the speech recognition decoder unit 13 includes a special token ([OOV] described later), the dictionary utilization unit 16 replaces the special token with a word in mixed kanji and kana notation (symbol string for outputting the recognition result). The dictionary utilization unit 16 acquires the word in mixed kanji and kana notation (symbol string for outputting the recognition result) by referring to the dictionary data stored in the unknown word dictionary storage unit 14.

[0028] The text output unit 17 outputs the text of the speech recognition result to the outside. The text of the speech recognition result is based on the output from the speech recognition decoder unit 13. However, if the output from the speech recognition decoder unit 13 contains a special token ([OOV]) corresponding to an unknown word, the text output unit 17 outputs the text after the dictionary utilization unit 16 has performed the above-mentioned replacement process.

[0029] The voice recognition device 1 having the above configuration can operate in three types of operation modes. These operation modes are a first learning mode, a second learning mode, and an estimation mode. In the first learning mode and the second learning mode, the voice recognition device 1 adjusts (optimizes) the values ​​of the internal parameters of the model using the learning data. In the first learning mode and the second learning mode, the voice recognition device 1 calculates the error (loss) between the correct answer and an estimation result calculated based on the internal parameter values ​​at that time, and updates the internal parameters using an error backpropagation method based on the error. In the estimation mode, the voice recognition device 1 processes an unknown input (voice) using a trained model (internal parameter values) and outputs an estimation result (text of the voice recognition result).

[0030] In the first learning mode, the voice recognition device 1 performs learning of only the voice recognition decoder unit 13 using only text data. Basically, the voice recognition device 1 performs learning in the first learning mode prior to learning in the second learning mode. In the first learning mode, the input data to the voice recognition decoder unit 13 is text, and the correct answer data output from the voice recognition decoder unit 13 is also text. On the other hand, in the second learning mode, the voice recognition device 1 performs learning of both the voice recognition encoder unit 12 and the voice recognition decoder unit 13 using input voice data and correct answer data of the output text. In the second learning mode, the input to the voice recognition encoder unit 12 is voice data (specifically, data of acoustic features), and the correct answer data output from the voice recognition decoder unit 13 is text (voice recognition result).

[0031] FIG. 2 is a schematic diagram showing an example of a more detailed functional configuration for realizing the speech recognition encoder unit 12 and the speech recognition decoder unit 13 in the speech recognition device 1. That is, the diagram shows the configuration of an encoder-decoder type end-to-end speech recognition model. The left side of the configuration shown in the diagram corresponds to the speech recognition encoder unit 12, and the right side corresponds to the speech recognition decoder unit 13. The configuration shown in FIG. 2 uses a Transformer encoder and decoder. However, other types of encoders and decoders may be used. The encoder and decoder are configured using, for example, a neural network, and their internal parameters are adjustable. In other words, the internal parameters can be updated based on learning data.

[0032] In the configuration shown in FIG. 2, data D1 is data of acoustic features input as a speech recognition target. The encoder inputs data D1 and outputs data D2. Data D2 is a vector representing encoded speech. Data D2 may be called a state vector. The encoder passes data D2 to the decoder. In addition to data D2, the decoder inputs data D3 (past text). The decoder in this embodiment further inputs an unknown word prompt passed from the unknown word prompt input unit 15. The decoder calculates and outputs data D4 based on these data. Data D4 is an estimated speech recognition result (next text). Data D4 may be, for example, information representing the likelihood (probability) of a symbol that is a speech recognition result. Data D4, which is the output from the decoder, is shifted (shifted to the right in the illustration) and input to the decoder itself as data D3. In other words, the decoder obtains and outputs data D4, which is the speech recognition result, based on data D3, which is the past output (speech recognition result), data D2 (a state vector representing the characteristics of the speech) passed from the encoder, and the unknown word prompt.

[0033] For example, when an acoustic feature corresponding to a speech uttered "It's sunny today" is input as data D1, data D3 is, for example, a past text output (y1, y2, . . . , y t-1 ), which corresponds to the text "It's sunny today." In this case, the data D4 output by the decoder is the following text output (y t ), which corresponds to the text "です" for example.

[0034] Since the transformer itself is a well-known existing technology, further detailed explanation of FIG. 2 is omitted here.

[0035] FIG. 3 is a schematic diagram showing an example of the configuration of dictionary data stored in the unknown word dictionary storage unit 14. As shown in the figure, the dictionary data stored in the unknown word dictionary storage unit 14 can be expressed as a set of pairs of a speech expression symbol string (e.g., Japanese kana) and a recognition result output symbol string (e.g., Japanese kanji or a kanji-kana mixed expression). In this example, the speech expression symbol string and the recognition result output symbol string are associated with each other by separating them with a colon ":". Also, in this example, the boundaries between entries of the dictionary data are separated with a comma ",". The beginning and end of the set of pairs are indicated by a left curly bracket and a right curly bracket, respectively. The curly brackets are also called "curly braces". However, the format of the data representing the set of pairs is not limited to the format of this example. The data of this example shows "'Yamada Taro':'Yamada Taro'" and "'Tanaka Jiro':'Tanaka Jiro'" among the multiple entries of the dictionary data. The words such as "Yamada Taro" and "Tanaka Jiro" contained in the dictionary data shown here are examples of names that may appear in video content. For example, names of politicians and cultural figures may appear in news content. The unknown word dictionary storage unit 14 may store, for example, pairs of speech expression symbol strings and symbol strings for outputting recognition results for such words.

[0036] When a speech to be recognized is a Japanese language utterance, the speech expression symbol string in the unknown word dictionary storage unit 14 may be expressed, for example, as a katakana text. The recognition result output symbol string may be expressed, for example, as a text containing both kanji and kana. The text containing both kanji and kana may be a text consisting of only kanji, a text consisting of only kana (katakana or hiragana), or a text containing both kanji and kana. When the speech to be recognized is a speech in a language other than Japanese, each of the speech expression symbol string and the recognition result output symbol string may be a symbol string expressed in an appropriate manner according to the language.

[0037] The number of dictionary data entries stored in the unknown word dictionary storage unit 14 may be any integer equal to or greater than 1. The unknown word dictionary storage unit 14 may store dictionary data of unknown words that are likely to appear in speech input to the speech recognition device 1. For example, the number of dictionary data entries stored in the unknown word dictionary storage unit 14 may be 10 or less.

[0038] FIG. 4 is a schematic diagram for explaining a pattern of pre-learning (first learning mode of the voice recognition device 1) of the voice recognition decoder unit 13 (decoder). The first learning mode is a mode in which learning is performed using input and output of only text, without using voice. As shown in the figure, the first learning mode includes three types of learning: pattern 1, pattern 2, and pattern 3. The learning of pattern 1 is learning using input data without a prompt. The learning of patterns 2 and 3 is learning using input data with a prompt. Of patterns 2 and 3, the learning of pattern 2 is learning when the character string of the specified prompt is included in the kana text (speech expression symbol string). On the other hand, the learning of pattern 3 is learning when the character string of the specified prompt is not included in the kana text (speech expression symbol string).

[0039] A more detailed explanation of Figure 4 is given below. In the first training mode, text data is used to train only the decoder (speech recognition decoder unit 13) of the encoder-decoder speech recognition model.

[0040] In learning of pattern 1 (without prompt), an example of input data is the text "[SOS] Yamada Taro desu [SEP]". Also, the correct output is the text "[EOS] Yamada Taro desu". Here, the [SOS] token is a dedicated token that represents the beginning of a sentence. "SOS" is an abbreviation for "start of sentence". Also, the [SEP] token is a dedicated token that represents the meaning of separation (delimitation of an expression). Also, the "EOS" token is a dedicated token that represents the meaning of end of sentence. In other words, in learning of pattern 1, the speech recognition decoder unit 13 performs learning to associate the speech expression symbol string "Yamada Taro desu" input without a prompt with the correct output "Yamada Taro desu" (symbol string for outputting recognition result). Note that the input / output data pair described here is merely an example.

[0041] In learning pattern 2 (with prompt, the prompted expression is included in the kana text), an example of input data is the text "[OOV] Yamada Taro [SOS] Yamada Taro desu [SEP]". Also, the correct output is the text "[OOV] desu [EOS]". Here, the [OOV] token is a dedicated token that has the meaning of "out of vocabulary". That is, in learning pattern 2, the correct output data has the [OOV] token at the location corresponding to the expression of the prompt. That is, in learning pattern 2, the speech recognition decoder unit 13 performs learning to associate the speech expression symbol string "Yamada Taro desu" input with the prompt "Yamada Taro" with the correct output (symbol string for outputting the recognition result) "[OOV] desu [EOS]". Note that the input / output data pair described here is merely an example.

[0042] In learning of pattern 3 (with prompt, the prompted expression is not included in the kana text), an example of input data is the text "[OOV] Tanaka Jiro [SOS] Yamada Taro desu [SEP]". Also, the correct output is the text "Yamada Taro desu [EOS]". In other words, in learning of pattern 3, the prompted expression is not included in the kana text, so the correct output does not have the [OOV] token. In other words, in learning of pattern 3, the speech recognition decoder unit 13 performs learning to associate the speech expression symbol string "Yamada Taro desu" input with the prompt "Tanaka Jiro" with the correct output "Yamada Taro desu [EOS]" (symbol string for outputting recognition result). In other words, when the speech recognition device 1 is actually operated, since it is not known when a word (unknown word) for which a dictionary is to be utilized will appear in the speech, the unknown word prompt input unit 15 continues to input the prompt (continues to pass it to the speech recognition decoder unit 13). The training of Pattern 3 is for the purpose of continuing correct speech recognition without outputting the [OOV] token even in such a situation if the word specified by the prompt is not present in the speech. Note that the input / output data pairs described here are merely examples.

[0043] FIG. 5 is a schematic diagram showing the flow of data when the encoder-decoder type end-to-end speech recognition model described in FIG. 2 is operated in the first learning mode (learning only by text). As shown in the figure, when operating in the first learning mode, the encoder (speech recognition encoder unit 12) side is not operated, and only the decoder (speech recognition decoder unit 13) side is operated. That is, the data of acoustic features is not input to the encoder. Also, the vector based on the data of acoustic features is not passed from the encoder to the decoder side. The data D3 input to the decoder in the first learning mode is prompt information, a speech expression symbol string, and a past output (data D4 output from the decoder in the past shifted). For example, when the expression of the prompt is "hare", the speech expression symbol string is "kyohaharedesu", and the past output is "kyo", the data D3 input to the decoder is "[OOV] hare [SOS] kyouhaharedesu [SEP] kyou". Here, [OOV], [SOS], and [SEP] are special tokens already explained.

[0044] When operating in the first learning mode (learning using text only), there is no output from the encoder (speech recognition encoder unit 12). That is, in this case, the Multi-Head Attention 132 in the decoder (speech recognition decoder unit 13) is not operated. That is, in this case, the output from Add&Norm·131 is not passed to Multi-Head Attention·132, and the output from Add&Norm·131 is passed only to Add&Norm·133, bypassing Multi-Head Attention·132. In this way, in the first learning mode, learning is performed only on the decoder (speech recognition decoder unit 13) side.

[0045] An example of the operation of the model shown in FIG. 5 in the first learning mode (learning by text only) is as follows. An example of the input to the decoder (data D3 in the figure) is "[OOV] hare [SOS] kyouha hare desu [SEP] today". In this example of data D3, "[OOV] hare" is a prompt for the unknown word "hare". Also, "[SOS] kyouha hare desu [SEP]" is a kana-only text (speech expression symbol string) corresponding to the speech recognition result. Also, "today" after the special token [SEP] is a past output from the decoder. The output from the decoder corresponding to such an input to the decoder (data D4 in the figure) is "wa". In other words, this "wa" is a symbol (word) in the mixed kanji and kana text (symbol string for outputting recognition results) that follows the above past output "today".

[0046] 6 is a schematic diagram showing examples of pairs of input data and correct answer data in the second learning mode (learning based on speech). The examples shown in the figure are two types, pattern 1 and pattern 2. In both examples of pattern 1 and pattern 2, the input to the encoder is acoustic features corresponding to the spoken speech "Yamada Taro desu."

[0047] In the learning example of pattern 1 shown in FIG. 6, the unknown word dictionary storage unit 14 does not have an unknown word "Yamada Taro". That is, in that case, the unknown word dictionary storage unit 14 has no information on unknown words at all, or has only unknown words other than "Yamada Taro". In this case, the unknown word prompt input unit 15 does not pass an unknown word prompt corresponding to "Yamada Taro" to the voice recognition decoder unit 13. Then, the correct output data is "Yamada Taro desu [SEP] Yamada Taro desu [EOS]". That is, in learning using this example, the voice recognition encoder unit 12 and the voice recognition decoder unit 13 adjust their internal parameters so as to output the text "Yamada Taro desu [SEP] Yamada Taro desu [EOS]" in response to the voice input of "Yamada Taro desu".

[0048] In the learning example of pattern 2 shown in FIG. 6, the unknown word dictionary storage unit 14 has an unknown word "Yamada Taro" registered therein. That is, in this case, the unknown word prompt input unit 15 passes an unknown word prompt corresponding to "Yamada Taro" to the speech recognition decoder unit 13. Then, the correct output data is "Yamada Taro desu [SEP] [OOV] desu [EOS]." That is, in learning using this example, the speech recognition encoder unit 12 and the speech recognition decoder unit 13 adjust their internal parameters so as to output the text "Yamada Taro desu [SEP] [OOV] desu [EOS]" in response to the speech input of "Yamada Taro desu," assuming the input of an unknown word prompt corresponding to "Yamada Taro."

[0049] 7 is a schematic diagram showing an example of input data and output data when inference is performed using a trained encoder (speech recognition encoder unit 12) and decoder (speech recognition decoder unit 13). In both examples of Pattern 1 and Pattern 2 shown in the figure, the input to the encoder is acoustic features corresponding to the spoken voice "Yamada Taro desu."

[0050] In pattern 1 shown in Fig. 7, the unknown word dictionary storage unit 14 does not store an entry for "Yamada Taro" as an unknown word. Therefore, the unknown word prompt input unit 15 does not pass an unknown word prompt corresponding to "Yamada Taro" to the speech recognition decoder unit 13. In such a case, an example of text output by the processing of the trained encoder and decoder is "Yamada Taro desu." In other words, the speech recognition decoder unit 13 outputs "Yamada Taro desu" from the output from its own neural network, "Yamada Taro desu [SEP] Yamada Taro desu [EOS]."

[0051] In pattern 2 shown in Fig. 7, the unknown word dictionary storage unit 14 stores an entry "Yamada Taro" as an unknown word. Therefore, the unknown word prompt input unit 15 passes an unknown word prompt corresponding to "Yamada Taro" to the speech recognition decoder unit 13. In such a case, an example of text output by the processing of the trained encoder and decoder is "It is [OOV]". In other words, the speech recognition decoder unit 13 outputs "It is [OOV]" from the output from its own neural network, "It is Yamada Taro desu [SEP] [OOV] desu [EOS]".

[0052] When the output from the speech recognition decoder unit 13 includes a special token [OOV], the dictionary utilization unit 16 replaces the special token [OOV] by referring to the unknown word dictionary storage unit 14. In this example (pattern 2), the unknown word corresponding to the special token [OOV] is "Yamada Taro", so the dictionary utilization unit 16 replaces the special token [OOV] included in the output text "It is [OOV]" passed from the speech recognition decoder unit 13 with the word "Yamada Taro" registered in the unknown word dictionary storage unit 14. In other words, the dictionary utilization unit 16 passes the replaced text "It is Yamada Taro" to the text output unit 17.

[0053] As in the above pattern 2 (FIG. 7), by registering a set of pairs of a speech expression symbol string of an unknown word that may be included in the input speech and a symbol string for outputting a recognition result in the unknown word dictionary storage unit 14, the speech recognition decoder unit 13 outputs a portion corresponding to the word as a special token [OOV]. In addition, the dictionary utilization unit 16 reliably replaces the special token [OOV] in the text output from the speech recognition decoder unit 13 with the symbol string for outputting a recognition result registered in the unknown word dictionary storage unit 14. In other words, the configuration of this embodiment improves the accuracy of the notation of the symbol string for outputting a recognition result of the speech recognition result, compared to speech recognition processing that does not use an unknown word dictionary. In other words, the speech recognition device 1 of this embodiment outputs text with more accurate notation of homonyms of unknown words, etc., compared to the prior art.

[0054] 8 is a flowchart showing the processing procedure of the voice recognition device 1. This flowchart basically explains the processing procedure when the voice recognition device 1 performs inference processing. However, even when the voice recognition device 1 operates in the second learning mode, the processing is basically performed according to the procedure shown in this flowchart, except for the calculation of errors and the updating of parameters. The processing procedure will be explained below according to this flowchart.

[0055] First, in step S1, the voice input unit 11 inputs voice data to the voice recognition encoder unit 12. The voice data here is, for example, data of acoustic features for one utterance. The voice recognition encoder unit 12 accepts the voice data passed from the voice input unit 11, performs calculations based on internal parameter values, and outputs the resulting data (state vector). The voice recognition encoder unit 12 passes this output to the voice recognition decoder unit 13.

[0056] Next, in step S2, the voice recognition decoder unit 13 obtains the output (the above-mentioned state vector) from the voice recognition encoder unit 12, a prompt (a prompt for an unknown word passed from the unknown word prompt input unit 15), and past outputs from the voice recognition decoder unit 13 (recognition results based on past voice data). The unknown word prompt input unit 15 passes a prompt for an unknown word to the voice recognition decoder unit 13 based on the entries of all unknown words stored in the unknown word dictionary storage unit 14. The voice recognition decoder unit 13 performs calculations based on the input data and internal parameter values.

[0057] Next, in step S3, the speech recognition decoder unit 13 obtains kana text (speech expression symbol string) and kana-kanji mixed text (post-kana-kanji conversion symbol string for outputting recognition results) as a result of performing the calculations described in step S2 above. The speech recognition decoder unit 13 passes the text obtained as a result of these calculations to the dictionary utilization unit 16. Note that the kana-kanji mixed text (post-kana-kanji conversion text, symbol string for outputting recognition results) here may contain a special token [OOV] corresponding to the input prompt.

[0058] Next, in step S4, the dictionary utilization unit 16 judges whether the kana-kanji mixed text (kana-kanji converted text, symbol string for outputting recognition results) passed from the speech recognition decoder unit 13 in step S3 includes the special token [OOV]. If the kana-kanji converted text includes the special token [OOV] (step S4: YES), the process proceeds to step S5. If the kana-kanji converted text does not include the special token [OOV] (step S4: NO), the process skips step S5 and proceeds to step S6.

[0059] Next, when the process proceeds to step S5, the dictionary utilization unit 16 searches the dictionary stored in the unknown word dictionary storage unit 14 to obtain a kanji expression (an expression that will become part of the symbol string for outputting the recognition result) corresponding to the special token [OOV] contained in the kana-kanji converted text. The dictionary utilization unit 16 then replaces the special token [OOV] in the kana-kanji converted text with the kanji expression (symbol string for outputting the recognition result) obtained from the unknown word dictionary storage unit 14. As a result, the kana-kanji converted text becomes a symbol string that does not have [OOV]. After the process of step S5 is completed, the process proceeds to step S6.

[0060] Next, in step S6, the dictionary utilization unit 16 passes the kana-kanji converted text to the text output unit 17. Even if the kana-kanji converted text output from the speech recognition decoder unit 13 contains a special token [OOV], the special token [OOV] has been converted to a word registered in the dictionary (a word written in kanji or the like) at the time of processing in this step. The text output unit 17 outputs the text (symbol string for outputting the recognition result) to the outside as the result of speech recognition.

[0061] Next, in step S7, the voice recognition device 1 determines whether or not unprocessed voice data (acoustic feature data) remains. If unprocessed voice data remains (step S7: YES), the process returns to step S1 to process the voice data. If unprocessed voice data does not remain (step S7: NO), the voice recognition device 1 ends the entire process of this flowchart.

[0062] In the process described above with reference to the flowchart, the voice data (acoustic feature data) may be input to the voice recognition encoder unit 12 one after another at a predetermined time interval, for example. The process of each step described in the flowchart does not necessarily have to be executed sequentially. In other words, at least a part of the process may be performed in parallel between steps. As long as the logical input / output relationship of information (the relationship in which the output is calculated based on the input) is maintained, there is no particular inconvenience even if the process is parallel.

[0063] The configuration and method of the speech recognition device 1 according to this embodiment can be summarized as follows.

[0064] The unknown word dictionary storage unit 14 stores dictionary data that indicates the correspondence between a phonetic expression symbol string (for example, kana) and a recognition result output symbol string (for example, a mixed kanji and kana expression) for an unknown word.

[0065] The unknown word prompt input unit 15 generates and outputs a prompt including a voice expression symbol string related to an unknown word by referring to the dictionary data stored in the unknown word dictionary storage unit 14. The unknown word prompt input unit 15 passes the generated prompt to the voice recognition decoder unit 13. The unknown word prompt input unit 15 may, for example, generate a plurality of prompts corresponding to all unknown words registered in the dictionary data and pass them to the voice recognition decoder unit 13.

[0066] The voice recognition encoder unit 12 receives a voice and outputs encoding result information (state vector) corresponding to the voice.

[0067] The speech recognition decoder unit 13 obtains and outputs recognition result text (text after kana-kanji conversion) based on the encoding result information output from the speech recognition encoder unit 12 and the prompt passed from the unknown word prompt input unit 15. The recognition result text may contain a special token that indicates an unknown word portion.

[0068] The dictionary utilization unit 16 receives the recognition result text from the speech recognition decoder unit 13. When the recognition result text output from the speech recognition decoder unit 13 contains a special token ([OOV]) indicating an unknown word portion, the dictionary utilization unit 16 refers to the dictionary data stored in the unknown word dictionary storage unit 14 to obtain a symbol string for outputting the recognition result corresponding to the special token, and replaces the portion of the special token in the recognition result text with the symbol string for outputting the recognition result.

[0069] Each of the voice recognition encoder unit 12 and the voice recognition decoder unit 13 is configured using a neural network. The voice recognition device 1 is configured so that the values ​​of internal parameters of each of the voice recognition encoder unit 12 and the voice recognition decoder unit 13 can be updated based on learning data. In other words, the voice recognition encoder unit 12 and the voice recognition decoder unit 13 are capable of machine learning.

[0070] The voice recognition device 1 performs learning in both the first learning mode and the second learning mode.

[0071] Among these, in the first learning mode, the voice recognition device 1 performs learning using a set of pairs of a voice expression symbol string (e.g., kana text) input to the voice recognition decoder unit 13 and a symbol string for outputting a recognition result (e.g., a text containing a mixture of kanji and kana) that is a correct answer output from the voice recognition decoder unit 13 as learning data. When there is no prompt passed from the unknown word prompt input unit 15 to the voice recognition decoder unit 13 as learning data in the first learning mode, the symbol string for outputting a correct recognition result does not include a special token indicating that it is an unknown word portion. When there is a prompt passed from the unknown word prompt input unit 15 to the voice recognition decoder unit 13 and the voice expression symbol string input to the voice recognition decoder unit 13 includes a word corresponding to the prompt, the symbol string for outputting a correct recognition result includes a special token indicating that it is an unknown word portion. In addition, when there is a prompt passed from the unknown word prompt input unit to the speech recognition decoder unit 13 and the voice expression symbol string input to the speech recognition decoder unit 13 does not contain a word corresponding to the prompt, the symbol string for outputting the correct recognition result does not contain a special token indicating that it is an unknown word portion.

[0072] In the second learning mode, the speech recognition device 1 performs learning using, as learning data, a set of pairs of speech input to the speech recognition encoder unit 12 and recognition result text, which is the correct answer output from the speech recognition decoder unit 13. The recognition result text, which is the correct answer output from the speech recognition decoder unit 13, is a pair of a speech expression symbol string and a symbol string for outputting the recognition result.

[0073] The dictionary data stored in the unknown word dictionary storage unit 14 may represent the correspondence between the speech expression symbol string and the recognition result output symbol string for a plurality of unknown words (see also FIG. 3). In this case, the unknown word prompt input unit 15 may generate and output a prompt associated with a special token specific to each of the plurality of unknown words. The unknown word prompt input unit 15 may generate prompts for all unknown words included in the dictionary data and pass them to the speech recognition decoder unit 13. In this case, the speech recognition decoder unit 13 obtains and outputs a recognition result text based on the plurality of prompts passed from the unknown word prompt input unit 15. In this case, when the recognition result text output from the speech recognition decoder unit 13 includes a special token indicating an unknown word portion, the dictionary utilization unit 16 identifies which unknown word out of the plurality of unknown words the special token corresponds to and refers to the dictionary data to obtain a recognition result output symbol string corresponding to the special token, and replaces the portion of the special token in the recognition result text with the recognition result output symbol string. The special token specific to each of the unknown words may include information on an index value that indicates the position of the specific unknown word in the dictionary data. For example, the special token specific to each of the unknown words may be [OOV1], [OOV2], ..., [OOVn], ..., etc. In this case, the numerical value (1, 2, ..., n, etc.) included in the special token is the index value.

[0074] The model held internally by the speech recognition device 1 is trained in both a first training mode (training using only text) and a second training mode (training based on a pair of speech and text).

[0075] In the first learning mode, the speech recognition decoder unit 13 is pre-trained using only text. At this time, the training data is a set of pairs of input data (text consisting of a prompt and kana) and output (correct answer) data (text after kana-kanji conversion). Based on the output error, the values ​​of the internal parameters of the speech recognition decoder unit 13 are adjusted. It is expected that by learning in the first learning mode, the speech recognition decoder unit 13 will be able to perform the task of kana-kanji conversion and the task of identifying the part of the kana text that corresponds to the prompt. Essentially, after the speech recognition decoder unit 13 is trained in the first learning mode, it moves on to learning in the second learning mode.

[0076] In the second learning mode, the speech recognition decoder unit 13 that has been previously trained is combined with the speech recognition encoder unit 12 to perform training. That is, in the second learning mode, an encoder-decoder type end-to-end speech recognition model is trained. As a configuration specific to this embodiment, the unknown word prompt input unit 15 uses data of the unknown word dictionary stored in the unknown word dictionary storage unit 14 to supply prompts to the speech recognition decoder unit 13. The training data in the second learning mode is a set of pairs of input data (speech data) and output (correct answer) data (output text). The output text includes text in kana (which may be called a speech expression symbol string) and text after kana-kanji conversion (which may be called a symbol string for outputting recognition results). Voice data is input to the speech recognition encoder unit 12. The speech recognition encoder unit 12 outputs numerical data (vector, state vector) corresponding to the input speech data. The speech recognition decoder unit 13 receives the output from the speech recognition encoder unit 12 and a prompt from the unknown word prompt input unit 15. Past outputs from the speech recognition decoder unit 13 are also input to the speech recognition decoder unit 13. The speech recognition decoder unit 13 outputs kana text (speech expression symbol string) and kana-kanji converted text (symbol string for outputting recognition result). The dictionary utilization unit 16 replaces special tokens [OOV] included in the kana-kanji converted text (symbol string for outputting recognition result) with expressions registered in a dictionary (part of the symbol string for outputting recognition result) and outputs the replaced text. The output from the dictionary utilization unit 16 is the final output as the speech recognition result.

[0077] In the above process, the prompt consists of a special token [OOV] that indicates an unknown word, and a kana notation (a string of phonetic symbols) of an expression (word) specified as the prompt. An example of a prompt + kana is "[OOV] Yamada Taro."

[0078] The kana text (speech expression symbol string) used in learning may be created based on a corpus of mixed kanji and kana text. For example, when the corpus contains the text "Yamada Taro is the Prime Minister," the kana text "Yamada Taro Souri Daijin Desu" can be generated in response to the mixed kanji and kana text. The kana text may be generated manually, or a part or all of the processing may be performed by a computer.

[0079] The input to the speech recognition decoder unit 13 in the first learning mode is a combination of a prompt and kana text. A special token [OOV] indicating a prompt, a special token [SOS] indicating the beginning of a sentence, and a special token [SEP] indicating a sentence break can be used. Using these special tokens, an example of the input to the speech recognition decoder unit 13 is data such as "[OOV] Yamada Tarou [SOS] Yamada Tarou Souri Dai Jin Desu [SEP]".

[0080] There are three types of patterns of learning data in the first learning mode (see also FIG. 4). In pattern 1, the input data to the voice recognition decoder unit 13 does not have a prompt. In pattern 2, the input data to the voice recognition decoder unit 13 has a prompt, and the word indicated by the prompt is included in the subsequent kana text. In pattern 3, the input data to the voice recognition decoder unit 13 has a prompt, but the word indicated by the prompt is not included in the subsequent kana text. By pre-training the voice recognition decoder unit 13 using these three patterns of learning data, the voice recognition decoder unit 13 is trained to be able to handle all patterns. That is, there are cases where there is no prompt, where there is a prompt and the word indicated by the prompt is included in the kana text, and where there is a prompt and the word indicated by the prompt is not included in the kana text.

[0081] Up to this point, an example in which only one prompt is input has been described, but the input data to the speech recognition decoder unit 13 may have multiple prompts. In this case, the special token [OOV] may have an index number. As an example, a prompt input to the speech recognition decoder unit 13 is "[OOV1] Yamada [OOV2] Tanaka [OOV3]... [OOVn] Takeda", etc., where [OOVn] is a special token representing the n-th prompt (n is a positive integer). In this way, by using special tokens with index numbers such as [OOV1], [OOV2], [OOV3],..., [OOVn], the index number can be used as an index of a position in the unknown word dictionary storage unit 14. In other words, if the output text passed from the speech recognition decoder unit 13 has a special token [OOV3], the dictionary utilization unit 16 can extract an expression to replace the special token [OOV3] by referring to an entry in the unknown word dictionary storage unit 14 corresponding to the index number "3".

[0082] It is preferable that the number of unknown word entries stored in the unknown word dictionary storage unit 14 is not too large. For example, it is preferable that the number of entries in the unknown word dictionary is about 1 to 10. However, the number of entries is not limited to 10 or less.

[0083] The contents of the dictionary data stored in the unknown word dictionary storage unit 14 may be changed as appropriate. For example, when the voice recognition device 1 recognizes speech in a broadcast news program, the contents of the dictionary data stored in the unknown word dictionary storage unit 14 may be changed for each news title (headline). In this case, a set of unknown words suitable for each news title can be stored in the unknown word dictionary storage unit 14. That is, the unknown word prompt input unit 15 can pass a prompt suitable for each news title to the voice recognition decoder unit 13. In addition, the contents of the dictionary data stored in the unknown word dictionary storage unit 14 can be changed according to the scene to be used, not limited to voice recognition in a news program, so that the voice recognition device 1 can output the optimum unknown word for each scene as the recognition result.

[0084] The structure of the speech recognition decoder unit 13 and its output are as follows. In other words, the inside of the speech recognition decoder unit 13 is realized by a functional configuration that is the same as or similar to that of a Transformer decoder. The inside of the speech recognition decoder unit 13 includes a self-attention process (Masked Multi-Head Attention·136 shown in FIG. 2) performed between past decoder outputs, and a source target attention process (Multi-Head Attention·132 shown in FIG. 2) performed between the output from the encoder (speech recognition encoder unit 12) and the decoder input that has undergone the above self-attention process.

[0085] In the first learning mode, the voice recognition decoder unit 13 is input with a prompt and kana. In the first learning mode, the voice recognition decoder unit 13 learns to skip the source target attention process (Multi-Head Attention·132 shown in FIG. 2) and output text after kana-kanji conversion. At this time, the voice recognition decoder unit 13 is trained to output the following three types of output depending on whether there is a prompt and whether a word corresponding to the prompt is included in the kana. Here, the case where the kana input to the voice recognition decoder unit 13 is "Yamada Tarou desu" will be described. In the first case (pattern 1 shown in FIG. 4), no prompt is passed to the voice recognition decoder unit 13, and the correct output text is "Yamada Tarou desu". In the second case (pattern 2 shown in FIG. 4), a prompt is passed to the voice recognition decoder unit 13, and the input kana contains a word corresponding to the prompt (here, "Yamada Tarou"), so the correct output text is "[OOV] desu". In the third case (pattern 3 shown in Figure 4), a prompt is passed to the speech recognition decoder unit 13, and the word corresponding to the prompt (here, "Tanaka") is not included in the input kana, so the correct output text is "Yamada Taro desu."

[0086] By sufficiently learning in the first learning mode as described above, the voice recognition decoder unit 13 acquires the kana-kanji conversion ability and the ability to identify the part of the text after kana-kanji conversion that corresponds to the word (kana string) that corresponds to the prompt.

[0087] In the second learning mode, learning is performed using appropriate learning data (a set of pairs of input voice data and output correct answer text). By performing this learning, the voice recognition encoder unit 12 and the voice recognition decoder unit 13 acquire the ability to perform the following processing. That is, based on the input voice data, the voice recognition encoder unit 12 outputs encoded voice data (state vector). Also, based on the vector output from the voice recognition encoder unit 12 and the prompt passed from the unknown word prompt input unit 15, the voice recognition decoder unit 13 outputs kana text and kana-kanji converted text based on the input voice data. However, if the kana specified by the prompt is included in the voice, the voice recognition decoder unit 13 outputs the corresponding part in the kana-kanji converted text as a special token [OOV].

[0088] The special token [OOV] in the kana-kanji converted text output from the speech recognition decoder unit 13 is replaced by the dictionary utilization unit 16. That is, the dictionary utilization unit 16 replaces the special token [OOV] with an appropriate kanji expression by referring to the unknown word dictionary storage unit 14. This creates the text to be finally output (text of the speech recognition result).

[0089] FIG. 9 is a block diagram showing an example of the internal configuration of a device for implementing the voice recognition device 1. The voice recognition device 1 can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. The RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with an external input / output device or the like. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data from and to the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port 903 via the bus 906.

[0090] At least some of the functions of the voice recognition device 1 in the above-mentioned embodiment can be realized by a computer and a program. In that case, a program for realizing the functions may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into a computer system and executed to realize the functions. Note that the term "computer system" here includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to a portable medium such as a flexible disk, an optical magnetic disk, a ROM, a CD-ROM, a DVD-ROM, a USB memory, and a storage device such as a hard disk built into a computer system. In other words, the term "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the term "computer-readable recording medium" may include a medium that temporarily and dynamically holds a program, such as a communication line when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in that case. Furthermore, the above-mentioned program may be a program for realizing some of the above-mentioned functions, and may further be a program that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0091] Although the embodiment has been described above, the present invention can also be embodied in the following modified examples.

[0092] In the embodiment, an example has been described in which the voice recognition device 1 recognizes a spoken voice in Japanese and outputs a text (recognition result) in Japanese. As a modified example, the voice recognition device 1 may perform voice recognition processing for any language, not limited to Japanese.

[0093] In the embodiment, the speech recognition device 1 performs learning in the first learning mode prior to learning in the second learning mode. As a modified example, part of the learning in the first learning mode (learning using only text) may be performed after learning in the second learning mode.

[0094] The above describes in detail an embodiment of the present invention (including modified examples) with reference to the drawings. However, the specific configuration is not limited to this embodiment, and also includes designs that do not deviate from the gist of the present invention.

[0095] According to the above-described embodiment (including the modified example), the speech recognition device 1 can output speech recognition results including unknown words that did not appear during learning. In addition, the unknown words at this time only need to be registered in the unknown word dictionary storage unit 14, and no additional learning is required, making the operation simple. Furthermore, when homonyms or spelling variations (especially in the case of proper nouns, etc.) exist, the speech recognition device 1 can output recognition results using the correct spelling registered in the unknown word dictionary storage unit 14. [Industrial Applicability]

[0096] The present invention can be used in any industry that uses voice recognition processing. For example, in the content distribution business (including broadcasting), the present invention can be used to recognize voice spoken in content and generate text. However, the scope of use of the present invention is not limited to the examples given here. [Explanation of symbols]

[0097] 1. Voice recognition device 11 Audio input section 12 Voice recognition encoder 13 Voice recognition decoder section 14 Unknown word dictionary memory section 15 Unknown word prompt input section 16 Dictionary Use Club 17 Text output section 131 Add&Norm 132 Multi-Head Attention 133 Add&Norm 136 Masked Multi-Head Attention 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus

Claims

1. an unknown word dictionary storage unit for storing dictionary data representing a correspondence relationship between a speech expression symbol string and a symbol string for outputting a recognition result for an unknown word; an unknown word prompt input unit that generates and outputs a prompt including a voice expression symbol string related to an unknown word by referring to the dictionary data stored in the unknown word dictionary storage unit; a voice recognition encoder unit which receives a voice and outputs encoded result information corresponding to the voice; a speech recognition decoder unit that obtains and outputs a recognition result text based on the encoding result information output from the speech recognition encoder unit and the prompt passed from the unknown word prompt input unit; a dictionary utilization unit that, when the recognition result text output from the speech recognition decoder unit includes a special token that indicates an unknown word portion, acquires a symbol string for outputting the recognition result corresponding to the special token by referring to the dictionary data stored in the unknown word dictionary storage unit, and replaces the portion of the special token in the recognition result text with the symbol string for outputting the recognition result; A voice recognition device comprising:

2. The voice recognition encoder unit and the voice recognition decoder unit are each configured using a neural network, and are configured so that values ​​of internal parameters of the voice recognition encoder unit and the voice recognition decoder unit can be updated based on learning data.

2. The speech recognition device according to claim 1.

3. In the first learning mode, a set of pairs of a voice expression symbol string to be input to the speech recognition decoder unit and a symbol string for outputting a recognition result, which is a correct answer output from the speech recognition decoder unit, is used as training data for learning; when there is no prompt passed from the unknown word prompt input unit to the speech recognition decoder unit, the symbol string for outputting the correct recognition result does not include a special token indicating that it is an unknown word portion; when there is a prompt passed from an unknown word prompt input unit to the speech recognition decoder unit and the speech expression symbol string input to the speech recognition decoder unit includes a word corresponding to the prompt, the symbol string for outputting a correct recognition result includes a special token indicating that the part is an unknown word; When there is a prompt passed from an unknown word prompt input unit to the speech recognition decoder unit and the speech expression symbol string input to the speech recognition decoder unit does not include a word corresponding to the prompt, the symbol string for outputting a correct recognition result does not include a special token indicating an unknown word portion.

3. The speech recognition device according to claim 2.

4. In the second learning mode, A set of pairs of a voice input to the speech recognition encoder unit and the recognition result text which is a correct answer output from the speech recognition decoder unit is used as training data for learning, The recognition result text is a pair of the speech expression symbol string and the recognition result output symbol string.

4. The speech recognition device according to claim 3.

5. the dictionary data stored in the unknown word dictionary storage unit represents a correspondence relationship between a speech expression symbol string and a recognition result output symbol string for a plurality of unknown words, the unknown word prompt input unit generates and outputs the prompt associated with the special token specific to each of the plurality of unknown words; the speech recognition decoder unit determines and outputs a recognition result text based on the prompt passed from the unknown word prompt input unit; When the recognition result text output from the speech recognition decoder includes a special token indicating an unknown word portion, the dictionary utilization unit identifies which unknown word among the plurality of unknown words the special token corresponds to and refers to the dictionary data to obtain a symbol string for outputting the recognition result corresponding to the special token, and replaces the portion of the special token in the recognition result text with the symbol string for outputting the recognition result.

2. The speech recognition device according to claim 1.

6. the special token specific to each of the plurality of unknown words includes information of an index value that indicates a position of the specific unknown word in the dictionary data; 6. The speech recognition device according to claim 5.

7. an unknown word dictionary storage unit for storing dictionary data representing a correspondence relationship between a speech expression symbol string and a symbol string for outputting a recognition result for an unknown word; an unknown word prompt input unit that generates and outputs a prompt including a voice expression symbol string related to an unknown word by referring to the dictionary data stored in the unknown word dictionary storage unit; a voice recognition encoder unit which receives a voice and outputs encoded result information corresponding to the voice; a speech recognition decoder unit that obtains and outputs a recognition result text based on the encoding result information output from the speech recognition encoder unit and the prompt passed from the unknown word prompt input unit; a dictionary utilization unit that, when the recognition result text output from the speech recognition decoder unit includes a special token that indicates an unknown word portion, acquires a symbol string for outputting the recognition result corresponding to the special token by referring to the dictionary data stored in the unknown word dictionary storage unit, and replaces the portion of the special token in the recognition result text with the symbol string for outputting the recognition result; A program for causing a computer to function as a speech recognition device comprising the above-mentioned.