Speech recognition device, speech recognition method, and speech recognition program
The 2-pass CB speech recognition system addresses the challenge of generating both character strings and pronunciations, improving the accuracy of speech recognition by applying contextual biasing to characters and words separately.
Patent Information
- Application Number
- JP2024099063
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-01-07
AI Technical Summary
Conventional speech recognition technologies face challenges in accurately generating information that includes both character strings and pronunciations, particularly for Japanese kanji characters, limiting their effectiveness in speech recognition applications.
A speech recognition system employing a 2-pass contextual biasing (CB) approach, where the first pass applies CB to characters and the second pass applies CB to words, generating both character strings and pronunciations of spoken words.
The system effectively generates both character strings and pronunciations of spoken words, enhancing the accuracy and appropriateness of speech recognition results.
Smart Images

Figure 2026001600000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice recognition device, a voice recognition method, and a voice recognition program. [Background technology]
[0002] Speech recognition that recognizes human speech is used in various services. For example, a speech recognition method using Contextual Biasing (hereinafter sometimes referred to as "CB") technology that takes into account the context of human speech has been proposed (see, for example, Patent Documents 1 and 2 and Non-Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7417634 [Patent Document 2] Patent No. 7092953 [Non-patent literature]
[0004] [Non-Patent Document 1] “Streaming End-to-end Speech Recognition For Mobile Devices”, Yanzhang He et al. <Internet> https: / / arxiv.org / abs / 1811.06621 (Retrieved June 8, 2024) [Non-patent document 2] “End-to-end ASR to jointly predict transcriptions and linguistic annotations”, Motoi Omachi, Yuya Fujita, Shinji Watanabe and Matthew Wiesner <Internet> https: / / aclanthology.org / 2021.naacl-main.149 / (Retrieved June 8, 2024) Summary of the Invention [Problem to be solved by the invention]
[0005] However, there is room for improvement in the above-mentioned conventional technology. It is difficult to apply the above-mentioned conventional technology to speech recognition that targets human speech and estimates (generates) information including a combination of words and their pronunciations, such as Japanese kanji characters and their pronunciations. Therefore, there is room for improvement in terms of performing appropriate speech recognition while using CB technology, and it is desired to appropriately generate information including both character strings indicating words included in an utterance and the pronunciation of the words.
[0006] The present application has been made in consideration of the above, and aims to provide a speech recognition device, a speech recognition method, and a speech recognition program that appropriately generate information including both a string of characters indicating words included in an utterance and the pronunciation of the words. [Means for solving the problem]
[0007] The speech recognition device according to the present application is characterized by comprising: an acquisition unit that acquires speech information, which is speech information indicating a user's speech; a generation unit that generates a recognition result including a recognition string indicating a spoken word recognized from the speech and recognition reading information indicating the reading of the spoken word, based on the result of a first speech recognition process that performs speech recognition by applying contextual biasing to a unit of components at least smaller than a word that is included in a character string that encodes the speech indicated by the speech information, and the result of a second speech recognition process that performs speech recognition by applying contextual biasing to a unit of spoken words that are included in the character string and constitute meaning; and an output unit that outputs the recognition result generated by the generation unit. [Effects of the Invention]
[0008] According to one aspect of the embodiment, it is possible to appropriately generate information including both character strings indicating words included in an utterance and the pronunciation of the words. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of voice recognition according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of a voice recognition system according to an embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of a voice recognition device according to an embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a model information storage unit according to the embodiment. [Figure 5] FIG. 5 is a flowchart illustrating an example of voice recognition according to the embodiment. [Figure 6] FIG. 6 is a conceptual diagram showing an example of a weighted finite state transducer. [Figure 7] FIG. 7 is a conceptual diagram showing an example of the 1st pass. [Figure 8] FIG. 8 is a diagram showing an example of the processing flow of the 1st pass. [Figure 9] FIG. 9 is a conceptual diagram showing an example of the 2nd pass. [Figure 10] FIG. 10 is a diagram illustrating an example of the processing flow of the 2nd pass. [Figure 11] FIG. 11 is a diagram showing another example of the processing flow of the 2nd pass. [Figure 12] FIG. 12 is a diagram showing another example of the flow of processing using beam search. [Figure 13] FIG. 13 is a hardware configuration diagram showing an example of a computer that realizes the functions of a voice recognition device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, a speech recognition device, a speech recognition method, and a speech recognition program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the speech recognition device, the speech recognition method, and the speech recognition program according to the present application are not limited to these embodiments. Furthermore, the same components in the following embodiments will be designated by the same reference numerals, and duplicated descriptions will be omitted.
[0011] (Embodiment) [1. Speech Recognition System] First, a processing flow of speech recognition processing (also simply referred to as "speech recognition") according to an embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of speech recognition according to an embodiment. Note that detailed description of the speech recognition performed in the speech recognition system 1 that is the same as conventional speech recognition processing will be omitted as appropriate. Note that, although Japanese will be described as an example below, the processing described below may be applied to various languages, not limited to Japanese, as long as the language is applicable. Also, although characters will be described as an example of components smaller than words, components smaller than words are not limited to characters and may be any components such as so-called subwords.
[0012] FIG. 1 illustrates an example in which user U1, an example of a user, performs voice recognition using his / her own terminal device 10. That is, FIG. 1 illustrates a case in which the terminal device 10 of user U1 is a voice recognition device that provides a voice recognition service to the user. For example, user U1 is a user identified by a user ID "U1." Note that FIG. 1 illustrates an example in which the terminal device 10 is a smartphone, but the terminal device 10 is not limited to a smartphone and may be any device (equipment) as long as it is a voice recognition device that provides a voice recognition service to the user; this point will be described in detail later.
[0013] An example of information processing will be described below with reference to Fig. 1. The terminal device 10 acquires a speech recognition model M1 for performing speech recognition from the server device 50 (see Fig. 2). For example, the speech recognition model M1 is a model that outputs a speech recognition result using CB (contextual biasing) for a speech (also referred to as "speech speech") uttered by a user. For example, the speech recognition model M1 generates a recognition result including a recognition character string indicating a word (also referred to as "spoken word") recognized from the speech of the user and reading information (also referred to as "recognized reading information") indicating the reading of the spoken word.
[0014] For example, the speech recognition model M1 estimates (generates) Kanji characters (also simply referred to as "Kanji") representing spoken words recognized from the speech and a recognized reading indicating the reading of the Kanji characters, based on the results of a first speech recognition process that performs speech recognition by applying CB to each character included in a string of characters that symbolizes Japanese speech, and the results of a second speech recognition process that performs speech recognition by applying CB to each word that constitutes meaning. As shown in FIG. 1 , the speech recognition model M1 includes components (processing) related to CB (also referred to as "2-pass CB"), which includes both a pass that applies CB to each character (also referred to as "1st pass") and a pass that applies CB to each word (also referred to as "2nd pass"). In this way, the speech recognition model M1 is a speech recognition model (also referred to as a "2-pass CB speech recognition model") that includes components (processing) related to 2-pass CB. As a result, the speech recognition model M1 estimates (generates) Kanji characters representing spoken words recognized from the speech by 2-pass CB, which is the first pass and the second pass of 2-pass CB, and a recognized reading indicating the reading of the Kanji characters. Details regarding the 1st pass and 2nd pass will be described later.
[0015] For example, the speech recognition model M1 is an end-to-end model for speech recognition. Note that the speech recognition model M1 is not limited to an end-to-end model, and may be a model of any format as long as it is possible to generate a recognition result including a recognition character string indicating an uttered word and recognition reading information indicating the reading of the uttered word by applying CB. Furthermore, when training the speech recognition model M1, the terminal device 10 may acquire data (training data) used in the training process of the speech recognition model M1 from the server device 50 or the like, and train the speech recognition model M1 using the training data; this point will be described later.
[0016] Note that a model for speech recognition that estimates the notation of kanji characters and their pronunciation corresponding to human speech without using CB technology (simultaneous reading and notation estimation model) is disclosed in Non-Patent Document 2, etc., and detailed explanation will be omitted here.
[0017] Now, the flow of the process shown in Fig. 1 will be described. In Fig. 1, user U1 utters "I want to go to Nipponbashi." The terminal device 10 detects the utterance PA of user U1 and accepts the utterance voice information of "I want to go to Nipponbashi," which is the utterance PA of user U1, as input (step S11).
[0018] The terminal device 10 then performs speech recognition processing using the speech information of "I want to go to Nipponbashi" received as input and the speech recognition model M1. The terminal device 10 inputs the speech information of "I want to go to Nipponbashi" as input information to the speech recognition model M1 (step S12). The speech recognition model M1 to which the input information has been input estimates (generates) a character string including kanji characters corresponding to the speech information and a recognized reading indicating the reading of the character string including the kanji characters, by speech recognition of the input speech information, and outputs them as a recognition result corresponding to the input speech information (step S13).
[0019] In Fig. 1, speech recognition model M1, to which speech information of "I want to go to Nipponbashi" is input, outputs "I want to go to Nipponbashi" as a recognition result OT1, which includes a character string including kanji characters corresponding to "I want to go to Nipponbashi" and a recognized reading indicating the reading of the character string including the kanji characters. As shown in Fig. 1, recognition result OT1 includes a character string formed by concatenating a combination of kanji characters of a spoken word recognized from the speech, a character indicating the reading of the kanji characters (also called "first character"), and a symbol indicating that the first character is the reading of the kanji characters (also called "second character").
[0020] For example, in the recognition result OT1, "Nihonbashi" is a character string (notation character string) indicating the kanji "Nihonbashi" estimated to be included in the spoken voice information. Also, in the recognition result OT1, "Ni_pp_pon_ba_shi_" is a character string (reading identification symbol combination character string) that combines a katakana character that is the first character indicating the reading of the kanji "Nihonbashi" in the spoken word with an underscore "_" that is a symbol indicating that the katakana character is the reading of the kanji "Nihonbashi." In this way, the terminal device 10 generates "Ni_pp_pon_ba_shi_Nihonbashini_ni_ni_iki_tai_itai_go" that includes a character string including the kanji character corresponding to "I want to go to Nipponbashi" and recognition reading information indicating the reading of the character string including the kanji character, as the recognition result OT1 corresponding to the spoken voice information of "I want to go to Nipponbashi" received as input.
[0021] Then, the terminal device 10 outputs the generated recognition result (step S14). For example, the terminal device 10 displays the generated recognition result. In FIG. 1, the terminal device 10 displays the recognition result OT1. In this case, the terminal device 10 displays the character string "I want to go to Nihonbashi-shi Nihonbashi-ni."
[0022] As described above, the terminal device 10 generates a character string including Kanji characters corresponding to utterance speech information and recognition reading information indicating the reading of the character string including the Kanji characters by using the speech recognition model M1 having a 2-pass CB including both a 1st pass that applies CB on a character-by-character basis and a 2nd pass that applies CB on a word-by-word basis. In this way, the terminal device 10 can appropriately generate information including both a character string indicating a word included in an utterance and the reading of the word by generating a recognition result including a recognized character string and recognition reading information based on a first speech recognition process that performs speech recognition by applying CB on a character-by-character basis and a second speech recognition process that performs speech recognition by applying CB on a word-by-word basis.
[0023] [2. Configuration of the speech recognition system] Next, the configuration of a speech recognition system 1 that realizes the following speech recognition will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of a speech recognition system according to an embodiment. As shown in Fig. 2, the speech recognition system 1 includes a plurality of terminal devices 10 and a server device 50. The terminal devices 10 and the server device 50 are connected to each other via a predetermined network N so as to be able to communicate with each other via wired or wireless communication. Note that the speech recognition system 1 shown in Fig. 2 may include a plurality of terminal devices 10 and a plurality of server devices 50.
[0024] The terminal device 10 is an example of a voice recognition device, and is a computer (information processing device) used by a user. For example, the terminal device 10 is realized by a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), etc. FIG. 1 shows an example in which the terminal device 10 is a smartphone.
[0025] The terminal device 10 performs speech recognition on the user's speech. For example, the terminal device 10 generates a recognized character string and recognized reading information based on the result of a first speech recognition process in which speech recognition is performed by applying CB to each character included in a character string that encodes the user's speech, and the result of a second speech recognition process in which speech recognition is performed by applying CB to each spoken word that is a word included in the character string and that constitutes the meaning.
[0026] Furthermore, the terminal device 10 outputs various types of information. For example, the terminal device 10 outputs various types of information using various applications. For example, the terminal device 10 outputs a recognition result generated by speech recognition. For example, the terminal device 10 displays content indicating the recognition result generated by speech recognition. The terminal device 10 displays content indicating the recognition result including a recognized character string and recognition reading information.
[0027] The terminal device 10 may execute various processes related to the display of content based on control information, etc., by appropriately using various conventional technologies related to the display of content. The terminal device 10 may execute various processes related to the display of content based on control information. The terminal device 10 may acquire, as the control information, a script to be executed on a predetermined application such as a web browser, and execute the acquired script. Such control information corresponds to a display program, etc. according to the embodiment, and is realized by, for example, CSS (Cascading Style Sheets), JavaScript (registered trademark), HTML (HyperText Markup Language), or any language capable of describing the above-described display processes, etc.
[0028] The server device 50 is a computer (information processing device) that provides various information used for processing by the terminal device 10. For example, the server device 50 is a server managed by an administrator of the voice recognition system 1 or the like.
[0029] The server device 50 provides various information to an external device that performs voice recognition. The server device 50 transmits various information to a terminal device 10 used by a user. For example, the server device 50 provides various information used for processing to the terminal device 10. The server device 50 may transmit various models, such as a voice recognition model, to the terminal device 10. Note that the distribution of the models may be performed by a device other than the server device 50.
[0030] The above-described configuration of the speech recognition system 1 is merely an example, and any device configuration and distribution of functions can be adopted for the speech recognition system 1. For example, the speech recognition system 1 may be configured such that a speech recognition model is installed on the server device 50 side, and speech recorded by the terminal device 10 is transmitted to perform speech recognition.
[0031] 3. Configuration of the speech recognition device Next, the configuration of a terminal device 10, which is an example of a voice recognition device, will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the configuration of a voice recognition device according to an embodiment. As shown in Fig. 3, the terminal device 10 has a communication unit 11, an input unit 12, a display unit 13, a storage unit 14, and a control unit 15. The terminal device 10 may also have a speaker that serves as an audio output interface. For example, the speaker may be connected to the terminal device 10 so as to be able to communicate with it via an external connection or the like.
[0032] (Communications Department 11) The communication unit 11 is realized by, for example, a communication circuit. The communication unit 11 is connected to a predetermined communication network (not shown) by wire or wirelessly, and transmits and receives information to and from an external information processing device. For example, the communication unit 11 is connected to a predetermined network N (see FIG. 2) by wire or wirelessly, and transmits and receives information to and from a server device 50.
[0033] (Input section 12) The input unit 12 functions as an interface that accepts input from a user. The input unit 12 accepts input from a user in various modalities. The input unit 12 receives the user's speech as input. For example, the input unit 12 has a microphone (sound sensor) that serves as a voice input / output interface. The input unit 12 may also have a configuration that allows various operations to be input from a user. For example, the input unit 12 may accept various operations from a user via a display surface (for example, the display unit 13) using a touch panel function. The input unit 12 may also accept various operations from buttons provided on the terminal device 10 or a keyboard or mouse connected to the terminal device 10.
[0034] (Display section 13) Display unit 13 is a screen that displays information. For example, display unit 13 is a display screen of a tablet terminal or the like realized by a liquid crystal display, an organic EL (Electro-Luminescence) display, or the like, and is a display device for displaying various types of information. Display unit 13 may also function as a touch panel screen.
[0035] (Storage unit 14) The storage unit 14 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 14 stores, for example, information about an application (e.g., a voice recognition application) installed in the terminal device 10, such as a program. As shown in FIG. 3 , the storage unit 14 according to the embodiment includes a model information storage unit 141, a data storage unit 142, and a processing result information storage unit 143. The storage unit 14 may store various information other than the above.
[0036] (Model information storage unit 141) The model information storage unit 141 according to the embodiment stores information about a model. For example, the model information storage unit 141 stores information (model data) about a trained model (model) trained (generated) through a learning process. The model information storage unit 141 shown in FIG. 4 stores data used in learning (trained data) in association with the trained model (model). FIG. 4 is a diagram illustrating an example of a model information storage unit according to the embodiment. In the example shown in FIG. 4, the model information storage unit 141 includes items such as "model ID," "purpose," "model data," and "trained data." In the example of FIG. 4, the model information storage unit 141 stores data used in learning (trained data) in association with the trained model (model).
[0037] "Model ID" indicates identification information for identifying a model. "Use" indicates the use of the corresponding model. "Model Data" indicates the data of the model. Figure 4 etc. shows an example in which conceptual information such as "MDT1" is stored in "Model Data," but in reality, various information that constitutes the model is included, such as information on the model configuration (network configuration) and information on parameters. For example, "Model Data" includes information including the nodes in each layer of the network, the functions employed by each node, the connection relationships between the nodes, and the connection coefficients set for the connections between the nodes.
[0038] "Learning data" refers to data used in training a trained model (model). "Learning data" stores information indicating a dataset used in training the corresponding model. For example, "Learning data" associates data (input information) with correct answer information (output information) corresponding to that data and stores the data as training data (also referred to as "training data"). While Figure 4 shows an example in which conceptual information such as "LDT1" is stored in "Learning data," in reality, various information related to the data used in training the corresponding model, such as data (input information) and correct answer information (output information) corresponding to that data, is included.
[0039] In FIG. 4, the model (speech recognition model M1) identified by the model ID "M1" indicates that its use is "speech recognition (notation + reading)." In other words, the speech recognition model M1 is a model that performs speech recognition of input spoken speech and outputs (generates) recognition results (such as character information) that include both notation and reading corresponding to the spoken speech. Also, the model data of the speech recognition model M1 is model data MDT1. Also, the training data used to train the speech recognition model M1 is training data LDT1.
[0040] Note that the model information storage unit 141 is not limited to the above and may store various types of information depending on the purpose. The model information storage unit 141 may store multiple voice recognition models. For example, the model information storage unit 141 may store voice recognition models depending on the application. For example, the model information storage unit 141 may store a voice recognition model used for a car navigation service (also referred to as a "car navigation model"). For example, the car navigation model is a voice recognition model that uses, as phrase information, information corresponding to phrases that a user is expected to utter when using a car navigation system. For example, the model information storage unit 141 may store a voice recognition model used for a search service. For example, the car navigation model is a voice recognition model that uses, as phrase information, information corresponding to phrases that a user is expected to utter when using a search system.
[0041] (Data storage unit 142) The data storage unit 142 according to the embodiment stores data used in processing. For example, the data storage unit 142 stores data that is the target of voice recognition. For example, the data storage unit 142 stores information related to the user's utterance that is the target of voice recognition. Note that the data storage unit 142 is not limited to the above, and may store various types of information depending on the purpose.
[0042] (Processing result information storage unit 143) The processing result information storage unit 143 according to the embodiment stores various information related to processing results. The processing result information storage unit 143 stores various information such as the processing results of speech recognition of the user who uses the terminal device 10. Note that the processing result information storage unit 143 is not limited to the above, and may store various information depending on the purpose.
[0043] (Control unit 15) The control unit 15 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in a storage device such as the storage unit 14 inside the terminal device 10 using RAM as a work area. For example, these various programs include an application program for performing voice recognition. The control unit 15 is also a controller, and is realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0044] 3, control unit 15 has an acquisition unit 151, a learning unit 152, a reception unit 153, a generation unit 154, an output unit 155, and a transmission unit 156, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of control unit 15 is not limited to the configuration shown in FIG. 3, and may be any other configuration that performs the information processing described below.
[0045] (Acquisition part 151) The acquisition unit 151 acquires various types of information. For example, the acquisition unit 151 acquires various types of information from the storage unit 14. For example, the acquisition unit 151 acquires various types of information from the model information storage unit 141, the data storage unit 142, the processing result information storage unit 143, etc. For example, the acquisition unit 151 acquires models such as the speech recognition model M1 from the model information storage unit 141.
[0046] The acquisition unit 151 may acquire various information from an external information processing device. For example, the acquisition unit 151 acquires various information from the server device 50. For example, the acquisition unit 151 acquires various information to be used for processing from the server device 50. For example, the acquisition unit 151 acquires models such as the speech recognition model M1 from the server device 50.
[0047] The acquisition unit 151 accepts various types of information. The acquisition unit 151 accepts various operations by the user. For example, the acquisition unit 151 accepts various operations by the user via the input unit 12. The acquisition unit 151 accepts operations by the user. The acquisition unit 151 accepts an operation in which the user moves a finger that has been placed in contact with the screen. The acquisition unit 151 accepts an operation to select content that is being displayed.
[0048] The acquisition unit 151 acquires speech information, which is speech information indicating a speech uttered by a user. The acquisition unit 151 acquires a speech recognition model M1 that outputs a recognition result corresponding to the speech information through a first speech recognition process and a second speech recognition process in response to input of the speech information.
[0049] (Learning Section 152) The learning unit 152 executes a learning process to learn a machine learning model (model). Note that, when the terminal device 10 acquires a machine learning model such as the speech recognition model M1 from the server device 50, the terminal device 10 does not need to include the learning unit 152.
[0050] For example, the learning unit 152 executes the learning process based on various information acquired by the acquisition unit 131. The learning unit 152 executes the learning process based on information from an external information processing device or information stored in the storage unit 14. The learning unit 152 executes the learning process based on information stored in the model information storage unit 141. The learning unit 152 stores the model generated by learning in the model information storage unit 141.
[0051] The learning unit 152 performs a learning process. The learning unit 152 performs various types of learning. The learning unit 152 learns various types of information based on the information acquired by the acquisition unit 131. The learning unit 152 learns (generates) a model. The learning unit 152 learns various types of information such as a model. The learning unit 152 generates a model through learning. The learning unit 152 learns the model using various machine learning techniques. For example, the learning unit 152 learns parameters of the model (network). The learning unit 152 learns the model using various machine learning techniques.
[0052] The learning unit 152 generates various machine learning models such as the speech recognition model M1. The learning unit 152 learns network parameters. For example, the learning unit 152 learns network parameters of various machine learning models such as the speech recognition model M1. The learning unit 152 performs a learning process using learning data (teacher data) stored in the data storage unit 142, thereby generating various machine learning models such as the speech recognition model M1. The learning unit 152 generates various machine learning models such as the speech recognition model M1 by learning network parameters of various machine learning models such as the speech recognition model M1.
[0053] For example, the learning unit 152 performs a learning process using a method such as backpropagation (error backpropagation) so that the recognition result (notation + reading) output by the speech recognition model M1 approaches the correct information associated with the speech information (input information) input to the speech recognition model M1. For example, the learning unit 152 adjusts the value of a weight (i.e., a connection coefficient) that is taken into account when a value is transmitted between nodes through the learning process. In this way, the learning unit 152 learns the speech recognition model M1 through a process such as backpropagation that corrects the parameters (connection coefficients) so as to reduce the error between the output of the speech recognition model M1 and the correct information corresponding to the input. For example, the learning unit 152 generates the speech recognition model M1 by performing a process such as backpropagation so as to minimize a predetermined loss function. This allows the learning unit 152 to perform a learning process to learn the parameters of the speech recognition model M1.
[0054] The models may be generated using a method based on a recurrent neural network such as an RNN (Recurrent Neural Network) or LSTM (Long Short-Term Memory Units), which is an extension of an RNN. The model learning method is not limited to the above-described method, and any known technique may be applied. Each model may be generated using various conventional machine learning techniques, as appropriate. For example, the models may be generated using deep learning techniques. The models may be generated using Transformer techniques. For example, the models may be generated using various deep learning techniques, such as a DNN (Deep Neural Network), an RNN, or a CNN (Convolutional Neural Network), as appropriate. The above descriptions regarding model generation are merely examples, and the models may be generated using a learning method appropriately selected depending on the obtainable information, etc. That is, the learning unit 152 may generate the speech recognition model M1 using any method as long as the speech recognition model M1 can be trained so that, when input information included in the training data is input, it outputs information corresponding to the correct answer information.
[0055] (Reception Department 153) The reception unit 153 receives various information. The reception unit 132 receives sounds uttered by the user. The reception unit 132 receives speech from the user via the input unit 12. The reception unit 132 receives speech from the user via the input unit 12.
[0056] (Generation unit 154) The generation unit 154 generates various information. The generation unit 154 generates various information based on information acquired by the acquisition unit 151. The generation unit 154 generates various information based on information stored in the storage unit 14. The generation unit 154 generates various information based on information accepted by the acceptance unit 153. The generation unit 154 generates various information based on information stored in the model information storage unit 141, the data storage unit 142, the processing result information storage unit 143, etc.
[0057] The generation unit 154 generates a recognition result including a recognized character string indicating a spoken word recognized from the spoken speech and recognition reading information indicating the reading of the spoken word, based on the result of a first speech recognition process in which speech recognition is performed by applying contextual biasing to units of components smaller than a word that are included in a character string obtained by encoding a spoken speech indicated by the spoken speech information, and the result of a second speech recognition process in which speech recognition is performed by applying contextual biasing to units of spoken words that are included in a character string and constitute a meaning. The generation unit 154 generates a recognition result including a recognized character string indicating a spoken word recognized from the spoken speech and recognition reading information indicating the reading of the spoken word, based on the result of a first speech recognition process in which speech recognition is performed by applying contextual biasing to units of characters that are included in a character string obtained by encoding a spoken speech indicated by the spoken speech information, and the result of a second speech recognition process in which speech recognition is performed by applying contextual biasing to units of spoken words that are included in a character string and constitute a meaning.
[0058] The generation unit 154 generates a recognition result including a recognition character string indicating the notation of a spoken word recognized from the spoken voice and recognition reading information. The generation unit 154 generates a recognition result including a recognition character string indicating the notation of a spoken word recognized from the spoken voice and recognition reading information.
[0059] The generation unit 154 generates recognition reading information including a first character that is a character indicating the reading of the spoken word. The generation unit 154 generates recognition reading information that combines the first character with a second character that is a symbol indicating that the first character is the reading of the spoken word. The generation unit 154 generates a recognition result that includes a character string that concatenates a recognition character string that indicates the kanji notation of the spoken word recognized from the spoken voice, and recognition reading information that combines the first character that is a character indicating the reading of the spoken word with the second character that is a symbol indicating that the first character is the reading of the spoken word.
[0060] The generation unit 154 inputs the speech information into the speech recognition model M1 and causes the speech recognition model M1 to output a recognition result corresponding to the speech information, thereby generating a recognition result corresponding to the speech information. The generation unit 154 generates a recognition result corresponding to the user's speech.
[0061] The generation unit 154 may execute a process for generating content to be provided to a user. The generation unit 154 generates content to be displayed on a screen (display unit 13). The generation unit 154 generates content by appropriately using various conventional technologies related to video generation for generating video. The generation unit 154 generates content indicating the results of speech recognition processing. For example, the generation unit 154 generates content indicating the results of speech recognition processing by appropriately using various technologies such as Java (registered trademark). Note that the generation unit 154 may generate content indicating the results of speech recognition processing based on the format of CSS, JavaScript (registered trademark), or HTML. Furthermore, for example, the generation unit 154 may generate content indicating the results of speech recognition processing in various formats such as JPEG (Joint Photographic Experts Group), GIF (Graphics Interchange Format), or PNG (Portable Network Graphics).
[0062] (output unit 155) The output unit 155 outputs various information. The output unit 155 outputs information generated by the generation unit 154. The output unit 155 outputs the recognition result generated by the generation unit 154. For example, the output unit 155 outputs various information in accordance with an operation input by the input unit 12 (also referred to as a "user operation").
[0063] For example, the output unit 155 outputs various pieces of information via the display unit 13. For example, the output unit 155 performs output by causing the display unit 13 to output the various pieces of information. The output unit 155 displays various pieces of information based on the information acquired by the acquisition unit 151. The output unit 155 displays various pieces of information based on the information stored in the storage unit 14. The output unit 155 displays various pieces of information based on the information stored in the model information storage unit 141, the data storage unit 142, the processing result information storage unit 143, etc. The output unit 155 displays various pieces of information generated by the generation unit 154.
[0064] (Transmitter 156) The transmitting unit 156 transmits various types of information. For example, the transmitting unit 156 transmits various types of information to an external information processing device in accordance with a user operation input via the input unit 12. The transmitting unit 156 transmits request information requesting various types of information from the external information processing device in accordance with the user operation to the external information processing device. The transmitting unit 156 may function as an output unit that outputs the information generated by the generating unit 154 to an external device.
[0065] Note that when the above-described information processing and other processes by the control unit 15 are performed by a predetermined application, each unit of the control unit 15 may be realized by, for example, the predetermined application. For example, the information processing and other processes by the control unit 15 may be realized by control information including JavaScript (registered trademark). Furthermore, when the above-described information processing and other processes are performed by a dedicated application, the control unit 15 may have, for example, an application control unit that controls the predetermined application (for example, a voice recognition application) or the dedicated application.
[0066] [4. Voice Recognition Flow] Next, a procedure for voice recognition by the terminal device 10 according to the embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing an example of voice recognition according to the embodiment.
[0067] 5, the terminal device 10 acquires speech information indicating a user's speech (step S101). Then, the terminal device 10 generates a recognized character string and recognized reading information based on the result of a first speech recognition process in which CB is applied to each character included in a character string obtained by symbolizing the speech, and the result of a second speech recognition process in which CB is applied to each spoken word included in the character string that constitutes the meaning (step S102). Then, the terminal device 10 outputs the recognition result (step S103).
[0068] [5. Voice Recognition Example] Based on the above content, an outline of the speech recognition process executed by the terminal device 10 will now be described. Note that explanations of points similar to those described above will be omitted as appropriate.
[0069] [5-1. Formulation] First, an example of the formulation of 2-pass CB will be described. The 2-pass CB in the above-mentioned speech recognition is formulated as shown in the following equation (1).
[0070]
number
[0071] Here, "y" in equation (1) corresponds to the recognition result hypothesis output by the speech recognition model. For example, "y" corresponds to the recognition result hypothesis before the spelling is replaced by 2-pass CB. Also, "y'" in equation (1) corresponds to the recognition result hypothesis after the spelling is replaced by 2-pass CB. "logPc(y'|y)" in equation (1) corresponds to the context score when "y" is converted to "y'". Also, "λ" in equation (1) is a hyperparameter that adjusts the strength of CB.
[0072] The above formulation is similar to the method in Patent Document 1 (also referred to as the "existing method"), but differs from the formulation in the existing method in that the context score in 2-pass CB is defined as "logPc(y'|y)" and argmax is taken for "y'". 2-pass CB uses a phrase to convert the notation label of "y" and then outputs a score corresponding to the conversion. For example, even if "hi_i_lung" is given as the input symbol, if "hi_i_hai" is given as the phrase, the output symbol will be "hi_i_hai". In this case, "y" will be "hi_i_lung" and "y'" will be "hi_i_hai".
[0073] [5-2. Speech Recognition Model] As described above, a simultaneous reading and spelling estimation model is used for two-pass CB speech recognition models such as speech recognition model M1. The simultaneous reading and spelling estimation model is constructed based on the technology disclosed in Non-Patent Document 2 and the like.
[0074] In a 2-pass CB speech recognition model such as speech recognition model M1, the recognition result includes both the reading and the spelling. An example of a recognition result is shown below.
[0075] Input voice (speech): I want to go to Nipponbashi Recognition result (generated information): I want to go to Nihonbashi
[0076] The reading information output by 2-pass CB speech recognition models such as speech recognition model M1 is output in katakana with underscores, and the recognition results are read morpheme by morpheme and output in the order of the spelling. Note that while a simultaneous reading and spelling estimation model can be constructed with any model structure, 2-pass CB speech recognition models such as speech recognition model M1 use, for example, an RNN-T model.
[0077] [5-3. Calculation of context score] [5-3-1.WFST] Hereinafter, the calculation of the context score will be described. Note that the explanation of the same points as those in the context score in the existing method will be omitted as appropriate.
[0078] Below, an example is shown in which a weighted finite-state transducer (sometimes referred to as "WFST") used in 2-pass CB is constructed for the phrase "Nihonbashi." A WFST (Weighted Finite-State Transducer) corresponding to the phrase "Nihonbashi" is constructed as shown in FIG. 6. FIG. 6 is a conceptual diagram showing an example of a weighted finite-state transducer. State transition diagram ST1 shown in FIG. 6 shows a WFST corresponding to the phrase "Nihonbashi." Each of the circles with numbers 1 to 15 written inside them in FIG. 6 corresponds to one of states state1 to state15. For example, the state in which the number inside the circle is 1 corresponds to state1.
[0079] The tags used in the WFST shown in Figure 6 have the following meanings: <eps>: A special symbol that indicates the absence of input or output, similar to ε in the existing method. <hyoki>: Any single character <unk>: Readings not included in the phrase (for example, in the case of the WFST in Figure 6, readings other than "Ni_, Ho_, N_, Ba_, Shi_") #second: disambiguation symbol to distinguish 2nd pass from 1st pass
[0080] Additionally, arcs (arrow lines) without a weight displayed represent a weight of 0. State 14 (when the number inside the circle is 14) and state 15 (when the number inside the circle is 15) each have the role of accepting any string and outputting a score of 0. This makes it possible to accept strings that do not contain phrases, or any strings that are included before or after a phrase. For example, in an input such as "I_K_SA_KI_YUKIKUSHIHA_HA_WA_NIHON_BA_SHI_NIHONBASHI_DES_DES", the "I_K_SA_KI_YUKIKUSHIHA_HA_WA" part is accepted by state 0 (when the number inside the circle is 0) and state 14, while the "DES_DES" part is accepted by state 13 (when the number inside the circle is 13) and state 15.
[0081] [5-3-2. Notation Cache] In 2-pass CB, the context score is calculated using a cache that temporarily stores the characters input to the WFST. The cache is used for the following two purposes:
[0082] Objective 1: To be able to output an appropriate context score even when the word boundaries of the string input to the WFST do not match the phrase. For example, to be able to output the same context score when the string "Nihonbashi_Hashi" is input to the above WFST as when the string "Nihonbashi_Hashi" is input. Objective 2: Replace the notation of the input string with the notation of the phrase on a word-by-word basis.
[0083] In addition, the cache is controlled according to rules 1 to 4 below.
[0084] Rule 1: When the following conditions 1-1 and 1-2 are both met, the input characters are cached. Condition 1-1: The input symbol to the WFST is a notation character ( <hyoki>) Condition 1-2: The output symbol is <eps>is.
[0085] Rule 2: When the following conditions 2-1, 2-2, and 2-3 are all met, the cached string is retrieved and input to the WFST. The contents of the cache are also deleted and the cache is left empty. Condition 2-1: The cache is not empty. Condition 2-2: The input symbol to the WFST is a notation character. Condition 2-3: The above written characters are the first characters of the word (for example, in the case of "Nihonbashibashi", the characters "Sun" and "Hashi" are applicable).
[0086] Rule 3: When both the following conditions 3-1 and 3-2 are met, the contents of the cache are deleted and discarded. Condition 3-1: The input symbol to the WFST is #second. Condition 3-2: The output symbol of the WFST is <eps>isn't it.
[0087] Rule 4: Under other conditions, do not perform any operation on the cache and maintain the contents. For example, Rules 1 and 2 correspond to Objective 1, and Rule 3 corresponds to Objective 2, but the specific processing will be described later.
[0088] [5-3-3. Overview of 1st pass and 2nd pass] As mentioned above, 2-passCB consists of two processes: 1st pass and 2nd pass. An overview of each process is given below.
[0089] (1st pass) This applies when the input to WFST and the phrase match exactly, including the spelling. The process is basically the same as existing methods, except that it can handle mismatched word boundaries by using a notation cache. -Context scores are output character by character.
[0090] (2nd pass) This is applied when the reading of the input to the WFST matches the phrase, but the spelling does not. -Converts the notation part of the input to WFST into phrase notation. Context scores are output on a word-by-word basis.
[0091] In this way, in a two-pass CB speech recognition model such as speech recognition model M1, if the input to the WFST and the phrase match perfectly, including the spelling, the first pass operates on a character-by-character basis, but if a spelling mismatch occurs, the second pass operates on a word-by-word basis.
[0092] For example, according to existing methods such as the above-mentioned Patent Document 1, it is desirable to output the context score on a character-by-character basis. This is because if the context score is output retroactively, such as on a word-by-word basis, there is a high possibility that hypotheses will be deleted by pruning during the beam search. On the other hand, it is desirable to perform the spelling replacement process on a word-by-word basis. Therefore, in a 2-pass CB speech recognition model such as speech recognition model M1, two types of processing, a 1st pass and a 2nd pass, can be performed to achieve both character-by-character context score output and word-by-word spelling replacement.
[0093] [5-3-4.1st pass] The 1st pass is executed by using path PS1, which is the area (region) indicated by the dashed line in FIG. 7, of the WFST shown in FIG. 6. FIG. 7 is a conceptual diagram showing an example of the 1st pass. Note that FIG. 7 is the same as FIG. 6 except that it shows path PS1. Also, FIG. 8 shows the processing of the 1st pass when "Ni_hon_n_nihonbashi_shi_bashi" is input. FIG. 8 is a diagram showing an example of the processing flow of the 1st pass.
[0094] As with existing methods, the context score is obtained character by character, so logPc(y'|y) can be calculated by adding up all the scores output at each processing step. In the above example, ·logPc(y'|y) = 0.1+0.1+0.1+0+0+0.1+0.1+0.51 = 1.01 This becomes:
[0095] Also, among the output symbols <eps>is a special tag that indicates that there is no output, so it is not included in y'. Therefore, in the above example, ·y = Nihonbashi Bridge ·y' = Nihonbashi Since 2-passCB is designed to ignore mismatches in word boundaries, "Nihonbashi" and "Nihonbashi" are considered to be the same recognition result.
[0096] As mentioned above, in the 1st pass, the input to the WFST and the phrase must match perfectly, including the spelling. In the 2-pass CB, however, any spelling can be used. <hyoki>As the input is a WFST, it may not be possible to satisfy this constraint. For this reason, a separate process may be added to compare the output symbols of the WFST with the input symbols and stop the search if they do not match. In this case, if the process stops, the calculation of the context score in the first pass may fail, but the method for executing the beam search in this case will be described later.
[0097] [5-3-5. 2nd pass] The second pass is executed using pass PS2, which is the area (region) indicated by the dashed line in Figure 9, in the WFST shown in Figure 6. Figure 9 is a conceptual diagram showing an example of the second pass. Note that Figure 9 is the same as Figure 6 except that it shows pass PS2. The second pass is processed by inputting the disambiguation symbol #second into the WFST. #second is input only when both of the following conditions 4-1 and 4-2 are met. Condition 4-1: The immediately preceding input symbol is the end of a word (for example, in the case of "Nihonbashibashi", "hon" and "bashi" are applicable). Condition 4-2: The notation cache is not empty.
[0098] Also, Figure 10 shows the processing of 2ndpass when "Ni_hon_n_nihonbashi_shi_bashi" is entered. Figure 10 is a diagram showing an example of the flow of processing 2ndpass. Note that Figure 10 shows an example of input with incorrect notation to give an idea of the processing.
[0099] As with the 1st pass, by integrating the output scores and symbols of all processing steps, the following results are obtained. logPc(y'|y) = 0.199 ·y = Nihonbashi Bridge ·y' = Nihonbashi
[0100] For example, the context score of the 2nd pass is set to be smaller than that of the 1st pass. For example, the context score of the 2nd pass may be set to about 0.2 times that of the 1st pass. For example, in the WFST, an ε transition is added from state 5 (the state where the number in the circle is 5) to state 6 (the state where the number in the circle is 6) and a negative score is output to adjust the ratio.
[0101] While the above is the basic processing of the second pass, applying this process can sometimes make it difficult to output homophones of a phrase. This can occur when adding a short phrase like "hai_i_hai." For example, "hai_i_hai ga_itai_i_itai" (it hurts), even in contexts where "lungs" clearly should be output, resulting in a decrease in recognition accuracy. To address this issue, the second pass can be enhanced with a decision process that references the ASR (Automatic Speech Recognition) score and does not perform spelling substitution if the score is above a threshold. Here, the ASR score is assumed to reflect the degree of confidence with which the speech recognition model outputs spellings. In this case, it is expected that the corresponding ASR score will be high in cases where the correct spelling is clear, such as the example above. Therefore, not performing spelling substitution when the ASR score is above a threshold can reduce the likelihood of the above issue occurring.
[0102] When adding the threshold processing described above, the ASR score is required to calculate the context score, so Equation (1) in the above formulation is modified to Equation (2) below.
[0103]
number
[0104] a in equation (2) is formulated as the following equation (3).
[0105]
number
[0106] In addition, when executing 2ndpass, the following processes 1 and 2 must be added. Process 1: Expand the notation cache to store not only notation characters but also the corresponding ASR scores. Process 2: During the notation replacement process, the average of the cached ASR scores is calculated, and if this value is equal to or greater than a threshold, the second pass process is stopped.
[0107] FIG. 11 shows the processing steps of the 2ndpass with threshold processing added. FIG. 11 is a diagram showing another example of the processing flow of the 2ndpass. In the example of FIG. 11, the input is "Ni_hon_n_nihonbashi_shi_bashi", the ASR scores corresponding to "Ni, Hon, Hashi" are "-0.1, -0.2, -0.3", respectively, and the threshold is set to -1. Also, in FIG. 11, processing steps before processing step 4 (e.g., processing steps 1 to 3) are not shown. For example, if the processing shown in FIG. 11 stops, the calculation of the context score will fail.
[0108] [5-4. Beam Search] Searching for speech recognition results is performed using beam search, as with existing methods. However, the WFST used in 2-pass CB may allow multiple transition patterns for the same input. For example, if "Ni_Hon_N_Go_Nihongo" is input into the WFST shown in Figure 6, the following transitions such as Pattern 1 and Pattern 2 are possible.
[0109] Pattern 1 state: 0→1→2→3→failed Output score: Failed to calculate (none)
[0110] Pattern 2 state: 0→14→14→14→14→0→14→0→14→0 Output score: 0
[0111] Note that in Pattern 1, the search failed midway because an input that did not match the phrase was entered. When multiple transition patterns are possible like this, the output symbols and output scores differ for each pattern. However, there is only one context score that can be input to the score combiner, which is a component that combines multiple scores (for example, the ASR score and the context score), and only one recognition result that can be output as the final result.
[0112] Therefore, in a 2-pass CB speech recognition model such as speech recognition model M1, the maximum value of the cumulative score at processing step t is calculated, and the difference between this and the maximum value at processing step t-1 is input to a score combiner. This process is then executed up to the final step, and the output symbol sequence of the transition pattern with the highest cumulative output score is finally output as the recognition result (y'). An example of the process is shown in FIG. 12. FIG. 12 is a diagram showing another example of the processing flow using beam search. Note that explanations of points similar to existing methods, such as the score combiner, will be omitted where appropriate.
[0113] In the process shown in Figure 12, the score of pattern 2 is selected as the maximum cumulative score in the final step. Therefore, y' becomes "Ni_Hon_N_Go_Nihongo", which is the output symbol sequence of pattern 2. As shown in processing step 4, if there is a pattern that has failed to transition along the way, the maximum cumulative score is selected from the remaining transition patterns to determine the input to the score combiner.
[0114] As mentioned above, the 1st pass and 2nd pass processing are forcibly stopped under certain conditions (such as when threshold processing is added), and the same processing is performed in this case. If both the 1st pass and 2nd pass processing fail, as in step 4 of the example above, the maximum cumulative score will be 0 and the input to the score combiner will be a negative value. Therefore, the same processing as the failure arc in existing methods is also performed in 2-pass CB.
[0115] [6. Effects] As described above, the speech recognition device according to the embodiment (the terminal device 10 in the embodiment) includes an acquisition unit (the acquisition unit 151 in the embodiment), a generation unit (the generation unit 154 in the embodiment), and an output unit (the output unit 155 in the embodiment). The acquisition unit acquires speech information, which is speech information indicating a user's speech. The generation unit generates a recognition result including a recognition string indicating a spoken word recognized from the speech and recognition reading information indicating the reading of the spoken word, based on the result of a first speech recognition process that performs speech recognition by applying contextual biasing to a unit of components smaller than a word that are included in a character string that encodes the speech indicated by the speech information, and the result of a second speech recognition process that performs speech recognition by applying contextual biasing to a unit of spoken words that are included in a character string and constitute meaning. The output unit outputs the recognition result generated by the generation unit.
[0116] In this way, the speech recognition device of the embodiment can appropriately generate information that includes both the character string indicating the word contained in the utterance and the pronunciation of the word by generating a recognition result that includes a recognized character string and recognized pronunciation information based on the result of a first speech recognition process that performs speech recognition by applying contextual biasing to units of components at least smaller than words, such as characters, and the result of a second speech recognition process that performs speech recognition by applying contextual biasing to units of spoken words, which are words contained in a character string and that constitute meaning.
[0117] In addition, in the speech recognition device according to the embodiment, the generation unit generates a recognition result including a recognized character string indicating the spelling of an uttered word recognized from an uttered speech, and recognized reading information.
[0118] In this way, the speech recognition device of the embodiment can appropriately generate information that includes both the character string indicating the word included in the utterance and the pronunciation of the word by generating a recognition result that includes the notation of the spoken word and its pronunciation.
[0119] In addition, in the speech recognition device according to the embodiment, the generation unit generates a recognition result including a recognized character string indicating the kanji notation of an uttered word recognized from an uttered speech, and recognized reading information.
[0120] In this way, the speech recognition device of the embodiment can appropriately generate information that includes both the character string indicating the word included in the utterance and the reading of the word by generating a recognition result that includes the kanji notation of the spoken word and the reading of the kanji.
[0121] In addition, in the voice recognition device according to the embodiment, the generation unit generates recognized reading information including a first character that is a character indicating the reading of the spoken word.
[0122] In this way, the speech recognition device of the embodiment can appropriately generate information including both a string of characters indicating the word included in the utterance and the reading of the word by generating recognition reading information including a first character that indicates the reading of the spoken word.
[0123] In the voice recognition device according to the embodiment, the generation unit generates recognized reading information that combines the first character with a second character that is a symbol indicating that the first character is the reading of a spoken word.
[0124] In this way, the speech recognition device of the embodiment can appropriately generate information that includes both a string of characters indicating a word included in an utterance and the pronunciation of the word by generating recognition pronunciation information that combines the first character with the second character, which is a symbol indicating that the first character is the pronunciation of the spoken word.
[0125] In addition, in the voice recognition device according to the embodiment, the generation unit generates a recognition result including a character string concatenating a recognition string indicating the kanji notation of a spoken word recognized from a spoken voice, and recognition reading information that combines a first character that is a character indicating the reading of the spoken word and a second character that is a symbol indicating that the first character is the reading of the spoken word.
[0126] In this way, the speech recognition device of the embodiment can appropriately generate information including both a string indicating the word contained in the utterance and the reading of the word by generating a string that concatenates recognition reading information that combines a first character that is a character indicating the kanji notation and the reading of the spoken word and a second character that is a symbol indicating that the first character is the reading of the spoken word.
[0127] In addition, in the speech recognition device according to the embodiment, the acquisition unit acquires a speech recognition model (speech recognition model M1 in the embodiment) that outputs a recognition result corresponding to the speech information through first and second speech recognition processes in response to input of speech information. The generation unit inputs the speech information to the speech recognition model and causes the speech recognition model to output the recognition result corresponding to the speech information, thereby generating the recognition result corresponding to the speech information.
[0128] In this way, the speech recognition device according to the embodiment inputs spoken speech information into a speech recognition model that outputs a recognition result corresponding to the speech information through a first speech recognition process and a second speech recognition process in response to input of speech information, and causes the speech recognition model to output a recognition result corresponding to the spoken speech information, thereby appropriately generating information that includes both character strings indicating words included in the speech and the pronunciation of the words.
[0129] Moreover, the speech recognition device according to the embodiment is a terminal device used by a user, and the generation unit generates a recognition result corresponding to an utterance of the user.
[0130] In this way, the speech recognition device according to the embodiment is a terminal device used by a user, and by generating recognition results corresponding to the user's speech, it is possible to appropriately generate information that includes both character strings indicating words included in the speech and the pronunciation of the words.
[0131] [7. Program] The above-described processes performed by the terminal device 10 and the server device 50 are realized by an information processing program according to the present application. For example, the generation unit 154 and the like of the terminal device 10 are realized by a CPU, an MPU, or the like of the terminal device 10, in which an information processing program included in, for example, a voice recognition application uses RAM as a working area and executes an information processing procedure according to the information processing program.
[0132] It should be noted that not all of the processes executed by the terminal device 10 or the server device 50 according to the present application need be realized by an information processing program. For example, information outside the terminal device 10 may be acquired by an OS (Operating System) of the terminal device 10. In other words, the information processing program itself may not execute the processes executed by the terminal device 10 as described above, but may instead realize the processes of the terminal device 10 described above by receiving data acquired by the OS (for example, data used to display (play) content such as images).
[0133] [8. Hardware Configuration] An information processing device such as the terminal device 10 according to the above-described embodiment is realized by, for example, a computer 1000 configured as shown in Fig. 13. Fig. 13 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0134] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.
[0135] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, etc. The communication interface 1500 receives data from other devices via a predetermined network N and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined network N.
[0136] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.
[0137] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0138] For example, when the computer 1000 functions as the terminal device 10 according to the embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to realize the functions of the control unit 15. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via a predetermined network N.
[0139] The above describes the embodiments of the present application in detail based on the drawings, but these are merely examples, and the present invention can be implemented in other forms that include various modifications and improvements based on the knowledge of those skilled in the art, including the aspects described in the Disclosure of the Invention.
[0140] [9. Other] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the information, including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0141] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0142] Furthermore, the processes described in the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the process contents.
[0143] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]
[0144] 1. Voice Recognition System 10 Terminal device (voice recognition device) 14 Storage section 141 Model information storage unit 142 Data storage unit 143 Processing result information storage unit 15 Control Unit 151 Acquisition Department 152 Learning Department 153 Reception Department 154 Generation part 155 Output section 156 Transmitter 50 Server device N Network< / hyoki> < / eps> < / eps> < / eps> < / hyoki> < / unk> < / hyoki> < / eps>
Claims
1. an acquisition unit that acquires speech information that is speech information indicating a speech of a user; a generation unit that generates a recognition result including a recognition character string indicating a spoken word recognized from the spoken speech and recognition reading information indicating the reading of the spoken word, based on a result of a first speech recognition process that performs speech recognition by applying contextual biasing to a unit of components smaller than a word that are included in a character string that encodes the spoken speech indicated by the spoken speech information, and a result of a second speech recognition process that performs speech recognition by applying contextual biasing to a unit of spoken words that are included in the character string and constitute meaning; an output unit that outputs the recognition result generated by the generation unit; A speech recognition device comprising:
2. The generation unit generating the recognition result including the recognized character string indicating the spelling of the spoken word recognized from the spoken voice and the recognized reading information; 2. The speech recognition device according to claim 1.
3. The generation unit generating the recognition result including the recognized character string indicating the kanji notation of the spoken word recognized from the spoken voice and the recognized reading information; 2. The speech recognition device according to claim 1.
4. The generation unit generating the recognized reading information including a first character that is a character indicating the reading of the spoken word; 2. The speech recognition device according to claim 1.
5. The generation unit generating the recognized reading information by combining the first character with a second character, which is a symbol indicating that the first character is the reading of the spoken word; 5. The speech recognition device according to claim 4.
6. The generation unit Generate the recognition result indicating a character string obtained by concatenating the recognized character string indicating the Kanji notation of the spoken word recognized from the spoken voice and the recognized reading information that combines a first character that is a character indicating the reading of the spoken word and a second character that is a symbol indicating that the first character is the reading of the spoken word.
2. The speech recognition device according to claim 1.
7. The acquisition unit acquiring a speech recognition model that outputs the recognition result corresponding to the speech information by the first speech recognition process and the second speech recognition process in response to the input of the speech information; The generation unit The speech information is input to the speech recognition model, and the speech recognition model outputs the recognition result corresponding to the speech information, thereby generating the recognition result corresponding to the speech information.
2. The speech recognition device according to claim 1.
8. A terminal device used by the user, The generation unit generating the recognition result corresponding to the user's utterance; 2. The speech recognition device according to claim 1.
9. 1. A computer-implemented method for speech recognition, comprising: an acquisition step of acquiring speech information that is speech information indicating a speech voice of a user; a generating step of generating a recognition result including a recognized character string indicating a spoken word recognized from the spoken voice and recognized reading information indicating the reading of the spoken word, based on a result of a first speech recognition process in which contextual biasing is applied to a unit of component smaller than a word included in a character string obtained by encoding the spoken voice indicated by the spoken voice information, and a result of a second speech recognition process in which contextual biasing is applied to a unit of spoken word which is a word included in the character string and constituting a meaning; an output step of outputting the recognition result generated by the generation step; A speech recognition method comprising:
10. an acquisition step of acquiring speech information that is speech information indicating a speech of a user; a generation step of generating a recognition result including a recognition character string indicating a spoken word recognized from the spoken speech and recognition reading information indicating the reading of the spoken word, based on a result of a first speech recognition process that performs speech recognition by applying contextual biasing to a unit of a component smaller than at least a word that is included in a character string that encodes the spoken speech indicated by the spoken speech information, and a result of a second speech recognition process that performs speech recognition by applying contextual biasing to a unit of a spoken word that is included in the character string and constitutes meaning; an output step of outputting the recognition result generated by the generation step; A speech recognition program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Phoneme-based context analysis for multilingual speech recognition using end-to-end models
JP7092953B2
Using context information in an end-to-end model for speech recognition
JP7417634B2