Speech recognition apparatus, method, and program
By generating and merging multiple extended speech data, the performance degradation caused by differences in noise, speech rate, speaker characteristics, and amplitude in speech recognition technology is solved, achieving a more efficient speech recognition effect.
Patent Information
- Application Number
- CN202210188336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-31
- Filing Date
- 2022-02-28
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2042-02-28
AI Technical Summary
Existing speech recognition technologies suffer from significant performance degradation when faced with environmental noise, varying speech rates, speaker characteristics, and differences in amplitude. Furthermore, robust speech recognition methods using multiple sound models or a single model suffer from time and memory consumption issues in terms of practicality.
By generating multiple extended speech data, the speech rate, volume, and tone quality are transformed using the data extension unit. Combined with voice score calculation, adjustment, merging, and grid generation, a merged grid is generated, thereby improving speech recognition performance.
Even under conditions of varying noise levels, speech rate, speaker characteristics, and amplitude, it can still improve the accuracy and efficiency of speech recognition, solving the problem of declining recognition performance in existing technologies.
Smart Images

Figure CN115482822B_ABST
Abstract
Description
[0001] This application is based on Japanese Patent Application 2021-091236 (filed on May 31, 2021), and enjoys priority based on that application. The entire contents of that application are incorporated herein by reference. Technical Field
[0002] The embodiments of the present invention relate to a speech recognition device, method, and program. Background Technology
[0003] There are techniques that use pre-learned sound models based on large amounts of general speech data to recognize spoken speech. Consider the following four reasons, for example, as major causes of decreased performance in spoken speech recognition.
[0004] (Main reason 1) The situation where environmental noise is mixed in with what is said.
[0005] (Main Reason 2) Speech rate significantly different from general speech data. This refers to situations where the person being recognized speaks significantly faster or slower.
[0006] (Main Reason 3) Cases where the speaker's characteristics are significantly different from those of general speech data. For example, general speech data consists of the voices of adults, while the target of recognition is the voice of a child.
[0007] (Main Reason 4) The amplitude of the input speech is significantly different from that of general speech data. For example, the gain of the microphone collecting the spoken speech is set to be significantly low.
[0008] If any of the above four main reasons occur, the features of the speech data being recognized as spoken speech will be inconsistent with the features of general speech data, resulting in a significant decrease in speech recognition performance.
[0009] One effective method to address the aforementioned problems is to use multiple voice models to recognize input speech and then merge the recognition results. By learning different noise levels, speech rates, speaker characteristics, and amplitudes of speech data for each voice model, it is possible to address the four main causes. However, learning multiple voice models is time-consuming. Furthermore, using multiple voice models for speech recognition via a computer consumes significant memory, making it impractical.
[0010] As another effective approach to solving this problem, there exists a method using a single sound model for robust speech recognition against noise. In this method, the input signal containing noise and the speech enhancement signal after noise suppression are combined and input into a single sound model. However, this method can address four main issues (Main Issue 1), but not (Main Issues 2) through (Main Issue 4). Furthermore, the sound model must be learned based on a pre-determined amount of input signal and speech enhancement signal, which is a significant limitation. Summary of the Invention
[0011] The technical problem to be solved by the present invention is to provide a speech recognition device, method and program that can improve speech recognition performance.
[0012] One embodiment of the speech recognition apparatus includes a data expansion unit, a voice score calculation unit, an adjustment unit, a voice score merging unit, a lattice generation unit, and a search unit. The data expansion unit generates multiple expanded voice data based on input speech data. The voice score calculation unit generates multiple voice scores based on each of the expanded voice data and a voice model. The adjustment unit generates multiple adjusted voice scores by resampling each of the multiple voice scores. The voice score merging unit generates a merged voice score by merging the multiple adjusted voice scores. The lattice generation unit generates a merged lattice based on the merged voice scores, a pronunciation dictionary, and a language model. The search unit searches for the speech recognition result with the highest likelihood from the merged lattice.
[0013] The speech recognition device with the above structure can improve speech recognition performance. Attached Figure Description
[0014] Figure 1 This is a block diagram illustrating the structure of the voice recognition device according to the first embodiment.
[0015] Figure 2 This is an example Figure 1 A block diagram of the structure of the merging processing unit.
[0016] Figure 3 This is an example Figure 1 A flowchart of the operation of the voice recognition device.
[0017] Figure 4 This is an example Figure 3 The flowchart for merging the flowcharts.
[0018] Figure 5 This is an explanation Figure 2 The resampling diagram in the adjustment section.
[0019] Figure 6This is a block diagram illustrating the structure of the merging processing unit of the voice recognition device according to a variation of the first embodiment.
[0020] Figure 7 This is a flowchart illustrating the merging process in the operation of a voice recognition device according to a variation of the first embodiment.
[0021] Figure 8 This is a table showing the experimental results of a speech recognition device according to a variation of the first embodiment.
[0022] Figure 9 This is a block diagram illustrating the structure of the voice recognition device according to the second embodiment.
[0023] Figure 10 This is an example Figure 9 The parameters automatically determine the structure of the part.
[0024] Figure 11 This is an example Figure 9 A flowchart of the operation of the voice recognition device.
[0025] Figure 12 This is an example Figure 9 The flowchart is a process for automatically estimating the parameters of the flowchart.
[0026] Figure 13 This is a block diagram illustrating the structure of the voice recognition device according to the third embodiment.
[0027] Figure 14 This is a block diagram illustrating the structure of the voice recognition device according to the fourth embodiment.
[0028] Figure 15 This is a block diagram illustrating the hardware structure of a computer according to one implementation method.
[0029] Figure 16 This is a block diagram illustrating the structure of a speech recognition system, including conventional speech recognition devices.
[0030] (Explanation of reference numerals in the attached diagram)
[0031] 10: Speech recognition device; 11: Voice score calculation unit; 12: Grid generation unit; 13: Search unit; 14: Voice model storage unit; 15: Pronunciation dictionary storage unit; 16: Language model storage unit; 20: Recording device; 30: Output device; 100: Speech recognition device; 110: Data expansion unit; 120: Merging processing unit; 121: Voice score calculation unit; 121-1: First calculation unit; 121-2: Second calculation unit; 121-N: Nth calculation unit; 122 122-1: First Adjustment Unit; 122-2: Second Adjustment Unit; 122-N: Nth Adjustment Unit; 123: Sound Score Merging Unit; 124: Mesh Generation Unit; 120A: Merging Processing Unit; 121A: Sound Score Calculation Unit; 121A-1: First Calculation Unit; 121A-2: Second Calculation Unit; 121A-N: Nth Calculation Unit; 122A: Mesh Generation Unit; 122A-1: First Generation Unit; 122A-2: Second Generation Unit; 122A-N 123A: Nth generation unit; 130: Mesh merging unit; 140: Search unit; 150: Sound model storage unit; 160: Pronunciation dictionary storage unit; 200: Language model storage unit; 210: Speech recognition device; 211: Automatic parameter determination unit; 212: Amplitude extraction unit; 213: General amplitude data storage unit; 214: Volume transformation parameter estimation unit; 215: Pitch extraction unit; 216: General pitch data storage unit; 217: Sound quality transformation parameter estimation unit; 217: Speech rate extraction unit; 218: General speech rate data storage unit; 219: Speech rate transformation parameter estimation unit; 300: Speech recognition device; 310: Adaptation unit; 320: Adapted sound model storage unit; 400: Speech recognition device; 410: Automatic parameter determination unit; 420: Adaptation unit; 430: Adapted sound model storage unit; 500: Computer; 530: Program memory; 540: Auxiliary storage device; 550: Input / output interface; 560: Bus. Detailed Implementation
[0032] First, let me give an overview of previous voice recognition devices.
[0033] Figure 16 This is a block diagram illustrating the structure of a voice recognition system including a conventional voice recognition device 10. The voice recognition system includes a voice recognition device 10, a recording device 20, and an output device 30.
[0034] The recording device 20 acquires speech data that is intended for speech recognition. The recording device 20 is, for example, a microphone. The recording device 20 outputs the acquired speech data to the speech recognition device 10. The speech data acquired by the recording device 20 will henceforth be referred to as input speech data.
[0035] The speech recognition device 10 includes a voice score calculation unit 11, a grid generation unit 12, a search unit 13, a voice model storage unit 14, a pronunciation dictionary storage unit 15, and a language model storage unit 16. Hereinafter, the voice model storage unit 14, the pronunciation dictionary storage unit 15, and the language model storage unit 16 will be described first.
[0036] The sound model storage unit 14 stores a sound model. The sound model is, for example, a learned model obtained through machine learning using pre-learned speech data. As machine learning, for example, a DNN (Deep Neural Network) is used. Specifically, the sound model is, for example, a single model learned by taking the waveform of speech data as input at least one unit—phoneme, syllable, character, word segment, or word—and outputting a posterior probability corresponding to the sound score, using the aforementioned DNN. Alternatively, the sound model may also be a model learned by taking features (or feature vectors) extracted from the waveform of the speech data, such as input power spectrum or Mel filter bank features.
[0037] The pronunciation dictionary storage unit 15 stores a pronunciation dictionary. A pronunciation dictionary is, for example, a dictionary that represents a word using a sequence of phonemes (phoneme sequence). The pronunciation dictionary is used to obtain words based on sound scores.
[0038] The language model storage unit 16 stores a language model. A language model is a model that describes the rules and constraints that make up sentences composed of word sequences. For example, language models may use rule-based grammar description methods or statistical methods such as N-grams. The language model is used to output the probabilities of multiple candidates for spoken sentences based on the recognition results composed of word sequences.
[0039] The sound score calculation unit 11 receives input speech data from the recording device 20 and a sound model from the sound model storage unit 14. The sound score calculation unit 11 generates a sound score based on the input speech data and the sound model. The sound score calculation unit 11 outputs the generated sound score to the grid generation unit 12.
[0040] Specifically, the sound score calculation unit 11, for example, divides the waveform data, which is the input speech data, into frames and generates a sound score for each frame. Alternatively, the sound score calculation unit 11 may also use feature vectors obtained from the waveform data divided into frames, such as those represented by Mel filter bank features, to generate the sound score. These can be appropriately changed depending on the type of sound model.
[0041] In other words, the sound score calculation unit 11 inputs the waveform data or feature vectors that are segmented according to each frame into the sound model and generates a sound score for each frame.
[0042] The mesh generation unit 12 receives voice scores from the voice score calculation unit 11, a pronunciation dictionary from the pronunciation dictionary storage unit 15, and a language model from the language model storage unit 16. The mesh generation unit 12 generates a mesh based on the voice scores, the pronunciation dictionary, and the language model. The mesh generation unit 12 outputs the generated mesh to the search unit 13.
[0043] Specifically, the grid generation unit 12 outputs higher-order candidates for the output word sequence based on sound scores, a pronunciation dictionary, and a language model. These higher-order candidates are output as a grid, where the higher-order candidates of the output word sequence are set as nodes, and the likelihood of each higher-order candidate word is set as an edge. In a broader sense, the grid is formed by setting candidate words based on speech recognition as nodes and the likelihood of each candidate word as an edge. Furthermore, the grid can also be referred to as a word grid.
[0044] The search unit 13 receives a grid from the grid generation unit 12. The search unit 13 searches the grid for the speech recognition result with the highest likelihood. The search unit 13 outputs the speech recognition result to the output device 30.
[0045] Furthermore, in the generation of upper-order candidates for the output word sequence in the grid generation unit 12 and the search for speech recognition results in the search unit 13, methods described in reference 1 (D. Rybach, J. Schalkwyk, M. Riley, “On Lattice Generation for Large Vocabulary Speech Recognition,” IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2017) can be used, for example.
[0046] The output device 30 receives the speech recognition result from the speech recognition device 10. The output device 30 is, for example, a display. The output device 30 converts the speech recognition result into the desired display format and presents it to the user.
[0047] Hereinafter, various embodiments of the voice recognition device will be described in detail with reference to the figures.
[0048] (First Implementation)
[0049] Figure 1This is a block diagram illustrating the structure of the speech recognition device 100 according to the first embodiment. The speech recognition device 100 includes a data expansion unit 110, a merging processing unit 120, a search unit 130, a voice model storage unit 140, a pronunciation dictionary storage unit 150, and a language model storage unit 160. Furthermore, the speech recognition device 100 may also include an acquisition unit for acquiring input speech data and an input speech data storage unit for storing the input speech data. Additionally, the voice model storage unit 140, the pronunciation dictionary storage unit 150, and the language model storage unit 160 may be combined into one or more storage units, or they may be provided separately or combined externally to the speech recognition device 100.
[0050] Furthermore, the sound model storage unit 140, the pronunciation dictionary storage unit 150, and the language model storage unit 160 are related to... Figure 16 The sound model storage unit 14, the pronunciation dictionary storage unit 15, and the language model storage unit 16 have roughly the same structure, so the description is omitted.
[0051] The data expansion unit 110 receives input voice data from a recording device (not shown). Based on the input voice data, the data expansion unit 110 generates multiple expanded voice data sets. The data expansion unit 110 outputs the multiple expanded voice data sets to the merging processing unit 120.
[0052] Specifically, the data expansion unit 110 performs at least one of the following transformation processes on the input speech data: speech rate transformation, volume transformation, and tone quality transformation, thereby generating at least one of multiple expanded speech data. Furthermore, the multiple expanded speech data may also include the input speech data. Hereinafter, the transformation processing will be explained separately for speech rate transformation, volume transformation, and tone quality transformation.
[0053] When the transformation process is speech rate transformation, the data expansion unit 110 generates expanded speech data by performing a speech rate transformation that sets the input speech data to a times the speed of speech data. The coefficient 'a' is a real number satisfying the condition that 'a>0' and 'a≠1', and is hereinafter referred to as the "speech rate transformation parameter". Regarding speech rate transformation, for example, it can be achieved by reproducing speech at a sampling rate different from the sampling rate of the input speech data, and then transforming the speech reproduced at the different sampling rate back to the original sampling rate. The speech rate transformation parameter 'a' can be any value satisfying the above conditions; commonly used values are 0.9 and 1.1.
[0054] When the transformation process is a volume transformation, the data expansion unit 110 generates expanded speech data by performing a volume transformation that sets the amplitude of the waveform of the input speech data to b times. Regarding the coefficient b, for example, when the input speech data is in 16-bit format, it is a real number that satisfies 0 < b < (32767 / the maximum value of the amplitude of the speech data), which will be referred to as the "volume transformation parameter" hereafter. Regarding the volume transformation parameter b, it can be randomly selected based on the above conditions.
[0055] When the transformation process is a voice quality transformation, the data expansion unit 110 generates expanded speech data by performing a voice quality transformation that sets the pitch of the input speech data to c times. Regarding the coefficient c, it is a real number greater than 0, which will be referred to as the "voice quality transformation parameter" hereafter. Voice quality transformation can be achieved, for example, by using Pitch Synchronous Overlap and Add (PSOLA).
[0056] In addition, PSOLA is described, for example, in Reference 2 (E. Moulines, and F. Charpentier, “Pitch synchronous waveform processing techniques for text-to-speech synthesis using diphones,” Speech Commn., 9: 453 - 467, 1990.) and so on.
[0057] In addition, the data expansion unit 110 can use one of the speech rate transformation, volume transformation, and voice quality transformation as the transformation process, or can use a combination of these multiple ones. Additionally, the data expansion unit 110 can also set the speech rate transformation parameter a, volume transformation parameter b, and voice quality transformation parameter c as transformation parameters to generate multiple expanded speech data. Regarding the type of transformation process, the number of generated expanded speech data, and the combination of transformation parameters, they can be arbitrarily set by the user.
[0058] The merging processing unit 120 receives multiple expanded speech data from the data expansion unit 110, receives the voice model from the voice model storage unit 140, receives the pronunciation dictionary from the pronunciation dictionary storage unit 150, and receives the language model from the language model storage unit 160. The merging processing unit 120 generates a merged grid by performing a merging process using the multiple expanded speech data. The merging processing unit 120 outputs the merged grid to the search unit 130. Next, use Figure 2 to illustrate the more specific structure of the merging processing unit 120.
[0059] Figure 2 is an illustration Figure 1A block diagram of the structure of the merging processing unit 120. The merging processing unit 120 includes a sound score calculation unit 121, an adjustment unit 122, a sound score merging unit 123, and a grid generation unit 124.
[0060] The voice score calculation unit 121 receives multiple extended speech data from the data extension unit 110 and a voice model from the voice model storage unit 140. The voice score calculation unit 121 generates multiple voice scores based on each of the extended speech data and the voice model. The voice score calculation unit 121 outputs the generated multiple voice scores to the adjustment unit 122. Furthermore, the specific generation of the voice scores is related to… Figure 16 The sound score calculation unit 11 is roughly the same.
[0061] The adjustment unit 122 receives multiple sound scores from the sound score calculation unit 121. The adjustment unit 122 generates multiple adjusted sound scores by resampling each of the multiple sound scores. The adjustment unit 122 outputs the generated multiple adjusted sound scores to the sound score merging unit 123.
[0062] Specifically, the adjustment unit 122 generates multiple adjusted sound scores by resampling each of the multiple sound scores in a manner that makes the number of time frames corresponding to each of the multiple sound scores consistent with the number of time frames of the input speech data. Furthermore, as the number of frames to make them consistent, the adjustment unit 122 can use either the input speech data as a reference or any extended speech data as a reference.
[0063] Furthermore, the voice score calculation unit 121 and the adjustment unit 122 may each have multiple calculation units and multiple adjustment units, matching the amount of extended voice data to be generated. For example, when the amount of extended voice data is N (N>1), the voice score calculation unit 121 has a first calculation unit 121-1, a second calculation unit 121-2, ..., an Nth calculation unit 121-N, and the adjustment unit 122 has a first adjustment unit 122-1, a second adjustment unit 122-2, ..., an Nth adjustment unit 122-N. Therefore, the voice score calculation unit 121 outputs N voice scores, and the adjustment unit 122 outputs N adjusted voice scores.
[0064] The sound score merging unit 123 receives multiple adjusted sound scores from the adjustment unit 122. The sound score merging unit 123 generates a merged sound score by merging the multiple adjusted sound scores. The sound score merging unit 123 outputs the generated merged sound score to the mesh generation unit 124.
[0065] Specifically, the sound score merging unit 123 generates a merged sound score by calculating at least one of the average, median, and maximum values of multiple adjusted sound scores. Furthermore, the sound score merging unit 123 can either combine the types of values to be calculated (average, median, and maximum) separately, or change the types of values to be calculated for each frame.
[0066] The mesh generation unit 124 receives merged sound scores from the sound score merging unit 123, a pronunciation dictionary from the pronunciation dictionary storage unit 150, and a language model from the language model storage unit 160. The mesh generation unit 124 generates a merged mesh based on the merged sound scores, the pronunciation dictionary, and the language model. The mesh generation unit 124 outputs the generated merged mesh to the search unit 130. Furthermore, the specific structure of the mesh generation unit 124 is similar to... Figure 16 The mesh generation unit 12 is roughly the same.
[0067] Search unit 130 receives merged grids from merge processing unit 120. Search unit 130 searches for the speech recognition result with the highest likelihood from the merged grids. Search unit 130 outputs the speech recognition result to an output device (not shown). Furthermore, the specific structure of search unit 130 is similar to... Figure 16 The search department 13 is roughly the same.
[0068] The structure of the voice recognition device 100 according to the first embodiment has been described above. Next, using... Figure 3 The flowchart is used to illustrate the operation of the voice recognition device 100.
[0069] Figure 3 This is an example Figure 1 A flowchart of the operation of the voice recognition device 100. Figure 3 The flowchart, for example, illustrates a series of steps to output speech recognition results from a grid corresponding to a sentence of input speech data.
[0070] (Step ST110)
[0071] The voice recognition device 100 acquires input voice data from the recording device.
[0072] (Step ST120)
[0073] After acquiring the input voice data, the data extension unit 110 generates multiple extended voice data based on the input voice data.
[0074] (Step ST130)
[0075] After generating multiple extended speech data, the merging processing unit 120 performs merging processing using the multiple extended speech data to generate a merged mesh. Hereafter, the processing in step ST130 will be referred to as "merging processing". Figure 4 The flowchart below illustrates a specific example of the merging process.
[0076] Figure 4 This is an example Figure 3 The flowchart for merging the flowcharts. Figure 4 The flowchart is transferred from step ST120.
[0077] (Step ST131)
[0078] After generating multiple extended speech data, the sound score calculation unit 121 generates multiple sound scores based on each extended speech data and the sound model of the multiple extended speech data.
[0079] (Step ST132)
[0080] After generating multiple sound scores, the adjustment unit 122 generates multiple adjusted sound scores by resampling each of the multiple sound scores. The following examples illustrate the processing of the adjustment unit 122.
[0081] For example, when merging multiple sound scores (hereinafter referred to as N sound scores) to generate a merged sound score, it is expected that the sound score merging unit 123 will perform processing on a per-frame basis, provided that the time frame number of each of the N sound scores is consistent.
[0082] However, for example, when generating extended speech data through speech rate transformation, the duration of the input speech data differs from the duration of the generated extended speech data, resulting in a problem where the number of time frames corresponding to each sound score is inconsistent. Because of this problem, the sound score merging unit 123 cannot perform processing on a frame-by-frame basis. Therefore, the above problem is addressed by the adjustment unit 122 performing processing to ensure consistency in the number of time frames corresponding to each of the multiple sound scores.
[0083] Let T be the number of time frames for the input speech data, and t be the index of each time frame (1 ≤ t ≤ T). Additionally, let T be the number of time frames for the N extended speech data. n (1≤n≤N). Let Y be the sound score of the t-th frame when the nth extended speech data has been input. t n The Y t n It is a K-dimensional (K is a natural number) vector.
[0084] If the nth extended speech data is generated through speech rate transformation and tone quality transformation, then Tn =T holds true. However, when the extended speech data is generated through speech rate variation, it becomes T. n ≠T, therefore it is necessary to start from T n The sound score corresponding to the frame is transformed into the sound score corresponding to the T frame.
[0085] The above transformation can be performed through the following process. First, the adjustment unit 122 adjusts the frame from frame 1 to frame T. n Frame sound score Y t n Extract the k-th dimension (1≤k≤K) from the sample (step 1). Next, the adjustment unit 122 will extract the T dimension. n Each score is considered T. n The time series data of the sample was used to create a T / T chart. n The score is obtained by resampling at a sampling rate of times. Therefore, the adjustment unit 122 can adjust T... n The scores are transformed into T scores (step 2). Then, the adjustment unit 122 repeats steps 1 and 2 above with respect to 1 ≤ k ≤ K, thereby enabling the transformation of the scores into T scores. n The corresponding amount of sound score Y in the frame t n The sound score is transformed into a quantity corresponding to the T-frame. This transformed sound score is the adjusted sound score described above.
[0086] Figure 5 This is an explanation Figure 2 The resampled image in the adjustment section 122. Figure 5 The text shows the process from T. n =5 downsampled to T=4. Adjustment unit 122, for example, from frame 1 to T... n The first dimension is extracted from the sound scores of 5 frames. Next, the adjustment unit 122 treats the extracted 5 scores as time-series data of 5 samples and generates a score by downsampling at a 4 / 5 times sampling rate. Thus, the adjustment unit 122 can transform the 5 scores into 4 scores.
[0087] (Step ST133)
[0088] After generating multiple adjusted voice scores, the voice score merging unit 123 generates a merged voice score by merging the multiple adjusted voice scores. The following examples illustrate the processing of the voice score merging unit 123.
[0089] The sound score merging unit 123 takes N adjusted sound scores, which are transformed into quantities corresponding to T frames, as input, and outputs a single merged sound score corresponding to T frames. The adjusted sound score of the t-th frame when the nth extended speech data is input is set as Z. t nAdditionally, the score of the merged audio is set to S. t Here, Z t n St and St are K-dimensional vectors, represented by the following equations (1) and (2), respectively.
[0090]
[0091] S t =[s t,1 s t,2 , ..., s t,k , ..., s t,K Equation (2)
[0092] In equations (1) and (2), the ' (apostrophe) indicates transposition. Furthermore, regarding the merged sound score S... t The elements S t,k For example, it can be obtained using any one of the following equations (3) to (5).
[0093]
[0094]
[0095]
[0096] For N adjusted voice scores, equation (3) calculates the average, equation (4) calculates the median, and equation (5) calculates the maximum. Furthermore, median(·) in equation (4) is a function that takes the median value with respect to 1 ≤ n ≤ N. Additionally, max(·) in equation (5) is a function that takes the maximum value with respect to 1 ≤ n ≤ N.
[0097] In summary, the sound score merging unit 123 can generate a merged sound score by calculating at least one of the average, median, and maximum values of N adjusted sound scores.
[0098] (Step ST134)
[0099] After generating the merged sound scores, the mesh generation unit 124 generates a merged mesh based on the merged sound scores, the pronunciation dictionary, and the language model. After the processing in step ST134, the process moves to step ST140.
[0100] (Step ST140)
[0101] After generating the merged grid, the search unit 130 searches for the speech recognition result with the highest likelihood from the merged grid.
[0102] (Step ST150)
[0103] After the speech recognition result is found, the speech recognition device 100 outputs the speech recognition result to the output device. After step ST150, the process ends.
[0104] Furthermore, as long as input voice data is continuously acquired, the voice recognition device 100 can follow... Figure 3 The flowchart processing continuously outputs the speech recognition results corresponding to the input speech data.
[0105] As explained above, the speech recognition device according to the first embodiment generates multiple extended speech data based on input speech data, generates multiple sound scores based on each extended speech data and a sound model, generates multiple adjusted sound scores by resampling each of the multiple sound scores, generates a merged sound score by merging the multiple adjusted sound scores, generates a merged grid based on the merged sound score, a pronunciation dictionary, and a language model, and searches for the speech recognition result with the highest likelihood from the merged grid. Therefore, the speech recognition device according to the first embodiment can improve speech recognition performance.
[0106] In the speech recognition device according to the first embodiment, even if there is only one voice model, it is possible to solve (main reason 2) to (main reason 4). Hereinafter, specific examples of solutions to (main reason 2) to (main reason 4) will be shown respectively.
[0107] Regarding (Main Reason 2), for example, if the input speech data is spoken at a fast pace, there may be parts where the correct recognition result cannot be obtained if the speech data is input as is, but the correct recognition result can be obtained if the speech rate is adjusted to 0.9x speed. In this case, the data expansion unit generates 0.9x speed data, and a merging process is applied to the input speech data and the 0.9x speed data. This allows the beneficial portions of the recognition results from both the input speech data and the 0.9x speed data to be obtained. As a result, the speech recognition device according to the first embodiment can improve speech recognition performance.
[0108] Regarding (Main Reason 3), for example, if the input speech data is a child's voice, there may be parts where the correct recognition result cannot be obtained if the input is presented as is, but the correct recognition result can be obtained if the pitch is changed to 0.95 times by modulating the tone. In this case, the data expansion unit generates speech data with a pitch of 0.95 times, and the input speech data and the speech data with a pitch of 0.95 times are combined. This allows the beneficial portions of the recognition result of the input speech data and the recognition result of the speech data with a pitch of 0.95 times to be obtained. As a result, the speech recognition device according to the first embodiment can improve speech recognition performance.
[0109] Regarding (Main Reason 4), for example, if the microphone collecting the input speech data has low gain, there may be areas where the correct recognition result cannot be obtained if the input is presented as is, but the correct recognition result can be obtained if the amplitude is doubled by adjusting the volume. In this case, the extension unit generates speech data with doubled amplitude, and the input speech data and the speech data with doubled amplitude are combined. This allows the beneficial portions of the recognition result of the input speech data and the recognition result of the speech data with doubled amplitude to be obtained. As a result, the speech recognition device according to the first embodiment can improve speech recognition performance.
[0110] (A variation of the first embodiment)
[0111] The speech recognition apparatus according to the first embodiment generates a merged grid by merging multiple voice scores in the merging processing unit. On the other hand, the speech recognition apparatus according to a variation of the first embodiment generates multiple grids based on multiple voice scores, merges the multiple grids, and thereby generates a merged grid.
[0112] The speech recognition device according to the modified example of the first embodiment includes a data expansion unit 110, a merging processing unit 120A, a search unit 130, a sound model storage unit 140, a pronunciation dictionary storage unit 150, and a language model storage unit 160.
[0113] Figure 6 This is a block diagram illustrating the structure of the merging processing unit 120A of the speech recognition device according to a modified example of the first embodiment. The merging processing unit 120A includes a voice score calculation unit 121A, a grid generation unit 122A, and a grid merging unit 123A. Furthermore, the voice score calculation unit 121A is... Figure 2 The sound score calculation unit 121 has a roughly the same structure, so the description is omitted.
[0114] The grid generation unit 122A receives multiple sound scores from the sound score calculation unit 121A, a pronunciation dictionary from the pronunciation dictionary storage unit 150, and a language model from the language model storage unit 160. The grid generation unit 122A generates multiple grids based on the individual sound scores, the pronunciation dictionary, and the language model. These multiple grids are, for example, word grids where candidate words based on speech recognition are set as nodes and the likelihood of the candidate words is set as edges. The grid generation unit 122A outputs the generated multiple grids to the grid merging unit 123A.
[0115] Furthermore, the voice score calculation unit 121A and the grid generation unit 122A can each have multiple calculation units and multiple generation units, respectively, to match the amount of extended speech data to be generated. For example, when the amount of extended speech data is N (N>1), the voice score calculation unit 121A has a first calculation unit 121A-1, a second calculation unit 121A-2, ..., an Nth calculation unit 121A-N, and the grid generation unit 122A has a first generation unit 122A-1, a second generation unit 122A-2, ..., an Nth generation unit 122A-N. Therefore, the voice score calculation unit 121A outputs N voice scores, and the grid generation unit 122A outputs N grids.
[0116] The mesh merging unit 123A receives multiple meshes from the mesh generation unit 122A. The mesh merging unit 123A generates a merged mesh by merging the multiple meshes. The mesh merging unit 123A outputs the generated merged mesh to the search unit 130.
[0117] Specifically, the grid merging unit 123A connects the starting points and ending points of multiple grids to each other, merging the common parts of candidate words to generate a merged grid. In merging multiple grids, methods described in reference 3 (V. Le, S. Seng, L. Besacier and B. Bigi, “Word / sub-word lattices decomposition and combination for speech recognition,” IEEE International Conference on Acoustics, Speech and Signal Processing, 2008) can be used.
[0118] The structure of the voice recognition device according to the modified example of the first embodiment has been described above. Next, using... Figure 7 The flowchart will be used to explain the operation related to the merging processing unit 120A of this embodiment. Furthermore, the operation of the voice recognition device according to the variation of the first embodiment involves... Figure 3 The process of step ST130 in the flowchart is replaced by the process of step ST130A.
[0119] Figure 7 This is a flowchart illustrating the merging process in the operation of a voice recognition device according to a variation of the first embodiment. Figure 7 The flowchart is equivalent to the processing of step ST130A, which is a transfer from step ST120.
[0120] (Step ST131A)
[0121] After generating multiple extended speech data, the sound score calculation unit 121A generates multiple sound scores based on each extended speech data and the sound model.
[0122] (Step ST132A)
[0123] After generating multiple voice scores, the grid generation unit 122A generates multiple grids based on each voice score, the pronunciation dictionary, and the language model.
[0124] (Step ST133A)
[0125] After generating multiple meshes, the mesh merging unit 123A generates a merged mesh by merging the multiple meshes. After the processing in step ST133A, the process moves to step ST140.
[0126] As explained above, the speech recognition device according to the variation of the first embodiment generates multiple extended speech data based on input speech data, generates multiple sound scores based on each extended speech data and a sound model, generates multiple grids based on each sound score, a pronunciation dictionary, and a language model, generates a merged grid by merging the multiple grids, and searches for the speech recognition result with the highest likelihood from the merged grid. Therefore, the speech recognition device according to the variation of the first embodiment can improve speech recognition performance.
[0127] Figure 8 This is a table showing the experimental results of a speech recognition device according to a variation of the first embodiment. Figure 8 This result compares the recognition performance of previous methods with that of a method using modified examples of the first embodiment, generated by the extension unit with input speech rates set to 0.9x and 1.1x. The evaluation metric is the Word Error Rate (WER), with lower values indicating better recognition performance. Furthermore, the voice model, pronunciation dictionary, and language model are learned based on the Japanese Spoken Language Corpus (CSJ), and the evaluation set of CSJ is used in the evaluation.
[0128] In addition, CSJ is documented in reference 4 (K. Maekawa, “Corpus of spontaneous Japanese: Its design and evaluation,” In Proceedings ISCA and IEEE workshop on spontaneous speech processing and recognition, SSPR 2003), etc.
[0129] exist Figure 8 In conventional methods, the input speech data (A: constant speed), data with a speech rate set to 0.9 times (B: 0.9x speed), and data with a speech rate set to 1.1 times (C: 1.1x speed) are input to the speech recognition device as is. On the other hand, in... Figure 8 The proposed method includes the WER of data obtained by generating and merging input speech data with data whose speech rate is set to 0.9 times (D: A+B), data obtained by generating and merging input speech data with data whose speech rate is set to 1.1 times (E: A+C), and data obtained by generating and merging input speech data with data whose speech rate is set to 0.9 times (F: A+B+C). Figure 8 The results from D to F show that the proposed method achieves better recognition performance.
[0130] (Second Implementation)
[0131] The speech recognition device according to the first embodiment and its variants uses pre-set transformation parameters in the data expansion section to generate expanded speech data. On the other hand, the speech recognition device according to the second embodiment determines the transformation parameters in real time based on the input speech data and uses these real-time determined transformation parameters to generate expanded speech data.
[0132] Figure 9 This is a block diagram illustrating the structure of the speech recognition device 200 according to the second embodiment. The speech recognition device 200 includes a data expansion unit 110, a merging processing unit 120, a search unit 130, a voice model storage unit 140, a pronunciation dictionary storage unit 150, a language model storage unit 160, and a parameter automatic determination unit 210.
[0133] In the second embodiment, the data expansion unit 110 receives input voice data from a recording device (not shown) and receives transformation parameters from the automatic parameter determination unit 210. The data expansion unit 110 generates multiple expanded voice data based on the input voice data and the transformation parameters.
[0134] Figure 10 This is an example Figure 9 A block diagram of the structure of the automatic parameter determination unit 210 is provided. The automatic parameter determination unit 210 includes an amplitude extraction unit 211, a general amplitude data storage unit 212, a volume transformation parameter estimation unit 213, a pitch extraction unit 214, a general pitch data storage unit 215, a tone quality transformation parameter estimation unit 216, a speech rate extraction unit 217, a general speech rate data storage unit 218, and a speech rate transformation parameter estimation unit 219. Hereinafter, the general amplitude data storage unit 212, the general pitch data storage unit 215, and the general speech rate data storage unit 218 will be described first.
[0135] Furthermore, the general amplitude data storage unit 212, the general tone data storage unit 215, and the general speech rate data storage unit 218 can be combined into one or more storage units, or they can be set separately or combined outside the speech recognition device 100.
[0136] The general amplitude data storage unit 212 stores general amplitude data. For example, the general amplitude data can be the average of the power spectrum obtained by performing a short-time Fourier transform on each general speech data. For example, the speech data used in learning a sound model can be used as general speech data.
[0137] The general pitch data storage unit 215 stores general pitch data. As general pitch data, for example, the pitch average of each voice is used based on various general speech data. Pitch averaging can be obtained by averaging the pitch of each time frame after obtaining pitch information for each time frame.
[0138] Furthermore, in obtaining the pitch average, for example, the method described in reference 5 (M. Lahat, R. Niederjohn and D. Krubsack, “A spectral autocorrelation method for measurement of the fundamental frequency of noise-corrupted speech,” in IEEE Transactions on Acoustics, Speech, and Signal Processing, vol.35, no.6, pp.741-750, June 1987, doi: 10.1109 / TASSP.1987.1165224.) can be used.
[0139] The general speech rate data storage unit 218 stores general speech rate data. This general speech rate data includes, for example, the number of mora per unit time (e.g., every 5 seconds) of each general speech data. A mora is a basic unit of rhythm in Japanese. The number of mora per unit time can be obtained, for example, based on the length and label (transcription) information of the general speech data.
[0140] The amplitude extraction unit 211 receives input speech data from the recording device. The amplitude extraction unit 211 extracts the amplitude of the input speech data. Specifically, the amplitude extraction unit 211 extracts the amplitude, for example, by averaging the power spectrum obtained by performing a short-time Fourier transform on the input speech data. The amplitude extraction unit 211 outputs the extracted amplitude information (amplitude information) to the volume transformation parameter estimation unit 213.
[0141] The volume shift parameter estimation unit 213 receives amplitude information from the amplitude extraction unit 211 and general amplitude data from the general amplitude data storage unit 212. Based on the amplitude information and the general amplitude data, the volume shift parameter estimation unit 213 estimates the volume shift parameters. The volume shift parameter estimation unit 213 outputs the estimated volume shift parameters to the data expansion unit 110.
[0142] The pitch extraction unit 214 receives input speech data from the recording device. The pitch extraction unit 214 extracts the pitch of the input speech data. Specifically, the pitch extraction unit 214 extracts the pitch by obtaining the average pitch of each voice sound from the input speech data. The pitch extraction unit 214 outputs the extracted pitch information (pitch information) to the sound quality transformation parameter estimation unit 216.
[0143] The tone quality transformation parameter estimation unit 216 receives tone information from the tone extraction unit 214 and general tone data from the general tone data storage unit 215. Based on the tone information and the general tone data, the tone quality transformation parameter estimation unit 216 estimates tone quality transformation parameters. The tone quality transformation parameter estimation unit 216 outputs the estimated tone quality transformation parameters to the data expansion unit 110.
[0144] The speech rate extraction unit 217 receives the speech recognition result generated by the speech recognition device 200. The speech rate extraction unit 217 extracts the speech rate from the speech recognition result. Specifically, the speech rate extraction unit 217 extracts the speech rate by obtaining the number of beats per unit time from the length of the input speech data corresponding to the speech recognition result and the speech recognition result. For the speech recognition result, there may also be a corresponding length of input speech data. The speech rate extraction unit 217 outputs the extracted speech rate information (speech rate information) to the speech rate transformation parameter estimation unit 219. Furthermore, the speech rate extraction unit 217 may also receive input speech data corresponding to the speech recognition result.
[0145] The speech rate conversion parameter estimation unit 219 receives speech rate information from the speech rate extraction unit 217 and general speech rate data from the general speech rate data storage unit 218. Based on the speech rate information and the general speech rate data, the speech rate conversion parameter estimation unit 219 estimates the speech rate conversion parameters. The speech rate conversion parameter estimation unit 219 outputs the estimated speech rate conversion parameters to the data expansion unit 110.
[0146] The structure of the voice recognition device 200 according to the second embodiment has been described above. Next, using... Figure 11 The flowchart is used to illustrate the operation of the voice recognition device 200.
[0147] Figure 11 This is an example Figure 9 A flowchart of the operation of the voice recognition device 200. Figure 11 The flowchart, for example, illustrates a series of steps to output speech recognition results from a grid corresponding to a sentence of input speech data.
[0148] (Step ST210)
[0149] The speech recognition device 100 acquires input speech data from a recording device. Furthermore, in Figure 11 After the flowchart processing loop is completed once, the speech recognition device 100 may further acquire (or retain) the output speech recognition results for use in the automatic parameter estimation processing described later.
[0150] (Step ST220)
[0151] After acquiring the input speech data, the automatic parameter determination unit 210 estimates the transform parameters related to the generation of extended speech data. In other words, the automatic parameter determination unit 210 automatically determines the transform parameters related to the transform processing based on the input speech data. Hereinafter, the processing in step ST220 will be referred to as "automatic parameter estimation processing." Figure 12 The flowchart below illustrates a specific example of automatic parameter estimation processing.
[0152] Figure 12 This is an example Figure 9 The flowchart is a process for automatically estimating the parameters of the flowchart. Figure 12 The flowchart is transferred from step ST220. Furthermore, as follows, the speech recognition device 100 outputs more than one speech recognition result.
[0153] (Step ST221)
[0154] After acquiring the input speech data, the amplitude extraction unit 211 extracts the amplitude of the input speech data.
[0155] (Step ST222)
[0156] After extracting the amplitude, the volume transformation parameter estimation unit 213 estimates the volume transformation parameters based on the extracted amplitude and general amplitude data. The processing of the volume transformation parameter estimation unit 213 will be illustrated below with specific examples.
[0157] The volume shift parameter can be estimated at the time point when the input speech data is acquired. The volume shift parameter estimation unit 213 sets the amplitude (average of the power spectrum) of the input speech data, which is the information of the extracted amplitude, as P, and sets the average of the general amplitude data as P', and uses the following equation (6) to estimate the volume shift parameter b.
[0158] b = (P′ / P) 1 / 2 Equation (6)
[0159] (Step ST223)
[0160] After estimating the volume transformation parameters, the pitch extraction unit 214 extracts the pitch of the input speech data.
[0161] (Step ST224)
[0162] After extracting the pitch, the sound quality transformation parameter estimation unit 216 estimates the sound quality transformation parameters based on the extracted pitch and general pitch data. The processing of the sound quality transformation parameter estimation unit 216 will be explained below with specific examples.
[0163] The sound quality transformation parameter can be estimated at the time point when the input speech data is acquired. The sound quality transformation parameter estimation unit 216 sets the pitch average of the input speech data, which is the information of the extracted pitch, to F, and sets the average of the general pitch data to F', and uses the following formula (7) to estimate the sound quality transformation parameter c.
[0164] c = F′ / F (Equation 7)
[0165] (Step ST225)
[0166] After estimating the sound quality transformation parameters, the speech rate extraction unit 217 extracts the speech rate based on the speech recognition result of the input speech data.
[0167] (Step ST226)
[0168] After extracting the speech rate, the speech rate transformation parameter estimation unit 219 estimates the speech rate transformation parameters based on the extracted speech rate and general speech rate data. The following examples illustrate the processing of the speech rate transformation parameter estimation unit 219.
[0169] The speech rate transformation parameter can only be estimated after speech recognition processing has been performed on at least one speech. The speech rate transformation parameter estimation unit 219 sets the number of beats per unit time of the input speech data, which serves as information about the extracted speech rate, to M, and sets the average of the general speech data to M', and uses the following equation (8) to estimate the speech rate transformation parameter a.
[0170] a = M′ / M (Equation 8)
[0171] (Step ST227)
[0172] After estimating the volume transformation parameters, tone quality transformation parameters, and speech rate transformation parameters, the automatic parameter determination unit 210 outputs the volume transformation parameters, tone quality transformation parameters, and speech rate transformation parameters. After the processing in step ST227, the processing proceeds to step ST230.
[0173] Furthermore, the processing of steps ST221 and ST222, steps ST223 and ST224, and steps ST225 and ST226 can be performed in different orders or simultaneously.
[0174] (Step ST230)
[0175] After estimating the transformation parameters, the data extension unit 110 generates multiple extended speech data based on the input speech data and the transformation parameters.
[0176] Furthermore, the processing of steps ST240 to ST260 and Figure 3 The processes in steps ST130 to ST150 are largely the same, so the description is omitted.
[0177] As explained above, the speech recognition apparatus according to the second embodiment can estimate transformation parameters in real time in accordance with the input speech data and apply them to the generation of extended speech data. Therefore, the speech recognition apparatus according to the second embodiment can generate extended speech data of an environment that closely approximates the learning dataset of the sound model, thereby improving speech recognition performance.
[0178] (Third Implementation)
[0179] The speech recognition device according to the first embodiment and the speech recognition device according to the modifications of the first embodiment perform speech recognition processing on input speech data and output speech recognition results. On the other hand, the speech recognition device according to the third embodiment further adapts the input speech data and the speech recognition results corresponding to the input speech data to a sound model to generate an adapted sound model.
[0180] Figure 13This is a block diagram illustrating the structure of the speech recognition device 300 according to the third embodiment. The speech recognition device 300 includes a data expansion unit 110, a merging processing unit 120, a search unit 130, a voice model storage unit 140, a pronunciation dictionary storage unit 150, a language model storage unit 160, an adaptation unit 310, and an adapted voice model storage unit 320. Furthermore, the voice model storage unit 140, the pronunciation dictionary storage unit 150, the language model storage unit 160, the adaptation unit 310, and the adapted voice model storage unit 320 can be merged into one or more storage units, or they can be provided separately or merged outside the speech recognition device 100.
[0181] The adaptation unit 310 receives input speech data from a recording device (not shown), receives a speech model from the speech model storage unit 140, and receives speech recognition results from the search unit 130. Based on the input speech data and the speech recognition results corresponding to the input speech data, the adaptation unit 310 generates an adapted speech model that adapts the speech model to the speaker of the input speech data. The adaptation unit 310 outputs the generated adapted speech model to the adapted speech model storage unit 320.
[0182] Specifically, the adaptive unit 310 uses the speech recognition result as the correct answer label and adaptive data that groups the input speech data and the correct answer label to adapt the sound model. The adaptation of the sound model can be performed, for example, by optimizing the parameters of the sound model using the adaptive data. More specifically, when a DNN is used in the sound model, the adaptive unit 310 optimizes the parameters of the sound model stored in the sound model storage unit 140 as initial values. As an optimization method, for example, the method described in reference 6 (PJ Werbos, “Backpropagation Through Time: What It Does and How to Do It,” Proceedings of the IEEE, vol. 78, no. 10, 1990.) can be used.
[0183] The adaptive voice model storage unit 320 receives the adapted voice model from the adaptation unit 310. The adaptive voice model storage unit 320 stores the adapted voice model. After certain conditions are met, the adaptive voice model storage unit 320 outputs the adapted voice model to the merging processing unit 120. The certain conditions are, for example, the elapsed time since the speech recognition device 300 began speech recognition.
[0184] A specific application example of the adaptive unit 310 and the adapted sound model storage unit 320 will be explained. When a user starts the speech recognition device 300, for an initial fixed period of time (e.g., at least 20 to 30 minutes), the speech recognition device 300 performs speech recognition processing using the sound model stored in the sound model storage unit 140 (hereinafter referred to as the initial sound model). Simultaneously with this processing, the adaptive unit 310 learns a sound model in the background based on the speech recognition result and the input speech data, and outputs the adapted sound model to the adapted sound model storage unit 320. Then, after a fixed period of time, the speech recognition device 300 switches from the initial sound model to the adapted sound model to perform speech recognition processing.
[0185] Furthermore, the speech recognition device 300 may also have a function that allows the user to choose whether to switch the voice model after a fixed period of time. Additionally, the speech recognition device 300 may also have the function of automatically determining whether to switch the voice model by comparing the reliability of the speech recognition result based on the initial voice model with the reliability of the speech recognition result based on the adapted voice model. In calculating the reliability, for example, methods described in reference 7 (A. Lee, et al., “Real-time word confidence scoring using local posterior probabilities on tree trellis search,” ICASSP 2004) and reference 8 (A. Kastanos, et al., “Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks,” ICASSP 2020) can be used.
[0186] As explained above, the speech recognition device according to the third embodiment can generate an adapted sound model based on the input speech data and the speech recognition result, making the sound model adaptable to the speaker of the input speech data. Therefore, the speech recognition device according to the third embodiment can generate a sound model adapted to the input speech data, thus improving speech recognition performance.
[0187] (Fourth Implementation)
[0188] The speech recognition device according to the second embodiment is formed by adding an automatic parameter determination unit to the speech recognition device according to the first embodiment (or the speech recognition device according to a variation of the first embodiment). On the other hand, the speech recognition device according to the third embodiment is formed by adding an adaptation unit and an adapted sound model storage unit to the speech recognition device according to the first embodiment (or the speech recognition device according to a variation of the first embodiment). The speech recognition device according to the fourth embodiment includes all of these.
[0189] Figure 14 This is a block diagram illustrating the structure of the speech recognition device 400 according to the fourth embodiment. The speech recognition device 400 includes a data expansion unit 110, a merging processing unit 120, a search unit 130, a voice model storage unit 140, a pronunciation dictionary storage unit 150, a language model storage unit 160, an automatic parameter determination unit 410, an adaptation unit 420, and an adapted voice model storage unit 430.
[0190] Automatic parameter determination unit 410 and Figure 9 The parameter automatic determination unit 210 is roughly the same, and the adaptive unit 420 is similar. Figure 13 The adaptive unit 310 is roughly the same as the adaptive sound model storage unit 430. Figure 13 The adaptive sound model storage unit 320 is roughly the same.
[0191] As explained above, the speech recognition device according to the fourth embodiment is expected to achieve the same effect as the speech recognition devices according to the above embodiments.
[0192] Figure 15 This is a block diagram illustrating the hardware structure of a computer according to one embodiment. As hardware, the computer 500 includes a CPU (Central Processing Unit) 510, RAM (Random Access Memory) 520, program memory 530, auxiliary storage device 540, and input / output interface 550. The CPU 510 communicates with the RAM 520, program memory 530, auxiliary storage device 540, and input / output interface 550 via a bus 560.
[0193] CPU 510 is an example of a general-purpose processor. RAM 520 is used by CPU 510 as working memory. RAM 520 includes volatile memory such as SDRAM (Synchronous Dynamic Random Access Memory). Program memory 530 stores various programs, including speech recognition processing programs. As program memory 530, for example, ROM (Read-Only Memory), part of auxiliary storage device 540, or a combination thereof is used. Auxiliary storage device 540 stores data non-transitorily. Auxiliary storage device 540 includes non-volatile memory such as HDD or SSD.
[0194] Input / output interface 550 is an interface used for connecting to other devices. For example, input / output interface 550 is used for connecting radio equipment to output devices.
[0195] The programs stored in the program memory 530 include computer-executable instructions. When a program (computer-executable instruction) is executed by the CPU 510, it causes the CPU 510 to perform specified processing. For example, when a speech recognition processing program is executed by the CPU 510, it causes the CPU 510 to perform processing related to speech recognition. Figure 1 , Figure 2 , Figure 6 , Figure 9 , Figure 10 , Figure 13 as well as Figure 14 The various parts are described in detail, outlining a series of processes.
[0196] The program can be provided to the computer 500 in a state where it is stored in a storage medium that can be read by the computer. In this case, for example, the computer 500 also includes a drive (not shown) for reading data from the storage medium and retrieving the program from the storage medium. Examples of storage media include magnetic disks, optical discs (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), optical disks (MO, etc.), and semiconductor memory. Alternatively, the program can be stored on a server on a communication network, and the computer 500 can download the program from the server using the input / output interface 550.
[0197] The processing described in the embodiments is not limited to being performed by executing a program using a general-purpose hardware processor such as a CPU 510, but may also be performed by a special-purpose hardware processor such as an ASIC (Application Specific Integrated Circuit). The term "processing circuit" (processing unit) includes at least one general-purpose hardware processor, at least one special-purpose hardware processor, or a combination of at least one general-purpose hardware processor and at least one special-purpose hardware processor. Figure 15In the example shown, CPU 510, RAM 520, and program memory 530 correspond to the processing circuit.
[0198] Therefore, the speech recognition performance can be improved according to the above implementation methods.
[0199] Several embodiments of the present invention have been described, but these embodiments are presented by way of example and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, and are included within the scope of the invention as described in the claims and its equivalents.
Claims
1. A speech recognition apparatus comprising: a parameter automatic decision section that automatically decides a transformation parameter based on input speech data; a data expansion section that generates at least one of a plurality of expanded speech data by performing at least one of a speech rate transformation, a volume transformation, and a voice quality transformation using the transformation parameter with respect to the input speech data, the speech rate transformation is a transformation by resampling at a sampling rate different from an original sampling rate of the input speech data and transforming to the original sampling rate, a transformation parameter of the speech rate transformation is based on a comparison of a number of beats per unit time and a reference speech rate, the volume transformation is a multiplication of an amplitude of a waveform of the input speech data by a coefficient, a transformation parameter of the volume transformation is based on a comparison of power spectra, the voice quality transformation is a change in a pitch of the input speech data using a pitch synchronous overlap-add method, a transformation parameter of the voice quality transformation is based on a comparison of pitch averages of the input speech data, a sound score calculation section that generates a plurality of sound scores based on each of the plurality of expanded speech data and a sound model; an adjustment section that generates a plurality of adjusted sound scores by resampling the plurality of sound scores respectively, the T is a number of time frames of the input speech data, whereby the adjustment section transforms a sequence of the sound scores into a sequence of the adjusted sound scores over the T time frames; a sound score merge section that generates a merged sound score by calculating at least one of an average, a median, and a maximum of the plurality of adjusted sound scores and merging; a lattice generation section that generates a merged lattice based on the merged sound score, a pronunciation dictionary, and a language model; and a search section that searches for a speech recognition result having a highest likelihood from the merged lattice. 2.The speech recognition apparatus according to claim 1, wherein the merged lattice is a word lattice that sets a candidate word based on speech recognition as a node and sets a likelihood of the candidate word as an edge. 3.The speech recognition apparatus according to claim 1 or 2, wherein the parameter automatic decision section estimates a speech rate transformation parameter related to the speech rate transformation based on a speech recognition result corresponding to the input speech data, and the data expansion section generates an expanded speech data using the speech rate transformation parameter. 4.The speech recognition apparatus according to any one of claims 1 to 3, wherein the parameter automatic decision section estimates a volume transformation parameter related to the volume transformation based on the input speech data, and the data expansion section generates an expanded speech data using the volume transformation parameter. 5.The speech recognition apparatus according to any one of claims 1 to 4, wherein the parameter automatic decision section estimates a voice quality transformation parameter related to the voice quality transformation based on the input speech data, and the data expansion section generates an expanded speech data using the voice quality transformation parameter. 6.The speech recognition apparatus according to any one of claims 1 to 5, wherein the plurality of expanded speech data includes the input speech data. The resampling is of each sound score column of K-dimensional vectors throughout T n time frames at a sampling rate of T / T n each independently resampled, The T n is the number of time frames of the plurality of extended speech data, 7. The speech recognition apparatus according to any one of claims 1 to 6, wherein the acoustic model is a single model learned in such a manner that speech data is inputted by at least one unit of a phoneme, a syllable, a character, a word piece, and a word, thereby outputting a posterior probability corresponding to an acoustic score.
8. The speech recognition apparatus according to any one of claims 1 to 7, wherein an adaptation unit that generates an adapted acoustic model adapted to a speaker of the input speech data based on the input speech data and the speech recognition result corresponding to the input speech data is further provided.
9. A speech recognition method comprising: automatically deciding a transformation parameter based on input speech data; generating at least one of a plurality of extended speech data by performing at least one of a speech rate transformation, a volume transformation, and a voice quality transformation using the transformation parameter for the input speech data, the speech rate transformation is a transformation by resampling at a sampling rate different from an original sampling rate of the input speech data and transforming to the original sampling rate, the transformation parameter of the speech rate transformation is based on a comparison of a number of beats per unit time and a reference speech rate, the volume transformation is multiplying an amplitude of a waveform of the input speech data by a coefficient, the transformation parameter of the volume transformation is based on a comparison of power spectra, the voice quality transformation is changing a pitch of the input speech data using a pitch synchronous overlap-add method, the transformation parameter of the voice quality transformation is based on a comparison of pitch averages of the input speech data, generating a plurality of acoustic scores based on each of the plurality of extended speech data and an acoustic model; generating a plurality of adjusted acoustic scores by resampling the plurality of acoustic scores respectively, The resampling is of each sound score column of K-dimensional vectors throughout T n time frames at a sampling rate of T / T n each independently resampled, The T n is the number of time frames of the plurality of extended speech data, the T is a number of time frames of the input speech data, thereby transforming a sequence of the acoustic scores into a sequence of the adjusted acoustic scores over the T time frames; generating an aggregated acoustic score by calculating at least one of a mean, a median, and a maximum of the plurality of adjusted acoustic scores and aggregating; generating an aggregated lattice based on the aggregated acoustic score, a pronunciation dictionary, and a language model; and searching for a speech recognition result having a highest likelihood from the aggregated lattice.
10. A computer program product including a computer program for causing a computer to function as: a unit that automatically decides a transformation parameter based on input speech data; a unit that generates at least one of a plurality of extended speech data by performing at least one of a speech rate transformation, a volume transformation, and a voice quality transformation using the transformation parameter for the input speech data, the speech rate transformation is a transformation by resampling at a sampling rate different from an original sampling rate of the input speech data and transforming to the original sampling rate, the transformation parameter of the speech rate transformation is based on a comparison of a number of beats per unit time and a reference speech rate, the volume transformation is multiplying an amplitude of a waveform of the input speech data by a coefficient, the transformation parameter of the volume transformation is based on a comparison of power spectra, a transformation parameter of the volume transformation is based on a comparison of power spectrums, the voice quality transformation is to change a pitch of the input speech data using a pitch synchronous overlap-add method, a transformation parameter of the voice quality transformation is based on a comparison of pitch averages of the input speech data, a unit of generating a plurality of sound scores based on each of the plurality of extended speech data and a sound model; a unit of generating a plurality of adjusted sound scores by resampling the plurality of sound scores respectively, The resampling is of each sound score column of K-dimensional vectors for sounds throughout T n time frames at a sampling rate of T / T n each independently resampled, The T n is the number of time frames of the plurality of extended speech data, the T is a number of time frames of the input speech data, thereby transforming a sequence of the sound scores into a sequence of the adjusted sound scores across the T time frames; a unit of generating a merged sound score by calculating at least one of an average, a median, and a maximum of the plurality of adjusted sound scores and merging; a unit of generating a merged lattice based on the merged sound score, a pronunciation dictionary, and a language model; and a unit of searching for a speech recognition result having a highest likelihood from the merged lattice.
Citation Information
Patent Citations
Liquid jet device
JP2021091236A
Voice conversion method, device and system, model training method, device and system and storage medium
CN112185342A
Apparatus, method, and medium for dialogue speech recognition using topic domain detection
US20070100618A1
Apparatus and method of acoustic score calculation and speech recognition
US20170025119A1