Speech recognition method, speech recognition device, and speech recognition program

The proposed speech recognition method addresses the challenge of weighting scores from multiple NLMs by iteratively updating lattice scores with coefficients adjusted for each NLM's performance, resulting in high-accuracy speech recognition.

JP7694805B2Active Publication Date: 2025-06-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024509551
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-06-18
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

Conventional speech recognition technologies using lattice rescoring face challenges in achieving high accuracy due to inadequate methods for weighting scores calculated by multiple neural language models (NLMs).

Method used

A speech recognition method that generates a lattice based on initial speech recognition results and iteratively updates lattice scores using outputs from multiple NLMs, with coefficients adjusted based on the number of repetitions or the performance of each NLM.

Benefits of technology

This approach enables high-accuracy speech recognition by effectively weighting and refining language scores through iterative lattice rescoring, leading to improved word prediction and reduced word error rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694805000004
    Figure 0007694805000004
  • Figure 0007694805000005
    Figure 0007694805000005
  • Figure 0007694805000006
    Figure 0007694805000006
Patent Text Reader

Abstract

A speech recognition device (10) according to an embodiment comprises a speech recognition unit (131) and a score calculation unit (132). The speech recognition unit (131) generates a lattice on the basis of the result of performing speech recognition of an utterance. The score calculation unit (132) updates a score of the lattice in each process, which is executed iteratively by a predetermined number of times, on the basis of an NLM output corresponding to each process and a coefficient based on the number of times of iterations or the performance of the NLM at the time of execution of each process (iterative lattice scoring).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a speech recognition method, a speech recognition apparatus, and a speech recognition program.

Background Art

[0002] Speech recognition is a technology for converting speech (utterance) made by a human into a word sequence (text) by a computer.

[0003] Normally, a speech recognition system outputs one word sequence (one-best hypothesis), which is the hypothesis (speech recognition result) with the highest speech recognition score, for one input utterance.

[0004] On the other hand, the accuracy of speech recognition processing by a speech recognition apparatus is not 100%. Conventionally, as a method for improving the accuracy of speech recognition processing, a method called lattice rescoring is known (see, for example, Non-Patent Document 1).

[0005] In lattice rescoring, instead of outputting only the one-best hypothesis for one input utterance, a lattice that efficiently represents a plurality of speech recognition hypotheses is output, and as post-processing, using some model, a hypothesis estimated to be the oracle hypothesis (the hypothesis with the highest accuracy, the hypothesis with the fewest errors) is selected from the lattice. Note that the oracle hypothesis may sometimes be the one-best hypothesis.

[0006] Also, a method of using a neural language model (NLM) based on a neural network for lattice rescoring is known (see, for example, Non-Patent Documents 2 and 3).

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

[0008] However, the conventional technology has a problem that it may not be possible to perform speech recognition by lattice rescoring with high accuracy.

[0009] For example, Non-Patent Document 4 describes a method of calculating scores in a lattice rescoring for a plurality of NLMs.

[0010] On the other hand, how to weight each of the scores calculated by the plurality of NLMs has not been sufficiently studied.

Means for Solving the Problems

[0011] In order to solve the above-described problems and achieve the object, the speech recognition method is a speech recognition method executed by a computer, including a generation step of generating a lattice based on a result of performing speech recognition of an utterance, and in each of the processes repeatedly executed a predetermined number of times, based on an output of an NLM corresponding to each process and a coefficient based on the number of repetitions at the time of executing each process or the performance of the NLM, a score calculation step of updating the score of the lattice.

Advantages of the Invention

[0012] According to the present invention, speech recognition by lattice rescoring can be performed with high accuracy.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0014] Hereinafter, embodiments of the speech recognition method, speech recognition apparatus, and speech recognition program according to the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to the embodiments described below.

[0015] [Configuration of the First Embodiment] First, with reference to FIG. 1, the configuration of the speech recognition apparatus according to the first embodiment will be described. FIG. 1 is a diagram showing an example of the configuration of the speech recognition apparatus according to the first embodiment. The speech recognition apparatus 10 receives an input of speech data, performs speech recognition, and outputs a word sequence as a speech recognition result.

[0016] As shown in FIG. 1, the speech recognition apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 13.

[0017] The communication unit 11 performs data communication with other devices via a network. For example, the communication unit 11 is a NIC (Network Interface Card).

[0018] The storage unit 12 is a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an optical disk. Note that the storage unit 12 may be a semiconductor memory capable of rewriting data, such as a RAM (Random Access Memory), a flash memory, or an NVSRAM (Non Volatile Static Random Access Memory). The storage unit 12 stores an OS (Operating System) and various programs executed by the speech recognition apparatus 10.

[0019] The storage unit 12 stores model information 121 and lattice information 122.

[0020] The model information 121 is information such as parameters for constructing each of a plurality of NLMs.

[0021] The lattice information 122 is information regarding the lattice. The lattice information 122 includes nodes, arcs, scores, and the like. Details of the lattice will be described later.

[0022] The control unit 13 controls the entire speech recognition device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0023] Also, the control unit 13 has an internal memory for storing programs and control data that define various processing procedures, and executes each process using the internal memory.

[0024] Also, the control unit 13 functions as various processing units when various programs operate. For example, the control unit 13 has a speech recognition unit 131 and a score calculation unit 132.

[0025] The speech recognition unit 131 performs speech recognition on an utterance. Also, the speech recognition unit 131 generates a lattice based on the result of performing speech recognition on the utterance. The speech recognition unit 131 stores the generated lattice in the storage unit 12 as the lattice information 122.

[0026] Here, the lattice will be described with reference to FIG. 2. FIG. 2 is a diagram for explaining the lattice.

[0027] As shown in FIG. 2, the lattice is composed of nodes and arcs. The nodes represent the word boundaries of the recognized result words (words obtained by speech recognition). The arcs are the recognized result words themselves.

[0028] The lattice shown in FIG. 2 is generated based on the utterance "I like speech recognition".

[0029] At this time, the 1-best hypothesis is "I like hot spring bathing is skiing" (dotted line in FIG. 2). Also, the oracle hypothesis is "I also like speech recognition" (dashed-dotted line in FIG. 2).

[0030] In this way, a plurality of word sequences are extracted from the lattice. Also, the extracted word sequences include the 1-best hypothesis and the oracle hypothesis. Note that the 1-best hypothesis may be the oracle hypothesis.

[0031] FIG. 3 is a diagram for explaining the acoustic score and the language score. As shown in FIG. 3, an acoustic score (log-likelihood) and a language score (log-probability) calculated by the speech recognition process are respectively assigned to the arcs.

[0032] The acoustic score is an estimated value representing how acoustically correct the recognized result word is. Also, the language score is an estimated value representing how linguistically correct the recognized result word is.

[0033] The speech recognition unit 131 can calculate the language score using an n-gram language model (n is usually about 3 to 5) that represents the n-chain probability of words. Also, the speech recognition unit 131 can calculate the acoustic score using a neural network for speech recognition that takes a speech signal as an input.

[0034] Note that the information for constructing the n-gram language model and the neural network for speech recognition is stored in the storage unit 12 as model information 121.

[0035] The score calculation unit 132 performs lattice rescoring. Lattice rescoring is performed using a rescoring model as post-processing of the speech recognition process.

[0036] According to lattice rescoring, as shown in FIG. 4, it is possible to assign a more accurate language score to an arc (recognized result word) than the language score assigned by an n-gram language model. FIG. 4 is a diagram for explaining the update of the language score. In the example of FIG. 4, the language score is updated using an NLM.

[0037] In recent years, an NLM that can capture a context longer than an n-gram language model and can perform word prediction with higher accuracy has been used as a rescoring model. High word prediction accuracy means that it is possible to predict with high accuracy the word to be generated next when a word history is given.

[0038] Regarding lattice rescoring using an NLM, it is described in, for example, Non-Patent Documents 2, 3, and 4.

[0039] Further, Non-Patent Document 1 describes a method of lattice rescoring based on the Push-forward algorithm, in which an NLM is used to perform a search (hypothesis expansion) on the lattice from the start node to the end node, and the language score recorded on the arc is updated.

[0040] In the method described in Non-Patent Document 1, among the hypotheses (word sequences) that reach the end node, the one with the highest score (the weighted sum score of the acoustic score and the updated language score) is taken as the final speech recognition result.

[0041] Here, focusing on the search process for an arc on the lattice as shown in FIG. 5, the iterative lattice rescoring by the score calculation unit 132 will be described.

[0042] w 1:t-1 Let be a hypothesis of length t-1. Hypothesis w 1:t-1The current score (log-likelihood) is log p(w 1:t-1 ) and it is assumed that we have reached an arc (recognized result word) w acou (w t ) with an acoustic score (log-likelihood) log p lang (w t ) and a language score (log-probability) log P t (w

[0043] The score calculation unit 132 calculates the score of the hypothesis w 1:t-1 of length t by expanding the hypothesis w t onto the arc w 1:t as shown in Equation (1).

[0044]

Equation

[0045] Here, log p resc (w t │w 1:t-1 is the language score of w 1:t-1 when w t is given, and it is calculated by the NLM for rescoring. β (0 < β < 1) is the interpolation coefficient between the original language score and the language score calculated by the NLM for rescoring. α (α > 0) is the weight of the language score with respect to the acoustic score.

[0046] The underlined term in Equation (1) corresponds to the updated language score. By the score calculation unit 132 performing the search process (score calculation for each reached arc) described here for all arcs on the lattice, a lattice with updated language scores can be obtained.

[0047] When using multiple NLMs, the score calculation unit 132 repeats the search process (iterative lattice rescoring). And each time the score calculation unit 132 repeats the search process, the language score (log-probability) log P lang w t is gradually updated (refined).

[0048] At this time, it is not obvious how to set β. If there are only a few NLMs to be used, it is possible to set β heuristically (manually), but when using more NLMs, it is necessary to design how to set β for each number of repetitions (i in FIG. 5). Hereinafter, a method for setting the interpolation coefficient β by the score calculation unit 132 will be described.

[0049] (Method for setting β 1) Let I be the number of repetitions (language score update) of the iterative lattice rescoring. That is, the number of NLMs used in the iterative lattice rescoring is I. When it can be assumed that the word prediction accuracies of these I NLMs are approximately the same, when the I repetitions are completed, the language scores output by the I NLMs may be equally evaluated (weighted). For this purpose, the score calculation unit 132 sets β in the i-th repetition as shown in Equation (2).

[0050] [Number]

[0051] In this way, the score calculation unit 132 sets a value that becomes smaller as the number of repetitions increases as a coefficient.

[0052] (Method for setting β 2) When the nature of the voice data used by the voice recognition unit 131 for voice recognition is clear and text data having the same nature as the voice data can be obtained, the score calculation unit 132 can set β using the word prediction accuracy of each NLM for the text data.

[0053] At this time, perplexity can be used as a measure of word prediction accuracy. Let PPL(i) be the perplexity of the text data for the NLM used in the i-th repetition. Then, the score calculation unit 132 sets β in the i-th repetition as shown in Equation (3).

[0054] [Number]

[0055] Here, PPL(0) is the perplexity of the n-gram language model with respect to the text data. Note that the above iterative lattice rescoring can also be applied to an N-best list which is a special shape of the lattice (iterative N-best list rescoring).

[0056] In this way, the score calculation unit 132 sets, as the coefficient β, a value that increases as the word prediction accuracy of the NLM corresponding to each process with respect to text data having the same nature as the utterance to be recognized is higher.

[0057] Note that perplexity is an example of an index representing the performance of the NLM. Also, PPL(i) decreases as the word estimation accuracy of the NLM is higher.

[0058] FIG. 6 is a flowchart showing the processing flow of the speech recognition apparatus according to the embodiment. As shown in FIG. 6, first, the speech recognition apparatus 10 receives an input of one utterance (step S11). The utterance is, for example, speech data representing a speech signal in a predetermined format.

[0059] Next, the speech recognition apparatus 10 performs speech recognition on the input utterance (step S12). Then, the speech recognition apparatus 10 generates a lattice based on the result of the speech recognition (step S13).

[0060] Here, the speech recognition apparatus 10 executes lattice rescoring (step S14). Then, the speech recognition apparatus 10 selects and outputs a hypothesis estimated as the oracle hypothesis from among the lattices whose scores have been updated by the lattice rescoring (step S15). For example, the speech recognition apparatus 10 outputs a word sequence based on the selected hypothesis.

[0061] FIG. 7 is a flowchart showing the processing flow of the lattice rescoring process. The process of FIG. 7 corresponds to the process of step S14 in FIG. 6.

[0062] As shown in FIG. 7, first, the speech recognition device 10 sets i to 1 (step S141). i is an index for identifying a model (for example, NLM) for calculating a score. Also, i can be said to be the current number of repetitions of lattice rescoring.

[0063] Also, information for constructing a plurality of models for calculating scores is included in the model information 121.

[0064] Here, the speech recognition device 10 sets a coefficient β(i) corresponding to the i-th NLM (step S142). For example, the speech recognition device 10 calculates the coefficient β(i) by the aforementioned β setting method 1 or β setting method 2.

[0065] Then, the speech recognition device 10 updates the score of the arc on the lattice based on the output of the i-th NLM and the coefficient β(i) (step S143).

[0066] Here, if i is not I (step S144, No), the speech recognition device 10 increments i by 1 (step S145), returns to step S142, and repeats the process.

[0067] On the other hand, if i is I (step S144, Yes), the speech recognition device 10 ends the process. I is the total number of repetitions of lattice rescoring and also the number of NLMs used.

[0068] In this way, the score calculation unit 132 updates the lattice score based on the output of the NLM corresponding to each process and the coefficient (β) based on the number of repetitions (for example, i) or the performance of the NLM at the time of executing each process, in each of the processes repeatedly executed a predetermined number of times (for example, I times).

[0069] [Effect of the First Embodiment] As described above, the speech recognition unit 131 generates a lattice based on the result of performing speech recognition on the utterance. The score calculation unit 132 updates the lattice score based on the output of the NLM corresponding to each process and a coefficient based on the number of repetitions or the performance of the NLM at the time of executing each process, in each of the processes repeatedly executed a predetermined number of times.

[0070] Thereby, weighting is performed based on the number of repetitions or a coefficient based on the performance of the NLM, and it becomes possible to perform speech recognition with high accuracy by lattice rescoring.

[0071] Also, the score calculation unit 132 sets a value that becomes smaller as the number of repetitions increases as a coefficient. Thereby, each NLM can be evaluated equally.

[0072] Also, the score calculation unit 132 sets, as a coefficient β, a value that becomes larger as the word prediction accuracy of the NLM corresponding to each process for text data having the same nature as the utterance to be recognized is higher. Thereby, the word prediction accuracy of each NLM can be reflected in the lattice score.

[0073] Here, the utterance to be recognized is the utterance on which speech recognition is performed by the speech recognition unit 131, and the recognition result (word sequence) of the utterance is unknown. On the other hand, when the nature of the utterance to be recognized is known, it is possible to calculate the perplexity in advance for text data having the same nature.

[0074] For example, when the utterance to be recognized is related to a weather forecast, the score calculation unit 132 can calculate the perplexity of the NLM for text data related to the weather forecast and set the coefficient β based on the calculated perplexity.

[0075] FIG. 8 shows the result of performing iterative lattice rescoring based on Expression (1) and Expression (2) by the method shown in the embodiment using eight NLMs. FIG. 8 is a diagram showing experimental results.

[0076] As can be seen from FIG. 8, the word error rate (the lower, the higher the accuracy) can be gradually reduced each time scoring is repeated. Finally, the word error rate of the single best hypothesis of the speech recognition process is reduced from 9.0% to 7.0%.

[0077] [System configuration, etc.] In addition, each component of each illustrated device is a functional concept, and it is not necessarily physically configured as shown in the figure. That is, the specific forms of distribution and integration of each device are not limited to those shown in the figure, and all or part of them can be functionally or physically distributed or integrated in any unit according to various loads, usage conditions, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic. Note that the program may be executed not only by the CPU but also by other processors such as a GPU.

[0078] Also, among the processes described in this embodiment, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified.

[0079] [Program] As one embodiment, the voice recognition device 10 can be implemented by installing a voice recognition program that executes the above-described voice recognition process as package software or online software on a desired computer. For example, by causing the information processing device to execute the above-described voice recognition program, the information processing device can function as the voice recognition device 10. The information processing device mentioned here includes desktop or notebook personal computers. In addition, other information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and further slate terminals such as PDA (Personal Digital Assistant) are included in this category.

[0080] In addition, the voice recognition device 10 can also be implemented as a voice recognition server device that uses the terminal device used by the user as a client and provides the above-described service related to the voice recognition process to the client. For example, the voice recognition server device is implemented as a server device that provides a voice recognition service that takes speech (voice data) as input and outputs a word sequence. In this case, the voice recognition server device may be implemented as a Web server, or may be implemented as a cloud that provides the above-described service related to the voice recognition process by outsourcing.

[0081] FIG. 9 is a diagram showing an example of a computer that executes a voice recognition program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0082] Memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0083] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the voice recognition device 10 is implemented as a program module 1093 in which computer-executable code is described. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration in the voice recognition device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0084] Also, the setting data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as needed, and executes the processing of the above-described embodiment.

[0085] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090. For example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a LAN (Local Area Network) or a WAN (Wide Area Network)). Then, the program module 1093 and the program data 1094 may be read by the CPU 1020 from another computer via the network interface 1070.

Explanation of Signs

[0086] 10 Voice recognition device 11 Communication unit 12 Storage unit 13 Control unit 121 Model information 122 Lattice information 131 Voice recognition section 132 Score calculation section

Claims

1. A speech recognition method executed by a computer, comprising: a generation step of generating a lattice based on a result of performing speech recognition of speech; in each of the processes repeatedly executed a predetermined number of times, based on the output of the NLM corresponding to each process and a coefficient based on the number of repetitions or the performance of the NLM at the time of execution of each process, a score calculation step of updating the score of the lattice; A speech recognition method, characterized by including the above.

2. The speech recognition method according to claim 1, characterized in that the score calculation step sets a value that becomes smaller as the number of repetitions increases as the coefficient.

3. The speech recognition method according to claim 1, characterized in that the score calculation step sets a value that becomes larger as the word prediction accuracy of the NLM corresponding to each process for text data having the same nature as the speech is higher as the coefficient.

4. A speech recognition unit that generates a lattice based on a result of performing speech recognition of speech; in each of the processes repeatedly executed a predetermined number of times, based on the output of the NLM corresponding to each process and a coefficient based on the number of repetitions or the performance of the NLM at the time of execution of each process, a score calculation unit that updates the score of the lattice; A speech recognition apparatus, characterized by including the above.

5. A speech recognition program for causing a computer to function as the speech recognition apparatus according to claim 4.

Citation Information

Patent Citations

  • Lattice encoding using recurrent neural networks

    US10176802B1

  • Method for estimating language model weight and system for the same

    US20120150539A1