Speech Recognition Method, Apparatus, Computer Device, and Storage Medium
By adjusting the probability suppression of the phoneme recognition results of the RNN-T model in the speech recognition system, the problem of error rate increase caused by the empty output is solved, and the accuracy of speech recognition is improved.
Patent Information
- Application Number
- CN202011536771.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-12-23
AI Technical Summary
The concept of empty output introduced by the RNN-T model in the phoneme recognition process leads to an increase in the error rate of subsequent decoding processes, especially the increase in the number of deleted errors, which affects the accuracy of speech recognition.
The speech signal is processed through the acoustic model, and the phoneme recognition results of each speech frame are obtained, and the probability of the empty output is suppressed and adjusted, reducing the ratio of the probability of the empty output to the probability of the individual phonemes, and improving the accuracy of the phoneme recognition results.
This reduces the possibility that the speech frame is mistakenly recognized as an empty output, reduces deletion errors, and thus improves the accuracy of speech recognition.
Smart Images

Figure CN113539242B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech recognition, and particularly to a speech recognition method, apparatus, computer device, and storage medium. Background Art
[0002] Speech recognition is a technology that recognizes speech into text, and it has a wide range of applications in various artificial intelligence (AI) scenarios.
[0003] A speech recognition framework usually includes an acoustic model part and a decoding part. Among them, the acoustic model part is used to phoneme each speech frame in the recognized input speech signal, and the decoding part outputs the text sequence of the speech signal through the phonemes of each recognized speech frame. In the related art, implementing the acoustic model through a Recurrent Neural Network Transducer (RNN-T) is one of the key research points in the industry.
[0004] However, the RNN-T model introduces the concept of null output in the phoneme recognition process, that is, it is predicted that a certain speech frame does not contain valid phonemes. The introduction of null output will increase the error rate of the subsequent decoding process in some application scenarios, especially resulting in an increase in deletion errors, affecting the accuracy of speech recognition. Summary of the Invention
[0005] Embodiments of this application provide a speech recognition method, apparatus, computer device, and storage medium, which can improve the accuracy of speech recognition. The technical solution is as follows:
[0006] On the one hand, a speech recognition method is provided. The method includes:
[0007] Processing the speech signal through an acoustic model to obtain the phoneme recognition result corresponding to each speech frame in the speech signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples;
[0008] Suppressing and adjusting the probability of the null output in the phoneme recognition results corresponding to each speech frame to reduce the ratio of the probability of the null output to the probability of each phoneme in the phoneme recognition results;
[0009] Inputting the adjusted phoneme recognition results corresponding to each speech frame into a decoding graph to obtain the recognition text sequence corresponding to the speech signal.
[0010] On the one hand, a speech recognition method is provided, and the method includes:
[0011] Obtain a speech signal, where the speech signal includes each speech frame obtained by segmenting the original speech;
[0012] Process the speech signal through an acoustic model to obtain a phoneme recognition result corresponding to each speech frame; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples;
[0013] Input the phoneme recognition results in which the probability of the null output in the phoneme recognition results corresponding to each speech frame satisfies a specified condition into a decoding graph to obtain a recognition text sequence corresponding to the speech signal.
[0014] On the other hand, a speech recognition device is provided, and the device includes:
[0015] A speech signal processing module, configured to process a speech signal through an acoustic model to obtain a phoneme recognition result corresponding to each speech frame in the speech signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples;
[0016] A probability adjustment module, configured to suppress and adjust the probability of the null output in the phoneme recognition results corresponding to each speech frame, so as to reduce the ratio of the probability of the null output in the phoneme recognition results to the probabilities of each phoneme;
[0017] A decoding module, configured to input the adjusted phoneme recognition results corresponding to each speech frame into a decoding graph to obtain a recognition text sequence corresponding to the speech signal.
[0018] In a possible implementation manner, the probability adjustment module is configured to adjust the phoneme recognition results corresponding to each speech frame through at least one of the following adjustment methods:
[0019] Reduce the probability of the null output in the phoneme recognition results corresponding to each speech frame;
[0020] And,
[0021] Increase the probabilities of each phoneme in the phoneme recognition results corresponding to each speech frame.
[0022] In a possible implementation, the probability adjustment module is configured to multiply the probability of null output in the phoneme recognition results corresponding to the respective speech frames by a first weight, where the first weight is less than 1 and greater than 0.
[0023] In a possible implementation, the probability adjustment module is configured to multiply the probabilities of the respective phonemes in the phoneme recognition results corresponding to the respective speech frames by a second weight, where the second weight is greater than 1.
[0024] In a possible implementation, the decoding module is configured to,
[0025] in response to the probability of null output in the target phoneme recognition result satisfying a specified condition, input the target phoneme recognition result into the decoding graph to obtain the recognition text corresponding to the target phoneme recognition result;
[0026] wherein the target phoneme recognition result is any one of the phoneme recognition results corresponding to the respective speech frames.
[0027] In a possible implementation, the specified condition includes:
[0028] the probability of null output in the target phoneme recognition result is less than a probability threshold.
[0029] In a possible implementation, the apparatus further includes:
[0030] a parameter acquisition module configured to acquire a threshold influence parameter, where the threshold influence parameter includes at least one of ambient sound intensity, the number of speech recognition failures within a specified time period, and user setting information;
[0031] a threshold determination module configured to determine the probability threshold based on the threshold influence parameter.
[0032] In a possible implementation, the speech signal processing module is configured to,
[0033] extract features from a target speech frame to obtain a feature vector of the target speech frame; the target speech frame is any one of the respective speech frames;
[0034] input the target speech frame into an encoder in the acoustic model to obtain an acoustic hidden layer representation vector of the target speech frame;
[0035] Input the phoneme information of the historical recognition text of the target speech frame into the predictor in the acoustic model to obtain the text hidden layer representation vector of the target speech frame; the historical recognition text of the target speech frame is the text obtained by the decoding graph recognizing the phoneme recognition results of the first n non-empty output speech frames of the target speech frame; n is an integer greater than or equal to 1.
[0036] Input the acoustic hidden layer representation vector of the target speech frame and the text hidden layer representation vector of the target speech frame into the joint network to obtain the phoneme recognition result of the target speech frame.
[0037] In a possible implementation manner, the encoder is a Forward Sequence Memory Network (FSMN).
[0038] In a possible implementation manner, the predictor is a one-dimensional convolutional network.
[0039] In a possible implementation manner, the decoding graph is composed of a phoneme dictionary and a language model.
[0040] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the above-mentioned speech recognition method.
[0041] On another hand, a computer-readable storage medium is provided. At least one computer instruction is stored in the storage medium, and the at least one computer instruction is loaded and executed by a processor to implement the above-mentioned speech recognition method.
[0042] On another hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned computer program method.
[0043] The technical solution provided by this application may include the following beneficial effects:
[0044] By suppressing the probability of the empty output in the phoneme recognition result before inputting the phoneme recognition result including the probability distribution of each phoneme and the empty output of the speech frame into the decoding graph, the probability of the speech frame being recognized as an empty output is reduced, thereby reducing the possibility of the speech frame being misrecognized as an empty output, that is, reducing the deletion error of the model, and thus improving the recognition accuracy of the model. Description of the Drawings
[0045] The accompanying drawings herein are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0046] Figure 1 It is a system configuration diagram of a speech recognition system related to various embodiments of the present application;
[0047] Figure 2 It is a schematic flowchart of a speech recognition method shown according to an exemplary embodiment;
[0048] Figure 3 It is a schematic flowchart of a speech recognition method shown according to an exemplary embodiment;
[0049] Figure 4 It is Figure 3 a schematic diagram of the alignment process related to the shown embodiment;
[0050] Figure 5 It is Figure 3 a schematic structural diagram of an acoustic model related to the shown embodiment;
[0051] Figure 6 It is Figure 3 a network structure diagram of a predictor related to the shown embodiment;
[0052] Figure 7 It is Figure 3 a flowchart of model training and application related to the shown embodiment;
[0053] Figure 8 It is a framework diagram of a speech recognition system shown according to an exemplary embodiment;
[0054] Figure 9 It is a structural block diagram of an object annotation device in a video shown according to an exemplary embodiment;
[0055] Figure 10 It is a structural block diagram of a computer device shown according to an exemplary embodiment. Detailed Description of the Embodiments
[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0057] Before describing the various embodiments shown in the present application, several concepts related to the present application will be introduced first:
[0058] 1) Artificial Intelligence (AI)
[0059] AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0060] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0061] 2) Speech Technology (ST)
[0062] The key technologies of speech technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0063] 3) Machine Learning (ML)
[0064] Machine learning is an interdisciplinary subject in multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0065] The solution provided in the embodiments of this application is applied to scenarios involving speech technology and machine learning technology in artificial intelligence to achieve accurate recognition of user speech as the corresponding text.
[0066] Please refer to Figure 1 , which shows a system configuration diagram of a speech recognition system involved in various embodiments of the present application. As Figure 1 shown, the system includes a sound collection component 120 and a speech recognition device 140.
[0067] Among them, the sound collection component 120 and the speech recognition device 140 are connected by a wired or wireless manner.
[0068] The sound collection component 120 can be implemented as a microphone, a microphone array, a pickup, etc. The sound collection component 120 is used to collect speech data when the user is speaking.
[0069] The speech recognition device 140 is used to recognize the speech data collected by the sound collection component 120 to obtain a recognized text sequence.
[0070] Optionally, the speech recognition device 140 can also perform natural semantic processing on the recognized text sequence to respond to the user's speech.
[0071] Among them, the sound collection component 120 and the speech recognition device 140 can be implemented as two independent hardware devices. For example, the sound collection component 120 is a microphone set on the vehicle steering wheel, and the speech recognition device 140 can be an in-vehicle intelligent device; or, the sound collection component 120 is a microphone set on the remote control, and the speech recognition device 140 can be a smart home device controlled by the remote control (such as a smart TV, a set-top box, an air conditioner, etc.).
[0072] Or, the sound collection component 120 and the speech recognition device 140 can be implemented as the same hardware device. For example, the speech recognition device 140 can be a smart device such as a smart phone, a tablet computer, a smart watch, a smart glasses, etc., and the sound collection component 120 can be a microphone built in the speech recognition device 140.
[0073] In a possible implementation manner, the above speech recognition system may further include a server 160.
[0074] Among them, the server 160 can be used to deploy and update the speech recognition model in the speech recognition device 140. Or, the server 160 can also provide cloud speech recognition services to the speech recognition device 140, that is, receive the speech data sent by the speech recognition device 140, perform speech recognition on the speech data, and then return the recognition result to the speech recognition device 140. Or, the server 160 can also cooperate with the speech recognition device 140 to complete operations such as recognition of speech data and response to speech data.
[0075] The server 160 is a server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center.
[0076] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0077] The server 160 is connected to the speech recognition device 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0078] Optionally, the system may further include a management device ( Figure 1 not shown), which is connected to the server 160 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0079] Optionally, the above wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, custom and / or proprietary data communication technologies can be used to replace or supplement the above data communication technologies.
[0080] Please refer to Figure 2, which is a schematic flowchart of a speech recognition method shown according to an exemplary embodiment. This speech recognition method can be used in a computer device. For example, the computer device can be the speech recognition device 140 or the server 160 in the above Figure 1 shown system, or the computer device can simultaneously include the speech recognition device 140 and the server 160 in the above Figure 1 shown system. As Figure 2 shown, the speech recognition method can include the following steps:
[0081] Step 21, process the speech signal through an acoustic model to obtain the phoneme recognition result corresponding to each speech frame in the speech signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space contains each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples.
[0082] A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing according to the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes are divided into two major categories: vowels and consonants. For example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, "dài" has three phonemes, etc.
[0083] A phoneme is the smallest unit or the smallest speech segment that constitutes a syllable, and is the smallest linear speech unit divided from the perspective of voice quality. A phoneme is a specific physical phenomenon. The phonetic symbols of the International Phonetic Alphabet (developed by the International Phonetic Association to uniformly mark the phonetic sounds of various countries. Also known as the "International Phonetic Alphabet" or "Universal Phonetic Alphabet") correspond one by one to the phonemes of all human languages.
[0084] In the embodiments of the present application, for each speech frame in the speech signal, the acoustic model can identify the phoneme corresponding to the speech frame to obtain the probabilities that the phoneme of the speech frame belongs to each preset phoneme and the null output.
[0085] For example, in a possible implementation manner, the above phoneme space contains 212 phonemes and a null output (indicating that there is no user pronunciation in the corresponding speech frame). That is to say, for an input speech frame, the acoustic model shown in the embodiments of the present application can output the probabilities that the speech frame corresponds to 212 phonemes and the null output respectively.
[0086] Step 22, suppress and adjust the probability of the null output in the phoneme recognition results corresponding to each speech frame to reduce the ratio of the probability of the null output to the probabilities of each phoneme in the phoneme recognition results.
[0087] Step 23: Input the phoneme recognition results corresponding to each adjusted speech frame into the decoding graph to obtain the recognition text sequence corresponding to the speech signal.
[0088] In the embodiment of the present application, after the phoneme recognition results are input into the decoding graph, the decoding graph determines whether the phoneme recognition results correspond to a certain phoneme or an empty output according to the probabilities of each phoneme and the empty output in the phoneme space of the phoneme recognition results, and determines the corresponding text according to the determined phoneme. If the phoneme recognition results correspond to an empty output, it is determined that the speech frame corresponding to the phoneme recognition results does not contain user pronunciation, that is, there is no corresponding text.
[0089] Since the above phoneme recognition results include empty outputs, it may lead to an increase in the recognition error rate. For example, it is possible that a speech frame with pronunciation is misrecognized as an empty output (this situation is also called a deletion error), thus affecting the accuracy of speech recognition. In response to this, the solution shown in the embodiment of the present application suppresses the probability of the empty output in the phoneme recognition results after the acoustic model outputs the phoneme recognition results. As the probability of the empty output in the phoneme recognition results is suppressed, the possibility that the phoneme recognition results are recognized as a certain phoneme also increases, thereby effectively reducing the situation where a speech frame with pronunciation is misrecognized as an empty output.
[0090] In summary, for the phoneme recognition results including the probability distribution of speech frames on each phoneme and the empty output, the solution shown in the embodiment of the present application suppresses the probability of the empty output in the phoneme recognition results before inputting the phoneme recognition results into the decoding graph, reduces the probability that the speech frame is recognized as an empty output, thereby reducing the possibility that the speech frame is misrecognized as an empty output, that is, reducing the deletion error of the model, and thus improving the recognition accuracy of the model.
[0091] Please refer to Figure 3 , which is a schematic flowchart of a speech recognition method shown according to an exemplary embodiment. The speech recognition method can be used in a computer device. For example, the computer device can be the speech recognition device 140 or the server 160 in the above Figure 1 shown system, or the computer device can include both the speech recognition device 140 and the server 160 in the above Figure 1 shown system. As Figure 3 shown, the speech recognition method can include the following steps:
[0092] Step 301: Obtain a speech signal, where the speech signal includes each speech frame obtained by segmenting the original speech.
[0093] In the embodiments of the present application, after the voice acquisition component collects the original voice during the user's speech, the collected original voice is sent to a computer device, for example, sent to a voice recognition device, and the voice recognition device segments the original voice to obtain a plurality of voice frames.
[0094] In a possible implementation manner, the voice recognition device may segment the original voice into short-time voice segments with overlap. For example, generally for a voice with a sampling rate of 16K, the length of one frame of voice after segmentation is 25ms, and the overlap between frames is 15ms. This process is also called "framing".
[0095] Step 302: Process the voice signal through an acoustic model to obtain the phoneme recognition result corresponding to each voice frame in the voice signal.
[0096] Wherein, the phoneme recognition result is used to indicate the probability distribution of the corresponding voice frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through voice signal samples and the actual phonemes of each voice frame in the voice signal samples.
[0097] In the embodiments of the present application, the acoustic model is an end-to-end machine learning model, and its input data includes voice frames in the voice signal (for example, the input includes the feature vector of the voice frame), and the output data is the distribution probability of the phoneme of the predicted voice frame in the phoneme space, that is, the phoneme recognition result.
[0098] For example, the above phoneme recognition result can be represented as a probability vector as shown below:
[0099] (p 0 , p 1 , p 2 , …… p 212 )
[0100] In the above probability vector, p 0 represents the probability that the voice frame is a null output, and p 1 represents the probability that the voice frame corresponds to the first phoneme. The entire phoneme space includes 212 phonemes plus a null output.
[0101] In a possible implementation manner, the process of processing the voice signal through the acoustic model to obtain the phoneme recognition result corresponding to each voice frame in the voice signal includes:
[0102] Extract features from the target voice frame to obtain the feature vector of the target voice frame; the target voice frame is any one of the various voice frames;
[0103] Input the target voice frame into the encoder in the acoustic model to obtain the acoustic hidden layer representation vector of the target voice frame;
[0104] Input the phoneme information of the historical recognition text of the target speech frame into the predictor in the acoustic model to obtain the text hidden layer representation vector of the target speech frame; the historical recognition text of the target speech frame is the text obtained by recognizing the phoneme recognition results of the first n non-empty output speech frames of the target speech frame in the decoding graph; n is an integer greater than or equal to 1.
[0105] Input the acoustic hidden layer representation vector of the target speech frame and the text hidden layer representation vector of the target speech frame into the joint network to obtain the phoneme recognition result of the target speech frame.
[0106] In the embodiments of the present application, the above acoustic model can be implemented through a Transducer model. The Transducer model is introduced as follows:
[0107] Given an input sequence:
[0108]
[0109] And an output sequence:
[0110]
[0111] Where, Represents the set of all input sequences, Represents the set of all output sequences, Are all real number vectors, And Represent the input and output spaces respectively. For example, in this solution, the Transducer model is used for phoneme recognition. The input sequence x is a sequence of feature vectors, such as Filter Bank (FBank) features, or Mel Frequency Cepstrum Coefficient (MFCC) features. x t Represents the feature vector at time t; the output sequence y is a sequence of phonemes, and y u Represents the phoneme at the u-th step.
[0112] Define an extended output space Represents the empty output symbol, indicating that the model has no output. After introducing the empty output symbol, the sequence Is equivalent to In this solution, due to the introduction of the empty output, the output sequence and the input sequence will have the same length. Therefore, the elements In the set Are called "alignments". Given any input sequence, the Transducer model defines a conditional distribution This conditional distribution will be used to calculate the probability of the output sequence y given the input sequence x:
[0113]
[0114] where denotes removing the empty outputs in the aligned sequence, denotes adding empty outputs to the output sequence to generate the aligned sequence. As can be seen from formula (1), to calculate the probability of the output sequence y, the conditional probabilities of all possible alignments a corresponding to the sequence y need to be summed up. Please refer to Figure 4 , which shows a schematic diagram of the alignment process involved in the embodiments of the present application. The Figure 4 gives an example to illustrate formula (1).
[0115] In Figure 4 , U = 3, T = 5, and all possible paths from the lower left corner to the upper right corner are an alignment. The bold arrow indicates one possible path. When the model takes a step forward vertically, a non-empty symbol (phoneme) will be output; when the model takes a step forward horizontally, an empty symbol (i.e., the above-mentioned empty output) will be output, indicating that no output is generated. At the same moment, the model allows multiple outputs to be generated.
[0116] To model generally three sub-networks are jointly modeled. Please refer to Figure 5 , which shows a schematic diagram of the structure of an acoustic model involved in the embodiments of the present application. As Figure 5 shown. The acoustic model includes an encoder 51, a predictor 52, and a joint network 53.
[0117] Among them, the encoder 51 (Encoder) can be a recurrent neural network, such as a Long Short-Term Memory (LSTM) network, which accepts the audio feature input at time t and outputs an acoustic hidden layer representation
[0118] The predictor 52 (Predictor) can be a recurrent neural network, such as an LSTM, which accepts the non-empty output labels of the model history and outputs a text hidden layer representation
[0119] The joint network 53 (Joint Network) can be a fully connected neural network, such as a linear layer plus an activation unit, which is used to and sum up after linear transformation and output a hidden unit representation z i ; finally, through a softmax function, it is converted into a probability distribution.
[0120] The above-mentioned Figure 5 in and aligned Finally, the calculation of formula (1) is as follows:
[0121]
[0122] For the calculation of formula (2), all possible alignment paths need to be traversed and calculated. Directly using this algorithm will result in a large amount of computational effort. During the model training process, a forward-backward algorithm can be used to calculate the probability of formula (2).
[0123] In a possible implementation, the encoder is a Feedforward Sequential Memory Networks (FSMN).
[0124] In a possible implementation, the predictor is a one-dimensional convolutional network.
[0125] The solution shown in the embodiments of this application can be applied to scenarios with limited computing power such as in-vehicle offline speech recognition systems. In-vehicle devices have high requirements for the number of model parameters and computational effort, and the computing power of the Central Processing Unit (CPU) is limited. Therefore, higher requirements are imposed on the number of model parameters and the model structure. In order to reduce the computational effort and adapt to such application scenarios with limited computing power, the solution shown in this application uses a fully feedforward neural network FSMN as the Encoder of the model, and uses a one-dimensional convolutional network to replace the commonly used Long Short-Term Memory network LSTM as the Predictor.
[0126] For the above-mentioned Transducer model, in order to characterize the historical information of the model, the Encoder and Predictor networks generally adopt a Recurrent Neural Network (RNN) structure, such as LSTM or Gated Recurrent Unit (GRU). However, on embedded devices with limited computing resources, the recurrent neural network will bring a large amount of computational effort and occupy a large amount of CPU resources. On the other hand, the content of in-vehicle offline speech recognition is mainly query and control instructions, and the sentences are relatively short, and there is no need for too long historical information. For this reason, this solution uses an Encoder based on FSMN and a Predictor network based on one-dimensional convolution. On the one hand, the model parameters can be compressed, and on the other hand, a large amount of computing resources can be saved, the computing speed can be greatly improved, and the real-time performance of speech recognition can be guaranteed.
[0127] In this solution, an FSMN-based Encoder structure is adopted. The FSMN network is applied to large vocabulary speech recognition tasks. The FSMN structure adopted in this solution can be a structure with a projection layer and residual connections.
[0128] For the Predictor network, a one-dimensional convolutional network is adopted in this solution to predict the output based on limited history to generate the current output. Please refer to Figure 6 , which shows the network structure diagram of the predictor involved in the embodiments of the present application. As Figure 6 shown, the Predictor network uses 4 non-empty historical outputs to predict the current output framework. That is, after the 4 non-empty historical outputs 61 corresponding to the current input are vector-mapped, they are input into a one-dimensional convolutional network 62 to obtain the text hidden layer representation vector.
[0129] In the embodiments of the present application, the above acoustic model can be trained through pre-set speech samples and the actual phonemes of each speech frame in the speech signal sample. For example, during the training process, a speech frame in the speech sample is input into the FSMN-based Encoder network of the acoustic model, and the actual phonemes of the first 4 non-empty speech frames of the speech frame (when there are no historical non-empty speech frames at the start of training, or when the historical non-empty speech frames are insufficient, pre-set phonemes can be used instead) are input into the Predictor network based on one-dimensional convolution. During the process of the acoustic model processing the input data, the parameters of the three parts (Encoder, Predictor, and joint network) in the acoustic model are updated to maximize the sum of the probabilities on all possible alignment paths, that is, the result of the above formula (2), thereby realizing the training of the acoustic model.
[0130] Step 303, suppress and adjust the probability of the empty output in the phoneme recognition result corresponding to each speech frame to reduce the ratio of the probability of the empty output in the phoneme recognition result to the probability of each phoneme.
[0131] In a possible implementation manner, the suppressing and adjusting the probability of the empty output in the phoneme recognition result corresponding to each speech frame includes:
[0132] Adjust the phoneme recognition result corresponding to each speech frame by at least one of the following adjustment methods:
[0133] Reduce the probability of the empty output in the phoneme recognition result corresponding to each speech frame;
[0134] And increase the probability of each phoneme in the phoneme recognition result corresponding to each speech frame.
[0135] In a possible implementation, reducing the probability of null output in the phoneme recognition results corresponding to each speech frame includes:
[0136] Multiplying the probability of null output in the phoneme recognition results corresponding to each speech frame by a first weight, where the first weight is less than 1 and greater than 0.
[0137] In the embodiments of the present application, to suppress the probability of null output in the phoneme recognition results, only the probability of null output in the phoneme recognition results can be reduced. For example, multiply a number between 0 and 1 by the probability of null output in the phoneme recognition results. In this way, when the probabilities of each phoneme in the phoneme recognition results remain unchanged, the ratio between the probability of null output and the probabilities of each phoneme can be reduced.
[0138] In a possible implementation, reducing the probability of null output in the phoneme recognition results corresponding to each speech frame includes:
[0139] Multiplying the probabilities of each phoneme in the phoneme recognition results corresponding to each speech frame by a second weight, where the second weight is greater than 1.
[0140] In the embodiments of the present application, to suppress the probability of null output in the phoneme recognition results, only the probability of null output in the lower phoneme recognition results can be increased. For example, multiply a number greater than 1 by the probabilities of each phoneme in the phoneme recognition results. In this way, when the probability of null output in the phoneme recognition results remains unchanged, the ratio between the probability of null output and the probabilities of each phoneme can be reduced.
[0141] In another exemplary solution, the computer device can also reduce the probability of null output in the phoneme recognition results while increasing the probabilities of each phoneme in the phoneme recognition results. For example, multiply a number between 0 and 1 by the probability of null output in the phoneme recognition results, and at the same time, multiply a number greater than 1 by the probabilities of each phoneme in the phoneme recognition results.
[0142] In this solution, in order to obtain the alignment before the input and output, the above acoustic model needs to insert a null output symbol in the input phoneme sequence, that is The symbol is predicted using the model like other phonemes. Assuming the total number of non-null phonemes is P, the output dimension of the final model is P + 1, and usually the 0th dimension represents the null output Experiments have found that the introduction of null output has significantly increased the deletion errors of the model, which indicates that a large number of phonemes are misrecognized as null output. To solve the problem of too high probability of null output, in the Transducer decoding process of the present application, the probability weight of null output is adjusted to reduce the generation of deletion errors.
[0143] Taking the example of multiplying the probability of null output in the phoneme recognition result corresponding to each speech frame by the first weight, assuming the probability of null output is To reduce the probability of null output, in this solution, on the basis of the original probability value of null output, divide by a weight α greater than 1, α > 1, and α is called the discount factor. The adjusted probability value of null output is:
[0144]
[0145] Generally speaking, the logarithmic probability is used as the final value to participate in the calculation of the final decoding score. Therefore, after taking the logarithm of both sides of formula (3), we can get:
[0146]
[0147] The result of the above formula (4) can be used as the adjusted probability of null output for subsequent decoding.
[0148] In a possible implementation manner, the above first weight or second weight is preset in the computer device by developers or managers. For example, the above first weight or second weight can be preset in the speech recognition model by developers.
[0149] Step 304, input the phoneme recognition results in which the probability of null output in the phoneme recognition results corresponding to each speech frame meets the specified condition into the decoding graph, and obtain the recognition text sequence corresponding to the speech signal.
[0150] In a possible implementation manner, inputting the adjusted phoneme recognition results corresponding to each speech frame into the decoding graph to obtain the recognition text sequence corresponding to the speech signal includes:
[0151] In response to the probability of null output in the target phoneme recognition result meeting the specified condition, input the target phoneme recognition result into the decoding graph to obtain the recognition text corresponding to the target phoneme recognition result;
[0152] Wherein, the target phoneme recognition result is any one of the phoneme recognition results corresponding to each speech frame.
[0153] In a possible implementation manner, the specified condition includes:
[0154] The probability of null output in the target phoneme recognition result is less than the probability threshold.
[0155] It is found in the experiment that, compared with the DNN-HMM model, the output of the Transducer model has an obvious spike effect, that is, at a certain moment, the model will output a certain prediction result with extremely high confidence. By using the spike effect of the model, we can skip the probability that the model predicts an empty output during the decoding process, that is, these probabilities will not participate in the decoding process of the decoding graph. Since this patent uses phonemes as the modeling unit, and at the same time, the empty output is skipped during decoding, the number of steps for decoding graph search is only related to the number of phonemes. This solution is called "Phone Synchronous Decoding (PSD)". The following figure shows the entire process of the PSD algorithm and the adjustment of the empty output weight proposed in this solution:
[0156] Algorithm 1: PSD algorithm;
[0157]
[0158]
[0159] Among them, the weight is adjusted in line 6 of the above algorithm, and β in the algorithm blank is 1 / α in formula (3). Lines 13-17 of the above algorithm are the PSD algorithm proposed in this solution, that is, only when the probability of the empty output is less than a certain threshold γ blank will the probability distribution of the network output participate in the decoding of the subsequent decoding graph.
[0160] In a possible implementation manner, the above probability threshold is preset in the computer device by developers or managers. For example, the above probability threshold can be preset in the speech recognition model by developers.
[0161] In a possible implementation manner, before inputting the phoneme recognition results corresponding to the adjusted speech frames into the decoding graph to obtain the recognition text sequence corresponding to the speech signal, it further includes:
[0162] Obtain a threshold influence parameter, where the threshold influence parameter includes at least one of ambient sound intensity, the number of speech recognition failures within a specified time period, and user setting information;
[0163] Determine the probability threshold based on the threshold influence parameter.
[0164] In the embodiments of the present application, the above probability threshold can also be adjusted by the computer device during the speech recognition process. That is to say, the computer device can obtain relevant parameters that may affect the value of the probability threshold and flexibly set the probability threshold through the relevant parameters.
[0165] For example, the ambient sound intensity may interfere with the speech uttered by the user. Therefore, when the ambient sound intensity is strong, the computer device can set a relatively high probability threshold, so that more phoneme recognition results are output to the decoding graph for decoding, thereby ensuring the accuracy of recognition; conversely, when the ambient sound intensity is weak, the computer device can set a relatively low probability threshold, so that more phoneme recognition results are skipped, thereby ensuring the efficiency of recognition.
[0166] For another example, the accuracy of the decoding graph in decoding the phoneme recognition results affects the success rate of speech recognition. When the number of speech recognition failures within a specified time period (such as a period of time before the current moment, for example, 5 minutes) is too large, the computer device can set a relatively high probability threshold, so that more phoneme recognition results are output to the decoding graph for decoding, thereby ensuring the accuracy of recognition; conversely, when the number of speech recognition failures within the specified time period is small or there is no failure, the computer device can set a relatively low probability threshold, so that more phoneme recognition results are skipped, thereby ensuring the efficiency of recognition.
[0167] In a possible implementation manner, the decoding graph is composed of a composite of a phoneme dictionary and a language model.
[0168] The decoding graph adopted in this solution is composed of a composite of two sub-weighted finite state transducer (WFST) graphs, namely a phoneme dictionary and a language model.
[0169] Phoneme dictionary WFST: The mapping from Chinese characters or words to phoneme sequences. Inputting a phoneme sequence string, the WFST can output the corresponding Chinese characters or words; generally, this WFST is independent of the text domain and is a common part in different recognition tasks;
[0170] Language model WFST: This WFST is usually converted from an n-gram language model. The language model is used to calculate the probability of a sentence appearing and is trained using training data and statistical methods. Generally, for texts in different domains, such as news and spoken dialogue texts, there are significant differences in common words and word collocations. Therefore, when performing speech recognition in different domains, adaptation can be achieved by changing the language model WFST.
[0171] Please refer to Figure 7 , which shows the model training and application flow chart involved in the embodiments of the present application. As Figure 7As shown in the figure, taking the application to in-vehicle devices as an example, after the model training shown in the embodiments of this application is completed, libtorch is used for model quantization and deployment. For the Android version of libtorch, the QNNPACK library is used for INT8 matrix calculations, greatly accelerating the matrix operation speed. The model is trained using pytorch in the Python environment 71, and then post-training quantization is performed on the model, that is, the model parameters are quantized to INT8, and INT8 matrix multiplication is used to accelerate the calculation. After exporting the quantized model, it is used for forward inference in the C++ environment 72 to be tested with test data.
[0172] Through the solution shown in this application, on the one hand, during the training process of the end-to-end model based on Transducer, frame-level alignment information is not required, greatly simplifying the modeling process; secondly, the decoding graph is simplified, reducing the search space. In the method proposed in this solution, due to phoneme modeling, the decoding graph only needs to be composed of L and G, and the search space is greatly reduced. Finally, by using phoneme modeling and combining with a custom decoding graph, flexible customization requirements can be achieved. According to different business scenarios, without changing the acoustic model, only the language model needs to be customized to adapt to their respective business scenarios.
[0173] Compared with the offline recognition system in the related technology, this solution has advantages in both recognition rate and CPU occupancy rate:
[0174] In terms of recognition rate, compared with the DNN combined with Hidden Markov Model (HMM) system model (DNN-HMM model), the system model shown in this solution has been greatly improved;
[0175] In terms of CPU occupancy, when the number of model parameters of the system model shown in this solution is 4 times that of the DNN-HMM system, it still has a similar CPU occupancy rate to the DNN-HMM system model.
[0176] The comparison of speech recognition rates is as follows:
[0177] Table 1 below shows the comparison of the Character Error Rate (CER) between the existing DNN-HMM system and the Transducer system proposed in this solution on 3 data sets.
[0178] Table 1
[0179] Model Number of parameters Test set 1 CER (%) Test set 2 CER (%) DNN-HMM 0.7M 14.88 19.77 Transducer1 0.8M 12.1 16.09 Tansducder2 1.9M 9.76 13.4 Tansducder3 2.1M 8.93 13.18
[0180] As can be seen from Table 1, with similar numbers of parameters, on the two test sets, the Transducer1 model achieved relative CER decreases of 18.7% and 18.6% respectively. At the same time, after increasing the number of model parameters, using Transducer3, word error rates of 8.93% and 13.18% were achieved respectively.
[0181] CPU Occupancy Rate Comparison:
[0182] Table 2
[0183] Model Number of parameters CPU occupancy (peak) DNN-HMM 0.7M 16% Transducer1 0.8M 18% Tansducder2 1.9M 20% Tansducder3 2.1M 20%
[0184] By comparing Transducer1 and DNN-HMM in Table 2, when the two models have the same number of parameters, the peak value of the Transducer1 model is 2% higher than that of the DNN-HMM model. However, when the number of model parameters increases, the peak value of the Transducer model does not change significantly. Under the conditions of significantly increasing the number of model parameters and reducing the recognition error rate, the CPU occupancy rate still remains at a low level.
[0185] In summary, for the phoneme recognition result including the probability distribution of speech frames on each phoneme and the empty output in the embodiment of the present application, before inputting the phoneme recognition result into the decoding graph, first suppress the probability of the empty output in the phoneme recognition result, reduce the probability that the speech frame is recognized as an empty output, thereby reducing the possibility that the speech frame is misrecognized as an empty output, that is, reducing the deletion error of the model, and thus improving the recognition accuracy of the model.
[0186] The above Figure 3 The solution in the embodiment shown in the present application is described by taking the simultaneous application of empty output weight adjustment (step 303) and decoding frame skipping (corresponding to step 304) as an example. In other implementation solutions, the empty output weight adjustment and decoding frame skipping can also be applied independently. For example, in an exemplary embodiment of the present application, when the above decoding frame skipping is applied independently, the solution shown in the present application can be as follows:
[0187] Obtain a speech signal, which includes each speech frame obtained by segmenting the original speech;
[0188] Process the speech signal through an acoustic model to obtain the phoneme recognition result corresponding to each speech frame; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and an empty output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples;
[0189] Among the phoneme recognition results corresponding to each of the voice frames, input the phoneme recognition results in which the probability of null output satisfies the specified condition into the decoding graph to obtain the recognition text sequence corresponding to the voice signal.
[0190] In summary, for the phoneme recognition results including the probability distribution of voice frames on each phoneme and null output, when inputting the phoneme recognition results into the decoding graph, the proposed solution in the embodiments of the present application can decode the phoneme recognition results in which the probability of null output meets the conditions, reducing the number of phoneme recognition results to be decoded and skipping unnecessary decoding steps, thereby effectively improving the speech recognition efficiency.
[0191] Please refer to Figure 8 , which is a framework diagram of a speech recognition system shown according to an exemplary embodiment. As Figure 8 shown, the audio acquisition device 81 is connected to the speech recognition device 82, and the speech recognition device 82 includes an acoustic model 82a, a probability adjustment unit 82b, a decoding graph input unit 82c, a decoding graph 82d, and a feature extraction unit 82e. Among them, the decoding graph 82d is composed of a phoneme dictionary and a language model.
[0192] During application, after the audio acquisition device 81 acquires the user's original speech, it transmits the original speech to the feature extraction unit 82e in the speech recognition device 82. After the feature extraction unit performs segmentation and feature extraction on each voice frame, the speech feature of a voice frame and the phonemes of the text recognized by the decoding graph 82d for the first 4 non-null voice frames of this voice frame are respectively input into the FSMN and one-dimensional convolutional network in the acoustic model 82a to obtain the phoneme recognition result of the acoustic model 82a for this voice frame.
[0193] The phoneme recognition result is input to the probability adjustment unit 82b for probability adjustment of null output to obtain the adjusted phoneme recognition result; the adjusted speech recognition result is judged by the decoding graph input unit 82c. When it is judged that the adjusted probability of null output is less than the threshold, it is determined that decoding is required, and the decoding graph input unit 82c inputs the adjusted phoneme recognition result into the decoding graph 82d, and the decoding graph 82d recognizes the text; otherwise, if it is judged that the adjusted probability of null output is not less than the threshold, it is determined that decoding is not required, and the adjusted speech recognition result is discarded.
[0194] After the above decoding graph recognizes the adjusted phoneme recognition results of each voice frame and outputs the text sequence, the text sequence can be output to the natural language processing component, and the natural language processing component responds to the user's input speech.
[0195] Figure 9It is a structural block diagram of a voice recognition device shown according to an exemplary embodiment. The voice recognition device can implement Figure 2 or Figure 3 all or part of the steps in the method provided by the illustrated embodiment. The voice recognition device may include:
[0196] A voice signal processing module 901, configured to process a voice signal through an acoustic model to obtain a phoneme recognition result corresponding to each voice frame in the voice signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding voice frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through voice signal samples and the actual phonemes of each voice frame in the voice signal samples;
[0197] A probability adjustment module 902, configured to suppress and adjust the probability of the null output in the phoneme recognition results corresponding to the respective voice frames, so as to reduce the ratio of the probability of the null output in the phoneme recognition results to the probabilities of the respective phonemes;
[0198] A decoding module 903, configured to input the adjusted phoneme recognition results corresponding to the respective voice frames into a decoding graph to obtain a recognition text sequence corresponding to the voice signal.
[0199] In a possible implementation manner, the probability adjustment module 902 is configured to adjust the phoneme recognition results corresponding to the respective voice frames by at least one of the following adjustment methods:
[0200] Reduce the probability of the null output in the phoneme recognition results corresponding to the respective voice frames;
[0201] And,
[0202] Increase the probabilities of the respective phonemes in the phoneme recognition results corresponding to the respective voice frames.
[0203] In a possible implementation manner, the probability adjustment module 902 is configured to multiply the probability of the null output in the phoneme recognition results corresponding to the respective voice frames by a first weight, where the first weight is less than 1 and greater than 0.
[0204] In a possible implementation manner, the probability adjustment module 902 is configured to multiply the probabilities of the respective phonemes in the phoneme recognition results corresponding to the respective voice frames by a second weight, where the second weight is greater than 1.
[0205] In a possible implementation manner, the decoding module 903 is configured to,
[0206] In response to the probability of the null output in the target phoneme recognition result satisfying a specified condition, input the target phoneme recognition result into the decoding graph to obtain the recognition text corresponding to the target phoneme recognition result;
[0207] Wherein, the target phoneme recognition result is any one of the phoneme recognition results corresponding to the respective speech frames.
[0208] In a possible implementation, the specified condition includes:
[0209] The probability of the null output in the target phoneme recognition result is less than a probability threshold.
[0210] In a possible implementation, the device further includes:
[0211] A parameter acquisition module, configured to acquire a threshold influence parameter, where the threshold influence parameter includes at least one of ambient sound intensity, the number of speech recognition failures within a specified time period, and user setting information;
[0212] A threshold determination module, configured to determine the probability threshold based on the threshold influence parameter.
[0213] In a possible implementation, the speech signal processing module 901 is configured to,
[0214] Extract features from a target speech frame to obtain a feature vector of the target speech frame; the target speech frame is any one of the respective speech frames;
[0215] Input the target speech frame into an encoder in the acoustic model to obtain an acoustic hidden layer representation vector of the target speech frame;
[0216] Input the phoneme information of the historical recognition text of the target speech frame into a predictor in the acoustic model to obtain a text hidden layer representation vector of the target speech frame; the historical recognition text of the target speech frame is the text obtained by the decoding graph recognizing the phoneme recognition results of the first n non-null output speech frames of the target speech frame; n is an integer greater than or equal to 1;
[0217] Input the acoustic hidden layer representation vector of the target speech frame and the text hidden layer representation vector of the target speech frame into a joint network to obtain the phoneme recognition result of the target speech frame.
[0218] In a possible implementation, the encoder is a forward sequence memory network FSMN.
[0219] In a possible implementation, the predictor is a one-dimensional convolutional network.
[0220] In a possible implementation, the decoding graph is composed of a compound of a phoneme dictionary and a language model.
[0221] In summary, for the phoneme recognition result including the probability distribution of speech frames on each phoneme and the null output, before inputting the phoneme recognition result into the decoding graph, the probability of the null output in the phoneme recognition result is first suppressed, reducing the probability that a speech frame is recognized as a null output, thereby reducing the possibility that a speech frame is misrecognized as a null output, that is, reducing the deletion error of the model, and thus improving the recognition accuracy of the model.
[0222] Figure 10 FIG. 7 is a schematic structural diagram of a computer device according to an exemplary embodiment. The computer device can be implemented as the computer device in each of the above method embodiments. The computer device 1000 includes a central processing unit 1001, a system memory 1004 including a random access memory (RAM) 1002 and a read-only memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 further includes a basic input / output system 1006 for facilitating information transfer between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0223] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is, the mass storage device 1007 can include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0224] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, flash memory or other solid-state storage technologies, CD-ROM, or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that the computer storage media is not limited to the above several types. The above-mentioned system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0225] The computer device 1000 can be connected to the Internet or other network devices through the network interface unit 1011 connected to the system bus 1005.
[0226] The memory further includes at least one computer instruction, and the at least one computer instruction is stored in the memory, and the processor realizes Figure 2 or Figure 3 all or part of the steps of the method shown.
[0227] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including a computer program (instructions), and the above program (instructions) can be executed by the processor of the computer device to complete the methods shown in various embodiments of the present application. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0228] In an exemplary embodiment, a computer program product or computer program is also provided, and the computer program product or computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods shown in the above various embodiments.
[0229] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the claims.
[0230] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A speech recognition method, characterized in that, the method includes: processing the speech signal through an acoustic model to obtain a phoneme recognition result corresponding to each speech frame in the speech signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples; suppressing and adjusting the probability of the null output in the phoneme recognition results corresponding to the respective speech frames to reduce the ratio of the probability of the null output in the phoneme recognition results to the probabilities of the respective phonemes; obtaining a threshold influence parameter, the threshold influence parameter including at least one of ambient sound intensity, the number of speech recognition failures within a specified time period, and user setting information; determining a probability threshold based on the threshold influence parameter; in response to the probability of the null output in the target phoneme recognition result being less than the probability threshold, inputting the target phoneme recognition result into a decoding graph to obtain a recognition text corresponding to the target phoneme recognition result; the target phoneme recognition result is any one of the adjusted phoneme recognition results corresponding to the respective speech frames.
2. The method according to claim 1, characterized in that, the suppressing and adjusting the probability of the null output in the phoneme recognition results corresponding to the respective speech frames includes: adjusting the phoneme recognition results corresponding to the respective speech frames by at least one of the following adjustment methods: reducing the probability of the null output in the phoneme recognition results corresponding to the respective speech frames; and, increasing the probabilities of the respective phonemes in the phoneme recognition results corresponding to the respective speech frames.
3. The method according to claim 2, characterized in that, the reducing the probability of the null output in the phoneme recognition results corresponding to the respective speech frames includes: multiplying the probability of the null output in the phoneme recognition results corresponding to the respective speech frames by a first weight, the first weight being less than 1 and greater than 0.
4. The method according to claim 2, characterized in that, the reducing the probability of the null output in the phoneme recognition results corresponding to the respective speech frames includes: multiplying the probabilities of the respective phonemes in the phoneme recognition results corresponding to the respective speech frames by a second weight, the second weight being greater than 1.
5. The method according to claim 1, characterized in that, the processing the speech signal through an acoustic model to obtain a phoneme recognition result corresponding to each speech frame in the speech signal includes: extracting features of a target speech frame to obtain a feature vector of the target speech frame; the target speech frame is any one of the respective speech frames; inputting the target speech frame into an encoder in the acoustic model to obtain an acoustic hidden layer representation vector of the target speech frame; Input the phoneme information of the historical recognition text of the target speech frame into the predictor in the acoustic model to obtain the text hidden layer representation vector of the target speech frame; the historical recognition text of the target speech frame is the text obtained by the decoding graph recognizing the phoneme recognition results of the first n non-empty output speech frames of the target speech frame; n is an integer greater than or equal to 1. Input the acoustic hidden layer representation vector of the target speech frame and the text hidden layer representation vector of the target speech frame into the joint network to obtain the phoneme recognition result of the target speech frame.
6. The method according to claim 5, wherein, the encoder is a forward sequence memory network FSMN.
7. The method according to claim 5, wherein, the predictor is a one-dimensional convolutional network.
8. The method according to any one of claims 1 to 7, wherein, the decoding graph is composed of a phoneme dictionary and a language model in combination.
9. A speech recognition device, wherein, the device includes: A speech signal processing module, configured to process a speech signal through an acoustic model to obtain a phoneme recognition result corresponding to each speech frame in the speech signal; the phoneme recognition result is used to indicate the probability distribution of the corresponding speech frame in the phoneme space; the phoneme space includes each phoneme and a null output; the acoustic model is trained through speech signal samples and the actual phonemes of each speech frame in the speech signal samples. A probability adjustment module, configured to suppress and adjust the probability of the null output in the phoneme recognition results corresponding to each speech frame, so as to reduce the ratio of the probability of the null output in the phoneme recognition results to the probabilities of each phoneme. A parameter acquisition module, configured to acquire a threshold influence parameter, where the threshold influence parameter includes at least one of ambient sound intensity, the number of speech recognition failures within a specified time period, and user setting information. A threshold determination module, configured to determine a probability threshold based on the threshold influence parameter. A decoding module, configured to input the target phoneme recognition result into a decoding graph in response to the probability of the null output in the target phoneme recognition result being less than the probability threshold, to obtain the recognition text corresponding to the target phoneme recognition result; the target phoneme recognition result is any one of the adjusted phoneme recognition results corresponding to each speech frame.
10. A computer device, wherein, the computer device includes a processor and a memory, and at least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the speech recognition method according to any one of claims 1 to 8.
11. A computer-readable storage medium, wherein, at least one computer instruction is stored in the storage medium, and the at least one computer instruction is loaded and executed by a processor to implement the speech recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Speech recognition with acoustic models
US20160372119A1