Speech Recognition Method, Apparatus and Medium

By using the historical decoding results of the Connectivist timing classification module in speech recognition technology for parallel decoding, the problem of low decoding dependence in the prior art is solved, and the real-time rate and efficiency of speech recognition are improved.

CN113889080BActive Publication Date: 2025-07-04BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111162913.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-07-04
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

In the existing speech recognition technology, the decoding of the current moment depends on the decoding results of the past moment, affecting the decoding efficiency, resulting in a low real-time rate of speech recognition.

Method used

The connectionist timing classification module is used to receive the first text sequence of the speech to be recognized and used as the historical decoding result of the decoding time, allowing the decoding of multiple decoding time to be performed in parallel, reducing the dependence on the decoding results of the past time.

Benefits of technology

The decoding efficiency and real-time rate of speech recognition are improved, and decoding can be performed without waiting for the decoding result of the past moment, thereby realizing parallel processing of multiple decoding moments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889080B_ABST
    Figure CN113889080B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a speech recognition method, apparatus, and medium. The method specifically includes: receiving a feature representation corresponding to the speech to be recognized from an encoder; receiving a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module; decoding the feature representation according to the first text sequence to obtain a corresponding second text sequence; and using the first text sequence as the historical decoding result corresponding to the decoding moment. The embodiment of the present invention can improve the real-time rate of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of speech recognition technology, and in particular, to a speech recognition method, apparatus, and medium. Background Art

[0002] Speech recognition technology is a technology that converts speech signals into corresponding texts through a computer. It converts the lexical content in human speech into actual text output, and is one of the main ways to achieve human-machine interaction. Speech recognition technology has been widely applied in various scenarios such as speech input methods, speech dialing, and in-vehicle navigation. With the continuous enrichment of the application scenarios of speech recognition technology, higher requirements are put forward for the real-time rate and accuracy of speech recognition technology.

[0003] Currently, for speech recognition technology, an attention-based end-to-end model can be adopted. The end-to-end model includes: an encoder and a decoder. Among them, the encoder encodes the input speech into a high-level feature representation. The decoder starts from the start symbol and, according to the output of the encoder and the decoding results of past moments, gradually decodes the text sequence corresponding to the current moment until the end marker is decoded.

[0004] The inventors found in the process of implementing the embodiments of the present invention that the decoding of the current moment depends on the decoding results of past moments, which affects the decoding efficiency and thus affects the real-time rate of speech recognition. Summary of the Invention

[0005] How to improve the real-time rate of speech recognition is a technical problem that needs to be solved by those skilled in the art. In view of the above problems, embodiments of the present invention propose a speech recognition method, apparatus, and medium that overcome the above problems or at least partially solve the above problems.

[0006] To solve the above problems, embodiments of the present invention disclose a speech recognition method, including:

[0007] Receiving a feature representation corresponding to the speech to be recognized from an encoder;

[0008] Receiving a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module;

[0009] Decoding the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0010] On the other hand, embodiments of the present invention disclose a speech recognition apparatus, the apparatus includes:

[0011] A first receiving module, configured to receive a feature representation corresponding to the speech to be recognized from an encoder;

[0012] A second receiving module, configured to receive a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module;

[0013] A decoding module, configured to decode the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as a historical decoding result corresponding to the decoding moment.

[0014] Optionally, the decoding module includes:

[0015] A historical decoding result determination module, configured to determine historical decoding results corresponding to multiple decoding moments according to the first text sequence;

[0016] A parallel decoding module, configured to perform parallel decoding of the feature representation at multiple decoding moments according to the historical decoding results corresponding to the multiple decoding moments.

[0017] Optionally, the apparatus further includes:

[0018] A first output module, configured to output the target text sequence obtained by the connectionist temporal classification module as a first speech recognition result; the target text sequence includes: the first text sequence, or a third text sequence;

[0019] A second output module, configured to output the second text sequence as a second speech recognition result; the second speech recognition result is used to replace the first speech recognition result.

[0020] Optionally, the feature representation corresponds to a data block included in the speech to be recognized; the first text sequence corresponds to the data block included in the speech to be recognized.

[0021] Optionally, the feature representation corresponds to a data block included in the speech to be recognized; the data block corresponds to a data block length;

[0022] The first text sequence is obtained according to a first feature representation corresponding to a first data block length; the third text sequence is obtained according to a second feature representation corresponding to a second data block length; the first data block length is greater than the second data block length.

[0023] Optionally, the apparatus is applied to a speech recognition model; the speech recognition model includes: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the encoder;

[0024] The encoder sends a second feature representation to the connectionist temporal classification module, and sends a first feature representation to the connectionist temporal classification module and the decoder;

[0025] The connectionist temporal classification module determines a third text sequence according to the second feature representation output by the encoder; the third text sequence is used as the first speech recognition result for output.

[0026] The connectionist temporal classification module determines a first text sequence according to the first feature representation output by the encoder, and sends the first text sequence to the decoder.

[0027] Optionally, the device is applied to a speech recognition model; the speech recognition model includes: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the encoder;

[0028] Wherein, during the training process, the decoder determines the historical decoding result corresponding to the current decoding moment according to the decoding result of the past moment; during the speech recognition process, the decoder determines the historical decoding results corresponding to multiple decoding moments according to the first text sequence.

[0029] On the other hand, an embodiment of the present invention discloses a device for speech recognition, including a memory, and one or more programs, wherein one or more programs are stored in the memory, and when the program is executed by one or more processors, the steps of the foregoing method are implemented.

[0030] On another aspect, an embodiment of the present invention discloses a machine-readable medium, on which instructions are stored, and when executed by one or more processors, the device is enabled to execute the foregoing method.

[0031] Embodiments of the present invention include the following advantages:

[0032] In the embodiments of the present invention, taking the first text sequence as the historical decoding result corresponding to the decoding moment can perform the decoding corresponding to the decoding moment without waiting for the decoding result of the past moment; further, it can enable the parallel execution of the decoding corresponding to multiple decoding moments respectively, so that the decoding efficiency can be improved, and further the real-time rate of speech recognition can be improved.

[0033] For example, the first text sequence includes: Then the first text sequence can be used as the historical decoding result corresponding to multiple decoding moments. Specifically, can be used as the historical decoding result corresponding to y2; and can be used as the historical decoding result corresponding to y3;... can be used as the historical decoding result corresponding to y i corresponding;... can be used as the historical decoding result corresponding to y nThe corresponding historical decoding results. Since multiple historical decoding results corresponding to multiple decoding moments can be provided via the first text sequence, it is possible to enable the parallel execution of decoding corresponding to multiple decoding moments respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic structural diagram of a speech recognition model according to an embodiment of the present invention;

[0035] Figure 2 is a schematic structural diagram of a speech recognition model according to an embodiment of the present invention;

[0036] Figure 3 is a flowchart of the steps of a speech recognition method according to an embodiment of the present invention;

[0037] Figure 4 is a schematic illustration of a speech recognition process according to an embodiment of the present invention;

[0038] Figure 5 is a flowchart of the steps of a speech recognition method according to an embodiment of the present invention;

[0039] Figure 6 is a flowchart of the steps of a speech recognition method according to an embodiment of the present invention;

[0040] Figure 7 is a flowchart of the steps of a speech recognition method according to an embodiment of the present invention;

[0041] Figure 8 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present invention;

[0042] Figure 9 is a block diagram of a device for speech recognition as a terminal according to an exemplary embodiment;

[0043] Figure 10 is a schematic structural diagram of a server in some embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0045] Embodiments of the present invention can be applied to speech recognition scenarios. The speech recognition scenario is used to convert speech into text, and the speech recognition scenario may include: speech input scenarios, intelligent chat scenarios, speech translation scenarios, etc.

[0046] Applying the current speech recognition technology, the decoding at the current moment usually depends on the decoding results of past moments, which affects the decoding efficiency and further affects the real-time rate of speech recognition.

[0047] For example, the decoding results include: y1, y2, y3, …… y i , …… y n , then the decoding of y i usually depends on the decoding results at past moments such as 1 to (i - 1). Therefore, the current decoding process generally includes: First, decode to obtain y1; then, decode y2 based on y1; next, decode y3 based on y1 and y2; …… Decode y i ; …… Decode y n based on the decoding results at past moments such as 1 to (n - 1). Since it is necessary to sequentially perform the decoding of y1, y2, y3, …… y i , …… y n , the decoding efficiency is thus relatively low, and further the real - time rate of speech recognition is relatively low.

[0048] Regarding the technical problem of how to improve the real - time rate of speech recognition, an embodiment of the present invention provides a speech recognition solution, which specifically includes: receiving a feature representation corresponding to the speech to be recognized from an encoder; receiving a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module; decoding the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0049] In the embodiment of the present invention, using the first text sequence as the historical decoding result corresponding to the decoding moment can perform the decoding corresponding to the decoding moment without waiting for the decoding results at past moments; further, it can enable the parallel execution of the decoding corresponding to multiple decoding moments respectively. Therefore, the decoding efficiency can be improved, and further the real - time rate of speech recognition can be improved.

[0050] For example, the first text sequence includes: Then this first text sequence can be used as the historical decoding result corresponding to multiple decoding moments. Specifically, can be used as the historical decoding result corresponding to y2; and can be used as the historical decoding result corresponding to y3; …… can be used as the historical decoding result corresponding to y i ; …… can be used as the historical decoding result corresponding to y n . Since the historical decoding results corresponding to multiple decoding moments can be provided simultaneously via the first text sequence, the parallel execution of the decoding corresponding to multiple decoding moments can be enabled.

[0051] The speech recognition method provided by the embodiments of the present invention can be applied to the application environment of a client and a server. The client and the server are located in a wired or wireless network, and through the wired or wireless network, the client and the server perform data interaction.

[0052] Optionally, the client can run on a terminal. For example, the client can be an APP (application program) running on the terminal, such as a speech transcription APP, or a speech translation APP, or an intelligent interaction APP, etc.

[0053] Taking the speech transcription APP as an example, the client can collect the speech to be recognized and send the speech to be recognized to the server. The server can use the solution of the embodiments of the present invention to process the speech to be recognized and return the speech recognition result to the client.

[0054] Taking the speech translation APP as an example, the client can collect the speech to be recognized and send the speech to be recognized to the server. The server can use the solution of the embodiments of the present invention to process the speech to be recognized, perform machine translation on the obtained speech recognition result to obtain a machine translation result, and return the machine translation result to the client.

[0055] Optionally, the above terminal may include: a conference terminal, a smart phone, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a vehicle-mounted computer, a desktop computer, a set-top box, a smart TV, a wearable device, a smart speaker, and so on. It can be understood that the embodiments of the present invention do not limit the specific terminal.

[0056] Method Embodiment 1

[0057] Method Embodiment 1 describes the training process of the speech recognition model. The speech recognition model can be trained according to training samples to improve the accuracy of speech recognition of the speech recognition model.

[0058] Referring to Figure 1 , a schematic structural diagram of a speech recognition model according to an embodiment of the present invention is shown. The speech recognition model may specifically include: an encoder 101, a decoder 102 and a connectionist temporal classification (CTC) module 103 respectively connected to the encoder 101.

[0059] Among them, the encoder 101 is used to encode the input speech into a high-level feature representation.

[0060] The connectionist temporal classification module 103 is used to decode the feature representation output by the encoder 101 to obtain a corresponding first text sequence.

[0061] The decoding methods adopted by the connectionist temporal classification module 103 may include, but are not limited to, the greedy search method, the beam search method, the prefix beam search method, etc. Among them, the greedy search method selects the output value with the highest probability at a certain moment. The beam search method calculates the probabilities of all possible hypotheses at a single moment and selects the highest several as a group; then, based on this group of hypotheses, the highest several with the highest probabilities are generated as a group of hypotheses, and so on until the last moment is reached. The prefix beam search records the label prefix corresponding to the output during the path search process.

[0062] The decoder 102 is used to decode the encoder output to obtain a corresponding second text sequence.

[0063] The training samples of the speech recognition model may include: N training samples, where N may be a positive integer greater than 1. The training samples may include: the speech samples corresponding to the complete sentences, and the text samples corresponding to the speech samples. In the embodiments of the present invention, it is not necessary to perform forced alignment on the speech samples and the text samples.

[0064] During the training process, the training samples can be used as the input sequence of the speech recognition model, and the input sequence is processed based on chunks. The processing based on chunks can enable the sharing of speech frames between different speech frames. Specifically, the processing of one chunk can utilize the information of the previous speech frames and can also utilize the information of the subsequent speech frames. Since more speech frame information can be utilized, the accuracy of speech recognition can be improved. Moreover, in the embodiments of the present invention, the processing results can be output in units of chunks. For example, the encoder can output the feature representation in units of chunks, so the training time and the speech recognition time can be reduced, and thus the real-time performance of speech recognition can be improved.

[0065] In one implementation, assuming that the number of speech frames included in a chunk is C, the number of speech frames that a chunk looks left is L, and the number of speech frames that a chunk looks right is R, then the length of the speech frames that a speech frame within a chunk can see is: L + C + R. Assuming that the number of layers of the encoder is Q, the delay between the input sequence and the output sequence can be: C + Q * L.

[0066] The training paths in the embodiments of the present invention specifically include: training path 1 and training path 2. Among them, training path 1 is specifically: encoder 101 + decoder 102. Training path 2 is specifically: encoder 101 + connectionist temporal classification module 103.

[0067] In the forward training stage of the speech recognition model, the loss 1 corresponding to training path 1 can be obtained, and the loss 2 corresponding to training path 2 can be obtained; and, the loss 1 and loss 2 can be fused to obtain a fused loss. Further, backward training can be performed according to the fused loss. The fusion methods adopted for fusing the loss 1 and loss 2 can include but are not limited to: weighted average method, product method, etc.

[0068] Refer to Figure 2 , which shows a schematic structural diagram of a speech recognition model according to an embodiment of the present invention. The speech recognition model may specifically include: an encoder 101, and a decoder 102 and a connectionist temporal classification module 103 respectively connected to the encoder 101.

[0069] The encoder 101 may include Q layers of neural network structures, and Q may be a positive integer greater than 1. The single-layer neural network structure of the encoder 101 may further include: a first attention unit 111, a first operation unit 112, a first neural network unit 113, and a second operation unit 114.

[0070] The first attention unit 111 may determine data blocks from the training samples and process the data blocks using the multi-head self-attention mechanism (SAN, multi-head self attention). Figure 2 In, the input sequence corresponding to the data block is represented by X.

[0071] The first operation unit 112 may perform operation processing such as normalization and summation on the output of the first attention unit 111.

[0072] The first neural network unit 113 may process the output of the first operation unit 112 using a neural network.

[0073] The second operation unit 114 may perform operation processing such as normalization and summation on the output of the first neural network unit 113.

[0074] The feature representation output by the encoder 101 may be an acoustic feature representation. The acoustic features may include but are not limited to: Linear Prediction Cepstral Coefficients (LPCC), Mel Frequency Cepstrum Coefficient (MFCC), etc.

[0075] The decoder 102 may also include a multi-layer neural network structure. The single-layer neural network structure of the decoder 102 may further include: a second attention unit 121, a third operation unit 122, a third attention unit 123, a fourth operation unit 124, a second neural network unit 125, and a fifth operation unit 126.

[0076] Among them, the second attention unit 121 may use a masked multi-head self-attention mechanism to process the input text sample Y. The mask often refers to using a brand-new layer of attention mechanism weights to represent and learn the key degree of some parts in the feature data.

[0077] The third operation unit 122 may perform operation processing such as normalization and summation on the output of the second attention unit 121.

[0078] The third attention unit 123 may receive the output of the encoder 101 and the output of the third operation unit 122, and use a multi-head self-attention mechanism to process the output of the encoder 101 and the output of the third operation unit 122.

[0079] The fourth operation unit 124 may perform operation processing such as normalization and summation on the output of the third attention unit 123.

[0080] The second neural network unit 125 may use a neural network to process the output of the fourth operation unit 124.

[0081] The fifth operation unit 126 may perform operation processing such as normalization and summation on the output of the second neural network unit 125.

[0082] In practical applications, a classification module 104 may be set after the decoder 102. The classification module 104 is used to classify the output of the decoder 102 to obtain a first decoding result. For example, the first decoding result may be the probability corresponding to the predicted text. The loss 1, denoted as L1, may be determined according to the error information between the first decoding result and the text sample.

[0083] Optionally, the classification module 104 may use an activation function for classification. Examples of activation functions may include: sigmoid function (S-shaped function), tanh function (hyperbolic tangent function), relu (Rectified Linear Unit) function, softmax function (normalized exponential function).

[0084] In a specific implementation, a loss function such as a cross-entropy function may be used to determine the loss 1. It can be understood that the specific determination method of the loss 1 in the embodiments of the present invention is not limited.

[0085] The connectionist temporal classification module 103 can decode the feature representation output by the encoder 101 to obtain a second decoding result. For example, the second decoding result can be the probability corresponding to the predicted text. The loss 2, denoted as L2, can be determined based on the error information between the second decoding result and the text sample.

[0086] The loss 1 and the loss 2 can be fused to obtain a fused loss. Further, based on the fused loss, backward training can be performed to update the parameters of each part in the speech recognition model according to the fused loss until the fused loss meets a preset condition. The preset condition can be that the value of the fused loss is less than a preset value, etc. It can be understood that the embodiments of the present invention do not limit the specific preset condition.

[0087] It should be noted that during the training process, the historical decoding result corresponding to the decoding moment can be provided by the decoder itself. Specifically, during the training process, the decoder determines the historical decoding result corresponding to the current decoding moment according to the decoding results of past moments.

[0088] For example, the decoding results include: y1, y2, y3,... y i 、... y n , then during the training process, the decoding process of the decoder can include: First, decode to obtain y1; then, decode to obtain y2 according to y1; then, decode to obtain y3 according to y1 and y2;... decode to obtain y i according to the decoding results of past moments such as 1 to (i - 1);... decode to obtain y n .

[0089] In summary, according to the loss 1 corresponding to the decoder 102 and the loss 2 corresponding to the connectionist temporal classification module 103, the embodiments of the present invention perform joint training, which can utilize the advantages of the fast convergence speed of the decoder 102 and the automatic alignment of the connectionist temporal classification module 103 for unaligned training samples. Therefore, the accuracy of the fused loss can be improved, and further, the recognition ability of the speech recognition model can be improved.

[0090] Method Embodiment Two

[0091] Method Embodiment Two describes the usage process of the speech recognition model (i.e., the speech recognition model). The speech recognition model can recognize the speech to be recognized and output the corresponding speech recognition result.

[0092] Referring to Figure 3 , a flowchart of the steps of a speech recognition method according to an embodiment of the present invention is shown, which may specifically include the following steps:

[0093] Step 301, receive the feature representation corresponding to the speech to be recognized from the encoder;

[0094] Step 302: Receive a first text sequence corresponding to the speech to be recognized from the connectionist temporal classification module;

[0095] Step 303: Decode the above feature representation according to the above first text sequence to obtain a corresponding second text sequence; the above first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0096] Figure 3 The method embodiment shown can be executed by a decoder to improve the decoding efficiency and thus improve the real-time performance of speech recognition. It can be understood that the specific execution entity of the method embodiment in the present invention is not limited.

[0097] In step 301, the encoder can encode the input speech to be recognized into a high-level feature representation. The above feature representation can include but is not limited to an acoustic feature representation.

[0098] In a specific implementation, the encoder can use the complete speech to be recognized as an input sequence and process the complete speech to be recognized.

[0099] Alternatively, the encoder can determine data blocks from the speech to be recognized and process the speech to be recognized based on the data blocks. In this case, the feature representation can correspond to the data blocks; in other words, the encoder can output the feature representation corresponding to the data blocks to the connectionist temporal classification module and the decoder. Based on the processing of the data blocks, it is possible to share speech frames between different speech frames, so the time consumption of speech recognition can be reduced, and thus the real-time performance of speech recognition can be improved.

[0100] In step 302, the connectionist temporal classification module can decode according to the feature representation output by the encoder to obtain a first text sequence corresponding to the speech to be recognized. Further, the connectionist temporal classification module can also output the first text sequence to the decoder.

[0101] In step 303, the encoder can decode the above feature representation according to the above first text sequence to obtain a corresponding second text sequence.

[0102] In practical applications, the first text sequence can be the text decoding result corresponding to a data block in the speech to be recognized, or the first text sequence can be the text decoding result corresponding to a sentence in the speech to be recognized. The feature representation can be the feature representation corresponding to the data blocks in the speech to be recognized.

[0103] In the embodiments of the present invention, the first text sequence is used as the historical decoding result corresponding to the decoding moment, so that decoding corresponding to the decoding moment can be performed without waiting for the decoding results of past moments; further, decoding corresponding to multiple decoding moments can be performed in parallel, so that the decoding efficiency can be improved, and thus the real-time rate of speech recognition can be improved.

[0104] In a specific implementation, the decoding of the feature representation may specifically include: determining historical decoding results corresponding to multiple decoding moments according to the first text sequence; and performing parallel decoding of the feature representation at multiple decoding moments according to the historical decoding results corresponding to the multiple decoding moments.

[0105] In a specific implementation, the multiple decoding moments may be decoding moments included in a data block or a sentence in the speech to be recognized. The embodiments of the present invention do not limit the specific decoding moments.

[0106] Referring to Figure 4 , a schematic diagram of a speech recognition process according to an embodiment of the present invention is shown. An encoder processes the speech to be recognized to obtain a corresponding feature representation, and sends the feature representation to a decoder and a connectionist temporal classification module. The connectionist temporal classification module obtains a first text sequence according to the feature representation, and sends the first text sequence to the decoder.

[0107] The first text sequence may specifically include: Figure 4 As shown in The first text sequence can be used as the historical decoding results corresponding to multiple decoding moments.

[0108] Figure 4 In , SOS can represent the start marker of a data block or a sentence, and EOS can represent the end marker of a data block or a sentence. Figure 4 In , can be used as the historical decoding result corresponding to y2; and can be used as the historical decoding result corresponding to y3. Since the historical decoding results corresponding to multiple decoding moments can be provided simultaneously via the first text sequence, parallel decoding corresponding to multiple decoding moments can be enabled.

[0109] It should be noted that the processing of the decoder in the training process is different from that in the speech recognition process. This difference is reflected in the different ways of determining the historical decoding results.

[0110] During the training process, the decoder adopts a first determination method. Specifically, according to the decoding results at past times, the historical decoding results corresponding to the current decoding time are determined. During the speech recognition process, the decoder adopts a second determination method and determines the historical decoding results corresponding to multiple decoding times according to the first text sequence.

[0111] In a specific implementation, according to a control parameter, the decoder can be controlled to execute either the first determination method or the second determination method. In other words, the first determination method and the second determination method can be switched by updating the control parameter.

[0112] For example, the control parameter includes a first control parameter and a second control parameter. Among them, the first control parameter corresponds to the training process and the first determination method, and the second control parameter corresponds to the speech recognition process and the second determination method. In this way, when the control parameter is the first control parameter, the decoder can execute the first determination method. Or, when the control parameter is the second control parameter, the decoder can execute the second determination method.

[0113] In summary, in the speech recognition method according to the embodiments of the present invention, by using the first text sequence as the historical decoding result corresponding to the decoding time, decoding corresponding to the decoding time can be performed without waiting for the decoding results at past times; further, decoding corresponding to multiple decoding times can be performed in parallel, so that the decoding efficiency can be improved, and thus the real-time rate of speech recognition can be improved.

[0114] Method Embodiment Three

[0115] Method Embodiment Three illustrates the process of the speech recognition model outputting a speech recognition result.

[0116] Referring to Figure 5 , a step flowchart of a speech recognition method according to an embodiment of the present invention is shown, which may specifically include the following steps:

[0117] Step 501: Receive a feature representation corresponding to the speech to be recognized from an encoder;

[0118] Step 502: Receive a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module;

[0119] Step 503: Decode the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding time;

[0120] Relative to Figure 3 the method embodiment shown, the method of this embodiment may further include:

[0121] Step 504: Use the target text sequence obtained by the connectionist temporal classification module as the first speech recognition result and output it; the target text sequence may specifically include: a first text sequence, or a third text sequence;

[0122] Step 505: Output the second text sequence as the second speech recognition result; the second speech recognition result is used to replace the first speech recognition result.

[0123] Steps 504 and 505 may be executed by the output module of the speech recognition module, and this output module is used to output the speech recognition module. It can be understood that the specific execution entity of the method embodiment in the present invention is not limited.

[0124] The embodiment of the present invention can output the speech recognition result in stages.

[0125] Among them, in the first stage, the first speech recognition result obtained by the connectionist temporal classification module is output. The connectionist temporal classification module has the advantage of fast decoding speed and can quickly output the first speech recognition result. Specifically, the connectionist temporal classification module decodes in the case of time slice independence (decoding moment independence), so it can output the decoding result in units of characters and can realize the real-time output of the first speech recognition result.

[0126] In the second stage, the second speech recognition result obtained by the decoder is output. The second speech recognition result is used to replace the first speech recognition result. For example, the first speech recognition result can be replaced with the second speech recognition result by means of refreshing the screen.

[0127] Among them, the decoder can decode in units of data blocks, and the historical decoding results are considered in the decoding process, so the accuracy of the second speech recognition result can be improved. Further, the decoder can decode according to the attention mechanism, which can further improve the accuracy of the second speech recognition result. Experimental results show that the accuracy of the second speech recognition result is higher than that of the first speech recognition result.

[0128] It should be noted that the embodiment of the present invention can perform decoding corresponding to multiple decoding moments in parallel, so the output speed and real-time rate of the second speech recognition result can be improved. Experimental results show that the decoding speed of the decoder in the embodiment of the present invention is more than 10 times higher than that of the traditional decoder. Therefore, the embodiment of the present invention can be applied to scenarios such as speech input and speech translation in streaming speech recognition.

[0129] Method Embodiment Four

[0130] Refer to Figure 6 , which shows the step flowchart of a speech recognition method according to an embodiment of the present invention, and specifically may include the following steps:

[0131] Step 601: The encoder sends the feature representation to the Connectionist Temporal Classification (CTC) module and the decoder.

[0132] Step 602: The CTC module determines the first text sequence based on the feature representation output by the encoder, and sends the first text sequence to the decoder.

[0133] Step 603: Output the first text sequence as the first speech recognition result.

[0134] Step 604: The decoder decodes the feature representation according to the first text sequence to obtain the corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0135] Step 605: Output the second text sequence as the second speech recognition result.

[0136] In practical applications, the data block may correspond to a data block length. In this embodiment, the encoder can process according to a data block length to obtain a feature representation corresponding to a data block length, and output the feature representation to the CTC module and the decoder.

[0137] Method Embodiment Five

[0138] Refer to Figure 7 , which shows the step flowchart of a speech recognition method according to an embodiment of the present invention, and may specifically include the following steps:

[0139] Step 701: The encoder sends the second feature representation to the CTC module, and sends the first feature representation to the CTC module and the decoder.

[0140] Step 702: The CTC module determines the third text sequence according to the second feature representation.

[0141] Step 703: Output the third text sequence as the first speech recognition result.

[0142] Step 704: The CTC module determines the first text sequence according to the first feature representation, and sends the first text sequence to the decoder.

[0143] Step 705: The decoder decodes the first feature representation according to the first text sequence to obtain the corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0144] Step 706: Output the second text sequence as the second speech recognition result.

[0145] In this embodiment, the encoder can process according to at least two data block lengths to obtain feature representations corresponding to at least two data block lengths. Suppose the feature representations corresponding to at least two data block lengths specifically include: a first feature representation and a second feature representation. Among them, the first feature representation corresponds to a first data block length, and the second feature representation corresponds to a second data block length. The first data block length can be greater than the second data block length.

[0146] In an embodiment of the present invention, the encoder can obtain a first feature representation according to the first data block length, and obtain a second feature representation according to the second data block length. Among them, the determination processes of the first feature representation and the second feature representation can be parallel or serial, and the present invention embodiment does not limit the sequence of the determination processes of the first feature representation and the second feature representation.

[0147] The second data block length is less than the first data block length, so the real-time performance of the first speech recognition result can be improved.

[0148] The first data block length can be greater than the second data block length, which can increase the amount of data used for decoding by the decoder, so the accuracy of the second speech recognition result can be improved.

[0149] In an embodiment of the present invention, the first text sequence may not be output as the first speech recognition result, but is provided to the decoder as the historical decoding result corresponding to the decoding moment.

[0150] Those skilled in the art can determine the first data block length and the second data block length according to actual application requirements, or determine the difference between the two. For example, the second data block length is 400 milliseconds, or the difference between the two is 200 milliseconds, etc. It can be understood that the present invention embodiment does not limit the specific values of the first data block length and the second data block length.

[0151] In summary, for the speech recognition method according to the embodiment of the present invention, the encoder can process according to at least two data block lengths to obtain feature representations corresponding to at least two data block lengths. Among them, the second data block length is less than the first data block length, and the second data block length is used to determine the first speech recognition result, so the real-time performance of the first speech recognition result can be improved. The first data block length can be greater than the second data block length, and the first data block length is used to determine the second speech recognition result, which can increase the amount of data used for decoding by the decoder, so the accuracy of the second speech recognition result can be improved.

[0152] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of combinations of motion actions. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described order of motion actions, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the involved motion actions are not necessarily essential for the embodiments of the present invention.

[0153] Device embodiments

[0154] Figure 8 FIG. is a schematic structural diagram of a voice recognition device provided by an embodiment of the present invention, and the voice recognition device is generally implemented in a hardware and / or software manner.

[0155] The voice recognition device may specifically include the following modules: a first receiving module 801, a second receiving module 802, and a decoding module 803.

[0156] Among them, the first receiving module 801 is configured to receive a feature representation corresponding to the voice to be recognized from an encoder;

[0157] The second receiving module 802 is configured to receive a first text sequence corresponding to the voice to be recognized from a connectionist temporal classification module;

[0158] The decoding module 803 is configured to decode the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

[0159] Optionally, the decoding module 803 may specifically include:

[0160] A historical decoding result determination module, configured to determine historical decoding results corresponding to multiple decoding moments according to the first text sequence;

[0161] A parallel decoding module, configured to perform parallel decoding of the feature representation at multiple decoding moments according to the historical decoding results corresponding to the multiple decoding moments.

[0162] Optionally, the above device may further include:

[0163] A first output module, configured to output the target text sequence obtained by the connectionist temporal classification module as a first voice recognition result; the above target text sequence includes: the first text sequence, or, a third text sequence;

[0164] A second output module, configured to output the second text sequence as a second voice recognition result; the above second voice recognition result is used to replace the above first voice recognition result.

[0165] Optionally, the above features may correspond to the data blocks included in the speech to be recognized; the above first text sequence may correspond to the data blocks included in the speech to be recognized.

[0166] Optionally, the above features represent corresponding to the data blocks included in the speech to be recognized; the above data blocks may correspond to a data block length;

[0167] The above first text sequence may be obtained according to the first feature representation corresponding to the first data block length; the above third text sequence may be obtained according to the second feature representation corresponding to the second data block length; the above first data block length may be greater than the above second data block length.

[0168] Optionally, the above device may be applied to a speech recognition model; the above speech recognition model may include: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the above encoder;

[0169] Wherein, the above encoder sends a second feature representation to the above connectionist temporal classification module, and sends a first feature representation to the above connectionist temporal classification module and the above decoder;

[0170] The above connectionist temporal classification module determines a third text sequence according to the second feature representation output by the above encoder; the above third text sequence is used as the first speech recognition result for output;

[0171] The above connectionist temporal classification module determines a first text sequence according to the first feature representation output by the above encoder, and sends the above first text sequence to the above decoder.

[0172] Optionally, the above device may be applied to a speech recognition model; the above speech recognition model may include: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the above encoder;

[0173] Wherein, during the training process, the above decoder determines the historical decoding result corresponding to the current decoding moment according to the decoding result of the past moment; during the speech recognition process, the above decoder determines the historical decoding results corresponding to multiple decoding moments according to the above first text sequence.

[0174] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, please refer to the partial description of the method embodiment.

[0175] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the embodiments, please refer to each other.

[0176] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0177] An embodiment of the present invention further provides a device for training, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include operation instructions for performing the operations included in the foregoing method.

[0178] Figure 9 It is a block diagram of a device for speech recognition as a terminal shown according to an exemplary embodiment. For example, the terminal 1100 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0179] Refer to Figure 9 , the terminal 1100 may include one or more of the following components: a processing component 1102, a memory 1104, a power supply component 1106, a multimedia component 1108, an audio component 1110, an input / output (I / O) interface 1112, a sensor component 1114, and a communication component 1116.

[0180] The processing component 1102 generally controls the overall operation of the terminal 1100, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing element 1102 may include one or more processors 1120 to execute download instructions to complete all or part of the steps of the above method. In addition, the processing component 1102 may include one or more modules to facilitate the interaction between the processing component 1102 and other components. For example, the processing component 1102 may include a multimedia module to facilitate the interaction between the multimedia component 1108 and the processing component 1102.

[0181] The memory 1104 is configured to store various types of data to support the operation of the terminal 1100. Examples of such data include download instructions for any application or method operating on the terminal 1100, contact data, phone book data, messages, pictures, videos, etc. The memory 1104 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0182] The power supply component 1106 provides power for various components of the terminal 1100. The power supply component 1106 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the terminal 1100.

[0183] The multimedia component 1108 includes a screen that provides an output interface between the terminal 1100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe motion actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1108 includes a front camera and / or a rear camera. When the terminal 1100 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0184] The audio component 1110 is configured to output and / or input audio signals. For example, the audio component 1110 includes a microphone (MIC) that is configured to receive external audio signals when the terminal 1100 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 1104 or transmitted via the communication component 1116. In some embodiments, the audio component 1110 further includes a speaker for outputting audio signals.

[0185] The I / O interface 1112 provides an interface between the processing component 1102 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0186] The sensor assembly 1114 includes one or more sensors for providing status assessments of various aspects for the terminal 1100. For example, the sensor assembly 1114 can detect the on / off state of the terminal 1100, the relative positioning of components, such as the display and keypad of the terminal 1100. The sensor assembly 1114 can also detect a change in the position of the terminal 1100 or a component of the terminal 1100, the presence or absence of user contact with the terminal 1100, the orientation or acceleration / deceleration of the terminal 1100, and the temperature change of the terminal 1100. The sensor assembly 1114 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1114 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1114 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0187] The communication component 1116 is configured to facilitate communication between the terminal 1100 and other devices in a wired or wireless manner. The terminal 1100 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1116 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1116 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0188] In an exemplary embodiment, the terminal 1100 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0189] In an exemplary embodiment, a non-transitory computer-readable storage medium including download instructions is also provided, such as a memory 1104 including download instructions, and the above download instructions can be executed by the processor 1120 of the terminal 1100 to complete the above methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0190] Figure 10It is a schematic structural diagram of a server in some embodiments of the present invention. The server 1900 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1922 (for example, one or more processors) and a memory 1932, and one or more storage media 1930 (for example, one or more mass storage devices) that store application programs 1942 or data 1944. Among them, the memory 1932 and the storage media 1930 may be transient storage or persistent storage. The program stored in the storage media 1930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1922 may be configured to communicate with the storage media 1930 and execute a series of instruction operations in the storage media 1930 on the server 1900.

[0191] The server 1900 may further include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input / output interfaces 1958, one or more keyboards 1956, and / or one or more operating systems 1941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0192] When the download instruction in the storage medium is executed by a processor of a device (terminal or server), the device is enabled to execute a speech recognition method, and the method includes: receiving a feature representation corresponding to the speech to be recognized from an encoder; receiving a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module; decoding the feature representation according to the first text sequence to obtain a corresponding second text sequence; and using the first text sequence as a historical decoding result corresponding to the decoding moment.

[0193] Those skilled in the art will readily think of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0194] It should be understood that the present invention is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

[0195] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

[0196] The above has introduced in detail a speech recognition method, a speech recognition device, a device for speech recognition, and a machine-readable medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A voice recognition method, characterized in that, The method includes: Receiving a feature representation corresponding to the speech to be recognized from an encoder; Receiving a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module; Decoding the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

2. The method according to claim 1, wherein The decoding of the feature representation includes: Determining historical decoding results corresponding to multiple decoding moments respectively according to the first text sequence; Performing parallel decoding of the feature representation at multiple decoding moments according to the historical decoding results corresponding to the multiple decoding moments.

3. The method according to claim 1, wherein The feature representation corresponds to a data block included in the speech to be recognized; the data block corresponds to a data block length, and the method further includes: Outputting the target text sequence obtained by the connectionist temporal classification module as the first speech recognition result; the target text sequence includes: the first text sequence, or, a third text sequence; the first text sequence is obtained according to a first feature representation corresponding to a first data block length; the third text sequence is obtained according to a second feature representation corresponding to a second data block length; the first data block length is greater than the second data block length; Outputting the second text sequence as the second speech recognition result; the second speech recognition result is used to replace the first speech recognition result.

4. The method according to claim 1, characterized in that The feature representation corresponds to a data block included in the speech to be recognized; the first text sequence corresponds to a data block included in the speech to be recognized.

5. The method according to claim 3, characterized in that, The method is applied to a speech recognition model; The speech recognition model includes: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the encoder; The method further includes: The encoder sends a second feature representation to the connectionist temporal classification module, and sends a first feature representation to the connectionist temporal classification module and the decoder; The connectionist temporal classification module determines a third text sequence according to the second feature representation output by the encoder; the third text sequence is used as the first speech recognition result for output; The connectionist temporal classification module determines a first text sequence according to the first feature representation output by the encoder, and sends the first text sequence to the decoder.

6. The method according to any one of claims 1 to 5, characterized in that, The method is applied to a speech recognition model; The speech recognition model includes: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the encoder; Wherein, during the training process, the decoder determines the historical decoding result corresponding to the current decoding moment according to the decoding results of past moments; during the speech recognition process, the decoder determines the historical decoding results corresponding to multiple decoding moments respectively according to the first text sequence.

7. A voice recognition device, characterized in that, The device includes: A first receiving module, configured to receive a feature representation corresponding to the speech to be recognized from an encoder; A second receiving module, configured to receive a first text sequence corresponding to the speech to be recognized from a connectionist temporal classification module; A decoding module, configured to decode the feature representation according to the first text sequence to obtain a corresponding second text sequence; the first text sequence is used as the historical decoding result corresponding to the decoding moment.

8. The device according to claim 7, wherein The decoding module includes: A historical decoding result determination module, configured to determine historical decoding results corresponding to multiple decoding moments according to the first text sequence; A parallel decoding module, configured to perform parallel decoding of the feature representation at multiple decoding moments according to the historical decoding results corresponding to the multiple decoding moments.

9. The device according to claim 7, characterized in that, The feature representation corresponds to a data block included in the speech to be recognized; the data block corresponds to a data block length, and the apparatus further includes: A first output module, configured to output the target text sequence obtained by the connectionist temporal classification module as the first speech recognition result; the target text sequence includes: a first text sequence, or a third text sequence; the first text sequence is obtained according to a first feature representation corresponding to a first data block length; the third text sequence is obtained according to a second feature representation corresponding to a second data block length; the first data block length is greater than the second data block length; A second output module, configured to output the second text sequence as the second speech recognition result; the second speech recognition result is used to replace the first speech recognition result.

10. The device according to claim 7, characterized in that, The feature representation corresponds to a data block included in the speech to be recognized; the first text sequence corresponds to a data block included in the speech to be recognized.

11. The device according to claim 9, characterized in that, The apparatus is applied to a speech recognition model; the speech recognition model includes: an encoder, and a decoder and a connectionist temporal classification module respectively connected to the encoder; The encoder sends a second feature representation to the connectionist temporal classification module, and sends a first feature representation to the connectionist temporal classification module and the decoder; The connectionist temporal classification module determines a third text sequence according to the second feature representation output by the encoder; the third text sequence is used as the first speech recognition result for output; The connectionist temporal classification module determines a first text sequence according to the first feature representation output by the encoder, and sends the first text sequence to the decoder.

12. A device for speech recognition, characterized in that, It includes a memory, and one or more programs, wherein one or more programs are stored in the memory, and when the programs are executed by one or more processors, the steps of the method according to any one of claims 1 to 6 are implemented.

13. A machine-readable medium, on which a download instruction is stored, and when executed by one or more processors, causes the apparatus to execute one or more of the methods according to claims 1 to 6.

Citation Information

Patent Citations

  • Chinese text overall recognition method in natural scene image

    CN108491836A

  • Speech recognition apparatus and method thereof

    US20090048839A1