Speech recognition method, training method of speech recognition model and related devices

By introducing the first word delay loss in speech recognition model training and optimizing model parameters, the first word delay problem in speech recognition technology is solved, and the user experience of real-time speech recognition is improved.

CN114495914BActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210135438.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-07-11
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

The existing voice recognition technology has the problem of first word delay, which causes users to wait for recognition results in the real-time voice recognition system for too long, reducing the user experience.

Method used

Introduce the first word delay loss during the speech recognition model training process, and optimize the model parameters to reduce the first word decoding time by adding the delay penalty term.

Benefits of technology

It effectively reduces the delay time of the speech recognition model when decoding the first word, and improves the user experience of the real-time speech recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495914B_ABST
    Figure CN114495914B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method, a training method of a speech recognition model, and related devices. The speech recognition method includes: obtaining a speech to be recognized; inputting the speech to be recognized into a trained speech recognition model to obtain an output text; wherein, the total loss used for training the speech recognition model is related to the first-word delay loss. By the above method, the present application can reduce the time of the first-word delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of speech recognition, and particularly relates to a speech recognition method, a training method of a speech recognition model, and related devices. Background Art

[0002] Speech recognition technology is a technology that recognizes the speech signals input by users and finally converts them into text / strings (that is, the recognition result is text), which provides convenience for natural human-computer interaction. Taking a mobile device adopting speech recognition technology as an example, with the support of speech recognition technology, as long as the user speaks to the mobile device, text will be automatically formed after being recognized by the speech recognition system, greatly improving the input efficiency of the user.

[0003] However, there is a problem of first-word delay in current speech recognition technology, that is, the first word of the transcription cannot be given within the expected short time. This defect will cause users to spend more time waiting for the recognition result when using applications with real-time speech recognition systems, greatly reducing the user experience. Summary of the Invention

[0004] This application provides a speech recognition method, a training method of a speech recognition model, and related devices to reduce the time of first-word delay.

[0005] To solve the above technical problems, a technical solution adopted by this application is: to provide a speech recognition method, including: obtaining the speech to be recognized; inputting the speech to be recognized into the trained speech recognition model to obtain an output text; wherein, the total loss used for training the speech recognition model is related to the first-word delay loss.

[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide a training method of a speech recognition model, including: obtaining a plurality of speech training samples, and each of the speech training samples is labeled with a text label; obtaining a plurality of frequency-domain features of the speech training samples in time series; inputting the plurality of frequency-domain features into the speech recognition model to obtain a predicted text; obtaining a first loss between the predicted text and the text label, and obtaining a first-word delay loss of the predicted text; taking the sum of the first loss and the first-word delay loss as the total loss, and adjusting the parameters of the speech recognition model according to the total loss.

[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide a speech recognition device, including: an obtaining module, configured to obtain the speech to be recognized; a recognition module, configured to input the speech to be recognized into the trained speech recognition model to obtain an output text; a training module, configured to train the speech recognition model, and the total loss used for training the speech recognition model is related to the first-word delay loss.

[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech recognition method described in any one of the above embodiments, or the training method of the speech recognition model.

[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide a storage device storing program instructions that can be run by a processor, and the program instructions are used to implement the speech recognition method described in any one of the above embodiments, or the training method of the speech recognition model.

[0010] Different from the prior art, the beneficial effect of this application is that: in the speech recognition method provided by this application, a speech recognition model will be used, and the total loss adopted during the training of the speech recognition model is related to the first-word delay loss; that is, in this case, a delay penalty term is added during model training to reduce the delay time when decoding the first word during model application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0012] Figure 1 It is a schematic flowchart of an embodiment of the speech recognition method of this application;

[0013] Figure 2 It is a schematic flowchart of an embodiment of the training method of the speech recognition model of this application;

[0014] Figure 3 is Figure 2 a schematic flowchart of an embodiment corresponding to step S203 in

[0015] Figure 4 It is a schematic structural diagram of an embodiment of the encoder;

[0016] Figure 5 is Figure 3 a schematic flowchart of an embodiment corresponding to step S302 in

[0017] Figure 6 It is a schematic diagram of an embodiment of the mask operation;

[0018] Figure 7Schematic diagram of an embodiment of the attention of the first character before and after adding the first character delay loss;

[0019] Figure 8 Schematic structural diagram of an embodiment of the voice recognition device of the present application;

[0020] Figure 9 Schematic structural diagram of an embodiment of the electronic device of the present application;

[0021] Figure 10 Schematic structural diagram of an embodiment of the storage device of the present application. Specific embodiments

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of an embodiment of the voice recognition method of the present application. The voice recognition method includes:

[0024] S101: Obtain the voice to be recognized.

[0025] S102: Input the voice to be recognized into the trained voice recognition model to obtain the output text; wherein, the total loss used in training the voice recognition model is related to the first character delay loss.

[0026] Specifically, the specific implementation process of the above step S102 can be: A. Extract features from the voice to be recognized to obtain a plurality of frequency domain features to be recognized (for example, Fbank features); for example, the number of frequency domain features to be recognized after feature extraction and conversion of a voice to be recognized with a duration of one second may be 100 frames. B. Input the plurality of frequency domain features to be recognized into the trained voice recognition model to obtain the corresponding output text.

[0027] The total loss used in training the voice recognition model used in the above application process is related to the first character delay loss; that is, in this case, a delay penalty term is added during model training to reduce the delay time when decoding the first character during model application.

[0028] In one embodiment, please refer to Figure 2 , Figure 2 which is a schematic flow chart of an embodiment of the training method of the voice recognition model of the present application. The training process specifically includes:

[0029] S201: Obtain a plurality of speech training samples, and each speech training sample is labeled with a text label.

[0030] Specifically, the speech training samples can be collected through networks or other means, and the text labels can be manually labeled. And the above-mentioned plurality of speech training samples can be in the same language, and the parameters of the speech recognition models corresponding to different languages are different.

[0031] S202: Obtain a plurality of frequency domain features of the speech training samples in time series.

[0032] Specifically, the above-mentioned frequency domain features can be fbank features, etc.; the specific implementation process of the above step S202 can be: successively perform pre-emphasis, framing, windowing, short-time Fourier transform (STFT), Mel filtering, mean removal, etc. on the speech training samples to obtain a plurality of frequency domain features.

[0033] S203: Input the plurality of frequency domain features into the speech recognition model to obtain a predicted text.

[0034] Specifically, please refer to Figure 3 , Figure 3 is Figure 2 a schematic flowchart of an implementation manner corresponding to step S203 in

[0035] S301: Encode the plurality of frequency domain features to obtain a plurality of first encoded features.

[0036] Specifically, the speech recognition model provided in this application includes an encoder. The principle of the encoder is to use a deep neural network such as a long short-term memory (LSTM), an encoding layer Transformer, etc. as the encoder to encode the frequency domain features of speech data with a time length of T to obtain its hidden space features; that is, the encoder is specifically used to implement the above step S301.

[0037] Optionally, please refer to Figure 4 , Figure 4 is a schematic structural diagram of an implementation manner of the encoder. The encoder 10 includes a plurality of branches 100. Each branch 100 includes a plurality of unidirectional long short-term memory (LSTM) modules (not labeled) and a bidirectional long short-term memory (LSTM) module (not labeled) connected in sequence. A frequency domain feature x i is input into one branch 100. In this embodiment, the number of unidirectional long short-term memory (LSTM) modules provided in each branch 100 can be two, three, four, etc., and the number of unidirectional long short-term memory (LSTM) modules provided in all branches 100 is the same.

[0038] Among them, for any two adjacent one-way long short-term memory (LSTM) modules located in the same layer, the hidden layer output h and the memory layer output c in the previous one-way LSTM module are input into the next one-way LSTM module. For example, as Figure 4 shown in the two one-way LSTM modules labeled 102 and 104 in the same layer, the hidden layer output h of the one-way LSTM module represented by 102 00 and the memory layer output c 00 will be used as the input of the one-way LSTM module represented by 104. Optionally, in this embodiment, taking any one-way LSTM module in the bottom layer as an example, its calculation process can be:

[0039] First, calculate h 0(i-1) , X i and process it using the sigmoid activation function, which is expressed by the formula as follows:

[0040] f i =σ(W f ·[X i , h 0(i-1) +b f );

[0041] Among them, h 0(i-1) is the hidden layer output of the (i - 1)-th one-way LSTM module in the bottom layer, X i is the frequency domain feature; W f , b f represent the parameters of the neural network, and σ represents the sigmoid activation function. The sigmoid function makes the output value between 0 and 1, so it is used to screen the memory retained in the (t - 1)-th calculation, and realizes the forgetting of information through vector multiplication, discarding the redundant memory retained in the previous calculation.

[0042] After calculating the forgotten information, the one-way LSTM module needs to extract the new memory to be retained from the current moment i, as shown in the following formula:

[0043] f info =σ(W info ·[X i , h 0(i-1) +b info );

[0044]

[0045] Among them, tanh represents the hyperbolic tangent activation function.

[0046] The unidirectional long short-term memory LSTM module updates the output C0(i-1) of the memory layer at this time, realizing the forgetting of old memories and the extraction of new memories, as shown in the following formula:

[0047]

[0048] Subsequently, the unidirectional long short-term memory LSTM module combines the latest memory to output a new hidden state h i0 information, as shown in the following formula:

[0049] h 0i =σ(W o ·[X i ,h 0(i-1) +b o )*tanh(C 0i );

[0050] Finally, the i-th unidirectional long short-term memory LSTM module in the bottom layer sends the latest memory layer output C 0i and the hidden state layer output h 0i to the next unidirectional long short-term memory LSTM module in the same layer, realizing the model's ability to model long-time series data.

[0051] In addition, for any two adjacent bidirectional long short-term memory LSTM modules in the same layer, the previous bidirectional long short-term memory LSTM module receives the hidden layer output h ni and the memory layer output c ni of the next bidirectional long short-term memory LSTM module, and the hidden layer output h n(i-1) and the memory layer output c n(i-1) of the previous bidirectional long short-term memory LSTM module will be input into the next bidirectional long short-term memory LSTM module. For example, for the two bidirectional long short-term memory LSTM modules labeled 106 and 108 in the same layer as shown in Figure 4 , the hidden layer output h n0 and the memory layer output c n0 of the bidirectional long short-term memory LSTM module represented by 106 will be used as the input of the bidirectional long short-term memory LSTM module represented by 108; and the hidden layer output h n1 and the memory layer output c n1 of the bidirectional long short-term memory LSTM module represented by 108 will also be used as the input of the bidirectional long short-term memory LSTM module represented by 106.

[0052] Speech data is usually long in the time dimension. After one second of audio is converted into frequency-domain features through feature extraction, it usually has 100 frames. To better model long sequences, the encoder 10 of the speech recognition model in this solution uses a multi-layer unidirectional long short-term memory (LSTM) module. This network model has the ability of long-term modeling, is not easy to forget long sequence information, and can be used in real-time scenarios. In addition, in this case, a bi-directional long short-term memory (BiLSTM) module is added after the multi-layer unidirectional long short-term memory (LSTM) module, that is, the BiLSTM calculates from the head and tail of the input as starting points respectively and combines the information for input. However, to ensure the real-time performance of speech recognition, this solution uses a truncated BiLSTM, that is, a window length is set, and there is a BiLSTM module corresponding to one branch of 100. When training the speech recognition model, the output of the previous layer of unidirectional LSTM module is divided into multiple subsequences, and BiLSTM calculations are performed on the subsequences respectively. This not only ensures that the BiLSTM does not need to obtain global speech features, but also ensures rich context information modeling.

[0053] Please refer to again Figure 4 , the encoder 10 provided in this application further includes a downsampling layer 101. After the output of the BiLSTM module enters the downsampling layer 101, the complexity of the encoded features can be reduced, which better serves the later decoding operation. At this time, the step S301 of encoding multiple frequency-domain features to obtain multiple first encoded features includes: encoding multiple frequency-domain features to obtain multiple second encoded features; where the number of multiple frequency-domain features is the same as the number of multiple second encoded features; performing downsampling processing on multiple second encoded features to obtain multiple first encoded features; where the number of multiple second encoded features is greater than the number of multiple first encoded features. For example, as Figure 4 shown, the downsampling layer 101 performs a 4-fold downsampling operation. The frequency-domain features with a time sequence length of T, after being calculated by a multi-layer unidirectional LSTM model, are sent into a BiLSTM model, and then a 4-fold downsampling is performed to obtain features with a time sequence length of T / 4.

[0054] S302: Obtain multiple context vectors based on multiple first encoded features.

[0055] S303: Decode multiple context vectors to obtain multiple words, where the multiple words form a predicted text.

[0056] Specifically, the speech recognition model provided by the present application further includes an attention layer connected to the output of the encoder and a decoder connected to the attention layer. Among them, the attention layer is mainly used to obtain multiple context vectors based on multiple first encoding features (i.e., to implement step S302). Currently, the main representatives of the commonly used Attention mechanism include soft-attention, monotonic attention, MoChA (monotonic chunkwise attention), etc. This solution improves the existing MoChA. The decoder, through the attention mechanism in the attention layer, pays attention to the information of the entire encoding feature during each decoding, finds the most important feature for the current decoding stage from the encoding features through the attention mechanism, and then classifies the extracted features. That is, the decoder layer is mainly used to decode multiple context vectors to obtain multiple words, and the multiple words form a predicted text (i.e., to implement step S303). Optionally, in the present application, the decoder includes a single-layer unidirectional long short-term memory LSTM structure.

[0057] In one embodiment, please refer to Figure 5 , Figure 5 which is Figure 3 a schematic flowchart of an implementation manner corresponding to step S302 in

[0058] S401: For each word, obtain the attention weight of the current word on each first encoding feature.

[0059] Specifically, taking Figure 4 as an example, through the encoder 10, multiple first encoding features can be obtained, which are h0, h1,..., h T / 4 . The specific implementation process of the above step S401 can be:

[0060] A. Calculate the monotonic energy value e i,j , and the specific calculation formula is as follows:

[0061] e i,j =monotonicEnergy(s i-1 ,h j )(1 ≤ j ≤ 4 / T);

[0062] Among them, e i,j represents the energy value that the i-th word pays attention to on the j-th first encoding feature h j , monotonicEnergy is the MoChA energy calculation function, and s i-1 represents the state value generated by the (i - 1)-th calculation of the unidirectional LSTM module in the decoder.

[0063] B. Calculate the selection probability p i,j , and the specific calculation formula is as follows:

[0064] p i,j = σ(e i,j );

[0065] where σ represents the sigmoid activation function.

[0066] C. Calculate the attention weight a j of the i-th word on the j-th first encoding feature h i,j , and the specific calculation formula is as follows:

[0067] a i,j = p i,j ((1 - p i,j-1 ) * a i,j-1 / p i,j-1 + a i-1,j ).

[0068] S402: Determine the sliding window of the current word according to the attention weights of the current word on each first encoding feature.

[0069] Specifically, the sliding window of the current word includes a start position and an end position. The start position corresponding to the current word is the same as the end position of the sliding window corresponding to the previous word. The sum of the second values of all attention weights at the start position, end position, and intermediate positions between the start position and the end position corresponding to the current word is greater than the threshold. The sum of the third values of all attention weights at the start position and intermediate positions between the start position and the end position corresponding to the current word is less than or equal to the threshold. Optionally, the threshold can be greater than or equal to 0.5 and less than or equal to 0.98; for example, the threshold can be 0.6, 0.7, 0.8, etc. The specific process of determining the sliding window of the current word can be: starting from the end position of the sliding window corresponding to the previous word, accumulate the attention weights corresponding to this position and the positions after it. If the sum of all attention weights from pos i-1 (the end position of the sliding window of word i - 1) to the j position is greater than the threshold, then stop sliding the window. It can be seen that the lengths of the sliding windows corresponding to different words may be different.

[0070] S404: Obtain the energy value based on the sliding window of the current word.

[0071] Specifically, the energy value β i,j can be obtained through the following formula:

[0072] μ i,j = ChunkEnergy(s i-1 , h j );

[0073]

[0074]

[0075] S405: Obtain the context vector corresponding to the current word according to the energy value.

[0076] Specifically, the context vector c i can be obtained through the following formula:

[0077]

[0078] S204: Obtain the first loss of the predicted text and the text label, and obtain the first-word delay loss of the predicted text.

[0079] Specifically, there are consecutive multiplication operations in the calculation process of MoChA, which makes the obtained attention weight a i,j relatively low. Sometimes it is difficult to even reach the threshold at the end of the sliding window proposed in MoChA, resulting in a large delay. And in the recognition of the first word, this situation is more prominent. To solve this problem, the present application proposes to increase the attention degree of MoChA during model training, so as to solve the MoChA delay problem and enhance the user experience of the real-time speech recognition system.

[0080] To reduce the degree of the first-word delay, a delay penalty term is added during the training of the speech recognition model. Specifically, the steps of obtaining the first-word delay loss of the predicted text in the above step S204 include: obtaining the attention weight of the first word on each first coding feature; setting all attention weights located after the expected delay size in time series to 0; obtaining the first sum value of the attention weights of the first word on each first coding feature; taking the absolute value of the first difference between the first sum value and a preset value as the first-word delay loss. Specifically, it is expressed by the formula as follows:

[0081] loss delay = |t - sum(mask(a0, b))|;

[0082] where, t is a preset value, and optionally, t takes the value of 1; loss delay represents the first-word delay loss; a0 is the attention weight of the first word calculated by mocha for each first coding feature, mask(a0, b) is the expected attention degree vector, b is the expected delay size, the mask operation sets the values after the position b in a0 to 0, and then sums the expected attention weights and hopes that the sum approaches 1.

[0083] For example, please refer to Figure 6 , Figure 6Schematic diagram of an implementation of the mask operation. The obtained attention weights related to the first character are: 0.04, 0.05, 0.05, 0.04, 0.15, 0.25, 0.13, 0.15, 0.05, 0.04, 0.02, 0.02, and 0.01; and the above attention weights are arranged in sequence from front to back according to the time sequence. When the expected delay size b is 6, the attention weights corresponding to the first 6 frames can be retained, and the attention weights corresponding to the subsequent frames are set to 0. At this time, loss delay =|1-(0.04 + 0.05 + 0.05 + 0.04 + 0.15 + 0.25)| = 0.42. The above design method can make the model parameters focus on the attention before b during optimization and make its value larger. This mask scheme will not cause a decrease in recognition accuracy for the acoustic model of real-time speech recognition because the model usually does not pay attention to the information in the speech after the expected delay size b when decoding the first character.

[0084] Of course, in other embodiments, the first character delay loss can also be obtained in other ways. For example, the length of the sliding window corresponding to the first character can be obtained, and the first character delay loss is obtained based on the length of the sliding window; among them, the larger the length of the sliding window, the greater the first character delay loss.

[0085] In addition, the process of obtaining the first loss between the predicted text and the text label in step S204 above can be: obtaining the cross-entropy loss between the predicted text and the text label.

[0086] S205: Obtain the total loss based on the first loss and the first character delay loss, and adjust the parameters of the speech recognition model according to the total loss.

[0087] Specifically, the first product of the first character delay loss and the adjustment parameter can be obtained, and the sum of the first product and the first loss is used as the total loss, which is expressed by the formula as follows:

[0088] loss = loss ASR + λ·loss delay ;

[0089] where loss represents the total loss, loss ASR represents the first loss, loss delay represents the first character delay loss, λ represents the adjustment parameter, and λ can be greater than 0 and less than or equal to 1.

[0090] In an application scenario, as Figure 7 shown, Figure 7 is a schematic diagram of an implementation of the attention of the first character before and after adding the first character delay loss. Optimizing the model through the above loss function makes the attention vector a output by the speech recognition model have a larger value within the expected delay size b, asFigure 7 As shown, it is easier to meet the conditions for the sliding window to stop. Figure 7 In the left figure in [reference], the attention weight of the first word calculated by MoChA without considering the loss of the first word delay is shown. Within the expected delay size b, the attention weight is small and it is difficult to meet the conditions for the sliding window to stop. Figure 7 In the right figure in [reference], by adding the loss of the first word delay, the speech recognition model can output a larger attention weight within the expected delay size b, and timely meet the stop conditions of the sliding window.

[0091] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the speech recognition device of the present application. The speech recognition device specifically includes an acquisition module 20, a recognition module 22, and a training module 24.

[0092] Among them, the acquisition module 20 is used to acquire the speech to be recognized; the recognition module 22 is connected to the acquisition module 20 and is used to input the speech to be recognized into the trained speech recognition model to obtain an output text. The training module 24 is connected to the recognition module 22 and is used to train the speech recognition model, and the total loss used for training the speech recognition model is related to the loss of the first word delay.

[0093] In one embodiment, the training module 24 includes a first sub-acquisition module, a second sub-acquisition module, a prediction sub-module, a third acquisition sub-module, and an adjustment sub-module. Among them, the first sub-acquisition module is used to acquire a plurality of speech training samples, and each speech training sample is labeled with a text label. The second sub-acquisition module is connected to the first sub-acquisition module and is used to acquire a plurality of frequency domain features of the speech training samples in time series. The prediction sub-module is connected to the second sub-acquisition module and is used to input the plurality of frequency domain features into the speech recognition model to obtain a predicted text. The third acquisition sub-module is connected to the prediction sub-module and is used to obtain a first loss between the predicted text and the text label, and to obtain the loss of the first word delay of the predicted text. The adjustment sub-module is connected to the third acquisition sub-module and is used to obtain the total loss based on the first loss and the loss of the first word delay, and to adjust the parameters of the speech recognition model according to the total loss.

[0094] Among them, the above-mentioned prediction sub-module is specifically used to encode a plurality of frequency domain features to obtain a plurality of first encoded features; obtain a plurality of context vectors based on the plurality of first encoded features; decode the plurality of context vectors to obtain a plurality of words, where the plurality of words constitute the predicted text.

[0095] Optionally, the step of obtaining a plurality of context vectors based on a plurality of first encoded features includes: for each word, obtaining an attention weight of the current word on each first encoded feature; determining a sliding window of the current word according to the attention weight of the current word on each first encoded feature; wherein, the sliding window of the current word includes a start position and an end position, the start position corresponding to the current word is the same as the end position corresponding to the previous word, the second sum value of all attention weights of the start position, the end position, and the intermediate positions between the start position and the end position corresponding to the current word is greater than a threshold value, and the third sum value of all attention weights of the start position corresponding to the current word and the intermediate positions between the start position and the end position is less than or equal to the threshold value; obtaining an energy value based on the sliding window of the current word, and obtaining a context vector corresponding to the current word according to the energy value.

[0096] Another optionally, the step of encoding a plurality of frequency domain features to obtain a plurality of first encoded features includes: encoding a plurality of frequency domain features to obtain a plurality of second encoded features; wherein, the number of the plurality of frequency domain features is the same as the number of the plurality of second encoded features; performing downsampling processing on the plurality of second encoded features to obtain a plurality of first encoded features; wherein, the number of the plurality of second encoded features is greater than the number of the plurality of first encoded features.

[0097] Further, the step of obtaining the first word delay loss of the predicted text in the third obtaining sub-module includes: obtaining the attention weight of the first word on each first encoded feature; setting all attention weights located after the expected delay size in time series to 0; obtaining the first sum value of the attention weights of the first word on each first encoded feature; taking the absolute value of the first difference between the first sum value and a preset value as the first word delay loss. Optionally, the preset value can be 1.

[0098] Optionally, in this embodiment, the speech recognition model includes an encoder, an attention layer, and a decoder connected in sequence. The encoder is used to encode multiple frequency-domain features to obtain multiple first encoded features; the attention layer is used to obtain multiple context vectors based on the multiple first encoded features; the decoder layer is used to decode the multiple context vectors to obtain multiple words, and the multiple words constitute a predicted text. Among them, the encoder includes multiple branches, each branch includes multiple single-direction long short-term memory (LSTM) modules connected in sequence and a bidirectional long short-term memory (LSTM) module, and one frequency-domain feature is input into one branch; and for any two adjacent single-direction long short-term memory (LSTM) modules in the same layer, the hidden layer output and the memory layer output in the previous single-direction long short-term memory (LSTM) module are input into the next single-direction long short-term memory (LSTM) module; for any two adjacent bidirectional long short-term memory (LSTM) modules in the same layer, the previous bidirectional long short-term memory (LSTM) module receives the hidden layer output and the memory layer output of the next bidirectional long short-term memory (LSTM) module, and the hidden layer output and the memory layer output of the previous bidirectional long short-term memory (LSTM) module are input into the next bidirectional long short-term memory (LSTM) module.

[0099] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device includes: a memory 32 and a processor 30 that are coupled to each other. The memory 32 stores program instructions, and the processor 30 is used to execute the program instructions to implement any of the above speech recognition methods or the training method of the speech recognition model. Specifically, the electronic device includes but is not limited to: desktop computers, laptop computers, tablet computers, servers, etc., which are not limited here. In addition, the processor 30 can also be called a CPU (Center Processing Unit, central processing unit). The processor 30 may be an integrated circuit chip with signal processing capabilities. The processor 30 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 30 may be implemented jointly by integrated circuit chips.

[0100] Please refer to Figure 10 , Figure 10FIG. 0 is a schematic structural diagram of an embodiment of the storage device of the present application. The storage device 40 stores program instructions 400 that can be run by a processor. The program instructions 400 are used to implement any of the above voice recognition methods or the training method of the voice recognition model.

[0101] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0102] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0103] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0104] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks, and other media that can store program codes.

[0105] The above are only embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.

Claims

1. A speech recognition method, characterized in that, Including: Obtaining the speech to be recognized; Inputting the speech to be recognized into the trained speech recognition model to obtain an output text; wherein, the total loss used for training the speech recognition model is related to the first-character delay loss; Wherein, the process of training the speech recognition model includes: obtaining a plurality of speech training samples, and each of the speech training samples is labeled with a text label; obtaining a plurality of frequency-domain features of the speech training samples in time sequence; inputting the plurality of frequency-domain features into the speech recognition model to obtain a predicted text; obtaining a first loss between the predicted text and the text label, and obtaining a first-character delay loss of the predicted text; obtaining the total loss based on the first loss and the first-character delay loss, and adjusting the parameters of the speech recognition model according to the total loss; Wherein, the step of inputting the plurality of frequency-domain features into the speech recognition model to obtain a predicted text includes: encoding the plurality of frequency-domain features to obtain a plurality of first encoding features; obtaining a plurality of context vectors based on the plurality of first encoding features; decoding the plurality of context vectors to obtain a plurality of words, wherein the plurality of words constitute the predicted text; The step of obtaining the first-character delay loss of the predicted text includes: obtaining the attention weights of the first character on each of the first encoding features; setting all the attention weights after the expected delay size in time sequence to 0; obtaining a first sum value of the attention weights of the first character on each of the first encoding features; taking the absolute value of the first difference between the first sum value and a preset value as the first-character delay loss.

2. The speech recognition method according to claim 1, wherein The preset value includes 1.

3. The voice recognition method according to claim 1, wherein The step of obtaining a plurality of context vectors based on the plurality of first encoding features includes: For each word, obtaining the attention weights of the current word on each of the first encoding features; Determining its sliding window according to the attention weights of the current word on each of the first encoding features; wherein, the sliding window of the current word includes a start position and an end position, the start position corresponding to the current word is the same as the end position corresponding to the previous word, the second sum value of all the attention weights at the start position, the end position, and the intermediate positions between the start position and the end position corresponding to the current word is greater than a threshold, and the third sum value of all the attention weights at the start position corresponding to the current word and the intermediate positions between the start position and the end position is less than or equal to the threshold; Obtaining an energy value based on the sliding window of the current word, and obtaining a context vector corresponding to the current word according to the energy value.

4. The voice recognition method according to claim 1, wherein The step of encoding the plurality of frequency-domain features to obtain a plurality of first encoding features includes: Encoding the plurality of frequency-domain features to obtain a plurality of second encoding features; wherein, the number of the plurality of frequency-domain features is the same as the number of the plurality of second encoding features; Performing downsampling processing on the plurality of second encoding features to obtain a plurality of first encoding features; wherein, the number of the plurality of second encoding features is greater than the number of the plurality of first encoding features.

5. The speech recognition method according to claim 1, wherein the speech recognition model includes an encoder, an attention layer, and a decoder connected in sequence. The encoder is configured to encode the multiple frequency domain features to obtain multiple first encoded features; the attention layer is configured to obtain multiple context vectors based on the multiple first encoded features; the decoder layer is configured to decode the multiple context vectors to obtain multiple words, and the multiple words constitute a predicted text; wherein, the encoder includes multiple branches, and each branch includes multiple single-direction long short-term memory (LSTM) modules and a bi-directional long short-term memory (LSTM) module connected in sequence. One frequency domain feature is input into one branch; For any two adjacent single-direction long short-term memory (LSTM) modules in the same layer, the hidden layer output and the memory layer output in the previous single-direction long short-term memory (LSTM) module are input into the next single-direction long short-term memory (LSTM) module; For any two adjacent bi-directional long short-term memory (LSTM) modules in the same layer, the previous bi-directional long short-term memory (LSTM) module receives the hidden layer output and the memory layer output of the next bi-directional long short-term memory (LSTM) module, and the hidden layer output and the memory layer output of the previous bi-directional long short-term memory (LSTM) module are input into the next bi-directional long short-term memory (LSTM) module.

6. A training method for a speech recognition model, characterized in that, including: obtaining multiple speech training samples, and each speech training sample is labeled with a text label; obtaining multiple frequency domain features of the speech training samples in time series; inputting the multiple frequency domain features into the speech recognition model to obtain a predicted text; obtaining a first loss between the predicted text and the text label, and obtaining a first word delay loss of the predicted text; using the sum of the first loss and the first word delay loss as the total loss, and adjusting the parameters of the speech recognition model according to the total loss; wherein, the step of inputting the multiple frequency domain features into the speech recognition model to obtain a predicted text includes: encoding the multiple frequency domain features to obtain multiple first encoded features; obtaining multiple context vectors based on the multiple first encoded features; decoding the multiple context vectors to obtain multiple words, wherein the multiple words constitute a predicted text; the step of obtaining the first word delay loss of the predicted text includes: obtaining the attention weights of the first word on each of the first encoded features; setting all attention weights after the expected delay size in time series to 0; obtaining a first sum value of the attention weights of the first word on each of the first encoded features; using the absolute value of the first difference between the first sum value and a preset value as the first word delay loss.

7. A voice recognition device, characterized in that, including: an obtaining module, configured to obtain the speech to be recognized; a recognition module, configured to input the speech to be recognized into the trained speech recognition model to obtain an output text; a training module, configured to train the speech recognition model, and the total loss used for training the speech recognition model is related to the first word delay loss; Among them, the training module includes a first sub-acquisition module, a second sub-acquisition module, a prediction sub-module, a third acquisition sub-module, and an adjustment sub-module. The first sub-acquisition module is used to acquire a plurality of speech training samples, and each of the speech training samples is labeled with a text label; the second sub-acquisition module is connected to the first sub-acquisition module and is used to acquire a plurality of frequency-domain features of the speech training samples in time series; the prediction sub-module is connected to the second sub-acquisition module and is used to input the plurality of frequency-domain features into the speech recognition model to obtain a predicted text; the third acquisition sub-module is connected to the prediction sub-module and is used to obtain a first loss between the predicted text and the text label, and to obtain a first-word delay loss of the predicted text. The adjustment sub-module is connected to the third acquisition sub-module and is used to obtain the total loss based on the first loss and the first-word delay loss, and to adjust the parameters of the speech recognition model according to the total loss. Among them, the prediction sub-module is further used to encode the plurality of frequency-domain features to obtain a plurality of first encoded features; obtain the plurality of context vectors based on the plurality of first encoded features; and decode the plurality of context vectors to obtain a plurality of words, where the plurality of words constitute the predicted text. The step of obtaining the first-word delay loss of the predicted text in the third acquisition sub-module includes: obtaining the attention weights of the first word on each of the first encoded features; setting all the attention weights after the expected delay size in time series to 0; obtaining the first sum value of the attention weights of the first word on each of the first encoded features; and taking the absolute value of the first difference between the first sum value and a preset value as the first-word delay loss.

8. An electronic device, characterized in that, It includes a memory and a processor that are mutually coupled. Program instructions are stored in the memory, and the processor is used to execute the program instructions to implement the speech recognition method according to any one of claims 1 to 5, or the training method of the speech recognition model according to claim 6.

9. A storage device, characterized in that, Program instructions that can be run by a processor are stored, and the program instructions are used to implement the speech recognition method according to any one of claims 1 to 5, or the training method of the speech recognition model according to claim 6.

Citation Information

Patent Citations

  • Speech recognition method and device, computer equipment and storage medium

    CN113539273A

  • Speech recognition model training method and device, equipment and medium

    CN113870845A