Speech recognition model training method, speech recognition method and related equipment

By fusing hot word tags and acoustic features of reference words to generate a fused feature vector, the problem of inaccurate hot word recognition in end-to-end speech recognition models is solved, thus improving the accuracy of speech recognition.

CN120932631APending Publication Date: 2025-11-11MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410569237.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing end-to-end speech recognition models suffer from inaccurate recognition due to a lack of training data when identifying hot words in specific scenarios.

Method used

By identifying hot word tags for reference words and fusing word features with acoustic features to generate a fused feature vector, a speech recognition model is used for hot word prediction and speech recognition. The model is then trained to improve its ability to capture hot words.

Benefits of technology

It improves the ability of speech recognition models to capture hot words in unknown scenarios, thereby enhancing the accuracy of speech recognition, especially in business scenarios where hot word recognition is insufficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932631A_ABST
    Figure CN120932631A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a speech recognition model, a speech recognition method and related equipment, and aims to accurately recognize hot words in various scenes so as to improve the speech recognition accuracy. The speech recognition model training method comprises the steps of determining a reference word of first speech data and a hot word tag of the reference word; performing fusion processing on the word features of the reference word and the acoustic features of the first voice data to obtain a fusion feature vector; through a voice recognition model, performing hot word prediction based on the fusion feature vector to obtain a hot word prediction result of the first voice data, and performing voice recognition based on the fusion feature vector to obtain a prediction text of the first voice data; and training a speech recognition model based on the hot word prediction result, the hot word tag and the prediction text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a training method for a speech recognition model, a speech recognition method, and related equipment. Background Technology

[0002] In existing end-to-end speech recognition, deeper networks often exhibit stronger generalization capabilities. However, end-to-end speech recognition targets are typically character-based, unlike hybrid models which use phonemes as modeling units. Therefore, it relies more heavily on training data, and the semantic information it carries is more readily absorbed by the training data. When the application scenario requires the recognition of certain specific words, such as trending words in certain business scenarios, the limited training data for these words prevents the trained speech recognition model from accurately identifying them, resulting in inaccurate speech recognition results. Summary of the Invention

[0003] The purpose of this application is to provide a training method for a speech recognition model, a speech recognition method, and related equipment, which can accurately identify hot words in various scenarios, thereby improving the accuracy of speech recognition.

[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0005] In a first aspect, embodiments of this application provide a method for training a speech recognition model, comprising:

[0006] Determine the reference words for the first speech data and the hot word tags for the reference words;

[0007] The word features of the reference words and the acoustic features of the first speech data are fused to obtain a fused feature vector;

[0008] Using a speech recognition model, hot word prediction results of the first speech data are obtained by performing hot word prediction based on the fused feature vector, and the predicted text of the first speech data is obtained by performing speech recognition based on the fused feature vector.

[0009] The speech recognition model is trained based on the hot word prediction results, the hot word tags, and the predicted text.

[0010] Secondly, embodiments of this application provide a speech recognition method, including:

[0011] Acquire target speech data and target hot words;

[0012] The word features of the target hot words and the acoustic features of the target speech data are fused to obtain a target fusion feature vector;

[0013] The target fusion feature vector is used to perform speech recognition by a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by the training method of the speech recognition model described in the first aspect.

[0014] Thirdly, embodiments of this application provide a training apparatus for a speech recognition model, comprising:

[0015] A determining unit is used to determine reference words in the first speech data and hot word tags for the reference words;

[0016] The fusion unit is used to fuse the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector;

[0017] The prediction unit is used to obtain the hot word prediction result of the first speech data by performing hot word prediction based on the fused feature vector through a speech recognition model, and to obtain the predicted text of the first speech data by performing speech recognition based on the fused feature vector.

[0018] The training unit is used to train the speech recognition model based on the hot word prediction results, the hot word tags, and the predicted text.

[0019] Fourthly, embodiments of this application provide a voice recognition device, including:

[0020] The acquisition unit is used to acquire target speech data and target hot words;

[0021] The fusion unit is used to fuse the word features of the target hot words and the acoustic features of the target speech data to obtain a target fusion feature vector;

[0022] The recognition unit is used to perform speech recognition on the target fusion feature vector through a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by the training method of the speech recognition model described in the first aspect.

[0023] Fifthly, embodiments of this application provide an electronic device, including:

[0024] processor;

[0025] Memory used to store the processor's executable instructions;

[0026] The processor is configured to execute the instructions to implement the training method for the speech recognition model as described in the first aspect; or, the processor is configured to execute the instructions to implement the speech recognition method as described in the second aspect.

[0027] In a sixth aspect, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a training method for a speech recognition model as described in the first aspect; or, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a speech recognition method as described in the second aspect.

[0028] In a seventh aspect, embodiments of this application provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the training method of the speech recognition model as described in the first aspect or the speech recognition method as described in the second aspect.

[0029] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:

[0030] The first speech data is used to identify reference words and hot word tags for those reference words. This guides the speech recognition model to capture similar hot word information as much as possible, resulting in better recognition performance even when the training data does not contain the hot words to be recognized in the application scenario. This improves the speech recognition model's ability to capture hot words in unknown scenarios. Next, the word features of the reference words and the acoustic features of the first speech data are fused together, effectively integrating the hot word information and acoustic features into a fused feature vector. Based on this, the speech recognition model performs hot word prediction on the fused feature vector to obtain the hot word prediction results for the first speech data. Speech recognition is then performed on the fused features to obtain the predicted text for the first speech data. The speech recognition model is then trained using the hot word prediction results, hot word tags, and predicted text. This not only helps improve the speech recognition model's ability to focus on hot word information in the first speech data, enabling it to learn how to combine acoustic features to capture hot word information in the given scenario during training, but also allows the model to apply the captured hot word information to speech recognition, thus possessing the ability to accurately capture hot words from speech data and improving speech recognition accuracy. Attached Figure Description

[0031] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0032] Figure 1 A flowchart illustrating a training method for a speech recognition model provided in one embodiment of this application;

[0033] Figure 2A schematic diagram illustrating an attention-based fusion processing procedure as provided in one embodiment of this application;

[0034] Figure 3 A flowchart illustrating a method for training a speech recognition model, provided as another embodiment of this application;

[0035] Figure 4 A flowchart illustrating a speech recognition method provided in one embodiment of this application;

[0036] Figure 5 A flowchart illustrating a speech recognition method provided for another embodiment of this application;

[0037] Figure 6 A schematic diagram of the structure of a speech recognition model training device provided in one embodiment of this application;

[0038] Figure 7 A schematic diagram of the structure of a speech recognition device provided in one embodiment of this application;

[0039] Figure 8 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] The terms "first," "second," etc., used in this specification and claims are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in this specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0042] Explanation of some concepts:

[0043] End-to-end speech recognition: The purpose of speech recognition is to convert the lexical content of human speech into text content. End-to-end speech recognition uses a pure neural network approach instead of the traditional method of training a mixture of alignment models, acoustic models, and language models separately.

[0044] Transformer: A temporal model based on a self-attention mechanism. Its encoder effectively encodes temporal information, demonstrating significantly better processing capabilities than LSTM (Long Short-Term Memory) networks, while also being faster. Transformers are widely used in natural language processing, computer vision, machine translation, speech recognition, and other fields.

[0045] Conformer: A model that combines transformer and CNN (Convolutional Neural Networks). Transformer excels at capturing content-based global interactions, while CNN effectively utilizes local features, making conformer a good modeler for both long-term global interaction information and local features.

[0046] CTC (Connectionist Temporal Classification) is a loss function used in time series labeling problems. Traditional sequence labeling algorithms require perfect alignment of input and output symbols at every time step, while CTC expands the label set by adding empty elements. After labeling the sequence using the expanded label set, all predicted sequences that can be converted into the true sequence through a mapping function are correct predictions. In other words, predicted sequences can be obtained without data alignment processing.

[0047] Hot words: These are words that frequently appear in the required speech recognition scenarios, but appear less frequently in the training data and often appear in specific scenarios.

[0048] As mentioned earlier, end-to-end speech recognition models rely heavily on training data. When the business scenario in which the model is applied involves hot words that need to be recognized, the training data containing such words is limited, resulting in the trained speech recognition model being unable to accurately identify these words and leading to inaccurate speech recognition results.

[0049] In view of this, this application proposes a training method for a speech recognition model. The method determines reference words and hot word tags for the first speech data to guide the speech recognition model to capture similar hot word information as much as possible. Even when the training data does not contain the hot words to be recognized in the application scenario, the recognition effect is still good, thereby improving the speech recognition model's ability to capture hot words in unknown scenarios. Next, the word features of the reference words and the acoustic features of the first speech data are fused, so that the hot word information and acoustic features of the first speech data are effectively fused into a fused feature vector. Based on this, the fused feature vector is further processed by the speech recognition model. Hot word prediction yields the hot word prediction results for the first speech data, and speech recognition is performed on the fused features to obtain the predicted text of the first speech data. Based on the hot word prediction results, hot word labels, and predicted text of the first speech data, the speech recognition model is trained. This not only helps improve the speech recognition model's ability to focus on hot word information in the first speech data, enabling the speech recognition model to learn how to combine acoustic features to capture hot word information in the given scene during training, but also enables the speech recognition model to apply the captured hot word information to speech recognition, thereby possessing the ability to accurately capture hot words from speech data and improving the accuracy of speech recognition.

[0050] Based on the speech recognition model trained using the above training method, this application also proposes a speech recognition method that fuses the acoustic features of the target speech data and the word features of the target hot words, so that the acoustic features and hot word information are effectively fused into the target fused feature vector. On this basis, since the trained speech recognition model has the ability to capture hot words and apply the captured hot word information to the speech recognition process, the speech recognition model can accurately capture the hot word information in the target speech data by performing speech recognition on the target fused feature vector, and thus the output predicted text is more accurate.

[0051] It should be understood that the training method and speech recognition method of the speech recognition model provided in the embodiments of this application can be executed by an electronic device. The electronic device referred to herein may include terminal devices, such as smartphones, tablets, laptops, desktop computers, intelligent voice interaction devices, smart home appliances, smartwatches, vehicle terminals, aircraft, etc.; or, the electronic device may also include a server, such as a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The electronic device is independent of the backend server.

[0052] It is worth noting that the speech recognition method provided in this application can be applied to various business scenarios. As an example, this speech recognition method can be applied to anti-fraud scenarios. In anti-fraud scenarios, training data containing hot words such as "fraud," "low interest," "fast payment," and "no collateral" is scarce. Since the recognition effect of the end-to-end speech recognition model depends on the training data, the speech recognition model trained using traditional training methods cannot accurately recognize and process such hot words, resulting in inaccurate speech recognition results.

[0053] The speech recognition model trained using the training method provided in this application embodiment can accurately capture information with similar hot words from a limited number of hot words, even if the training data containing the aforementioned hot words is limited. This allows for accurate capture of hot word information from speech data in anti-fraud scenarios, and the captured hot word information can be applied to the speech recognition process to output highly accurate predicted text.

[0054] Subsequently, during the dialogue with the user, the system receives voice data input from the user in real time. The word features of hot words in the anti-fraud scenario and the acoustic features of the voice data are input into the trained speech recognition model to obtain predicted text. Then, intent recognition is performed on the predicted text to obtain the predicted intent, determining whether the user has fraudulent intent. Further, based on the predicted intent, a corresponding response text is generated and converted into response voice data, which is then returned to the user, thus completing one round of response. This process is repeated until the dialogue with the user ends.

[0055] It should be understood that the application of this method in anti-fraud scenarios is merely an illustrative example and should not be construed as limiting the scope of its use.

[0056] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0057] Please see Figure 1 The following is a flowchart illustrating a method for training a speech recognition model, provided as an embodiment of this application. The method includes the following steps:

[0058] S102, determine the reference words of the first speech data and the hot word tags of the reference words.

[0059] The first speech data may include a portion of the speech data extracted from a speech dataset. The speech dataset contains multiple speech data points, each with corresponding labeled text. The labeled text can be obtained through manual recognition of the corresponding speech data. The labeled text is used to provide supervision signals to the speech recognition model during training, guiding the model to learn how to convert the speech data into correct text.

[0060] In practice, the training process for the speech recognition model involves multiple rounds of training. As an example, in each round of training, a specified number of speech data are randomly extracted from the speech dataset to obtain multiple first speech data for that round of training.

[0061] The hot word tag for reference words is used to indicate the type of reference words, such as hot words and non-hot words.

[0062] In one embodiment, the reference words may include a subset of words determined from the annotated text of the first speech data.

[0063] In another embodiment, the reference words may include pre-collected words that are associated with the application scenarios of the speech recognition model.

[0064] In another embodiment, S102 includes the following steps: S121, extracting text from the labeled text of the speech data contained in the speech dataset to obtain reference words of the first speech data; S122, determining hot word tags of the reference words based on the labeled text of the first speech data.

[0065] As an example, in S121, words are randomly extracted from the labeled text of the speech data contained in the speech dataset as reference words for the first speech data; in S122, hot word tags of the reference words are determined based on the matching status between the reference words and the words in the labeled text of the first speech data.

[0066] For example, the starting index and number of words to be extracted are preset as the text extraction range; then, within the text extraction range, words are randomly extracted from the labeled text of all speech data contained in the speech dataset as reference words; then, for each reference word, the hot word is matched with the words in the labeled text of the first speech data. If they match, the reference word is determined as a hot word, and a hot word label representing the hot word is set for the reference word; if they do not match, the reference word is determined as a non-hot word, and a hot word label representing the non-hot word is set for the reference word.

[0067] In practical applications, in order to prevent the length of reference words from exceeding the limit and affecting the hot word capture capability of the speech recognition model, a length threshold for the extracted words can be set within the text extraction range to ensure that the length of the extracted words is within the length threshold range.

[0068] In this embodiment, by extracting reference words from the annotated text of all the speech data in the original speech dataset and determining the hot word tags of the reference words, the hot word information contained in these data becomes more universal, which can guide the speech recognition model to capture similar hot word information as much as possible, thereby improving the hot word capture capability of the speech recognition model in various scenarios.

[0069] This application embodiment illustrates a partial implementation of S102 described above. It should be understood that S102 can also be implemented in other ways, and this application embodiment does not limit this implementation.

[0070] S104, the word features of the reference words and the acoustic features of the first speech data are fused to obtain a fused feature vector.

[0071] The word features of the reference words may include, but are not limited to, the index of each character in the reference words in the dictionary, the position of each character in the reference words, etc., and this application embodiment does not limit this. The word features of the reference words can be obtained by feature extraction of the reference words.

[0072] The acoustic features of the first speech data can be obtained using various acoustic features such as fbank features. These features can be extracted using various acoustic feature extraction techniques. For example, by processing the first speech data through pre-emphasis, framing, windowing, discrete Fourier transform, and Mel filtering, fbank features can be obtained and used as the acoustic features of the first speech data.

[0073] By fusing the word features of the reference words and the acoustic features of the first speech data, the hot word information and acoustic features of the first speech data are effectively integrated into the fused feature vector. This fusion process can be implemented in various appropriate ways, and the specific method can be selected according to actual needs.

[0074] In one embodiment, S104 includes the following steps: concatenating the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector.

[0075] In another embodiment, S104 includes the following steps: determining a query matrix based on the word features of reference words, determining a key matrix and a value matrix based on the acoustic features of the first speech data, and performing attention operations on the query matrix, key matrix, and value matrix based on an attention mechanism to obtain a fused feature vector.

[0076] In this embodiment, an attention mechanism is used to fuse word features and acoustic features, similar to a dictionary lookup. Using word features as a reference, acoustic information corresponding to hot word information is searched in the acoustic features, thereby enabling word features and acoustic features to be better integrated, which helps to improve the hot word prediction effect of the speech recognition model.

[0077] In another embodiment, S104 above includes the following steps:

[0078] S141, perform block attention operation on the word features of the reference words and the acoustic features of the first speech data to obtain the first-level attention operation result.

[0079] As an example, in S141 above, a first query matrix is ​​determined based on the word features of the reference words, a first key matrix and a first value matrix are determined based on the acoustic features of the first speech data, and block attention operation is performed based on the first query matrix, the first key matrix and the first value matrix to obtain the first-level attention operation result.

[0080] Specifically, the above-mentioned block attention operation based on the first query matrix, the first key matrix, and the first value matrix to obtain the first-level attention operation result includes the following steps:

[0081] Step A1: Divide the first query matrix into n subquery matrices. n is an integer greater than 1.

[0082] For example, such as Figure 2 As shown, after encoding the word features of the reference words, word representation vectors are obtained. Then, the word representation vectors are used as the first query matrix, and the first query matrix is ​​divided into n sub-query matrices, Q1 to Qn, according to a specified block size. In practical applications, the block size can be used as a hyperparameter and trained together with the speech recognition model.

[0083] Step A2: Divide the first key matrix into n sub-key matrices and the first value matrix into n sub-value matrices.

[0084] For example, such as Figure 2 As shown, after encoding the acoustic features, an acoustic representation vector is obtained. This vector is then used as the first key matrix and the first value matrix. The first key matrix is ​​divided into n sub-key matrices (K1 to Kn) according to a specified block size, and the first value matrix is ​​divided into n sub-value matrices (V1 to Vn). In practical applications, the block size can be used as a hyperparameter and trained together with the speech recognition model.

[0085] Step A3: Combine the n subquery matrices, n subkey matrices, and n subvalue matrices to obtain m combinations of the first matrix.

[0086] Each first matrix combination consists of a subquery matrix, a subkey matrix, and a subvalue matrix, where m is an integer greater than n.

[0087] For example, suppose there are n subquery matrices including Q1 and Q2, n subkey matrices including K1 and K2, and n subvalue matrices including V1 and V2. By combining them, we can obtain the following 8 first matrix combinations: {Q1,K1,V1}, {Q1,K1,V2}, {Q1,K2,V1}, {Q1,K2,V2}, {Q2,K1,V1}, {Q2,K1,V2}, {Q2,K2,V1}, {Q2,K2,V2}.

[0088] Step A4: For each first matrix combination, perform block attention operation on the first matrix combination to obtain the first attention feature of the first matrix combination, and add the first attention feature of the first matrix combination to the first-level attention operation result.

[0089] The result of the first-level attention operation includes the first attention features of the combination of m first matrices.

[0090] For example, for each first matrix combination, the first attention feature of the first matrix combination can be obtained by block attention calculation as follows (1).

[0091]

[0092] Where Attention(Q,K,V) represents the first attention feature of the first matrix combination, Q represents the subquery matrix in the first matrix combination, K represents the subkey matrix in the first matrix combination, and K... T Let V be the transpose of the key matrix, V be the subvalue matrix in the first matrix combination, and d be the subvalue matrix. k The dimension of the subvalue matrix is ​​represented by Softmax, which represents the normalization exponential function used to implement attention operations.

[0093] S142, based on the result of the first-level attention operation, performs a cross-attention operation to obtain the fused feature vector. Specifically, as follows: Figure 2 As shown, based on the first attention features of m combinations of first matrices, m second query matrices, m second key matrices, and m second value matrices are determined. The m second query matrices, m second key matrices, and m second value matrices are combined to obtain k combinations of second matrices. Each combination of second matrices includes a second query matrix, a second key matrix, and a second value matrix, where k is an integer greater than m. For each combination of second matrices, a cross-attention operation is performed to obtain the second attention features of the second matrix combinations. The second attention features of the k combinations of second matrices are concatenated to obtain a fused feature vector.

[0094] For example, for each first attention feature, it is treated as a second query matrix, a second key matrix, and a second value matrix. Thus, based on the first attention features of m combinations of first matrices, m second query matrices, m second key matrices, and m second value matrices are obtained. After combining these second query matrices, second key matrices, and second value matrices to obtain k combinations of second matrices, for each combination, a cross-attention operation is performed on the second query matrix, second key matrix, and second value matrix to obtain the second attention features of that combination. The second attention features of all second matrix combinations are then concatenated to obtain the fused feature vector.

[0095] In this embodiment, considering that when the word features of the reference words are too large, directly fusing the word features and acoustic features using a traditional attention mechanism would lead to attention dispersion in the speech recognition model, resulting in poor hot word prediction performance, a block expansion technique is adopted. First, the first query matrix, first key matrix, and first value matrix are divided into blocks. Attention operations are then performed on the resulting sub-query matrices, sub-key matrices, and sub-value matrices. Since the size of each matrix is ​​reduced during the attention operation, attention can be more focused. This allows for more accurate capture of hot word information from the word features when processing excessive word features, thus preventing attention dispersion. Furthermore, considering that the block division might lead to incomplete hot word information capture due to loose contextual information, affecting hot word prediction performance, all the first attention features obtained from the block division are fused. This makes the connection between contextual information in the word features closer, helping the speech recognition model capture comprehensive hot word information and improving hot word prediction performance.

[0096] This application embodiment illustrates a portion of specific implementations of S104 described above. It should be understood that S104 can also be implemented in other ways, and this application embodiment does not limit this implementation.

[0097] In S104 above, the fusion of word features and acoustic features can be performed by directly fusing the original word features and acoustic features, or by first encoding the word features and acoustic features and then fusing the resulting word representation vector and acoustic representation vector. The latter method yields better fusion results, further increasing the focus on hot word information in the fused feature vector.

[0098] For the latter fusion method, in order to better integrate hot word information into the encoding process of the speech recognition model and enhance the generalization effect of the speech recognition model, in S104 above, word features are encoded by the speech recognition model to obtain word representation vectors; acoustic features are encoded at multiple levels by the speech recognition model to obtain the first acoustic representation vector of each level of encoding; and the first target acoustic representation vector and word representation vector are fused to obtain a fused feature vector. Here, the first target acoustic representation vector refers to the first acoustic representation vector of the target-level encoding that meets the first screening condition in the multi-level encoding. The first screening condition can be set according to actual needs, such as selecting the last level encoding as the target-level encoding and using its encoded first acoustic representation vector as the first target acoustic representation vector, etc. This application embodiment does not limit this.

[0099] For example, word features are encoded at two levels using a speech recognition model, and the vector obtained from the second level encoding is used as the word representation vector. Acoustic features are also encoded at eight levels using the same model, resulting in an eighth-level first acoustic representation vector. A fused feature vector is obtained by fusing the eighth-level first acoustic representation vector and the word representation vector.

[0100] In practical applications, multi-level encoding of acoustic features can be achieved through a speech recognition network within a speech recognition model. As an example, such as... Figure 3 As shown, the speech recognition network includes an acoustic coding sub-network. The acoustic features are encoded at level t through the conformer layer of the acoustic coding sub-network to obtain the first acoustic representation vector at level t, where t is an integer greater than 1.

[0101] The encoding and fusion of word features can be achieved through a hot word prediction network in a speech recognition model. As an example, such as... Figure 3 As shown, the hot word prediction network includes a hot word encoding subnetwork and a fusion subnetwork. The word features are input into the hot word encoding subnetwork and encoded at level s through the transformer layer of the hot word encoding subnetwork to obtain the word representation vector, where s is a positive integer. Then, the word representation vector and the first acoustic representation vector at level t are fused through the fusion subnetwork to obtain the fused feature vector.

[0102] It is worth noting that, under the above fusion method, the fusion processing of the first acoustic representation vector and the word representation vector at level t can be carried out by the implementation of steps S141 to S145 above, or the two can be directly fused based on the attention mechanism, or the two can be fused by splicing, etc.

[0103] S106, using a speech recognition model, hot word prediction results of the first speech data are obtained by performing hot word prediction based on fused feature vectors, and predicted text of the first speech data is obtained by performing speech recognition based on fused feature vectors.

[0104] Hot word prediction results can be obtained through various methods based on fused features. As one example, a linear mapping of the fused feature vector using a speech recognition model can yield the hot word prediction result for the first speech data. As another example, a first hot word prediction result can be obtained by linearly mapping the fused feature vector using a speech recognition model based on the CTC mechanism, and a second hot word prediction result can be obtained by linearly mapping the fused feature vector using a softmax-based linear mapping mechanism. Finally, based on the first and second hot word prediction results, the final hot word prediction result is determined. The hot word prediction result can include the probability that a reference word belongs to the hot word category.

[0105] In practical applications, hot word prediction results can be obtained by using a hot word prediction network of a speech recognition model to predict hot words from fused features. As an example, such as... Figure 3 As shown, the hot word prediction network includes not only a hot word encoding subnetwork and a fusion subnetwork, but also a fusion subnetwork and a hot word decoding subnetwork. The hot word encoding subnetwork's transformer layer performs two-level encoding of the word features of the reference words to obtain a word representation vector. The fusion subnetwork fuses the t-th level first acoustic representation vector with the word representation vector to obtain a fused feature vector. A CTC mechanism is then used to perform a linear mapping on the fused feature vector to obtain the first hot word prediction result. The hot word decoding subnetwork's transformer layer performs multi-level decoding on the fused feature vector based on the hot word labels of the reference words, and a softmax linear mapping mechanism is used to perform a linear mapping on the decoding result to obtain the second hot word prediction result. Finally, based on the first and second hot word prediction results, the final hot word prediction result is determined.

[0106] Predicted text can be obtained through speech recognition based on fused feature vectors in various ways. As one example, linear mapping is applied to the fused feature vectors to obtain the predicted text. As another example, when the fused feature vector is obtained by fusing a first target acoustic representation vector and a word representation vector, the fused feature vector is then encoded at multiple levels to obtain a second acoustic representation vector for each level; speech recognition is then performed based on the second target acoustic representation vector to obtain the predicted text of the first speech data. Here, the second target acoustic representation vector refers to the second acoustic representation vector of the target-level code that satisfies the second filtering condition in the multi-level encoding. The second filtering condition can be set according to actual needs, such as selecting the last level code as the target-level code and using its encoded second acoustic representation vector as the second target acoustic representation vector, etc. This application does not limit this.

[0107] More specifically, the second target acoustic representation vector is subjected to a first linear mapping process to obtain first semantic information; the second target acoustic representation vector is decoded to obtain a semantic vector; the semantic vector is subjected to a second linear mapping process to obtain second semantic information; and speech recognition is performed based on the first and second semantic information to obtain the predicted text of the first speech data.

[0108] It is understandable that multi-level encoding of the fused feature vector is similar to an acoustic model modeling from speech to text, achieving an acoustic mapping from speech to text. Therefore, the second acoustic representation vector encoded at each level implicitly contains the mapping relationship between speech and text. Then, the second target acoustic representation vector is selected, and a first linear mapping process is performed on it, mapping it to higher-level first semantic information. Decoding the second target acoustic representation vector is similar to a language model predicting the next character from the first n characters; the resulting semantic vector carries richer semantic information. Based on this, a second linear mapping process is performed on the second target acoustic representation vector, mapping to higher-level second semantic information. This second semantic information is richer in meaning than the first semantic information, helping to identify homophones and reducing the probability of misidentifying homophones. Therefore, speech recognition based on both first and second semantic information yields more accurate predicted text.

[0109] In practical applications, predicted text can be obtained by performing speech recognition on fused features using a speech recognition network within a speech recognition model. As an example, such as... Figure 3As shown, the speech recognition network includes both an acoustic coding subnetwork and an acoustic decoding subnetwork. The acoustic coding subnetwork performs p-level encoding on the fused features through a transformer layer to obtain a p-level second acoustic representation vector. Based on the CTC mechanism, it uses this p-level second acoustic representation vector as a second target acoustic representation vector and performs a first linear mapping process to obtain first semantic information, where p is an integer greater than 1. The acoustic decoding subnetwork, through a transformer layer, performs q-level decoding on the p-level second acoustic representation vector based on the labeled text of the first speech data to obtain a semantic vector. It then performs a second linear mapping process on the semantic vector to obtain second semantic information. Finally, the first speech information and the second semantic information are weighted and fused for speech recognition to obtain the predicted text of the first speech data.

[0110] S108, the speech recognition model is trained based on the hot word prediction results, hot word labels and predicted text of the first speech data.

[0111] In one embodiment, S108 includes the following steps: determining a first prediction loss of the hot word prediction network based on the hot word prediction result and hot word label of the first speech data; determining a second prediction loss of the speech recognition network based on the predicted text and labeled text of the first speech data; and then updating the network parameters of the hot word prediction network and the network parameters of the speech recognition network based on the first prediction loss and the second prediction loss.

[0112] In another embodiment, the speech recognition network can be pre-trained first, and then the hot word prediction network can be fine-tuned to achieve the training effect of the speech recognition model. Specifically, the speech recognition network is a pre-trained network, for example, obtained by pre-training based on the second speech data and the labeled text of the second speech data. In this case, S108 above includes the following steps: determining the first prediction loss of the hot word prediction network based on the hot word prediction results and hot word labels; determining the second prediction loss of the speech recognition network based on the predicted text of the first speech data and the labeled text of the first speech data; then, freezing the network parameters of the speech recognition network, and fine-tuning the network parameters of the hot word prediction network based on the first prediction loss and the second prediction loss.

[0113] In practical applications, the second voice data can be a portion of the voice data extracted from the voice dataset, or it can be other voice data outside the voice dataset.

[0114] The first prediction loss can include prediction losses caused by two linear mapping processes, namely, as follows: Figure 3As shown, the prediction loss caused by performing a first linear mapping on the p-th level second acoustic representation vector (referred to as the first CTC loss) and the prediction loss caused by performing a second linear mapping on the semantic vector (referred to as the first Attention loss) are shown. The first CTC loss can be calculated based on the difference between the reference semantic information of the annotated text of the first speech data and the first semantic information, and the first Attention loss can be calculated based on the difference between the reference semantic information of the annotated text of the first speech data and the second semantic information.

[0115] Similarly, the second prediction loss can also include the prediction loss caused by the two linear mapping processes, i.e., as... Figure 3 As shown, the prediction loss caused by linear mapping of the fused feature vector (referred to as the second CTC loss) and the prediction loss caused by linear mapping of the decoding result of the fused feature vector (referred to as the second Attention loss) are calculated. The second CTC loss can be calculated based on the difference between the hot word label of the reference word and the prediction result of the first hot word, and the second Attention loss can be calculated based on the difference between the hot word label of the reference word and the prediction result of the second hot word.

[0116] Furthermore, during training, an improved Adam optimizer can be used. This involves adding an L2 regularization term only during parameter updates, replacing the L2 regularization term added during gradient updates in the original Adam. This reduces the drastic parameter oscillations caused by gradient changes, improving the training speed and optimization performance of speech recognition. In this approach, the L2 regularization term is added only during the final parameter update. In the original Adam, adding L2 regularization during gradient updates would cause gradients to accumulate, amplifying the effect of regularization and leading to significant parameter oscillations. The current solution, by adding regularization only during the final parameter update in each gradient update cycle, allows the regularization term to better serve parameter updates, thus reducing drastic parameter oscillations and improving optimization speed and performance.

[0117] In the above implementation, the speech recognition network is pre-trained first, and then the network parameters of the speech recognition network are frozen. The network parameters of the hot word prediction network are fine-tuned, which can make the network parameters of each network in the speech recognition model more stable, prevent the training oscillation problem that is easy to occur after adding reference words and hot word labels. In addition, it can also reduce overfitting and enhance the guiding role of reference words and their hot word labels in the hot word capture ability of the speech recognition model.

[0118] It is worth noting that S102 to S108 above only constitutes one round of model training. In actual training, the speech recognition model will undergo multiple rounds of training, that is, S102 to S108 will be executed multiple times until the training stopping condition is met, and the speech recognition model after the last round of training will be used as the final speech recognition model. The training stopping condition can be set according to actual needs, such as including but not limited to: reaching a preset threshold number of training iterations, or the convergence of the weighted sum of the first and second prediction losses.

[0119] One or more embodiments of this application provide a training method for a speech recognition model, which determines reference words and hot word tags of the reference words in the first speech data. This guides the speech recognition model to capture similar hot word information as much as possible, resulting in better recognition performance even when the training data does not contain the hot words to be recognized in the application scenario, thereby improving the speech recognition model's ability to capture hot words in unknown scenarios. Next, the word features of the reference words and the acoustic features of the first speech data are fused, effectively integrating the hot word information and acoustic features of the first speech data into a fused feature vector. Based on this, the fused feature vector is further processed by the speech recognition model. The hot word prediction results of the first speech data are obtained by performing hot word prediction, and the predicted text of the first speech data is obtained by performing speech recognition on the fused features. Based on the hot word prediction results, hot word labels and predicted text of the first speech data, the speech recognition model is trained. This not only helps to improve the speech recognition model's ability to pay attention to hot word information in the first speech data, but also enables the speech recognition model to learn how to combine acoustic features to capture hot word information in the given scene during the training process. Furthermore, it enables the speech recognition model to apply the captured hot word information to speech recognition, thereby having the ability to accurately capture hot words from speech data and improve the accuracy of speech recognition.

[0120] Based on the speech recognition model trained using the above-described training method, this application also provides a speech recognition method. Please refer to [link to relevant documentation]. Figure 4 The following is a flowchart illustrating a speech recognition method according to an embodiment of this application. The method includes the following steps:

[0121] S402, acquire target speech data and target hot words.

[0122] Target keywords can be set according to specific application scenarios, and this application embodiment does not limit this. For example, in an anti-fraud scenario, target keywords may include, but are not limited to, words such as "fraud," "low interest," "fast payment," and "no collateral."

[0123] S404: The word features of the target hot words and the acoustic features of the target speech data are fused to obtain the target fused feature vector.

[0124] The specific implementation of S404 described above is the same as described above. Figure 1 The specific implementation of S104 in the illustrated embodiment is similar and will not be described again.

[0125] S406, the target fusion feature vector is used to perform speech recognition through a speech recognition model to obtain the predicted text of the target speech data.

[0126] The specific implementation of S406 described above is the same as described above. Figure 1 The specific implementation of S106 in the illustrated embodiment is similar and will not be described again.

[0127] As an example, such as Figure 5 As shown, the word features of the target hot words are encoded at two levels through the transformer layer of the hot word encoding subnetwork to obtain a word representation vector, and the acoustic features of the target speech data are encoded at level t through the conformer layer of the acoustic encoding subnetwork to obtain a level t third acoustic representation vector. Then, the level t third acoustic representation vector and the word representation vector are fused through the fusion subnetwork to obtain a target fused feature vector. Further, the target fused feature vector is encoded at level p through the conformer layer of the acoustic encoding subnetwork to obtain a level p fourth acoustic representation vector. Further, the level p fourth acoustic representation vector is processed by the CTC mechanism to obtain the third semantic information, and the level p fourth acoustic representation vector is decoded at level q through the acoustic decoding subnetwork to obtain a semantic vector. The semantic vector is then processed by the second linear mapping to obtain the fourth semantic information. Finally, the third and fourth semantic information are weighted and fused for speech recognition to obtain the predicted text of the target speech data.

[0128] One or more embodiments of this application provide a speech recognition method that fuses the acoustic features of the target speech data and the word features of the target hot words, thereby effectively integrating the acoustic features and hot word information into the target fused feature vector. Based on this, since the trained speech recognition model has the ability to capture hot words and apply the captured hot word information to the speech recognition process, the speech recognition model can accurately capture the hot word information in the target speech data by performing speech recognition on the target fused feature vector, resulting in more accurate predicted text output.

[0129] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0130] Based on the same inventive concept, this application also provides a training device for a speech recognition model. Please refer to... Figure 6 The following is a schematic diagram of the structure of a speech recognition model training device 600 provided in an embodiment of this application. The device 600 includes: a determination unit 610, a fusion unit 620, a prediction unit 630, and a training unit 640.

[0131] The determining unit 610 is used to determine the reference words of the first speech data and the hot word tags of the reference words.

[0132] The fusion unit 620 is used to fuse the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector.

[0133] The prediction unit 630 is used to perform hot word prediction on the fused feature vector using a speech recognition model to obtain the hot word prediction result of the first speech data, and to perform speech recognition on the fused feature vector to obtain the predicted text of the first speech data.

[0134] The training unit 640 is used to train the speech recognition model based on the hot word prediction results, the hot word tags, and the predicted text.

[0135] In another embodiment, when the fusion unit fuses the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector, it performs the following steps:

[0136] Block attention operations are performed on the word features of the reference words and the acoustic features of the first speech data to obtain the first-level attention operation results;

[0137] Based on the results of the first-level attention operation, a cross-attention operation is performed to obtain the fused feature vector.

[0138] In another embodiment, when the fusion unit performs block attention operations on the word features of the reference words and the acoustic features of the first speech data to obtain the first-level attention operation result, it performs the following steps:

[0139] The first query matrix is ​​determined based on the word features of the reference words;

[0140] The first bond matrix and the first value matrix are determined based on the acoustic features;

[0141] Block attention operations are performed based on the first query matrix, the first key matrix, and the first value matrix to obtain the first-level attention operation result.

[0142] In another embodiment, when the fusion unit performs block attention operations based on the first query matrix, the first key matrix, and the first value matrix to obtain the first-level attention operation result, it executes the following steps:

[0143] Divide the first query matrix into n sub-query matrices, where n is an integer greater than 1;

[0144] The first key matrix is ​​divided into n sub-key matrices, and the first value matrix is ​​divided into n sub-value matrices;

[0145] The n subquery matrices, the n subkey matrices, and the n subvalue matrices are combined to obtain m first matrix combinations. Each first matrix combination includes a subquery matrix, a subkey matrix, and a subvalue matrix, where m is an integer greater than n.

[0146] For each first matrix combination, block attention operation is performed on the first matrix combination to obtain the first attention feature of the first matrix combination, and the first attention feature of the first matrix combination is added to the first-level attention operation result.

[0147] In another embodiment, the first-level attention operation result includes a first attention feature composed of m first matrices;

[0148] When the fusion unit performs cross-attention operation based on the first-level attention operation result to obtain the fused feature vector, it performs the following steps:

[0149] Based on the first attention features of the combination of the m first matrices, determine m second query matrices, m second key matrices, and m second value matrices;

[0150] The m second query matrices, the m second key matrices, and the m second value matrices are combined to obtain k second matrix combinations. Each second matrix combination includes a second query matrix, a second key matrix, and a second value matrix, where k is an integer greater than m.

[0151] For each combination of second matrices, a cross-attention operation is performed on the combination of second matrices to obtain the second attention feature of the combination of second matrices;

[0152] The second attention features of the k combinations of second matrices are concatenated to obtain the fused feature vector.

[0153] In another embodiment, when the fusion unit fuses the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector, it performs the following steps:

[0154] The word features are encoded using the speech recognition model to obtain word representation vectors;

[0155] The acoustic features are encoded in multiple levels using the speech recognition model to obtain the first acoustic representation vector of each level of encoding. The first target acoustic representation vector and the word representation vector are then fused to obtain the fused feature vector. The first target acoustic representation vector refers to the first acoustic representation vector of the target level encoding that meets the first screening condition in the multi-level encoding.

[0156] In another embodiment, when the prediction unit performs speech recognition based on the fused feature vector to obtain the predicted text of the first speech data, it performs the following steps:

[0157] The fused feature vector is encoded in multiple levels to obtain the second acoustic representation vector for each level of encoding;

[0158] Speech recognition is performed based on the second target acoustic representation vector to obtain the predicted text of the first speech data; the second target acoustic representation vector refers to the second acoustic representation vector of the target-level coding that meets the second screening condition in the multi-level coding.

[0159] In another embodiment, when the prediction unit performs speech recognition based on the second target acoustic representation vector to obtain the predicted text of the first speech data, it performs the following steps:

[0160] The second target acoustic representation vector is subjected to a first linear mapping process to obtain first semantic information; and the second target acoustic representation vector is decoded to obtain a semantic vector.

[0161] The semantic vector is subjected to a second linear mapping process to obtain the second semantic information;

[0162] Speech recognition is performed based on the first semantic information and the second semantic information to obtain the predicted text of the first speech data.

[0163] In another embodiment, the determining unit is used to:

[0164] Text extraction is performed on the labeled text of the speech data contained in the speech dataset to obtain the reference words of the first speech data;

[0165] Based on the annotated text of the first speech data, hot word tags for the reference words are determined.

[0166] In another embodiment, when the determining unit extracts text from the labeled text of the speech data contained in the speech dataset to obtain reference words of the first speech data, it performs the following steps:

[0167] Words are randomly extracted from the annotated text of the speech data contained in the speech dataset and used as reference words for the first speech data;

[0168] When determining the hot word tags of the reference words based on the annotated text of the first speech data, the determining unit performs the following steps:

[0169] Based on the matching status between the reference words and the words in the annotated text of the first speech data, hot word tags for the reference words are determined.

[0170] In another embodiment, the speech recognition model includes a speech recognition network and a hot word prediction network, wherein the speech recognition network is a trained network;

[0171] When the training unit trains the speech recognition model based on the hot word prediction results, the hot word tags, and the predicted text, it performs the following steps:

[0172] Based on the hot word prediction results and the hot word tags, the first prediction loss of the hot word prediction network is determined;

[0173] Based on the predicted text of the first speech data and the labeled text of the first speech data, the second prediction loss of the speech recognition network is determined.

[0174] The network parameters of the speech recognition network are frozen, and the network parameters of the hot word prediction network are fine-tuned based on the first prediction loss and the second prediction loss.

[0175] Obviously, the speech recognition model training device provided in this application embodiment can be used as... Figure 1 The entity that performs the training method for the speech recognition model shown, for example Figure 1 In the training method of the speech recognition model shown, step S102 can be performed by... Figure 6 The determination unit 610 in the training device for the speech recognition model shown executes step S104, which can be performed by... Figure 6 The fusion unit 620 in the training device for the speech recognition model shown executes step S106, which can be performed by... Figure 6 The prediction unit 630 in the training device of the speech recognition model shown executes step S108, which can be performed by... Figure 6The training unit 640 in the training device for the speech recognition model shown is executed.

[0176] According to another embodiment of this application, Figure 6 The units in the training device of the speech recognition model shown can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effect of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In the embodiments of this application, the training device of the speech recognition model may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0177] According to another embodiment of this application, a general-purpose computing device, such as a computer, including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), can run an application capable of performing tasks such as... Figure 1 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 6 The document describes a training apparatus for a speech recognition model and a training method for implementing the speech recognition model according to embodiments of this application. The computer program may be recorded on, for example, a computer-readable storage medium, and may be transferred to and run in an electronic device via such a medium.

[0178] Based on the same inventive concept, embodiments of this application also provide a voice recognition device. Please see below. Figure 7 The following is a schematic diagram of the structure of a speech recognition device 700 provided in an embodiment of this application. The device 700 includes: an acquisition unit 710, a fusion unit 720, and a recognition unit 730.

[0179] The acquisition unit 710 is used to acquire target speech data and target hot words.

[0180] The fusion unit 720 is used to fuse the word features of the target hot words and the acoustic features of the target speech data to obtain a target fusion feature vector.

[0181] The recognition unit 730 is used to perform speech recognition on the target fused feature vector through a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained based on the training method of the speech recognition model provided in the embodiments of this application.

[0182] Obviously, the speech recognition device provided in this application embodiment can serve as... Figure 4 The entity performing the speech recognition method shown, for example Figure 4 In the speech recognition method shown, step S402 can be performed by... Figure 7 The acquisition unit 710 in the voice recognition device shown executes step S404, which can be performed by... Figure 7 The fusion unit 720 in the speech recognition device shown executes step S406, which can be performed by... Figure 7 The recognition unit 730 in the voice recognition device shown is executed.

[0183] According to another embodiment of this application, Figure 7 The various units in the illustrated speech recognition device can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one module. In the embodiments of this application, the speech recognition device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0184] According to another embodiment of this application, a general-purpose computing device, such as a computer, including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), can run an application capable of performing tasks such as... Figure 4 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 7 The document describes a speech recognition device and a speech recognition method for implementing embodiments of this application. The computer program may be recorded on, for example, a computer-readable storage medium, and may be transferred to and run in an electronic device via such a medium.

[0185] Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 8At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0186] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0187] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0188] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a training device for a speech recognition model at the logical level. The processor executes the program stored in memory and specifically performs the following operations:

[0189] Determine the reference words for the first speech data and the hot word tags for the reference words;

[0190] The word features of the reference words and the acoustic features of the first speech data are fused to obtain a fused feature vector;

[0191] Using a speech recognition model, hot word prediction results of the first speech data are obtained by performing hot word prediction based on the fused feature vector, and the predicted text of the first speech data is obtained by performing speech recognition based on the fused feature vector.

[0192] The speech recognition model is trained based on the hot word prediction results, the hot word tags, and the predicted text.

[0193] Alternatively, the processor reads the corresponding computer program from non-volatile memory into memory and runs it, forming a speech recognition device at the logical level. The processor executes the program stored in memory and specifically performs the following operations:

[0194] Acquire target speech data and target hot words;

[0195] The word features of the target hot words and the acoustic features of the target speech data are fused to obtain a target fusion feature vector;

[0196] The target fusion feature vector is used to perform speech recognition by a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by the training method of the speech recognition model described in the first aspect.

[0197] The above is as stated in this application. Figure 1 The method executed by the speech recognition model training apparatus disclosed in the illustrated embodiment, or as described in this application. Figure 4 The method executed by the voice recognition device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0198] The electronic device can also perform Figure 1 The method, and the implementation of a training device for the speech recognition model in Figure 1 , Figure 2 , Figure 3 The illustrated embodiment may also perform the functions of the electronic device, or the electronic device may also perform the functions of the embodiment. Figure 4 The method, and implement the speech recognition device in Figure 4 , Figure 5 The functions of the embodiments shown are not described in detail here.

[0199] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0200] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 The method of the illustrated embodiment is specifically used to perform the following operations:

[0201] Determine the reference words for the first speech data and the hot word tags for the reference words;

[0202] The word features of the reference words and the acoustic features of the first speech data are fused to obtain a fused feature vector;

[0203] Using a speech recognition model, hot word prediction results of the first speech data are obtained by performing hot word prediction based on the fused feature vector, and the predicted text of the first speech data is obtained by performing speech recognition based on the fused feature vector.

[0204] The speech recognition model is trained based on the hot word prediction results, the hot word tags, and the predicted text.

[0205] Alternatively, when executed by a portable electronic device that includes multiple applications, the instruction can enable the portable electronic device to perform... Figure 1 The method of the illustrated embodiment is specifically used to perform the following operations:

[0206] Acquire target speech data and target hot words;

[0207] The word features of the target hot words and the acoustic features of the target speech data are fused to obtain a target fusion feature vector;

[0208] The target fusion feature vector is used to perform speech recognition by a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by the training method of the speech recognition model described in the first aspect.

[0209] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the training method of the speech recognition model or the speech recognition method provided in this application.

[0210] In summary, the above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0211] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0212] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0213] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0214] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

Claims

1. A method for training a speech recognition model, characterized in that, include: Determine the reference words of the first speech data and the hot word tags of the reference words; The word features of the reference words and the acoustic features of the first speech data are fused to obtain a fused feature vector; Using a speech recognition model, hot word prediction results of the first speech data are obtained by performing hot word prediction based on the fused feature vector, and the predicted text of the first speech data is obtained by performing speech recognition based on the fused feature vector. The speech recognition model is trained based on the hot word prediction results, the hot word tags, and the predicted text.

2. The method according to claim 1, characterized in that, The process of fusing the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector includes: Block attention operations are performed on the word features of the reference words and the acoustic features of the first speech data to obtain the first-level attention operation results; Based on the results of the first-level attention operation, a cross-attention operation is performed to obtain the fused feature vector.

3. The method according to claim 2, characterized in that, The step of performing block attention operations on the word features of the reference words and the acoustic features of the first speech data to obtain the first-level attention operation result includes: The first query matrix is ​​determined based on the word features of the reference words; The first bond matrix and the first value matrix are determined based on the acoustic features; Block attention operations are performed based on the first query matrix, the first key matrix, and the first value matrix to obtain the first-level attention operation result.

4. The method according to claim 3, characterized in that, The step of performing block attention operations based on the first query matrix, the first key matrix, and the first value matrix to obtain the first-level attention operation result includes: The first query matrix is ​​divided into n sub-query matrices, where n is an integer greater than 1; and the first key matrix is ​​divided into n sub-key matrices, and the first value matrix is ​​divided into n sub-value matrices. The n subquery matrices, the n subkey matrices, and the n subvalue matrices are combined to obtain m first matrix combinations. Each first matrix combination includes a subquery matrix, a subkey matrix, and a subvalue matrix; m is an integer greater than n. For each first matrix combination, block attention operation is performed on the first matrix combination to obtain the first attention feature of the first matrix combination, and the first attention feature of the first matrix combination is added to the first-level attention operation result.

5. The method according to any one of claims 2-4, characterized in that, The result of the first-level attention operation includes first attention features composed of m first matrix combinations; The step of performing cross-attention operation based on the first-level attention operation result to obtain the fused feature vector includes: Based on the first attention features of the combination of the m first matrices, determine m second query matrices, m second key matrices, and m second value matrices; The m second query matrices, the m second key matrices, and the m second value matrices are combined to obtain k second matrix combinations. Each second matrix combination includes a second query matrix, a second key matrix, and a second value matrix, where k is an integer greater than m. For each combination of second matrices, perform a cross-attention operation on the combination of second matrices to obtain the second attention feature of the combination of second matrices; The second attention features of the k combinations of second matrices are concatenated to obtain the fused feature vector.

6. The method according to claim 1, characterized in that, The process of fusing the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector includes: The word features are encoded using the speech recognition model to obtain word representation vectors; The acoustic features are encoded in multiple levels using the speech recognition model to obtain the first acoustic representation vector of each level of encoding. The first target acoustic representation vector and the word representation vector are then fused to obtain the fused feature vector. The first target acoustic representation vector refers to the first acoustic representation vector of the target level encoding that meets the first screening condition in the multi-level encoding.

7. The method according to claim 6, characterized in that, The step of performing speech recognition based on the fused feature vector to obtain the predicted text of the first speech data includes: The fused feature vector is encoded in multiple levels to obtain the second acoustic representation vector for each level of encoding; Speech recognition is performed based on the second target acoustic representation vector to obtain the predicted text of the first speech data; the second target acoustic representation vector refers to the second acoustic representation vector of the target-level coding that meets the second screening condition in the multi-level coding.

8. The method according to claim 7, characterized in that, The step of performing speech recognition based on the second target acoustic representation vector to obtain the predicted text of the first speech data includes: The second target acoustic representation vector is subjected to a first linear mapping process to obtain first semantic information; and the second target acoustic representation vector is decoded to obtain a semantic vector. The semantic vector is subjected to a second linear mapping process to obtain the second semantic information; Speech recognition is performed based on the first semantic information and the second semantic information to obtain the predicted text of the first speech data.

9. The method according to claim 1, characterized in that, The process of determining reference words for the first speech data and hot word tags for the reference words includes: Text extraction is performed on the labeled text of the speech data contained in the speech dataset to obtain the reference words of the first speech data; Based on the annotated text of the first speech data, hot word tags for the reference words are determined.

10. The method according to claim 1, characterized in that, The step of extracting text from the labeled text of the speech data contained in the speech dataset to obtain reference words for the first speech data includes: Words are randomly extracted from the annotated text of the speech data contained in the speech dataset and used as reference words for the first speech data; The step of determining the hot word tags of the reference words based on the annotated text of the first speech data includes: Based on the matching status between the reference words and the words in the annotated text of the first speech data, hot word tags for the reference words are determined.

11. The method according to claim 1, characterized in that, The speech recognition model includes a speech recognition network and a hot word prediction network, wherein the speech recognition network is a trained network. The step of training the speech recognition model based on the hot word prediction results, the hot word tags, and the predicted text includes: Based on the hot word prediction results and the hot word tags, the first prediction loss of the hot word prediction network is determined; Based on the predicted text of the first speech data and the labeled text of the first speech data, the second prediction loss of the speech recognition network is determined. The network parameters of the speech recognition network are frozen, and the network parameters of the hot word prediction network are fine-tuned based on the first prediction loss and the second prediction loss.

12. A speech recognition method, characterized in that, include: Acquire target speech data and target hot words; The word features of the target hot words and the acoustic features of the target speech data are fused to obtain a target fusion feature vector; The target fusion feature vector is subjected to speech recognition by a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by any one of the weights 1-11.

13. A training device for a speech recognition model, characterized in that, include: A determining unit is used to determine reference words in the first speech data and hot word tags for the reference words; The fusion unit is used to fuse the word features of the reference words and the acoustic features of the first speech data to obtain a fused feature vector; The prediction unit is used to obtain the hot word prediction result of the first speech data by performing hot word prediction based on the fused feature vector through a speech recognition model, and to obtain the predicted text of the first speech data by performing speech recognition based on the fused feature vector. The training unit is used to train the speech recognition model based on the hot word prediction results, the hot word tags, and the predicted text.

14. A voice recognition device, characterized in that, include: The acquisition unit is used to acquire target speech data and target hot words; The fusion unit is used to fuse the word features of the target hot words and the acoustic features of the target speech data to obtain a target fusion feature vector; The recognition unit is used to perform speech recognition on the target fusion feature vector through a speech recognition model to obtain the predicted text of the target speech data; the speech recognition model is trained by any one of the methods of weight 1 to weight 11.

15. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the training method of the speech recognition model as described in any one of claims 1 to 11; or, the processor is configured to execute the instructions to implement the speech recognition method as described in claim 12.

16. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the speech recognition model as described in any one of claims 1 to 11; or, when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the speech recognition method as described in any one of claims 12.

17. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of the training method for the speech recognition model as claimed in any one of claims 1 to 11 or the speech recognition method as claimed in claim 12.