Speech Recognition Method, Speech Recognition Device and Computer Readable Storage Medium
By introducing reference text features into the speech recognition method, the recognition of speech features is solved, and the problem of low accuracy in recognition text in the prior art is solved, and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202210400143.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-15
AI Technical Summary
The recognition text obtained by existing speech recognition methods is not very accurate, especially when dealing with words with the same or similar pronunciations, it is easy to cause confusion.
Reference text features are introduced to assist in the recognition of speech features, and the global and local features of the reference text are extracted, and combined with speech features are combined to generate recognition text. The context of the reference text is related to the context of the speech to be recognized, and the speech time of the reference speech is preceded by the speech time of the speech to be recognized.
By considering the connection between the reference text and the context of the text to be identified, the accuracy of the recognition text is improved and the recognition errors of the same or similar words are reduced.
Smart Images

Figure CN114944149B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech processing, and in particular, to a speech recognition method, a speech recognition device, and a computer-readable storage medium. Background Art
[0002] There are many application scenarios for speech recognition technology, such as audio-visual subtitle generation, automatic meeting minutes transcription, recording transcription, intelligent voice assistants, in-vehicle human-machine interaction systems, and so on.
[0003] The speech recognition method can generally be described as extracting the speech features of the speech to be recognized, recognizing the speech features, and obtaining the recognized text. However, the accuracy of the recognized text obtained by this method is not high. Summary of the Invention
[0004] This application provides a speech recognition method, a speech recognition device, and a computer-readable storage medium, which can solve the problem that the accuracy of the recognized text obtained by the existing speech recognition method is not high.
[0005] To solve the above technical problems, a technical solution adopted by this application is: to provide a speech recognition method. The method includes: extracting speech features based on the speech to be recognized to obtain speech features, and extracting text features based on a reference text to obtain reference text features, where the reference text is obtained by recognizing a reference speech, the context of the reference text is related to the context of the speech to be recognized, and the speaking time of the reference speech is earlier than the speaking time of the speech to be recognized; recognizing the recognized text of the speech to be recognized based on the reference text features and the speech features.
[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide a speech recognition device, which includes a feature extraction module and a recognition module. The feature extraction module is used to extract speech features based on the speech to be recognized to obtain speech features, and extract text features based on a reference text to obtain reference text features, where the reference text is obtained by recognizing a reference speech whose context is related to the context of the speech to be recognized, and the speaking time of the reference speech is earlier than the speaking time of the speech to be recognized; the recognition module is used to recognize the recognized text of the speech to be recognized based on the reference text features and the speech features.
[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide a speech recognition device, which includes a processor and a memory connected to the processor, where the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the above method.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium storing program instructions, which can implement the above method when executed.
[0009] In the above manner, this application additionally introduces reference text features to assist in the recognition of speech features. Since the reference speech is related to the context of the speech to be recognized, the reference text is related to the context of the text expressed by the speech to be recognized. The reference text features can express the context of the reference text to a certain extent. Therefore, based on the reference text features to assist in the recognition of speech features, the connection between the reference text and the context of the text expressed by the speech to be recognized can be considered, and the closeness between the recognized text and the text expressed by the speech to be recognized can be improved, that is, the accuracy of the recognized text can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of a speech recognition method in the related art;
[0011] Figure 2 is a schematic flowchart of an embodiment of the speech recognition method of this application;
[0012] Figure 3 is a schematic flowchart of global text feature extraction;
[0013] Figure 4 is a schematic flowchart of local text feature extraction;
[0014] Figure 5 is a schematic structural diagram of an RNN;
[0015] Figure 6 is a schematic flowchart of another embodiment of the speech recognition method of this application;
[0016] Figure 7 is a schematic flowchart of yet another embodiment of the speech recognition method of this application;
[0017] Figure 8 is Figure 7 a specific flowchart of S32 in
[0018] Figure 9 is Figure 7 another specific flowchart of S32 in
[0019] Figure 10 is Figure 7 a specific flowchart of S33 in
[0020] Figure 11 is a schematic structural diagram of a Transformer model;
[0021] Figure 12It is another schematic structural diagram of the Transformer model;
[0022] Figure 13 It is another schematic structural diagram of the Transformer model;
[0023] Figure 14 It is a schematic diagram of an application scenario of the speech recognition method of this application;
[0024] Figure 15 It is a schematic flow diagram of a specific example of the speech recognition method of this application;
[0025] Figure 16 It is a schematic structural diagram of a speech recognition model;
[0026] Figure 17 It is a schematic flow diagram of the training method of the speech recognition model of this application;
[0027] Figure 18 It is a schematic structural diagram of a speech recognition device;
[0028] Figure 19 It is a schematic structural diagram of an embodiment of the speech recognition device of this application;
[0029] Figure 20 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of this application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0031] The terms "first", "second", and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0032] References to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein can be combined with other embodiments.
[0033] Before introducing the speech recognition method provided by the present application, first, in combination with Figure 1 the process of the speech recognition method in the related art will be described as follows:
[0034] As Figure 1 shown, in the time direction, the source speech to be recognized is sequentially segmented into consecutive speech segments to be recognized 1, speech segment to be recognized 2, speech segment to be recognized 3,... in sequence of sentences, and each speech segment to be recognized expresses a sentence of a fixed length. Each speech segment to be recognized is recognized one by one using a speech recognition model to obtain recognition text 1, recognition text 2, recognition text 3,...
[0035] In the above speech recognition method, the recognition processes for different speech segments to be recognized are independent. Therefore, during the recognition of a single speech segment to be recognized, only the information within the sentence expressed by that single speech segment to be recognized is considered.
[0036] Since there are different characters and words with the same or similar pronunciations in the vocabulary, when recognizing the speech segment to be recognized, it is very easy to confuse some characters and words with the same or similar pronunciations, resulting in errors in the recognition text obtained from speech recognition.
[0037] For the convenience of understanding, some specific application scenarios are listed: when the sentence expressed by the speech segment to be recognized has a character with the pronunciation of "TA", it may be recognized as "she", "he", "it", etc. during speech recognition. Another example is that when the sentence expressed by the speech segment to be recognized has a phrase with the pronunciation of "QINGYUANZI", it may be recognized as "hydrogen atom", "green garden", etc. during speech recognition.
[0038] In order to avoid the problem of errors in the recognition text caused by confusing words with the same or similar pronunciations, during the speech recognition process of the present application, reference text features are introduced to assist in the recognition of the speech segment to be recognized, specifically as follows:
[0039] Figure 2 is a schematic flowchart of an embodiment of the speech recognition method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 2 the flowchart order shown. As Figure 2 shown, this embodiment may include:
[0040] S11: Extract speech features based on the speech to be recognized to obtain speech features, and extract text features based on the reference text to obtain reference text features.
[0041] Among them, the reference text is obtained by recognizing the reference speech. The context of the reference text is related to the context of the speech to be recognized, and the speaking time of the reference speech is prior to the speaking time of the speech to be recognized.
[0042] The source speech can be segmented according to a preset rule in the time direction (generally segmented by sentence) to obtain a number of speech segments, and each speech segment is used as a speech to be recognized. During the speech recognition process, the several speeches to be recognized are recognized one by one. The embodiments of the present application will be described by taking the recognition of one of the speeches to be recognized as an example.
[0043] Since the speaking time of the reference speech is prior to the speaking time of the speech to be recognized, from the perspective of speaking time, the reference speech is a historical speech relative to the speech to be recognized, and the reference text is a historical text relative to the recognized text. The context of the reference speech / speech to be recognized mentioned in the present application refers to the context of the text expressed by the reference speech / speech to be recognized. The context of the reference speech being related to the context of the speech to be recognized means that the context of the text expressed by the reference speech is related to the context of the text expressed by the speech to be recognized.
[0044] It can be understood that according to speaking habits, in the same passage, the context of the subsequent sentence is related to the context of the previous sentence (that is, the context of different sentences is contextually related). Or, in different passages with the same speaking topic, the context of the subsequent passage is related to the context of the previous passage (that is, the context of different passages is contextually related).
[0045] Based on this, the reference speech and the speech to be recognized can come from the same source speech, and the speech to be recognized is the speech segment expressing the subsequent sentence in the same source speech, and the reference speech is the speech segment expressing the previous sentence in the same source speech. For example, the same source speech expresses a passage, which includes 5 sentences and is segmented into speech segments 1-5 respectively expressing the 5 sentences. During the speech recognition process, speech segments 1-5 are sequentially used as the speeches to be recognized for recognition. When speech segment 3 is used as the speech to be recognized, speech segments 1 and 2 that have been speech recognized are used as the reference speech.
[0046] Alternatively, the reference speech and the speech to be recognized may come from two different source speeches. The two different source speeches may be collected by the same or different speech collection devices used by the same speaker at different time periods, or may be collected by the speech collection devices used by different speakers participating in the speech at different time periods. For example, the reference speech comes from source speech 1, and the speech to be recognized comes from source speech 2. Source speech 1 is collected by the speech collection device A used by speaker A in time period 1, and source speech 2 is collected by the speech collection device B used by speaker A in time period 2 after time period 1. Another example is that source speech 1 and source speech 2 are respectively collected by the speech collection device A used by speaker A in time period 1 and time period 2. Still another example is that source speech 1 is collected by the speech collection device A used by speaker A in time period 1, and source speech 2 is collected by the speech collection device C used by speaker B in time period 3 after time period 1.
[0047] For the convenience of understanding, a specific application scenario is listed: A recently wants to buy a car. On the first day, A talks with friend A about the car purchase topic, and the relevant speech a1 is collected by the speech collection device A used by A. On the second day, A talks with friend B about the car purchase topic, and the relevant speech a2 is collected by the speech collection device A used by A. The speech to be recognized comes from speech a2, and the reference speech comes from speech a1.
[0048] Generally speaking, the smaller the time interval between the reference speech and the speech to be recognized, the higher the degree of context relevance of the text expressed by the reference speech and the speech to be recognized, and vice versa. Correspondingly, the longer the duration of the reference speech, the more additional processing overhead is required to introduce the reference text features in the recognition process of the speech to be recognized. In some embodiments, the time interval threshold can be used to balance the relevance degree between the reference speech and the speech to be recognized and the additional processing overhead. The time interval threshold can be fixed or can adaptively change according to the application scenario. For example, in the case where the speech topic has a strong unfolding property, a longer time interval can be set; in the case where the speech topic has a weak unfolding property, a shorter time interval can be set. The limitation of the time interval can be achieved through a sliding window.
[0049] For the convenience of understanding, a specific application scenario is given: The speech to be recognized and the reference speech come from the same source speech, and the same source speech is a conference recording, which is collected from the speeches of each speaker around a certain conference topic in the conference. The conference recording is segmented into 100 speech segments by sentences, and the 100 speech segments are sequentially used as the speeches to be recognized for recognition. When the speech segment 50 is used as the speech to be recognized, compared with the speech segments 1 - 39 that do not meet the time interval threshold requirement, the speech segments 40 - 49 that meet the time interval threshold requirement have a higher degree of context relevance with the speech segment 50, and the speech segments 1 - 39 are used as the reference speech.
[0050] The reference text features may include at least one of the global text features of the reference text and the local text features of each keyword in the reference text.
[0051] The reference text includes several sentences. For the global text features: several sentences in the reference text can be concatenated to obtain a concatenated sentence; global feature extraction is performed on the concatenated sentence to obtain global text features.
[0052] The model for global text feature extraction includes but is not limited to the BERT model. The BERT model has a full-field attention mechanism, so it can extract global text features for expressing the global information of the reference text. When the length of the concatenated sentence is N, the computational complexity of the model is O(N 2 ), so the longer the concatenated sentence (the longer the duration of the reference speech), the higher the complexity of the model to achieve global text feature extraction.
[0053] Combined with Figure 3 For example, as Figure 3 shown, sentences 1 to t + w are all historical texts that have been recognized, sentence t + w + 1 is the text to be recognized for the speech representation, and sentences 1 to t + w within the adjacent sliding window (length w) are the reference text. The sentences within the sliding window can be concatenated to obtain a concatenated sentence; the concatenated sentence is input into the BERT model to obtain global text features.
[0054] For the local text features of each keyword: several sentences in the reference text can be concatenated to obtain a concatenated sentence; several keywords are extracted from the concatenated sentence; local feature extraction is performed on each keyword to obtain the local text features of each keyword.
[0055] Among them, the steps for keyword extraction may include:
[0056] The concatenated sentence is tokenized to obtain several phrases in the concatenated sentence; tokenization is to add boundary markers between words in the concatenated sentence, and the models used include but are not limited to the N-gram model (Ngram) and the neural network language model (MLM).
[0057] Furthermore, several candidate keywords are extracted from several phrases; the models for extracting candidate keywords include but are not limited to TFIDF, Topic-model, and RAKE. Taking TFIDF as an example, 2N candidate keywords can be extracted according to the pre-set number of keywords N.
[0058] Further, candidate keywords are filtered to obtain the final keywords. Among them, redundant candidate keywords can be merged. For example, "iFlytek Co., Ltd." and "iFlytek" are merged, and only the longer keyword is retained. Redundant phrases caused by inconsistent word segmentation can also be deleted. For example, "total lunar eclipse" and "total lunar eclipse" are regarded as the same word. Numbers, units, etc. are deleted, and single characters are deleted. Still taking the TFIDF model as an example, the scores of 2N candidate keywords can be output according to the TFIDF model, and the N candidate keywords with the highest scores are selected from the 2N candidate keywords as the final keywords.
[0059] Combined with Figure 4 Taking an example for illustration, the sentences within the sliding window can be concatenated to obtain a concatenated sentence. Based on the Ngram speech model, the concatenated sentence is segmented, and 2N candidate keywords are extracted from the segmentation results based on the TFIDF model. The candidate keywords are filtered to obtain keywords 1 to keyword N. Each keyword is encoded to obtain the local text feature H = [h 1 , h 2 , …, h N . Among them, h i (i = 1, …, N) represents the local text feature of the i-th keyword.
[0060] S12: Based on the reference text feature and the speech feature, the recognized text of the speech to be recognized is obtained.
[0061] The speech feature is a sequence composed of multiple speech sub-features, and each speech sub-feature represents a character or a word composed of multiple characters. The recognition based on the reference text feature and the speech feature is divided into multiple rounds (multiple time steps), and in each round, the recognition is performed based on different speech sub-features and reference text features in turn. More specifically, before performing the recognition of a specific round, it is necessary to first determine the speech sub-feature (specific speech sub-feature) to be recognized in this specific round from the speech feature. The method for determining the specific speech sub-feature can be to assign attention weights to each speech sub-feature in the speech feature, and the attention weight assigned to the specific speech sub-feature is much higher than that of other speech sub-features. Or, the method for determining the specific speech sub-feature can be to define the position of the specific speech sub-feature through position indication information. Thus, in the recognition process of a specific round, the information of other speech sub-features is masked to exclude the interference of other speech sub-features on the recognition of the specific sub-feature.
[0062] The recognition process in S12 can also be called a decoding process, and the decoder used can be any type of recurrent neural network (RNN), such as LSTM, or a Transformer network, or a derivative network of the networks listed above, etc.
[0063] Recognition based on speech features and reference text features essentially uses reference text features to assist in the recognition of speech features. The way of assisting with reference text features can be to splice the reference text features with the speech features, and subsequent decoding is performed on the splicing result of the reference text features and the speech features; or it can be to fuse the reference text features with the intermediate decoding results (such as hidden states) of the speech features, and subsequent decoding is performed on the fusion result of the reference text features and the intermediate decoding results; or it can be to perform attention processing on the speech features based on the reference text features, and subsequent decoding is performed based on the result of the attention processing, and so on.
[0064] It can be understood that the semantics of different characters and words with the same or similar pronunciations are different, and semantics is related to the context. Therefore, combining the context where the words are located for recognition can reduce the possibility of confusion of different words with the same or similar pronunciations during the speech recognition process.
[0065] However, if recognition is only based on the speech features of the speech to be recognized, the context considered is limited to the context within the text expressed by the speech to be recognized. Therefore, in this embodiment, reference text features are additionally introduced to assist in the recognition of speech features. Since the reference speech is related to the context of the speech to be recognized, the reference text is related to the context of the text expressed by the speech to be recognized. The reference text features can express the context of the reference text to a certain extent. Therefore, based on the reference text features to assist in the recognition of speech features, the connection between the context of the reference text and the context of the text expressed by the speech to be recognized can be considered, and the closeness between the recognized text and the text expressed by the speech to be recognized can be improved, that is, the accuracy of the recognized text can be improved.
[0066] To facilitate understanding of the technical effects brought by introducing reference text features, an example is given in combination with an actual application scenario:
[0067] When the text expressed by the speech to be recognized has a character with the pronunciation "TA", it may be recognized as "she", "he", "it", etc. during speech recognition. If the reference text contains the obvious gender-directed word "female student", then the character with the pronunciation "TA" is more likely to be recognized as "she". If the reference text contains the obvious gender-directed word "male student", then the character with the pronunciation "TA" is more likely to be recognized as "he".
[0068] When the sentence expressed by the speech to be recognized has a word with the pronunciation "QINGYUANZI", it may be recognized as "hydrogen atom", "green garden", etc. during speech recognition. If the reference text contains chemical knowledge such as organic substances, then the word with the pronunciation "QINGYUANZI" is more likely to be recognized as "hydrogen atom". If the reference text contains parks, scenery, etc., then the word with the pronunciation "QINGYUANZI" is more likely to be recognized as "green garden".
[0069] Further, when the decoder on which the recognition process is based in S12 is an RNN, S12 is described as follows:
[0070] Figure 5 It is a schematic structural diagram of an RNN. As Figure 5 shown, the RNN includes an input layer X, a hidden layer S, and an output layer O. t represents a time step (also referred to as a round hereinafter), and O t represents the output of the t-th round. U, W, and V represent weights. X t represents the input of the t-th round, and S t represents the value of the hidden layer (hidden state) of the t-th round. S t not only depends on X t , but also depends on S t-1 (the hidden state of the (t - 1)-th round), that is, the input of the hidden layer of the t-th round includes X t and S t-1 , and the output is S t .
[0071] The original processing logic of the RNN for X t can be expressed as the following formula:
[0072] O t = g(V * S t );
[0073] S t = f(U * X t + W * S t-1 );
[0074] where f(.) represents the function for processing X t and S t-1 , and g(.) represents the function for processing S t . The initial hidden state S 0 of the hidden layer is a feature of all zeros.
[0075] Based on the structure of the RNN, S12 is further expanded:
[0076] Figure 6 It is a schematic flowchart of another embodiment of the speech recognition method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 6 the process sequence shown. As Figure 6 shown, this embodiment may include:
[0077] S21: Obtain the decoding state of the front-wheel decoding.
[0078] The decoding state, which is the hidden state mentioned above, is obtained by decoding speech features during the previous round of decoding. In the case where this round of decoding is the first round of decoding, the decoding state of the previous round of decoding is all-zero features (the feature values of each dimension are all 0), or the decoding state of the previous round of decoding is obtained based on the global text features. Among them, the global text features can be transformed, and the transformed global text features are adapted to the decoder state, and used as the decoding state of the previous round of decoding (i.e., the hidden state of the hidden layer, the initial hidden state), which can be reasonably utilized by the decoder.
[0079] S22: Decode based on the decoding state of the previous round of decoding and the reference features to obtain the decoded character and decoding state of this round of decoding.
[0080] Among them, the reference features at least include speech features.
[0081] In other embodiments, the reference features may further include the fused text features of the local text features of each keyword in this round of decoding. Among them, the fused text features of the local text features of each keyword in different rounds of decoding may be the same or different.
[0082] In the same case, according to the same first weight, the local text features of each keyword are weighted to obtain the fused text features, and the fused text features are applied to each round of decoding.
[0083] In different cases, based on the decoding state of this round of decoding, the first weight of each keyword in this round of decoding can be obtained; based on the first weight of each keyword in this round of decoding, the first text features of each keyword are weighted to obtain the fused text features. Among them, the features of each keyword can be matched with the decoding state of this round of decoding to obtain the first weight of each keyword in this round of decoding. For example, denote the local text features of each keyword as H = [h 1 ,h 2 ,…,h N , h i (i = 1,..., N) represents the local text feature of the i-th keyword. By matching with the decoding state s of this round of decoding, the first weight b of each keyword is obtained as b = softmax(v T tanh(Ws + VH); based on the first weight of each keyword, the local text features of each keyword are weighted to obtain the fused text features
[0084] In the case where the reference text features include global text features, S22 may include: Decode based on the decoding state corresponding to the previous round of decoding and the global text features to obtain the decoded character and decoding state of this round of decoding.
[0085] When the reference text features further include fused text features, S22 may include: concatenating the speech features and the fused text features decoded in the current round to obtain the first concatenated feature decoded in the current round; performing decoding based on the first concatenated feature decoded in the current round and the decoding state decoded in the previous round to obtain the decoded character and decoding state decoded in the current round.
[0086] S23: Combine the decoded characters in each round of decoding to obtain the recognized text.
[0087] The following uses several examples to elaborate in detail on the t-th round of decoding in S21 - S23:
[0088] Example 1: The reference text features include global text features, and the reference features include speech features.
[0089] 1) In the t-th round, the attention mechanism is used to process the speech features, and the attention of the processed speech features Xt focuses on the speech sub-features to be recognized in the t-th round.
[0090] 2) The processing logic of the RNN for Xt can be described as:
[0091] S 0 = y(M);
[0092] O t = g(V * S t );
[0093] S t = f(U(X t + N t ) + W * S t-1 );
[0094] Among them, M represents the global text feature, S 0 represents the initial state of the hidden layer (initial hidden state, initial decoding state), N t represents the fused text feature of the global text features of each keyword in the t-th round of decoding, X t + N t represents the concatenated feature of X t and N t , and S 0 is a feature of all zeros.
[0095] Example 2: The reference text features include local text features of each keyword, and the reference features include speech features and fused text features.
[0096] 1) In the t-th round, the attention mechanism is used to process the speech features, and the attention of the processed speech features X t focuses on the speech sub-features to be recognized in the t-th round.
[0097] 2) The RNN processes Xt The processing logic can be described as:
[0098] O t = g(V * S t );
[0099] S t = f(U(X t + N t ) + W * S t-1 );
[0100] Wherein, N t represents the fused text feature of the local text features of each keyword in the t-th round of decoding, and X t + N t represents the concatenated feature of X t and N t , and S 0 is a feature of all zeros.
[0101] Example 3: The reference text feature includes the global text feature and the local text features of each keyword, and the reference feature includes the speech feature and the fused text feature.
[0102] 1) In the t-th round, the attention mechanism is used to process the speech feature, and the attention of the processed speech feature Xt is concentrated on the speech sub-feature to be recognized in the t-th round.
[0103] 2) The processing logic of the RNN for Xt can be described as:
[0104] S0 = y(M);
[0105] Ot = g(v * St);
[0106] St = f(U(Xt + Nt) + W * St-1).
[0107] It can be understood that in this embodiment, when the reference text feature includes the local text features of each keyword and the reference feature includes the fused text feature, compared with the original processing logic of the RNN, the concatenated feature of the fused text feature and the speech feature is used as the input of this round, so that the fused text feature will be referred to in each step of processing the speech feature until the decoding of this round is completed. In this way, the fused text feature is used to assist in the recognition of the speech feature, improving the accuracy of the recognized text.
[0108] In the case where the reference text features include global text features, compared with the original processing logic of the RNN, the decoding state required for the first-round decoding is obtained based on the global text features, so that the global text features assist in the recognition of speech features by integrating into the hidden state, which is equivalent to providing a bias to the decoder, making the hidden state of the decoder more inclined to decode in the direction of the text context that conforms to the speech expression to be recognized, masking the information that does not conform to the text context of the speech expression to be recognized, thereby improving the accuracy of the recognized text.
[0109] Furthermore, with reference to Figure 7 , in the case where the decoder relied on in the recognition process in S12 is a Transformer model, S12 may include the following sub-steps:
[0110] S31: Obtain the character features corresponding to each historical round of decoding.
[0111] Among them, the character features corresponding to the historical round of decoding are extracted based on the decoded characters of the historical round of decoding.
[0112] Each historical round is relative to the current round, including the previous round, the round before the previous round,..., the first round of the current round. For example, if the current round is the 2nd round, each historical round includes the 0th round and the 1st round.
[0113] The character features may include a first character query feature Q, a first character key feature K, and a first character value feature V. Among them, the first character query feature is obtained by converting the decoded characters of the previous round of decoding through a query mapping parameter, the first character key feature is obtained by converting the decoded characters of the previous round of decoding through a key mapping parameter, and the first character value feature is obtained by converting the decoded characters of the previous round of decoding through a value mapping parameter. The query mapping parameter, the key mapping parameter, and the value mapping parameter are three matrices of the same size, which are multiplied by the decoded characters obtained from the previous round of decoding respectively to obtain Q, K, and V.
[0114] S32: Based on the character features corresponding to each historical round of decoding, perform attention processing on the decoded characters obtained from the previous round of decoding to obtain a first attention processing result.
[0115] In some embodiments, the global text features are also referred to during the process of performing attention processing on the decoded characters of the previous round of decoding.
[0116] In the case of not referring to the global text features, with reference to Figure 8 , S32 may include the following sub-steps:
[0117] S321: Calculate a second weight based on the first character query feature corresponding to the previous round of decoding and the first character key features corresponding to each historical round of decoding.
[0118] S322: Obtain a first attention processing result based on the second weight and the first character value features corresponding to each historical round of decoding.
[0119] In some embodiments, S321 - S322 can be implemented in the following manner: Concatenate the first character key features corresponding to each historical round of decoding to obtain a first concatenated feature; match the first character query feature corresponding to the previous round of decoding with the second concatenated feature to obtain a second weight; concatenate the first character value features corresponding to each historical round of decoding to obtain a third concatenated feature; multiply the second weight by the third concatenated feature to obtain a first attention processing result.
[0120] In some embodiments, S321 - S322 can also be implemented in the following manner: Match the first character query feature corresponding to the previous round of decoding with the first character key features corresponding to each historical round of decoding respectively to obtain the second weight of the first character value features corresponding to each historical round of decoding in this round; Weight the first character value features corresponding to each historical round of decoding according to the second weight to obtain a first attention processing result.
[0121] In the case of referring to the global text feature, in combination with referring to Figure 9 , S32 can include the following sub - steps:
[0122] S323: Calculate a second weight based on the first character query feature corresponding to the previous round of decoding, the first character key features corresponding to each historical round of decoding, and the global text feature.
[0123] S324: Obtain a first attention processing result based on the second weight, the first character value features corresponding to each historical round of decoding, and the global text feature.
[0124] In some embodiments, S323 - S324 can be implemented in the following manner: Concatenate the first character key features corresponding to each historical round of decoding and the global text feature to obtain a second concatenated feature; match the first character query feature corresponding to the previous round of decoding with the second concatenated feature to obtain a second weight; concatenate the first character value features corresponding to each historical round of decoding and the global text feature to obtain a third concatenated feature; multiply the second weight by the third concatenated feature to obtain a first attention processing result.
[0125] In some embodiments, S323 - S324 can also be implemented in the following manner: Match the first character query feature corresponding to the previous round of decoding with the first character key features corresponding to each historical round of decoding and the global text feature respectively to obtain the second weight of the first character value features corresponding to each historical round of decoding and the global text feature in this round; Weight the first character value features corresponding to each historical round of decoding and the global text feature according to the second weight to obtain a first attention processing result.
[0126] It can be understood that S321 to S322 can be regarded as a process of performing self-attention processing on the decoded characters of the front-wheel decoding. The process of performing self-attention processing on the decoded characters of the front-wheel decoding can extract information related to the current round of decoding from the decoded characters of the front-wheel decoding and mask the information unrelated to the current round of decoding in the decoded characters of the front-wheel decoding. Compared with S321 to S322, the process of performing self-attention processing on the decoded characters of the front-wheel decoding by S323 to S324 additionally considers the global text features. Therefore, the first attention processing result obtained can more accurately locate the information related to the current round of decoding in the decoded characters of the front-wheel decoding, and the first attention processing result obtained is more accurate.
[0127] S33: Decode based on the first attention processing result and the reference features to obtain the decoded characters of the current round of decoding.
[0128] Among them, the reference features at least include speech features.
[0129] In some embodiments, the reference features may further include the local text features of each keyword in the reference text.
[0130] In the case where the reference features include speech features, the speech features can be enhanced based on the first attention processing result to obtain enhanced features; decode based on the enhanced features to obtain the decoded characters of the current round of decoding. The way to enhance the speech features can be to convert the first attention processing result to obtain a third character query feature, convert the speech features to obtain a third character key feature and a third character value feature, and perform attention processing based on the third character query feature, the third character key feature, and the third character value feature (the method is similar to the process of performing attention processing on the first character query feature, the first character key feature, and the first character value feature described above) to obtain enhanced features.
[0131] In the case where the reference features further include local text features, with reference to Figure 10 , S33 may include the following sub-steps:
[0132] S331: Enhance the speech features based on the first attention processing result to obtain enhanced features.
[0133] S332: Decode based on the enhanced features and the local text features of the current round of decoding to obtain the decoded characters of the current round of decoding.
[0134] The enhanced feature can be transformed through query mapping parameters to obtain a second character query feature, and the local text feature can be transformed through the mapping parameter and the value mapping parameter respectively to obtain a second character key feature and a second character value feature; attention processing is performed based on the second character query feature, the second character key feature, and the second character key feature to obtain a second attention processing result; decoding is performed based on the second attention processing result to obtain the decoded character of this round of decoding.
[0135] That is to say, when the reference feature further includes the local text feature, after obtaining the enhanced feature, instead of directly decoding the enhanced feature to obtain the decoded character of this round of decoding, the local text feature is also referred to, so the accuracy of the decoded character of this round of decoding can be improved.
[0136] S34: Combine the decoded characters obtained from each round of decoding to obtain the recognized text.
[0137] The following combines the structure of the Transformer model and illustrates S31 to S34 in the form of several examples:
[0138] Example 4: During the attention processing of the decoded characters obtained from the previous round of decoding, the global text feature and the character features corresponding to each historical round of decoding are referred to; the reference feature includes the speech feature.
[0139] Figure 11 is a schematic structural diagram of the Transformer model, as Figure 11 shown, the Transformer model includes N Transformer block structures, and each Transformer block structure includes a self-attention module, an encoder-decoder attention module, and a fully connected layer.
[0140] The self-attention module processes the decoded characters obtained from the previous round of decoding by using the self-attention mechanism (which can be a single-head self-attention mechanism or a multi-head self-attention mechanism. In the following of this application, the multi-head self-attention mechanism is taken as an example for illustration).
[0141] 1) In the 0th round, the self-attention module processes to obtain the first attention processing result of the 0th round of decoding. Specifically, the self-attention module processes the start flag " <s>The word embedding vector a of " 0 is transformed to obtain the first character query feature Q 10 , the first character key feature K 10 and the first character value feature V 10 ; K 10 is concatenated with the global text feature L to obtain the second concatenated feature of the 0th round, and V 10 is concatenated with the global text feature L to obtain the third concatenated feature of the 0th round; Q 1 is matched with the second concatenated feature of the 0th round to obtain the second weight α 10 of the first-round decoding; α 10 is multiplied by the third concatenated feature of the 0th round to obtain the first attention processing result P10 of the 0th round decoding.
[0142] 2) The encoder-decoder attention module enhances the part X 10 to be decoded in the 0th round in the speech features based on P 0 to obtain the enhanced feature X 0 '. Specifically, P 10 is transformed into the third character query feature Q 30 , and X 0 is respectively transformed into the third character key feature K 30 and the third character value feature V 30 ; attention processing is performed based on Q 30 , K 30 and V 30 to obtain the enhanced feature X 0 '.
[0143] 3) The fully connected layer performs a non-linear transformation based on the enhanced feature X 0 ' to obtain the decoded character "ke" of the 0th round decoding. For specific instructions, please refer to the related technology and will not be elaborated here.
[0144] 4) In the 1st round, the self-attention module processes to obtain the first attention processing result of the 1st round decoding. Specifically, the self-attention module transforms the word embedding vector a 1 of "ke" to obtain the first character query feature Q 11 , the first character key feature K 11 and the first character value feature V 11 ; K 11 , K 10 are concatenated with L to obtain the second concatenated feature of the 1st round, and V 11 , V 10 are concatenated with L to obtain the third concatenated feature of the 1st round; Q 11 is matched with the second concatenated feature of the 1st round to obtain the second weight α 11 of the 1st round decoding, and α 11 Multiply it by the third concatenated feature of the first round to get the first attention processing result P of the first round of decoding 11 .
[0145] 5) Encoder-decoder attention module based on P 11 For the first round of decoding part X in the speech feature 1 Enhance and obtain enhanced feature X 1 '. The specific process is the same as obtaining X 0 'The process is similar and will not be described here.
[0146] 6) The fully connected layer is based on the enhanced feature X 1 'Perform nonlinear transformation to obtain the decoded character "大" in the first round of decoding. For detailed description, please refer to the relevant technology, which will not be repeated here.
[0147] The processing flow of subsequent rounds is similar.
[0148] Example 5: In the process of performing attention processing on the decoded characters obtained from the previous round of decoding, the character features corresponding to each historical round of decoding are referred to; the reference features include speech features and local text features.
[0149] Figure 12 This is another structural diagram of the Transformer model, such as Figure 12 As shown in the figure, the Transformer model includes N Transformer block structures, each of which includes a self-attention module, an encoder-decoder attention module, a keyword-decoder attention module and a fully connected layer.
[0150] 1) In round 0, the self-attention module processes the first attention processing result of round 0 decoding. Specifically, the self-attention module processes the start marker " <s>The word embedding vector a of " 0 is transformed to obtain the first character query feature Q 10 , the first character key feature K 10 and the first character value feature V 10 ; Q 10 and K 10 are matched to obtain the second weight α for the first-round decoding 10 ; α 10 and V 10 are multiplied to obtain the first attention processing result P for the 0th-round decoding 10 .
[0151] 2) The encoder-decoder attention module enhances the part X 10 to be decoded in the 0th round in the speech features based on P 0 to obtain the enhanced feature X 0 '.
[0152] 3) The keyword-decoder attention module obtains the second attention processing result P for the 0th-round decoding based on X 0 ' and the local text features H of each keyword. Specifically, X 20 ' is transformed to obtain the second character query feature Q 0 , H is transformed to obtain the second character key feature K and the second character value feature V; attention processing is performed based on Q 20 , K and V to obtain the second attention processing result P 20 , K and V to obtain the second attention processing result P 20 .
[0153] 4) The fully connected layer performs a non-linear transformation based on the enhanced feature P 20 to obtain the decoded character "Ke" for the 0th-round decoding. For specific descriptions, please refer to the related technology and will not be elaborated here
[0154] The subsequent processing flow is similar to the previous one and will not be elaborated here
[0155] Example 6: In the process of performing attention processing on the decoded characters obtained from the previous-round decoding, the global text features and the character features corresponding to each historical-round decoding are referred to; the reference features include speech features
[0156] Figure 13 is another structural schematic diagram of the Transformer model. As Figure 13 shown, the Transformer model includes N Transformer block structures, and each Transformer block structure includes a self-attention module, an encoder-decoder attention module, and a fully connected layer
[0157] 1) In the 0th round, the self-attention module processes to obtain the first attention processing result of the 0th round decoding. Specifically, the self-attention module processes the start identifier " <s>The word embedding vector a of "" 0 is transformed to obtain the first character query feature Q 10 , the first character key feature K 10 and the first character value feature V 10 ; K 10 is concatenated with the global text feature L to obtain the second concatenated feature in the 0th round, and V 10 is concatenated with the global text feature L to obtain the third concatenated feature in the 0th round; Q 1 is matched with the second concatenated feature in the 0th round to obtain the second weight α 10 in the first-round decoding; α 10 is multiplied by the third concatenated feature in the 0th round to obtain the first attention processing result P 10 in the 0th round decoding.
[0158] 2) The encoder-decoder attention module enhances the part X 10 to be decoded in the 0th round in the speech features based on P 0 to obtain the enhanced feature X 0 '.
[0159] 3) The keyword-decoder attention module obtains the second attention processing result P 0 in the 0th round decoding based on X 20 ' and H. Specifically, X 0 ' is transformed to obtain the second character query feature Q 20 , and H is transformed to obtain the second character key feature K 20 and the second character value feature V 20 ; attention processing is performed based on Q 20 , K 20 and V 20 to obtain the enhanced feature X 0 '.
[0160] 4) The fully connected layer performs a non-linear transformation based on the enhanced feature X 0 ' to obtain the decoded character "Ke" in the 0th round decoding. For specific descriptions, please refer to the related technology and will not be elaborated here.
[0161] The subsequent processing flow is similar to the previous one and will not be elaborated here.
[0162] Combined with Figure 14 , list a specific application scenario of the speech recognition method of this application: The speaker holds the end - side 1, end - side 2, and end - side 3. The end - sides 1 - 3 can perform speech collection and recognition. The text (with timestamp information) obtained by the end - sides 1 - 3 after collecting and recognizing the speech will be synchronized to the speech assistance device (the speech assistance device can be one of the end - sides 1 - 3 or other devices). The speech assistance device sorts the texts sent by the end - sides 1 - 3 according to the timestamp information and extracts the text features after sorting. When the speaker uses one of the end - sides (for example, end - side 1), the speech assistance device sends the text features as reference text features to end - side 1 for reference when end - side 1 recognizes newly collected speech.
[0163] Further, the recognized text is obtained based on a speech recognition model. The following is an illustration of the speech recognition method provided in this application in the form of an example: Figures 15 - 16
[0164] Example 7: As Figure 15 shown, the source speech is segmented into sentences to obtain the speech to be recognized 1, the speech to be recognized 2, and the speech to be recognized 3. Since the speech to be recognized 1 represents the first sentence, its reference text feature is all - zero feature (indicating an empty history). The reference text feature and the speech to be recognized 1 are input into the speech recognition model to obtain the recognition result 1; the recognition result 1 is used as the reference text feature (historical representation 1) of the speech to be recognized 2, and the historical representation 1 and the speech to be recognized 2 are input into the speech recognition model to obtain the recognition result 2; the recognition result 2 is used as the historical representation 2 and the speech to be recognized 3 are input into the speech recognition model to obtain the recognition result 3, and so on.
[0165] Figure 16 is the structure of the speech recognition model (it should be noted that this structure does not limit the speech recognition model, that is, the speech recognition model can also be other structures in related technologies), as Figure 16 shown, the speech recognition model includes a speech encoder, an attention module, and a decoder. Based on this, the processing flow of the speech recognition model can include:
[0166] 1) The speech encoder encodes the speech to obtain speech features.
[0167] 2) Before each round of decoding, the attention module masks the part of the speech features that is irrelevant to the corresponding round of decoding;
[0168] 3) The decoder, based on the reference text features and the speech features, recognizes the recognition result / recognized text of the speech to be recognized.
[0169] Further, before applying the speech recognition model as described above, the speech recognition model needs to be trained to the desired level. The speech recognition model is trained based on sample data, which includes sample speech to be recognized, sample reference text, and the sample actual text expressed by the sample speech to be recognized. The sample reference text is obtained by recognizing a sample reference speech whose context is related to the context of the sample speech to be recognized, and the speaking time of the sample reference speech precedes that of the sample speech to be recognized. The training of the speech recognition model includes several rounds.
[0170] In the case of insufficient sample data, if there is only the sample speech to be recognized, the sample reference text can be obtained by methods such as manual annotation and ASR engine transcribing; if there is only the sample reference text, the sample speech to be recognized can be obtained by methods such as manual annotation and TTS engine synthesis.
[0171] Combined with reference to Figure 17 , in some embodiments, the training steps of the speech recognition model may include:
[0172] S41: Extract speech features based on the sample speech to be recognized to obtain sample speech features, and extract text features based on the sample reference text to obtain sample text features.
[0173] S42: Based on the sample text features and sample speech features, recognize the sample recognition text of the sample speech to be recognized.
[0174] S43: Adjust the network parameters of the speech recognition model based on the difference between the sample recognition text and the sample actual text.
[0175] The principle of this embodiment is similar to that of the previous embodiments and will not be elaborated here.
[0176] Further, considering that in actual applications, there may be a situation where the speech to be recognized has no reference speech (for example, when the speech to be recognized expresses the first sentence of a paragraph, refer to Figure 15 the relevant description), therefore, in order to take into account the performance of the speech recognition model in the case of no reference speech, during the training process, the sample text features of some rounds or some sample speeches to be recognized are discarded, that is, set to all-zero features with the same size as the sample text features.
[0177] Thus, in some embodiments, before S42, it may also be determined whether the current round of training meets the training conditions for discarding the sample text features; in response to the current round of training not meeting the training conditions, S42 is executed; in response to the current round of training meeting the training conditions, the sample text features are replaced with all-zero features, and based on the all-zero features and the sample speech features, the sample recognition text of the sample speech to be recognized is obtained. It can be understood that a certain proportion of the sample text features of the sample speech to be recognized in the training dataset can be set as the speech meeting the training conditions, or a certain proportion of the training rounds can be set as the rounds meeting the training conditions, etc. Correspondingly, the training conditions may include that the current round meets the training conditions, the sample speech meets the training conditions, etc.
[0178] It can be understood that in the related art, for the training of a speech recognition model, the sample data used includes the sample speech to be recognized and the sample actual text expressed by the sample speech to be recognized. Correspondingly, the training process includes: extracting speech features from the sample speech to be recognized to obtain speech features; recognizing the speech features to obtain the sample recognition text of the sample speech to be recognized; and adjusting the network parameters of the speech recognition model based on the difference between the sample recognition text and the sample actual text.
[0179] Compared with the training method of the speech recognition model in the related art, in the embodiments of the present application, a sample text feature is additionally introduced to assist in the recognition of the sample speech features. Since the sample reference speech is related to the context of the sample speech to be recognized, the sample reference text is related to the context of the text expressed by the sample speech to be recognized. The sample text features can express the context of the sample reference text to a certain extent. Therefore, based on the sample text features to assist in the recognition of the speech features, the connection between the sample reference text and the context of the text expressed by the sample speech to be recognized can be considered, improving the closeness between the obtained sample recognition text and the text expressed by the sample speech to be recognized, that is, improving the accuracy of the sample recognition text and the training effect of the speech recognition model.
[0180] Figure 18 is a schematic structural diagram of a speech recognition device, as Figure 18 shown, the speech recognition device may include a feature extraction module 11 and an identification module 12.
[0181] The feature extraction module 11 can be used to extract speech features based on the speech to be recognized to obtain speech features, and extract text features based on the reference text to obtain reference text features. Among them, the reference text is obtained by recognizing the reference speech whose context is related to the context of the speech to be recognized, and the speaking time of the reference speech is prior to the speaking time of the speech to be recognized. The recognition module 12 can be used to recognize the recognized text of the speech to be recognized based on the reference text features and the speech features. The reference text features include at least one of the global text features of the reference text and the local text features of each keyword in the reference text.
[0182] Through the implementation of this embodiment, in the processing process of the recognition module 12 of this embodiment, a reference text feature is additionally introduced to assist in the recognition of speech features. Since the reference speech is related to the context of the speech to be recognized, the reference text is related to the context of the text expressed by the speech to be recognized. The reference text features can express the context of the reference text to a certain extent. Therefore, based on the reference text features to assist in the recognition of speech features, the connection between the reference text and the context of the text expressed by the speech to be recognized can be considered, and the closeness between the obtained recognized text and the text expressed by the speech to be recognized can be improved, that is, the accuracy of the recognized text can be improved.
[0183] In some embodiments, the recognition module 12 can specifically be used to: obtain the decoding state of the previous round of decoding; perform decoding based on the decoding state of the previous round of decoding and the reference features to obtain the decoded characters and decoding state of the current round of decoding; where the reference features at least include the speech features; combine the decoded characters of each round of decoding to obtain the recognized text. Among them, when the current round of decoding is the first round of decoding, the decoding state of the previous round of decoding is obtained based on the global text features, and / or the reference features further include the fused text features of the local text features of each keyword in the current round of decoding.
[0184] Since when the current round of decoding is the first round of decoding, the decoding state of the previous round of decoding is obtained based on the global text features, and / or the reference features further include the fused text features of the local text features of each keyword in the current round of decoding. Therefore, the accuracy of the recognized text can be improved.
[0185] In some embodiments, the recognition module 12 can specifically be used to: based on the decoding state of the current round of decoding, obtain the first weight of each keyword in the current round of decoding; weight the first text features of each keyword based on the first weight of each keyword in the current round of decoding to obtain the fused text features.
[0186] It can be understood that according to the method of obtaining the first weight of each keyword in the current round of decoding based on the decoding state of the current round of decoding, keywords related to the current round of decoding can be given a greater weight, and keywords that are not relevant can be given a smaller weight, so that the obtained fused text features can better assist in the recognition of speech features.
[0187] In some embodiments, the recognition module 12 may specifically be configured to: when the reference feature further includes the fused text feature, splice the speech feature and the fused text feature decoded in the current round to obtain a first spliced feature decoded in the current round; and perform decoding based on the first spliced feature decoded in the current round and the decoding state decoded in the previous round to obtain the decoded character and the decoding state decoded in the current round.
[0188] It can be understood that by participating in splicing the fused text feature to obtain the first spliced feature and the second spliced feature, the fused text feature can be incorporated into the speech feature and the decoding state, and then decoding based on the first spliced feature and the second spliced feature can provide the accuracy of recognition and decoding.
[0189] In some embodiments, the recognition module 12 may specifically be configured to: obtain the character features corresponding to each historical round of decoding; wherein, the character features corresponding to the historical round of decoding are extracted based on the decoded characters of the historical round of decoding; perform attention processing on the decoded characters obtained by the previous round of decoding based on the character features corresponding to each historical round of decoding to obtain a first attention processing result; perform decoding based on the first attention processing result and the reference feature to obtain the decoded character decoded in the current round; wherein, the reference feature at least includes the speech feature; combine the decoded characters obtained in each round of decoding to obtain the recognized text; wherein, the global text feature is also referred to during the process of performing attention processing on the decoded characters of the previous round of decoding, and / or, the reference feature further includes the local text features of each keyword in the reference text.
[0190] By the way of also referring to the global text feature during the process of performing self-attention processing on the decoded characters of the previous round of decoding, and / or, the reference feature further includes the local text features of each keyword in the reference text, the recognition of the speech feature by the reference text feature can be realized, and the accuracy of the recognized text can be improved.
[0191] In some embodiments, the character feature includes a first character query feature, a first character key feature, and a first character value feature. The first character query feature is obtained by converting the decoded characters of the previous round of decoding through a query mapping parameter, the first character key feature is obtained by converting the decoded characters of the previous round of decoding through a key mapping parameter, and the first character value feature is obtained by converting the decoded characters of the previous round of decoding through a value mapping parameter.
[0192] When the global text feature is also referred to during the process of performing self-attention processing on the decoded characters of the previous round of decoding, the recognition module 12 may specifically be configured to: calculate a second weight based on the first character query feature corresponding to the previous round of decoding, the first character key features corresponding to each historical round of decoding, and the global text feature; and obtain a first attention processing result based on the second weight, the first character value features corresponding to each historical round of decoding, and the global text feature.
[0193] The above process can be regarded as a process of performing self-attention processing on the characters corresponding to the previous-round decoding. Since the global text features are referred to in the process of performing self-attention processing on the character features corresponding to the previous-round decoding, the information related to the current-round decoding in the character features corresponding to the previous-round decoding can be more accurately located, so that the first attention processing result can more accurately retain the information related to the current-round decoding and mask the information unrelated to the current-round decoding.
[0194] In some embodiments, the recognition module 12 can specifically be configured to: splice the first character key features corresponding to each historical-round decoding and the global text features to obtain a second spliced feature; match the first character query feature corresponding to the previous-round decoding with the second spliced feature to obtain a second weight; splice the first character value features corresponding to each historical-round decoding and the global text features to obtain a third spliced feature; multiply the second weight by the third spliced feature to obtain a first attention processing result.
[0195] Thus, the global text features can be incorporated into the process of self-attention processing on the character features corresponding to the previous-round decoding by being spliced into the first character key features and the first character value features.
[0196] In some embodiments, the recognition module 12 can specifically be configured to: when the reference features further include the local text features of each keyword in the reference text, enhance the speech features based on the first attention processing result to obtain enhanced features; perform decoding based on the enhanced features and the local text features to obtain the decoded character of the current-round decoding.
[0197] Since the local text features are also referred to when decoding the enhanced features, the accuracy of the decoded character of the current-round decoding can be improved.
[0198] In some embodiments, the recognition module 12 can specifically be configured to: convert the enhanced features through query mapping parameters to obtain second character query features, and convert the local text features through key mapping parameters and value mapping parameters respectively to obtain second character key features and second character value features; perform attention processing based on the second character query features, the second character key features, and the second character key features to obtain a second attention processing result; perform decoding based on the second attention processing result to obtain the decoded character of the current-round decoding.
[0199] Since, after obtaining the enhanced features, the decoded character of the current-round decoding is not directly obtained based on the enhanced features, but further attention processing is performed through the local text features and the enhanced features, the information unrelated to the current-round decoding can be further masked, and the accuracy of the subsequent recognized text can be improved.
[0200] In some embodiments, the recognition module 12 obtains recognition text based on a speech recognition model. The speech recognition model is trained based on sample data, which includes sample speech to be recognized, sample reference text, and sample actual text expressed by the sample speech to be recognized. The sample reference text is obtained by recognizing a sample reference speech whose context is related to the context of the sample speech to be recognized, and the speaking time of the sample reference speech is prior to the speaking time of the sample speech to be recognized.
[0201] In some embodiments, the training steps of the speech recognition model include: extracting speech features based on the sample speech to be recognized to obtain sample speech features, and extracting text features based on the sample reference text to obtain sample text features; recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features; and adjusting the network parameters of the speech recognition model based on the difference between the sample recognition text and the sample actual text.
[0202] In this embodiment, sample text features are additionally introduced to assist in the recognition of sample speech features. Since the sample reference speech is contextually related to the sample speech to be recognized, the sample reference text is contextually related to the text expressed by the sample speech to be recognized. The sample text features can express the context of the sample reference text to a certain extent. Therefore, based on the sample text features to assist in the recognition of speech features, the connection between the context of the sample reference text and the text expressed by the sample speech to be recognized can be considered, improving the closeness between the obtained sample recognition text and the text expressed by the sample speech to be recognized, that is, improving the accuracy of the sample recognition text and providing the training effect of the speech recognition model.
[0203] In some embodiments, the speech recognition model is obtained through several rounds of training. The recognition module 12 can specifically be used to: before recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features, determine whether the current round of training meets the training conditions for discarding the sample text features; in response to the current round of training not meeting the training conditions, perform the step of recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features; in response to the current round of training meeting the training conditions, replace the sample text features with all-zero features, and recognize the sample recognition text of the sample speech to be recognized based on the all-zero features and the sample speech features.
[0204] It can be understood that considering that there may be a situation where there is no reference speech for the speech to be recognized in actual applications (such as when the speech to be recognized expresses the first sentence of a paragraph), therefore, during the training process, by using training conditions to limit the replacement of the sample text features of certain sample speeches to be recognized with all-zero features (discarding), the performance of the speech recognition model in the case of no reference speech can be taken into account, improving the robustness of the trained speech recognition model.
[0205] Figure 19 This is a schematic structural diagram of an embodiment of the voice recognition device of the present application. The voice recognition device can be any device with voice recognition function, such as mobile phones, walkie-talkies, computers, and so on. As Figure 19 shown, the voice recognition device includes a processor 21 and a memory 22 which are coupled to each other.
[0206] Among them, the memory 22 stores program instructions for implementing the method of any of the above embodiments; the processor 21 is configured to execute the program instructions stored in the memory 22 to implement the steps of the above method embodiments. Among them, the processor 21 can also be called a CPU (Central Processing Unit, central processing unit). The processor 21 may be an integrated circuit chip with signal processing capabilities. The processor 21 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0207] Figure 20 This is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. As Figure 20 shown, the computer-readable storage medium 30 of the embodiment of the present application stores program instructions 31, and when the program instructions 31 are executed, the method provided in the above embodiments of the present application is implemented. Among them, the program instructions 31 can form a program file and be stored in the above computer-readable storage medium 30 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of various embodiments of the present application. And the aforementioned computer-readable storage medium 30 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.
[0208] Among them, and / or, in addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only the implementation manner of the present application, and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.< / s> < / s> < / s>
Claims
1. A speech recognition method, characterized in that, it includes: Performing speech feature extraction based on the speech to be recognized to obtain speech features, and performing text feature extraction based on a reference text to obtain reference text features, wherein the reference text is obtained by recognizing a reference speech, the context of the reference text is related to the context of the speech to be recognized, the speaking time of the reference speech is earlier than the speaking time of the speech to be recognized, and the reference text features include at least one of the global text features of the reference text and the local text features of each keyword in the reference text; Based on the reference text features and the speech features, recognizing the recognized text of the speech to be recognized, including: Obtaining the decoding state of the previous round of decoding; decoding based on the decoding state of the previous round of decoding and a first reference feature to obtain the decoded character and decoding state of the current round of decoding; wherein the first reference feature includes at least the speech features; combining the decoded characters of each round of decoding to obtain the recognized text; wherein, when the current round of decoding is the first round of decoding, the decoding state of the previous round of decoding is obtained based on the global text features, and / or, the first reference feature further includes the fused text features of the local text features of each keyword in the current round of decoding; Or, Obtaining the character features corresponding to each historical round of decoding; wherein the character features corresponding to the historical round of decoding are extracted based on the decoded characters of the historical round of decoding; performing attention processing on the decoded characters obtained by the previous round of decoding based on the character features corresponding to each historical round of decoding to obtain a first attention processing result; decoding based on the first attention processing result and a second reference feature to obtain the decoded character of the current round of decoding; wherein the second reference feature includes at least the speech features; combining the decoded characters obtained by each round of decoding to obtain the recognized text; wherein the global text features are also referred to during the process of performing attention processing on the decoded characters of the previous round of decoding, and / or, the second reference feature further includes the local text features of each keyword.
2. The method according to claim 1, characterized in that, the step of obtaining the fused text features includes: Based on the decoding state of the current round of decoding, obtaining the first weight of each keyword in the current round of decoding; Weighting the first text features of each keyword based on the first weight of each keyword in the current round of decoding to obtain the fused text features.
3. The method according to claim 1, characterized in that, when the first reference feature further includes the fused text features, the decoding based on the decoding state of the previous round of decoding and the first reference feature to obtain the decoded character and decoding state of the current round of decoding includes: Concatenating the speech features and the fused text features of the current round of decoding to obtain the first concatenated feature of the current round of decoding; Decoding based on the first concatenated feature of the current round of decoding and the decoding state of the previous round of decoding to obtain the decoded character and decoding state of the current round of decoding.
4. The method according to claim 1, characterized in that, The character features include a first character query feature, a first character key feature, and a first character value feature; When the global text feature is also referred to during the attention processing of the decoded characters of the front wheel decoding, the attention processing of the decoded characters of the front wheel decoding based on the character features corresponding to each of the historical round decodings to obtain a first attention processing result includes: Calculating a second weight based on the first character query feature corresponding to the front wheel decoding, the first character key features corresponding to each of the historical round decodings, and the global text feature; Obtaining the first attention processing result based on the second weight, the first character value features corresponding to each of the historical round decodings, and the global text feature.
5. The method according to claim 1, wherein, When the second reference feature further includes local text features of each keyword in the reference text, the decoding based on the first attention processing result and the second reference feature to obtain the decoded character of the current round decoding includes: Enhancing the speech feature based on the first attention processing result to obtain an enhanced feature; Decoding based on the enhanced feature and the local text feature to obtain the decoded character of the current round decoding.
6. The method according to claim 1, wherein, The recognized text is recognized based on a speech recognition model, the speech recognition model is trained based on sample data, the sample data includes sample speech to be recognized, sample reference text, and sample actual text expressed by the sample speech to be recognized, the sample reference text is recognized from a sample reference speech whose context is related to the context of the sample speech to be recognized, and the speaking time of the sample reference speech is earlier than the speaking time of the sample speech to be recognized; The training steps of the speech recognition model include: Performing speech feature extraction based on the sample speech to be recognized to obtain sample speech features, and performing text feature extraction based on the sample reference text to obtain sample text features; Recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features; Adjusting the network parameters of the speech recognition model based on the difference between the sample recognition text and the sample actual text.
7. The method according to claim 6, wherein, The speech recognition model is obtained through several rounds of training. Before recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features, the method further includes: Determining whether the current round of training meets the training condition for discarding the sample text features; In response to the current round of training not meeting the training condition, performing the step of recognizing the sample recognition text of the sample speech to be recognized based on the sample text features and the sample speech features; In response to the current round of training meeting the training condition, replacing the sample text features with all-zero features, and recognizing the sample recognition text of the sample speech to be recognized based on the all-zero features and the sample speech features.
8. A speech recognition device, wherein, Including: A feature extraction module, configured to perform speech feature extraction based on the speech to be recognized to obtain speech features, and perform text feature extraction based on a reference text to obtain reference text features, where the reference text is obtained by recognizing a reference speech whose context is related to the context of the speech to be recognized, the speaking time of the reference speech is earlier than the speaking time of the speech to be recognized, and the reference text features include at least one of the global text features of the reference text and the local text features of each keyword in the reference text; An identification module, configured to identify the recognized text of the speech to be recognized based on the reference text features and the speech features, including: obtaining the decoding state of the previous-round decoding; decoding based on the decoding state of the previous-round decoding and a first reference feature to obtain the decoded character and decoding state of the current-round decoding; where the first reference feature includes at least the speech features; combining the decoded characters of each round of decoding to obtain the recognized text; where, when the current-round decoding is the first-round decoding, the decoding state of the previous-round decoding is obtained based on the global text features, and / or, the first reference feature further includes the fused text features of the local text features of each keyword in the current-round decoding; Or, Obtaining the character features corresponding to each historical-round decoding; where the character features corresponding to the historical-round decoding are extracted based on the decoded characters of the historical-round decoding; performing attention processing on the decoded characters obtained by the previous-round decoding based on the character features corresponding to each historical-round decoding to obtain a first attention processing result; decoding based on the first attention processing result and a second reference feature to obtain the decoded character of the current-round decoding; where the second reference feature includes at least the speech features; combining the decoded characters obtained by each round of decoding to obtain the recognized text; where the global text features are also referred to in the process of performing attention processing on the decoded characters of the previous-round decoding, and / or, the second reference feature further includes the local text features of each keyword.
9. A speech recognition device Characterized in that It includes a memory and a processor that are mutually coupled, and the processor is configured to execute the program instructions stored in the memory to implement the speech recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium Characterized in that The computer-readable storage medium stores program instructions that can be run by a processor, and the program instructions are used to implement the speech recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition method and device, electronic equipment and storage medium
CN112562659A