Speech recognition methods, devices, equipment and storage media
By preprocessing speech data and fusing style feature vectors, the impact of different speaking styles and emotional changes on speech recognition is resolved, thereby improving the accuracy and reliability of speech recognition.
Patent Information
- Application Number
- CN202111101833.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-09-18
AI Technical Summary
How to improve the accuracy of speech recognition, especially when faced with different speaking styles and emotional changes.
By preprocessing the speech data to be recognized, speech feature parameters are determined, and a style feature vector is generated using a pre-trained style recognition model. This vector is then combined with the speech recognition model for recognition, and the style feature vector and speech feature parameters are fused to generate the recognition result.
It effectively reduces the impact of speaking style on speech recognition, improving the accuracy and reliability of speech recognition results.
Smart Images

Figure CN115841813B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of speech technology and artificial intelligence technology such as deep learning, and particularly to a speech recognition method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of internet technology, emerging industries such as short videos and online education are injecting new vitality into economic development. Voice recognition technology, as a fundamental service across various industries, has broad application prospects in new internet business areas.
[0003] Therefore, improving the accuracy of speech recognition is an urgent problem that needs to be solved. Summary of the Invention
[0004] This disclosure provides a speech recognition method, apparatus, device, and storage medium.
[0005] According to one aspect of this disclosure, a speech recognition method is provided, comprising:
[0006] The speech data to be recognized is preprocessed to determine the speech feature parameters corresponding to the speech data;
[0007] The pre-trained style recognition model is used to identify the speech feature parameters in order to determine the style feature vector corresponding to the speech data.
[0008] Based on the style feature vector, the speech data is recognized using a pre-trained speech recognition model to generate the recognition result corresponding to the speech data.
[0009] Optionally, the step of recognizing the speech data using a pre-trained speech recognition model based on the style feature vector to generate a recognition result corresponding to the speech data includes:
[0010] According to preset rules, the style feature vector and the speech feature parameters are fused to generate a vector to be recognized;
[0011] The speech recognition model is used to recognize the vector to generate a recognition result corresponding to the speech data.
[0012] Optionally, the step of recognizing the speech data using a pre-trained speech recognition model based on the style feature vector to generate a recognition result corresponding to the speech data includes:
[0013] The speech feature parameters are encoded using the encoder in the speech recognition model to generate the audio vector corresponding to the speech data;
[0014] According to preset rules, the style feature vector and the audio vector are fused to generate a first vector;
[0015] The first vector is identified using the recognition module in the speech recognition model to generate the recognition result corresponding to the speech data.
[0016] Optionally, the step of recognizing the speech data using a pre-trained speech recognition model based on the style feature vector to generate a recognition result corresponding to the speech data includes:
[0017] According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector.
[0018] The second vector is encoded using the remaining layers in the encoder to obtain the third vector;
[0019] The recognition module in the speech recognition model is used to recognize the third vector in order to generate the recognition result corresponding to the speech data.
[0020] Optionally, the style feature vector contains one first sub-vector, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameters to be fused. The preset rule includes any one of the following:
[0021] The style feature vector is placed after the audio vector to be fused;
[0022] The style feature vector is placed after each second sub-vector in the audio vector to be fused;
[0023] The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and,
[0024] The style feature vector is added to each second sub-vector in the audio vector to be fused.
[0025] Optionally, the first sub-vectors included in the style feature vector have the same number of second sub-vectors as the audio vector to be fused, and the dimensions of the first sub-vectors are the same as the dimensions of the second sub-vectors. The preset rule includes any one of the following:
[0026] The style feature vector is placed after the audio vector to be fused;
[0027] Each of the first sub-vectors is placed after the corresponding second sub-vector;
[0028] Each of the first sub-vectors is concatenated with its corresponding second sub-vector; and...
[0029] Each of the first sub-vectors is added to its corresponding second sub-vector.
[0030] Optionally, the dimension of the style feature vector is different from the dimension of the audio vector to be fused, and the preset rule is to concatenate the style feature vector with the audio vector.
[0031] According to a second aspect of this disclosure, a voice recognition device is provided, comprising:
[0032] The first determining module is used to preprocess the speech data to be recognized in order to determine the speech feature parameters corresponding to the speech data;
[0033] The second determining module is used to identify the speech feature parameters using a pre-trained style recognition model, so as to determine the style feature vector corresponding to the speech data.
[0034] The recognition module is used to recognize the speech data based on the style feature vector using a pre-trained speech recognition model, so as to generate a recognition result corresponding to the speech data.
[0035] Optionally, the identification module is specifically used for:
[0036] According to preset rules, the style feature vector and the speech feature parameters are fused to generate a vector to be recognized;
[0037] The speech recognition model is used to recognize the vector to generate a recognition result corresponding to the speech data.
[0038] Optionally, the identification module is specifically used for:
[0039] The speech feature parameters are encoded using the encoder in the speech recognition model to generate the audio vector corresponding to the speech data;
[0040] According to preset rules, the style feature vector and the audio vector are fused to generate a first vector;
[0041] The first vector is identified using the recognition module in the speech recognition model to generate the recognition result corresponding to the speech data.
[0042] Optionally, the identification module is specifically used for:
[0043] According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector.
[0044] The second vector is encoded using the remaining layers in the encoder to obtain the third vector;
[0045] The recognition module in the speech recognition model is used to recognize the third vector in order to generate the recognition result corresponding to the speech data.
[0046] Optionally, the style feature vector contains one first sub-vector, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameters to be fused. The preset rule includes any one of the following:
[0047] The style feature vector is placed after the audio vector to be fused;
[0048] The style feature vector is placed after each second sub-vector in the audio vector to be fused;
[0049] The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and,
[0050] The style feature vector is added to each second sub-vector in the audio vector to be fused.
[0051] Optionally, the first sub-vectors included in the style feature vector have the same number of second sub-vectors as the audio vector to be fused, and the dimensions of the first sub-vectors are the same as the dimensions of the second sub-vectors. The preset rule includes any one of the following:
[0052] The style feature vector is placed after the audio vector to be fused;
[0053] Each of the first sub-vectors is placed after the corresponding second sub-vector;
[0054] Each of the first sub-vectors is concatenated with its corresponding second sub-vector; and...
[0055] Each of the first sub-vectors is added to its corresponding second sub-vector.
[0056] Optionally, the dimension of the style feature vector is different from the dimension of the audio vector to be fused, and the preset rule is to concatenate the style feature vector with the audio vector.
[0057] A third aspect of this disclosure provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect of this application.
[0058] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as described in the first aspect of this application.
[0059] A fifth aspect of this disclosure provides a computer program product that, when executed by an instruction processor, performs the method proposed in a first aspect of this disclosure.
[0060] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Finally, based on the style feature vector, the pre-trained speech recognition model is used to recognize the speech data to generate the recognition result corresponding to the speech data. Therefore, by recognizing the speech data based on the style feature vector corresponding to the speech recognition data during the speech recognition process, the influence of speaking style on speech recognition is avoided, improving the accuracy and reliability of the speech recognition results.
[0061] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0062] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0063] Figure 1 A flowchart illustrating a speech recognition method provided in an embodiment of this disclosure;
[0064] Figure 2 A flowchart illustrating another speech recognition method provided in this embodiment of the present disclosure;
[0065] Figure 2a This is a schematic diagram of a first fusion method provided in an embodiment of the present disclosure;
[0066] Figure 2b This is a schematic diagram of a second fusion method provided in an embodiment of the present disclosure;
[0067] Figure 2c This is a schematic diagram of a third fusion method provided in the embodiments of this disclosure;
[0068] Figure 2d This is a schematic diagram of the fourth fusion method provided in the embodiments of this disclosure;
[0069] Figure 2e This is a schematic diagram of the fifth fusion method provided in the embodiments of this disclosure;
[0070] Figure 2f This is a schematic diagram of the sixth fusion method provided in the embodiments of this disclosure;
[0071] Figure 2g This is a schematic diagram of the seventh fusion method provided in the embodiments of this disclosure;
[0072] Figure 2h This is a schematic diagram of the eighth fusion method provided in the embodiments of this disclosure;
[0073] Figure 2i This is a schematic diagram of the ninth fusion method provided in the embodiments of this disclosure;
[0074] Figure 3 This diagram illustrates the overall architecture of a speech recognition model.
[0075] Figure 4 A flowchart illustrating yet another speech recognition method provided in this disclosure embodiment;
[0076] Figure 5 This diagram illustrates the overall architecture of yet another speech recognition model.
[0077] Figure 6 A flowchart illustrating yet another speech recognition method provided in this disclosure embodiment;
[0078] Figure 7 This diagram illustrates the overall architecture of another speech recognition model.
[0079] Figure 8 A structural block diagram of a speech recognition device provided in an embodiment of this disclosure;
[0080] Figure 9 This is a block diagram of an electronic device used to implement the speech recognition method of the embodiments of this disclosure. Detailed Implementation
[0081] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0082] This disclosure provides a speech recognition method, which can be executed by a speech recognition device provided by this disclosure or by an electronic device provided by this disclosure. The electronic device can be a terminal device, such as a user device, mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant, handheld device, computing device, vehicle device, wearable device, etc., without limitation, and can also be a server.
[0083] The following describes the implementation of a speech recognition method provided in this disclosure using a speech recognition device, and is not intended to limit the scope of this disclosure.
[0084] The speech recognition method, apparatus, computer equipment, and storage medium provided in this disclosure are described in detail below with reference to the accompanying drawings.
[0085] Figure 1 This is a flowchart illustrating a speech recognition method provided according to an embodiment of the present disclosure.
[0086] like Figure 1 As shown, the speech recognition method may include the following steps:
[0087] Step 101: Preprocess the speech data to be recognized to determine the speech feature parameters corresponding to the speech data.
[0088] The speech data to be identified can be speech data in any language, and can be any form of data recorded or transmitted by speech. It can be audio files, such as songs or audiobooks, or even everyday conversation data, etc. There are no restrictions here.
[0089] Among them, the speech feature parameters can be any data that can reflect the frequency domain characteristics of speech data, such as Mel Frequency Cepstrum Coefficient (MFCC), etc., without limitation.
[0090] It should be noted that although the grammar and rhythm of different languages may vary greatly, the rules of human pronunciation are consistent. That is, as a single species, humans share universal rules regarding vocal cord vibration and oral cavity closure, and these rules can be directly reflected in the frequency domain characteristics of sound. Therefore, in this disclosure, any data that can characterize the frequency domain characteristics corresponding to speech data, such as MFCC, can be used as a speech feature parameter to characterize the frequency domain characteristics of sound.
[0091] Understandably, since this speech feature parameter conforms to the auditory perception characteristics of human ears regarding sound frequencies, it can enhance speech features to a certain extent and suppress non-speech features.
[0092] Optionally, the speech feature parameters can also be linear predictive cepstral coefficients (LPCC) or perceptual linear predictive coefficients (PLP), etc., which are not limited here.
[0093] Specifically, taking MFCC as an example, when preprocessing the speech data to be recognized, it can be processed by pre-emphasis, framing, windowing, etc., followed by discrete Fourier transform and Mel filtering, and finally cepstral and energy difference to obtain speech feature parameters.
[0094] Step 102: Using a pre-trained style recognition model, identify the speech feature parameters to determine the style feature vector corresponding to the speech data.
[0095] It should be noted that because different people have different speaking styles, the speech data generated by different people for the same text may have different styles; or, the same person may interpret the same text differently depending on their mood, resulting in different styles of speech data generated by the same person for the same text under different circumstances. To reduce the difficulties caused by such problems to speech recognition, this disclosure introduces a style recognition model to predict the style features corresponding to the speech data, so that the style feature vector output by the style recognition model can be used to further improve the accuracy of speech recognition.
[0096] The style recognition model can be a pre-trained model, that is, a pre-trained model.
[0097] It should be noted that the style feature vector can be composed of one or more sub-vectors, and each sub-vector can implicitly represent any speaking style, such as tone, speaking speed, intonation, emphasis, etc., or it can be an organic combination of one or more of the above speaking style factors, which is not limited here.
[0098] Optionally, in this disclosure, the style recognition model can be trained in an unsupervised manner. For example, a style vector dictionary can be generated, and then the model can be trained using multiple unsupervised data, so that each sub-vector in the style vector dictionary can autonomously absorb different speaking styles.
[0099] Step 103: Based on the style feature vector, the speech recognition model generated in the pre-trained manner is used to recognize the speech data in order to generate the recognition result corresponding to the speech data.
[0100] This disclosure introduces style feature vectors, meaning that the style feature vectors output by the style recognition model can be used to guide speech data recognition, thereby improving recognition accuracy. For example, style feature vectors can be fused with speech feature parameters, and then a pre-trained speech recognition model can recognize the fused result to generate a recognition result corresponding to the speech data. Thus, the speech recognition model can better identify the content information in the speech data based on the style feature vectors, thereby making the recognition results more accurate and reliable.
[0101] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Finally, based on the style feature vector, the pre-trained speech recognition model is used to recognize the speech data to generate the recognition result corresponding to the speech data. Therefore, by recognizing the speech data based on the style feature vector corresponding to the speech recognition data during the speech recognition process, the influence of speaking style on speech recognition is avoided, improving the accuracy and reliability of the speech recognition results.
[0102] Figure 2 This is a flowchart illustrating another speech recognition method provided according to an embodiment of the present disclosure.
[0103] like Figure 2 As shown, the speech recognition method may include the following steps:
[0104] Step 201: Preprocess the speech data to be recognized to determine the speech feature parameters corresponding to the speech data.
[0105] Step 202: Using a pre-trained style recognition model, identify the speech feature parameters to determine the style feature vector corresponding to the speech data.
[0106] It should be noted that the specific implementation methods of steps 201 and 202 can refer to the above embodiments, and will not be repeated here.
[0107] Step 203: According to preset rules, the style feature vector and speech feature parameters are fused to generate the vector to be recognized.
[0108] In this disclosure, the style recognition model can output a single overall style feature vector for speech data, or it can output a corresponding style feature vector for each frame of audio in the speech data. That is, the style feature vector output by the style recognition model can include a first sub-vector representing the frequency domain features of the overall speech data, or it can include multiple first sub-vectors representing the frequency domain features of each frame of audio. Accordingly, in this disclosure, the style feature vector can be fused with speech feature parameters using different rules based on the number and dimension of the first sub-vectors in the style feature vector.
[0109] Optionally, if the style feature vector contains one first sub-vector, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameter to be fused, where the audio vector is the vector corresponding to the speech feature parameter, then the style feature vector and the speech feature parameter can be fused using any of the following preset rules:
[0110] Place the style feature vector after the audio vector to be fused;
[0111] The style feature vectors are placed after each second sub-vector in the audio vector to be fused;
[0112] The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and,
[0113] The style feature vector is added to each second sub-vector in the audio vector to be fused.
[0114] It is understood that in this disclosure, the audio vector to be fused contains the vector corresponding to each frame of audio in the speech data, that is, it can be regarded as a time series, and this time series consists of multiple second sub-vectors of a certain dimension. It is understood that each second sub-vector corresponds to a frame of audio in the speech data.
[0115] Example 1: If the dimensions of the first and second subvectors are the same, the first subvector has a quantity of 1, denoted as K1, and the second subvector has a quantity of 5, denoted as A1, A2, A3, A4, A5 respectively in chronological order. Figure 2a As shown, the style feature vector can be placed after the audio vector to be fused to generate the current fused vector to be identified: [A1, A2, A3, A4, A5, K1].
[0116] Example 2: If the dimensions of the first and second subvectors are the same, the first subvector has a quantity of 1, denoted as K1, and the second subvector has a quantity of 3, denoted as A1, A2, and A3 respectively in chronological order. Figure 2bAs shown, the style feature vector can be placed after each second sub-vector in the audio vector to be fused, thereby generating the current fused vector to be identified: [A1, K1, A2, K1, A3, K1].
[0117] Example 3: If the dimensions of the first and second subvectors are the same, the first subvector has a quantity of 1, denoted as K1, and the second subvector has a quantity of 3, denoted as A1, A2, and A3 respectively in chronological order. Figure 2c As shown, the style feature vector can be concatenated with each second sub-vector in the audio vector to be fused to generate the current fused vector to be identified: [A1K1, A2K1, A3K1].
[0118] Example 4: If the dimensions of the first and second subvectors are the same, the first subvector has a quantity of 1, denoted as K1, and the second subvector has a quantity of 3, denoted as A1, A2, and A3 respectively in chronological order. Figure 2d As shown, the style feature vector can be added to each second sub-vector in the audio vector to be fused, thereby generating the current fused vector to be identified: [A1+K1, A2+K1, A3+K1].
[0119] It should be noted that the above examples are merely illustrative of this disclosure and are not intended to limit this disclosure.
[0120] Optionally, if the number of first sub-vectors in the style feature vector is the same as the number of second sub-vectors in the audio vector to be fused, and the dimensions of the first sub-vector and the second sub-vector are the same, then the style feature vector and speech feature parameters can be fused according to any of the following preset rules:
[0121] Place the style feature vector after the audio vector to be fused;
[0122] Place each first sub-vector after its corresponding second sub-vector;
[0123] Each first subvector is concatenated with its corresponding second subvector; and...
[0124] Add each first subvector to its corresponding second subvector.
[0125] Based on the aforementioned pre-defined rules, the following examples are provided for illustrative purposes and are not intended to limit the scope of this disclosure.
[0126] Example 5: If the dimensions of the first and second subvectors are the same, the number of the first subvectors is 3, and they are denoted as M1, M2, M3 in chronological order. The number of the second subvectors is also 3, and they are denoted as N1, N2, N3 in chronological order. Figure 2eAs shown, the style feature vector can be placed after the audio vector to be fused to generate the current fused vector to be identified: [N1, N2, N3, M1, M2, M3].
[0127] Example 6: If the dimensions of the first and second subvectors are the same, the number of the first subvectors is 3, and they are denoted as M1, M2, M3 in chronological order. The number of the second subvectors is also 3, and they are denoted as N1, N2, N3 in chronological order. Figure 2f As shown, the style feature vector can be placed after the audio vector to be fused to generate the current fused vector to be identified: [N1, M1, N2, M2, N3, M3].
[0128] Example 7: If the dimensions of the first and second subvectors are the same, the number of the first subvectors is 3, and they are denoted as M1, M2, M3 in chronological order. The number of the second subvectors is also 3, and they are denoted as N1, N2, N3 in chronological order. Figure 2g As shown, the style feature vector can be placed after the audio vector to be fused to generate the current fused vector to be identified: [N1M1, N2M2, N3M3].
[0129] Example 8: If the dimensions of the first and second subvectors are the same, the number of the first subvectors is 3, and they are denoted as M1, M2, M3 in chronological order. The number of the second subvectors is 3, and they are denoted as N1, N2, N3 in chronological order. Figure 2h As shown, the style feature vector can be placed after the audio vector to be fused, thereby generating the current fused vector to be identified: [N1+M1, N2+M2, N3+M3].
[0130] Optionally, if the dimension of the style feature vector is different from the dimension of the audio vector to be fused, the style feature vector and the audio vector can be concatenated to generate the vector to be identified.
[0131] Example 9: If the dimension of the first subvector is h1, the dimension of the second subvector is h2, and h1 is different from h2, the number of the first subvector is 3, and they are denoted as P1, P2, P3 in chronological order; the number of the second subvector is 3, and they are denoted as Q1, Q2, Q3 in chronological order. Figure 2i As shown, the style feature vector can be concatenated with the audio vector to generate the current fused vector to be identified: [P1Q1, P2Q2, P3Q3].
[0132] It should be noted that, through the above fusion methods, the vector to be identified can simultaneously contain both speech style features and audio features.
[0133] Step 204: Use a speech recognition model to recognize the vectors to generate the recognition results corresponding to the speech data.
[0134] It should be noted that since the vector is generated by fusing style feature vectors and audio vectors, the speaking style information corresponding to the speech data can be incorporated into the recognition process as a reference. This allows the speech recognition model to better and more accurately recognize speech data based on the information in the style feature vectors. Optionally, an attention mechanism and a decoder module can be combined to decode the encoded information corresponding to the vectors, thereby improving the performance of the speech recognition model.
[0135] Figure 3 This diagram illustrates the overall architecture of a speech recognition model, as follows: Figure 3 As shown, audio, i.e., speech data, can be input into the recognition encoder and the speech style encoder respectively. Then, an attention module is used to calculate the speech style features. The speech style features (style embedding) and audio are then input into the recognition encoder together. After that, the encoder state module and the attention module are used for calculation. Finally, the recognition result is output through the recognition decoder.
[0136] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Next, the style feature vector and the speech feature parameters are fused according to preset rules to generate a vector to be recognized. Finally, the speech recognition model is used to recognize the vector to generate the recognition result corresponding to the speech data. Therefore, by fusing the style feature vector and speech feature parameters corresponding to the speech data according to preset rules, the speech recognition model can recognize and decode the speech data based on the speaking style as much as possible, thereby improving the accuracy and effectiveness of speech recognition.
[0137] Figure 4 This is a flowchart illustrating another speech recognition method provided according to an embodiment of the present disclosure.
[0138] like Figure 4 As shown, the speech recognition method may include the following steps:
[0139] Step 301: Preprocess the speech data to be recognized to determine the speech feature parameters corresponding to the speech data.
[0140] Step 302: Using a pre-trained style recognition model, identify the speech feature parameters to determine the style feature vector corresponding to the speech data.
[0141] It should be noted that the specific implementation methods of steps 301 and 302 can refer to the above embodiments, and will not be repeated here.
[0142] Step 303: The speech feature parameters are encoded using the encoder in the speech recognition model to generate the audio vector corresponding to the speech data.
[0143] It should be noted that in this disclosure, the speech feature parameters can be downsampled in the temporal domain using a convolutional neural network in the encoder, thereby reducing computational complexity. The encoder can consist of an input layer, an output layer, and multiple hidden state layers.
[0144] Optionally, a transformer structure can be used in the encoder, which can improve the parallelism of model computation and alleviate the problem of information loss during sequential computation.
[0145] It is understandable that encoding speech feature parameters to generate audio vectors corresponding to speech data can provide data support for the subsequent fusion of style feature vectors and audio vectors.
[0146] Step 304: According to preset rules, the style feature vector and the audio vector are fused to generate the first vector.
[0147] The first vector can be the vector to be identified, which is obtained by fusing the style feature vector and the audio vector.
[0148] It should be noted that the process of fusing style feature vectors and audio vectors according to preset rules in this embodiment can refer to the above embodiment, and will not be repeated here.
[0149] Step 305: Use the recognition module in the speech recognition model to recognize the first vector to generate the recognition result corresponding to the speech data.
[0150] It should be noted that the process of using the recognition module in the speech recognition model to recognize the first vector to generate the recognition result corresponding to the speech data in this embodiment can refer to the above embodiment, and will not be repeated here.
[0151] Figure 5 This shows the overall architecture diagram of another speech recognition model, such as Figure 5As shown, audio, i.e., speech data, can be input into the recognition encoder and the speech style encoder respectively. Then, the attention module is used to calculate the speech style features, i.e., the style feature vector. This vector, along with the data output from the recognition encoder, is then input into the encoding state module. Finally, the recognition result is output through the calculation of the attention module and the recognition decoder.
[0152] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Next, the encoder in the speech recognition model encodes the speech feature parameters to generate an audio vector corresponding to the speech data. Then, according to preset rules, the style feature vector and the audio vector are fused to generate a first vector. Finally, the recognition module in the speech recognition model recognizes the first vector to generate the recognition result corresponding to the speech data. Therefore, by fusing the style feature vector and the audio vector according to preset rules before using the speech recognition model for recognition, the style feature vector can guide the entire recognition process of the speech recognition model, resulting in higher accuracy and reliability of the recognition results.
[0153] Figure 6 This is a flowchart illustrating another speech recognition method provided according to an embodiment of the present disclosure.
[0154] like Figure 6 As shown, the speech recognition method may include the following steps:
[0155] Step 401: Preprocess the speech data to be recognized to determine the speech feature parameters corresponding to the speech data.
[0156] Step 402: Using a pre-trained style recognition model, identify the speech feature parameters to determine the style feature vector corresponding to the speech data.
[0157] It should be noted that the specific implementation methods of steps 401 and 402 can refer to the above embodiments, and will not be repeated here.
[0158] Step 403: According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector.
[0159] The second vector can be a vector generated by fusing the style feature vector with the hidden state vector output from any hidden state layer in the encoder. The hidden state vector can be a vector output from any hidden state layer.
[0160] The hidden state layer can consist of a series of convolutional layers, pooling layers, and fully connected layers. In this disclosure, the style feature vector can be fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model, such as a pooling layer or a fully connected layer, without limitation.
[0161] It should be noted that the process of fusing the style feature vector with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model according to the preset rules in this embodiment can refer to any of the above embodiments, and will not be repeated here.
[0162] In this embodiment, since the hidden state vector is a vector that has undergone downsampling, by fusing the style feature vector with any hidden state vector, not only can the style feature vector guide the recognition process of the speech recognition model, but the amount of data processed by the speech recognition model is also minimized.
[0163] Step 404: Encode the second vector using the remaining layers in the encoder to obtain the third vector.
[0164] The remaining layers can be the network layers outside the hidden state layers in the encoder. By using the remaining layers to encode the second vector, the third vector to be identified can be obtained.
[0165] Step 405: Use the recognition module in the speech recognition model to recognize the third vector to generate the recognition result corresponding to the speech data.
[0166] It should be noted that the process of using the recognition module in the speech recognition model to recognize the third vector to generate the recognition result corresponding to the speech data in this embodiment can refer to any of the above embodiments, and will not be repeated here.
[0167] Figure 7 This shows an overall architecture diagram of another speech recognition model, such as Figure 7 As shown, audio, i.e., speech data, can be input into the recognition encoder and the speech style encoder respectively. Then, the attention module is used to calculate the speech style features, i.e., the style feature vector, and then input it together with the speech data into the recognition encoder. After passing through the encoding state module and the attention module for calculation, the recognition result is finally output through the recognition decoder.
[0168] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Next, according to preset rules, the style feature vector is fused with the hidden state vector output from any hidden state layer in the encoder of the speech recognition model to generate a second vector. The remaining layers in the encoder are used to encode the second vector to obtain a third vector. Finally, the recognition module in the speech recognition model is used to recognize the third vector to generate the recognition result corresponding to the speech data. Therefore, by fusing the style feature vector with the hidden state vector output from any hidden state layer in the encoder of the speech recognition model, not only can the style feature vector guide the recognition process of the speech recognition model, improving the accuracy and reliability of the speech recognition result, but the amount of data processed by the speech recognition model is also minimized.
[0169] To implement the above embodiments, this disclosure also proposes a voice recognition device.
[0170] Figure 8 This is a schematic diagram of the structure of a voice recognition device provided in an embodiment of the present disclosure.
[0171] like Figure 8 As shown, the voice recognition device 800 includes a first determining module 810, a second determining module 820, and a recognition module 830:
[0172] The first determining module 810 is used to preprocess the speech data to be recognized in order to determine the speech feature parameters corresponding to the speech data.
[0173] The second determining module 820 is used to identify the speech feature parameters using a pre-trained style recognition model, so as to determine the style feature vector corresponding to the speech data.
[0174] The recognition module 830 is used to recognize the speech data based on the style feature vector using a pre-trained speech recognition model, so as to generate a recognition result corresponding to the speech data.
[0175] Optionally, the identification module is specifically used for:
[0176] According to preset rules, the style feature vector and the speech feature parameters are fused to generate a vector to be recognized;
[0177] The speech recognition model is used to recognize the vector to generate a recognition result corresponding to the speech data.
[0178] Optionally, the identification module is specifically used for:
[0179] The speech feature parameters are encoded using the encoder in the speech recognition model to generate the audio vector corresponding to the speech data;
[0180] According to preset rules, the style feature vector and the audio vector are fused to generate a first vector;
[0181] The first vector is identified using the recognition module in the speech recognition model to generate the recognition result corresponding to the speech data.
[0182] Optionally, the identification module is specifically used for:
[0183] According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector.
[0184] The second vector is recognized using the recognition module in the speech recognition model to generate the recognition result corresponding to the speech data.
[0185] Optionally, the style feature vector contains one first sub-vector, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameters to be fused. The preset rule includes any one of the following:
[0186] The style feature vector is placed after the audio vector to be fused;
[0187] The style feature vector is placed after each second sub-vector in the audio vector to be fused;
[0188] The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and,
[0189] The style feature vector is added to each second sub-vector in the audio vector to be fused.
[0190] Optionally, the first sub-vectors included in the style feature vector have the same number of second sub-vectors as the audio vector to be fused, and the dimensions of the first sub-vectors are the same as the dimensions of the second sub-vectors. The preset rule includes any one of the following:
[0191] The style feature vector is placed after the audio vector to be fused;
[0192] Each of the first sub-vectors is placed after the corresponding second sub-vector;
[0193] Each of the first sub-vectors is concatenated with its corresponding second sub-vector; and...
[0194] Each of the first sub-vectors is added to its corresponding second sub-vector.
[0195] Optionally, the dimension of the style feature vector is different from the dimension of the audio vector to be fused, and the preset rule is to concatenate the style feature vector with the audio vector.
[0196] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Finally, based on the style feature vector, the pre-trained speech recognition model is used to recognize the speech data to generate the recognition result corresponding to the speech data. Therefore, by recognizing the speech data based on the style feature vector corresponding to the speech recognition data during the speech recognition process, the influence of speaking style on speech recognition is avoided, improving the accuracy and reliability of the speech recognition results.
[0197] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0198] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0199] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0200] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0201] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as speech recognition methods. For example, in some embodiments, the speech recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the speech recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform speech recognition methods by any other suitable means (e.g., by means of firmware).
[0202] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0203] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0204] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0205] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0206] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0207] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0208] In this embodiment, the speech data to be recognized is first preprocessed to determine the speech feature parameters corresponding to the speech data. Then, a pre-trained style recognition model is used to recognize the speech feature parameters to determine the style feature vector corresponding to the speech data. Finally, based on the style feature vector, the pre-trained speech recognition model is used to recognize the speech data to generate the recognition result corresponding to the speech data. Therefore, by recognizing the speech data based on the style feature vector corresponding to the speech recognition data during the speech recognition process, the influence of speaking style on speech recognition is avoided, improving the accuracy and reliability of the speech recognition results.
[0209] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0210] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A speech recognition method, characterized in that, include: The speech data to be recognized is preprocessed to determine the speech feature parameters corresponding to the speech data; The pre-trained style recognition model is used to identify the speech feature parameters in order to determine the style feature vector corresponding to the speech data. Based on the style feature vector, the speech data is recognized using a pre-trained speech recognition model to generate the recognition result corresponding to the speech data; Based on the style feature vector, a pre-trained speech recognition model is used to recognize the speech data to generate a recognition result corresponding to the speech data, including: According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector. The second vector is encoded using the remaining layers in the encoder to obtain the third vector; The recognition module in the speech recognition model is used to recognize the third vector in order to generate the recognition result corresponding to the speech data.
2. The method as described in claim 1, characterized in that, The style feature vector contains a first sub-vector with a quantity of 1, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameters to be fused. The preset rule includes any one of the following: The style feature vector is placed after the audio vector to be fused; The style feature vector is placed after each second sub-vector in the audio vector to be fused; The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and, The style feature vector is added to each second sub-vector in the audio vector to be fused.
3. The method as described in claim 1, characterized in that, The first sub-vector in the style feature vector has the same number of sub-vectors as the second sub-vector in the audio vector to be fused, and the dimensions of the first sub-vector and the second sub-vector are the same. The preset rule includes any one of the following: The style feature vector is placed after the audio vector to be fused; Each of the first sub-vectors is placed after the corresponding second sub-vector; Each of the first sub-vectors is concatenated with its corresponding second sub-vector; and... Each of the first sub-vectors is added to its corresponding second sub-vector.
4. The method as described in claim 1, characterized in that, The dimension of the style feature vector is different from the dimension of the audio vector to be fused. The preset rule is to concatenate the style feature vector with the audio vector.
5. A voice recognition device, characterized in that, include: The first determining module is used to preprocess the speech data to be recognized in order to determine the speech feature parameters corresponding to the speech data; The second determining module uses a pre-trained style recognition model to identify the speech feature parameters in order to determine the style feature vector corresponding to the speech data. The recognition module is used to recognize the speech data based on the style feature vector using a pre-trained speech recognition model, so as to generate a recognition result corresponding to the speech data. The identification module is specifically used for: According to preset rules, the style feature vector is fused with the hidden state vector output by any hidden state layer in the encoder of the speech recognition model to generate a second vector. The second vector is encoded using the remaining layers in the encoder to obtain the third vector; The recognition module in the speech recognition model is used to recognize the third vector in order to generate the recognition result corresponding to the speech data.
6. The apparatus according to claim 5, characterized in that, The style feature vector contains a first sub-vector with a quantity of 1, and the dimension of the first sub-vector is the same as the dimension of the second sub-vector in the audio vector corresponding to the speech feature parameters to be fused. The preset rule includes any one of the following: The style feature vector is placed after the audio vector to be fused; The style feature vector is placed after each second sub-vector in the audio vector to be fused; The style feature vector is concatenated with each second sub-vector in the audio vector to be fused; and, The style feature vector is added to each second sub-vector in the audio vector to be fused.
7. The apparatus according to claim 5, characterized in that, The first sub-vector in the style feature vector has the same number of sub-vectors as the second sub-vector in the audio vector to be fused, and the dimensions of the first sub-vector and the second sub-vector are the same. The preset rule includes any one of the following: The style feature vector is placed after the audio vector to be fused; Each of the first sub-vectors is placed after the corresponding second sub-vector; Each of the first sub-vectors is concatenated with its corresponding second sub-vector; and... Each of the first sub-vectors is added to its corresponding second sub-vector.
8. The apparatus according to claim 5, characterized in that, The dimension of the style feature vector is different from the dimension of the audio vector to be fused. The preset rule is to concatenate the style feature vector with the audio vector.
9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.
11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
Text representation method and device
CN111581335A
Voice recognition device and method, electronic equipment and computer readable storage medium
CN111862944A