Speech recognition method, terminal device and storage medium
Patent Information
- Application Number
- CN202310688637.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-06-09
AI Technical Summary
[0003]然而,常见的语音识别方法在一些场景复杂多样、音频信噪比低等音频场景中,并不能精准地识别出说话人的年龄信息,存在语音识别准确性低的问题,从而降低了终端设备的智能性
[0014]本申请实施例提供了一种语音识别方法,终端设备及存储介质,终端设备对待识别语音进行声纹特征提取,确定待识别语音对应的第一声纹信息;将待识别语音和第一声纹信息输入至语音识别模型中,获得待识别语音对应的识别结果;其中,语音识别模型是通过增强后的语音数据对自注意力网络进行训练获得的。也就是说,在本申请的实施例中,终端设备在进行语音识别的过程中融合了语音的声纹特征,提升了语音识别的准确率,同时,终端设备进行语音识别处理所使用的语音识别模型是通过增强后的语音数据训练的,能够有效提升该语音识别模型的鲁棒性,使得语音识别模型能够更好的应用于复杂的音频场景中,且语音识别模型可以通过自注意力网络来充分利用语音的上下文信息,进一步提升了语音识别的准确率。由此可见,在本申请的实施例中,基于声纹特征和语音识别模型的结合使用,能够精准地识别出说话人的年龄信息,大大提升了语音识别准确性,从而提高了终端设备的智能性。
Smart Images

Figure CN119107955B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a speech recognition method, terminal device, and storage medium. Background Technology
[0002] Speaker Age Estimation (SAE) is widely considered a subproblem of speech attribute recognition, which estimates the speaker's age using acquired audio data.
[0003] However, common speech recognition methods cannot accurately identify the speaker's age in some complex and diverse audio scenarios with low signal-to-noise ratios, resulting in low speech recognition accuracy and thus reducing the intelligence of terminal devices. Summary of the Invention
[0004] This application provides a speech recognition method, a terminal device, and a storage medium, which can improve the accuracy of speech recognition and thus enhance the intelligence of the terminal device.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a speech recognition method, the method comprising:
[0007] Voiceprint features are extracted from the speech to be recognized to determine the first voiceprint information corresponding to the speech to be recognized;
[0008] The speech to be recognized and the first voiceprint information are input into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data.
[0009] Secondly, embodiments of this application provide a terminal device, the terminal device comprising: a determining unit, an acquiring unit, and...
[0010] The determining unit is used to extract voiceprint features from the speech to be recognized and determine the first voiceprint information corresponding to the speech to be recognized.
[0011] The acquisition unit is used to input the speech to be recognized and the first voiceprint information into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data.
[0012] Thirdly, embodiments of this application provide a terminal device, which includes a processor and a memory storing processor-executable instructions. When the instructions are executed by the processor, the method described in the first aspect above is implemented.
[0013] Fourthly, embodiments of this application provide a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.
[0014] This application provides a speech recognition method, a terminal device, and a storage medium. The terminal device extracts voiceprint features from the speech to be recognized to determine the first voiceprint information corresponding to the speech. The speech to be recognized and the first voiceprint information are input into a speech recognition model to obtain the recognition result corresponding to the speech. The speech recognition model is obtained by training a self-attention network with enhanced speech data. In other words, in this application, the terminal device integrates voiceprint features during speech recognition, improving the accuracy of speech recognition. Furthermore, the speech recognition model used by the terminal device is trained with enhanced speech data, effectively improving its robustness and enabling it to be better applied to complex audio scenarios. The speech recognition model can also fully utilize the contextual information of the speech through a self-attention network, further improving the accuracy of speech recognition. Therefore, in this application, the combined use of voiceprint features and a speech recognition model can accurately identify the speaker's age information, greatly improving speech recognition accuracy and thus enhancing the intelligence of the terminal device. Attached Figure Description
[0015] Figure 1 A schematic diagram of the implementation process of the speech recognition method. Figure 1 ;
[0016] Figure 2 Schematic diagram of the voiceprint recognition network Figure 1 ;
[0017] Figure 3 Schematic diagram of the voiceprint recognition network Figure 2 ;
[0018] Figure 4 Schematic diagram of the speech recognition model Figure 1 ;
[0019] Figure 5 Schematic diagram of the speech recognition model Figure 2 ;
[0020] Figure 6 This is a schematic diagram of a dot product attention mechanism network;
[0021] Figure 7 A schematic diagram of a multi-head self-attention mechanism network;
[0022] Figure 8 This is a schematic diagram illustrating the process of using a speech recognition model;
[0023] Figure 9 This is a schematic diagram of the training process of a speech recognition model;
[0024] Figure 10 A schematic diagram illustrating the application scenarios of a speech recognition solution;
[0025] Figure 11 This is a diagram illustrating the implementation framework of a speech recognition method.
[0026] Figure 12 This is a schematic diagram of the structure of the speech recognition model proposed in the embodiments of this application;
[0027] Figure 13 This is a schematic diagram of the composition structure of the terminal device proposed in the embodiments of this application. Figure 1 ;
[0028] Figure 14 This is a schematic diagram of the composition structure of the terminal device proposed in the embodiments of this application. Figure 2 . Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant application and not for limiting the application. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant application are shown in the accompanying drawings.
[0030] With the rapid development of mobile terminal technology, consumers' demands for mobile photo albums are also shifting towards intelligence. Therefore, how to fully utilize the audio information in photo album videos to make photo album editing technology more intelligent is a direction worth exploring.
[0031] Speaker Age Estimation (SAE) is widely considered a subproblem of speech attribute recognition, which estimates the age (child, middle-aged, elderly) of a speaker using acquired audio data.
[0032] Common SAE (Speaker Age Estimation) techniques are primarily based on traditional machine learning (non-deep learning) methods. For example, using Support Vector Machines (SVM) for speaker age estimation; borrowing the ivector method from Speaker Recognition (SR) is also a commonly used technique. Additionally, since speakers of different ages exhibit variations in vocal cord coefficients and formant frequencies, these acoustic features can also be utilized for classification. These methods require less training data, offer fast recognition speeds, and demonstrate good accuracy in scenarios with high signal-to-noise ratios.
[0033] However, in the scenario of user photo album videos, the presence of a large number of low signal-to-noise ratio scenes will cause a significant drop in the accuracy of age estimation. Therefore, the above method cannot be applied to creating wonderful memories of children in user videos.
[0034] With the development of deep learning technology, in the field of SAE (Speech Engineering), methods based on deep neural networks (DNNs) have shown superior performance compared to traditional machine learning methods. Currently, mainstream SAE systems use speech acoustic features (such as FBANK (Filter Bank) features) as input, and then train a neural network model. This generally achieves good results in well-matched test scenarios. For example, generative adversarial networks (GANs) can be used for age recognition, or residual networks (ResNets) can be used for speaker age estimation.
[0035] However, the following problems remain unresolved in user photo album videos: 1. User video scenes vary widely, and some user videos have very low signal-to-noise ratios, which places high demands on the robustness of the SAE model; 2. Existing technologies directly extract acoustic features from speech and then train neural network models, ignoring the voiceprint characteristics of speech, while voiceprint information is helpful for SAE estimation; 3. Existing SAEs based on deep neural network technology have difficulty in modeling the temporal information of a speech segment well, which also leads to a decrease in the accuracy of age estimation.
[0036] In other words, common traditional machine learning or deep learning technologies like SAE (Search Engine Effect) solutions cannot be effectively applied to scenarios such as user photo albums. This is because user photo album videos are characterized by complex and diverse scenes and low audio signal-to-noise ratios. Therefore, SAE technology needs to be highly robust, able to handle various scenarios and low signal-to-noise ratios. Furthermore, the SAE module's framework for creating memorable childhood voices requires extremely high accuracy in determining whether a speech segment is a child's voice. Current mainstream SAE technologies struggle to meet the accuracy requirements for age estimation in creating memorable childhood voices.
[0037] It is evident that traditional machine learning-based SAE methods suffer from a significant drop in recognition rate in low signal-to-noise ratio (SNR) scenarios, making it difficult to achieve good results in user photo album video scenarios. Furthermore, current mainstream deep learning-based SAE methods also fail to obtain accurate age estimation results in user photo album video scenarios, as well as in scenarios with diverse speech signal variations and user videos exhibiting very low SNR.
[0038] In other words, common speech recognition methods cannot accurately identify the speaker's age in some complex and diverse audio scenarios with low signal-to-noise ratios, resulting in low speech recognition accuracy and thus reducing the intelligence of terminal devices.
[0039] To address the aforementioned issues, in the embodiments of this application, the terminal device extracts voiceprint features from the speech to be recognized, determining the first voiceprint information corresponding to the speech; the speech to be recognized and the first voiceprint information are input into a speech recognition model to obtain the recognition result corresponding to the speech; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data. In other words, in the embodiments of this application, the terminal device integrates voiceprint features during speech recognition, improving the accuracy of speech recognition. Simultaneously, the speech recognition model used by the terminal device for speech recognition processing is trained with enhanced speech data, effectively improving the robustness of the speech recognition model and enabling it to be better applied to complex audio scenarios. Furthermore, the speech recognition model can fully utilize the contextual information of the speech through a self-attention network, further improving the accuracy of speech recognition. Therefore, in the embodiments of this application, the combined use of voiceprint features and a speech recognition model can accurately identify the speaker's age information, greatly improving speech recognition accuracy and thus enhancing the intelligence of the terminal device.
[0040] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0041] One embodiment of this application provides a speech recognition method that can be applied to terminal devices. Figure 1 A schematic diagram of the implementation process of the speech recognition method. Figure 1 ,like Figure 1 As shown, the method for terminal devices to perform voice recognition may include the following steps:
[0042] Step 101: Extract voiceprint features from the speech to be recognized and determine the first voiceprint information corresponding to the speech to be recognized.
[0043] In the embodiments of this application, the terminal device can first extract voiceprint features from the speech to be recognized, thereby determining the first voiceprint information corresponding to the speech to be recognized.
[0044] It should be noted that, in the embodiments of this application, the speech to be recognized can be speech information of any format, and this application does not impose any specific limitations.
[0045] It should be noted that, in the embodiments of this application, the terminal device can acquire the speech to be recognized in any way, and this application does not impose any specific limitations.
[0046] For example, in some embodiments, the speech to be recognized can be speech information extracted from video by the terminal device, speech information acquired by the terminal device through a speech acquisition device, or speech information sent by other devices received by the terminal device.
[0047] Furthermore, in the embodiments of the application, the terminal device can use a voiceprint recognition network to perform voiceprint recognition processing on the speech to be recognized, thereby determining the first voiceprint information corresponding to the speech to be recognized.
[0048] It should be noted that, in the embodiments of this application, the voiceprint recognition network is used to extract voiceprint features from the speech to be recognized, and finally obtain the corresponding first voiceprint information used to characterize the voiceprint features.
[0049] It is understood that, in the embodiments of this application, the voiceprint recognition network can be understood as a voiceprint recognition system based on a neural network. The voiceprint recognition network can be a ResNet network structure, wherein the voiceprint recognition network includes a frame-level feature learning layer, a feature aggregation layer, and a sentence-level feature learning layer.
[0050] It should be noted that, in the embodiments of this application, the first voiceprint information obtained based on the voiceprint recognition network can be a feature vector that can characterize the voiceprint features of the speech to be recognized. For example, the first voiceprint information can be an r-vector.
[0051] Exemplary, in some embodiments, Figure 2 Schematic diagram of the voiceprint recognition network Figure 1 ,like Figure 2 As shown, the voiceprint recognition network can adopt a ResNet-based network structure, which can be divided into a frame-level deep feature learning layer, a feature aggregation layer, and a segment-level embedding learning layer. The voiceprint recognition network is trained using a speaker-discriminatory training criterion. Speaker embedding vectors are extracted at the segment-level layer, and the final extracted speaker embedding vector is simply called the r-vector, which represents the voiceprint information.
[0052] Exemplary, in some embodiments, Figure 3 Schematic diagram of the voiceprint recognition network Figure 2 ,like Figure 3 As shown, taking a 34-layer ResNet network structure for voiceprint recognition as an example, the first part is a Conv2d layer, which increases the number of input feature channels while keeping the time and frequency dimensions unchanged. The second part is a stacked 32-layer Conv2d layer (ResNetBlock-1, ResNetBlock-2, ResNetBlock-3, ResNetBlock-4), which performs convolutions on the time and frequency dimensions to extract high-dimensional speech features. The third part consists of StatsPooling and Flatten layers, which are used to extract the mean and variance at the sentence level, converting the frame-level feature representation into a sentence-level feature representation. The fourth part consists of fully connected neural networks Dense1 and Dense2. Dense1 reduces the sentence-level feature representation to 256 dimensions, which is the vector representation r-vector of the voiceprint features. Dense2 maps the r-vector to the final number of speakers N and uses the softmax function to obtain the posterior probability representation of speaker classification.
[0053] Step 102: Input the speech to be recognized and the first voiceprint information into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training the self-attention network with the enhanced speech data.
[0054] In the embodiments of this application, after the terminal device extracts the voiceprint features of the speech to be recognized and determines the first voiceprint information corresponding to the speech to be recognized, it can further input the speech to be recognized and the first voiceprint information into the speech recognition model, thereby obtaining the recognition result corresponding to the speech to be recognized.
[0055] It should be noted that, in the embodiments of this application, the speech recognition model can be obtained by training a self-attention based network (SAN) using enhanced speech data. The enhanced speech data enables the trained speech recognition model to achieve higher robustness, and the self-attention network has a strong ability to model the temporal information of speech, thus improving the speech recognition accuracy of the model.
[0056] In other words, in the embodiments of this application, the terminal device can use the enhanced speech data to train the initial recognition model composed of the self-attention network, and finally obtain the corresponding speech recognition model.
[0057] Furthermore, in the embodiments of this application, the trained speech recognition model may include a first self-attention network and a second self-attention network. Accordingly, when the terminal device inputs the speech to be recognized and the first voiceprint information into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized, it can first extract attribute features of the speech to be recognized through the first self-attention network to obtain the first attribute information corresponding to the speech to be recognized; then, based on the first attribute information and the first voiceprint information, determine the fused feature information corresponding to the speech to be recognized; finally, through the second self-attention network and the fused feature information, determine the recognition result corresponding to the speech to be recognized.
[0058] In other words, in the embodiments of this application, the first self-attention network included in the speech recognition model can be used to extract attribute features corresponding to the speech to be recognized. Specifically, the first attribute information corresponding to the speech to be recognized can determine the age attribute corresponding to the speech.
[0059] It should be noted that, in the embodiments of this application, the first attribute information obtained based on the first self-attention network can be a feature vector that can characterize the age features of the speech to be recognized. For example, the first attribute information can be an a-vector.
[0060] It is understood that, in the embodiments of this application, the second self-attention network included in the speech recognition model can be used to further extract features from the attribute information of the speech to be recognized after the fusion of the first attribute information and the first voiceprint information.
[0061] It should be noted that, in the embodiments of this application, the information input into the speech recognition model includes, in addition to the first attribute information representing the age characteristics of the speech to be recognized, the first voiceprint information representing the voiceprint characteristics of the speech to be recognized. By fusing the voiceprint feature information into the speech recognition model, the speech recognition model can learn the voiceprint features in the speech at the same time, thereby improving the accuracy of speech recognition.
[0062] Exemplary, in some embodiments, Figure 4 Schematic diagram of the speech recognition model Figure 1 ,like Figure 4 As shown, the speech recognition model includes two self-attention networks, namely the first self-attention network Self-Attention1 and the second self-attention network Self-Attention2. Self-Attention1 can extract the age attribute information a-vector from the speech to be recognized, that is, obtain the first attribute information, and then add it with the first voiceprint information r-vector representing the voiceprint features to obtain the fused feature information. Then, the second self-attention network Self-Attention2 is used to learn the fused feature information obtained by fusing a-vector and r-vector, and finally obtain the recognition result corresponding to the speech to be recognized.
[0063] Furthermore, in the embodiments of this application, the trained speech recognition model may include a fully connected layer. Accordingly, when the terminal device determines the recognition result corresponding to the speech to be recognized through the second self-attention network and the fused feature information, it can first learn the fused feature information through the second self-attention network to obtain the feature information corresponding to the speech to be recognized; then, it can classify the feature information corresponding to the speech to be recognized through the fully connected layer to determine the recognition result corresponding to the speech to be recognized.
[0064] It is understood that, in the embodiments of this application, the fully connected layer included in the speech recognition model can perform classification processing based on the learned feature information. That is, the fully connected layer can further map the feature information corresponding to the speech to be recognized to the labeled sample space, thereby obtaining the final recognition result.
[0065] Exemplary, in some embodiments, Figure 5 Schematic diagram of the speech recognition model Figure 2 ,like Figure 5As shown, the speech recognition model also includes fully connected layers (FC). For the feature information of the speech to be recognized output by the second self-attention network Self-Attention2, which is learned from the fused feature information, the fully connected layer can map it to the output node (such as age category), that is, obtain the recognition result corresponding to the speech to be recognized.
[0066] It should be noted that, in the embodiments of this application, both the first self-attention network and the second self-attention network include multiple self-attention layers; wherein, the self-attention layer includes a multi-head self-attention mechanism network and a feedforward network.
[0067] It is understood that in the embodiments of this application, the network structures of the first attention network and the second attention network can be the same, wherein the attention network can be composed of multiple self-attention layers. Each self-attention layer can include two sub-layers: a multi-head self-attention mechanism network and a feedforward network.
[0068] It should be noted that in the embodiments of this application, the multi-head self-attention mechanism network is the same as the multi-head self-attention mechanism (MHA), and the feedforward network is the same as the simple feedforward network (FFN).
[0069] It should be noted that, in the embodiments of this application, for each sub-layer, residual connections (RC) can be used, followed by layer normalization (LN). To maintain consistency with the dimension of the voiceprint feature r-vector, all sub-layers in the model produce outputs with the same dimension as the r-vector. For example, if the dimension of the voiceprint feature r-vector is 256, then each sub-layer can also produce an output with a dimension of 256.
[0070] Furthermore, in embodiments of this application, the multi-head self-attention mechanism can be described as mapping a query (Q) and a set of key-value pairs to an output, where the query (Q), key (K), value (V), and output are all vectors. The output is calculated by a weighted sum of the values, where the weight assigned to each value is calculated by Q and the corresponding value of K. The core idea of MHA is to project Q, K, and V through h different linear transformations, and finally concatenate the different attention results.
[0071] Exemplary, in some embodiments, Figure 6This is a schematic diagram of a dot product attention mechanism network, such as... Figure 6 As shown, for the Scaled Dot-Product Attention (SDPA) mechanism, a set of queries is packaged into matrix Q, and the keys and values are also packaged into matrices K and V, respectively. The formula for calculating the output matrix is:
[0072]
[0073] Where, d k For example, scaling factor, such as d k Set it to 32.
[0074] Exemplary, in some embodiments, Figure 7 A schematic diagram of a multi-head self-attention mechanism network, such as... Figure 7 As shown, in contrast, the FFN layer in the Self-Attention layer is essentially a two-layer fully connected neural network. The activation function of the first layer is generally a Rectified Linear Unit (ReLU), and the second layer is a linear activation function. If the output of the MHA layer is represented as Z, then the FFN can be represented as:
[0075] FFN(Z)=max(0,ZW1+b1)W2+b2 (2)
[0076] It should be noted that, in the embodiments of this application, when determining the fused feature information corresponding to the speech to be recognized based on the first attribute information and the first voiceprint information, the terminal device can perform fusion processing on the first attribute information and the first voiceprint information according to a preset fusion strategy, thereby obtaining the fused feature information.
[0077] It is understood that, in the embodiments of this application, the preset fusion strategy can be used to fuse different information. Based on the preset fusion strategy, the terminal device can select any fusion algorithm to fuse the first attribute information and the first voiceprint information. For example, the preset fusion strategy includes additive fusion, or dot product fusion, or concatenation fusion. This application does not impose specific limitations.
[0078] For example, in some embodiments, the terminal device may use an additive fusion algorithm to fuse the first attribute information and the first voiceprint information; or, the terminal device may use a dot product fusion algorithm to fuse the first attribute information and the first voiceprint information; or, the terminal device may use a splicing fusion algorithm to fuse the first attribute information and the first voiceprint information.
[0079] It should be noted that, in the embodiments of this application, the speech recognition model can be used to predict the age corresponding to the speech to be recognized. Therefore, the final recognition result corresponding to the speech to be recognized may include age-related information. For example, the recognition result of the speech to be recognized may be a specific age range or an age type classified based on age. This application does not impose any specific limitations on this.
[0080] For example, in some embodiments, the recognition result corresponding to the speech to be recognized determined by the terminal based on the speech recognition model can be under 10 years old, or 10 to 16 years old, or 16 to 25 years old, or 25 to 45 years old, or over 45 years old.
[0081] For example, in some embodiments, the recognition result corresponding to the speech to be recognized determined by the terminal based on the speech recognition model can be child, middle-aged, or elderly.
[0082] Exemplary, in some embodiments, Figure 8 This is a diagram illustrating the process of using a speech recognition model, such as... Figure 8 As shown, the model's process is mainly based on two parts, including the SAE model (i.e., speech recognition model) and the speaker recognition network. For the input speech to be recognized, on the one hand, the speaker recognition network can extract speaker features to obtain the corresponding speaker feature r-vector (i.e., the first speaker information); on the other hand, the speech to be recognized can be input into the first self-attention network in the SAE model, and the corresponding age attribute feature a-vector (i.e., the first attribute information) can be obtained through the first self-attention network. Then, the age attribute feature and speaker feature can be fused, and the fused information (i.e., the fused feature information) can be input into the second self-attention network for learning to obtain the corresponding feature information. Finally, the input feature information can be mapped to the output node through a fully connected layer to obtain the corresponding recognition result.
[0083] Furthermore, in the real-time implementation of this application, during the training process of the speech recognition model, the terminal device needs to first perform data augmentation processing on the speech data in the speech dataset to obtain augmented speech data; then the augmented speech data can be used to train the initial recognition model to obtain the speech recognition model.
[0084] It should be noted that, in the embodiments of this application, in order to make the speech recognition model more robust, the speech data in the original speech dataset can be subjected to data augmentation processing before model training. The data augmentation processing can include at least one or more of the following: speed adjustment processing, noise addition processing, and reverberation processing.
[0085] It is understood that, in the embodiments of this application, the speed-changing processing can be used to change the speech rate of the speech data according to a preset speech rate parameter; wherein, the preset speech rate parameter can be any value greater than 0, and this application does not impose any specific limitation.
[0086] For example, in some embodiments, when processing speech data at varying speeds, the principle of changing speed without changing pitch can be adopted. For each speech data point, the speed can be changed to 0.9, 1.0, or 1.1 times the original speech rate, respectively, while ensuring that the pitch remains unchanged. This speed-changing-without-pitch-changing processing can simulate the changes in a speaker's speech rate in real-world scenarios, thereby improving the robustness of the model.
[0087] It is understood that, in the embodiments of this application, noise addition processing can be used to add noise to speech data according to noise data; wherein, the noise data can be obtained from any speech dataset, and this application does not impose any specific limitations.
[0088] For example, in some embodiments, when adding noise to speech data, the open-source MUSAN dataset can be used. The MUSAN dataset contains three types of noise: speech, music, and noise. For speech noise, the signal-to-noise ratio (SNR) is randomly selected from 0dB to 15dB; for music noise, the SNR is randomly selected from 5dB to 15dB; and for noise, the SNR is randomly selected from 10dB to 20dB.
[0089] It should be noted that, in the embodiments of this application, the MUSAN dataset covers a wide variety of noise types, and it is possible to randomly select a type of noise for each speech data to add noise. Moreover, the noise addition process covers different signal-to-noise ratios, so it can simulate complex scenes in various types of videos to a certain extent, thereby improving the accuracy of the model.
[0090] It is understood that, in the embodiments of this application, the reverberation processing can be used to add reverberation to the speech data according to the reverberation data; wherein, the reverberation data can be obtained from any speech dataset, and this application does not make any specific limitation.
[0091] For example, in some embodiments, when performing reverberation processing on speech data, the simulated_rirs reverberation data from the open-source RIRS dataset can be used to add reverberation to the speech signal. Specifically, a reverberation data point can be randomly selected to add reverberation to the input speech data.
[0092] It should be noted that in the embodiments of this application, the RIRS dataset covers a large amount of reverberation data, and the addition of reverberation can better simulate real reverberation scenes, thereby improving the accuracy of the model.
[0093] It is understood that in the embodiments of this application, the initial recognition model and the speech recognition model obtained after training have the same model structure, that is, the initial recognition model may include a first self-attention network, a second self-attention network and a fully connected layer.
[0094] Furthermore, in the embodiments of this application, after completing the data augmentation processing of the speech dataset, when training the initial recognition model using the augmented speech data to obtain the speech recognition model, a speaker recognition network can first be used to perform speaker recognition processing on the speech data in the speech dataset to determine the second speaker information corresponding to the augmented speech data; then, a first self-attention network can be used to extract attribute features from the augmented speech data to obtain the second attribute information corresponding to the augmented speech data; next, a second self-attention network can be used to learn the fusion features to determine the feature information corresponding to the augmented speech data; wherein, the fusion features are obtained by fusing the second attribute information and the second speaker information; then, a fully connected layer can be used to classify the feature information corresponding to the augmented speech data to determine the training result corresponding to the augmented speech data; finally, a loss function can be calculated based on the training result, and the initial recognition model can be corrected according to the loss function to obtain the speech recognition model.
[0095] Exemplary, in some embodiments, Figure 9 This is a schematic diagram illustrating the training process of a speech recognition model, as shown below. Figure 9 As shown, the model training process is mainly based on three parts, including data augmentation, the SAE model (i.e., the initial recognition model), and the speaker recognition network. For the speech data in the input speech dataset, on the one hand, data augmentation processing such as speed adjustment, noise addition, and reverberation processing can be used to obtain the corresponding augmented speech data. On the other hand, the speaker recognition network can extract speaker features from the speech data in the speech dataset to obtain the corresponding speaker feature r-vector (i.e., the second speaker information). Then, the augmented speech data and speaker features can be input into the SAE model. Specifically, the augmented speech data can be input into the first self-attention network to obtain the corresponding age attribute feature a-vector (i.e., the second attribute information). Then, the age attribute feature and speaker features can be fused, and the fused information (i.e., the fused feature) can be input into the second self-attention network for learning to obtain the corresponding feature information. Then, the input feature information can be mapped to the output node through a fully connected layer to obtain the corresponding training results. Subsequently, the loss function can be calculated based on the training results, and the initial recognition model can be corrected according to the loss function to finally obtain the speech recognition model.
[0096] In summary, the speech recognition method proposed in steps 101 to 102 has several advantages. First, using data-augmented speech data for model training effectively increases the robustness of the model. Second, combining the voiceprint information corresponding to the speech to be recognized during the speech recognition process supplements the speech recognition model, thereby improving the accuracy of speech recognition. Third, the speech recognition model, which includes a self-attention network, can fully utilize the contextual information in the speech, further improving the accuracy of speech recognition.
[0097] In other words, the speech recognition method proposed in this application applies a self-attention network structure to SAE technology. The self-attention network has a strong ability to model the temporal information of speech, which enables the model to learn the contextual information in the speech sequence better, thereby improving the model's accuracy in estimating the age of the speaker in the user's video.
[0098] Furthermore, the speech recognition method proposed in this application integrates voiceprint feature information into the speech recognition model, thereby enabling the speech recognition model to learn the voiceprint features in the speech simultaneously, which can improve the accuracy of age estimation.
[0099] This application provides a speech recognition method. A terminal device extracts voiceprint features from the speech to be recognized to determine the first voiceprint information corresponding to the speech. The speech to be recognized and the first voiceprint information are input into a speech recognition model to obtain the recognition result corresponding to the speech. The speech recognition model is obtained by training a self-attention network with enhanced speech data. In other words, in this application, the terminal device integrates voiceprint features during speech recognition, improving the accuracy of speech recognition. Furthermore, the speech recognition model used by the terminal device is trained with enhanced speech data, effectively improving its robustness and enabling it to be better applied to complex audio scenarios. The speech recognition model can also fully utilize the contextual information of the speech through a self-attention network, further improving the accuracy of speech recognition. Therefore, in this application, the combined use of voiceprint features and a speech recognition model can accurately identify the speaker's age information, greatly improving speech recognition accuracy and thus enhancing the intelligence of the terminal device.
[0100] Based on the above embodiments, this application provides a speech recognition method that can improve and optimize the accuracy of speech recognition from three aspects: data augmentation, fusion of voiceprint features, and the use of a self-attention based network (SAN). Specifically, during the training process of the SAE model (speech recognition model), data augmentation processing is performed on the training data (speech data in the speech dataset), which may include speed variation without pitch change, noise addition, and reverberation addition, thereby improving the model's robustness to complex scenarios. During the training process of the SAE model, voiceprint feature vectors are added, enabling the SAE model to fuse voiceprint features to improve the accuracy of speech recognition. Finally, an SAE model based on a self-attention network is built, allowing the model to fully utilize the contextual information of the speech, thereby achieving better speech recognition results.
[0101] Understandably, in scenarios where accurate speaker age identification is crucial, such as creating memorable video clips based on children's voices from photo albums, age estimation is performed on audio segments extracted from the videos. Video clips identified as having children's voices are then used as material for creating these memorable clips. Therefore, the accuracy of speech recognition directly impacts the final effectiveness of the memorable clips.
[0102] Exemplary, in some embodiments, Figure 10 This is a diagram illustrating application scenarios for speech recognition solutions, such as... Figure 10 As shown, in the children's voice detection module of the multimodal intelligent editing and search engine project for photo albums, wonderful memories in the user's photo album are created based on children's video clips. The overall technical framework includes the following parts: 1. Acquiring voice data, i.e., acquiring photo album videos and extracting audio from the videos; 2. Performing active voice monitoring on the voice to be recognized, i.e., performing active voice detection on the audio and removing silent and noisy segments; 3. Acquiring the voice to be recognized, i.e., using speaker segmentation and clustering technology to segment the voice clips, and then clustering them according to the voiceprint information of the voice clips, thereby separating the voices of each individual speaker; 4. Performing age recognition processing on the voice to be recognized, i.e., using SAE to determine whether the voice of each separated speaker belongs to a child; 5. Generating children's voice memories, i.e., creating wonderful video memories of children's voices based on the detected children's voice clips.
[0103] As can be seen, SAE technology is directly responsible for the video footage of children's memorable voices in the aforementioned application scenarios. Therefore, the accuracy and effectiveness of speech recognition play a crucial role in creating compelling memories of children's voices.
[0104] To improve the accuracy of speech recognition, this application proposes a speech recognition method, which is a speaker age estimation technique that integrates voiceprint features. Figure 11 The following is a framework diagram for implementing a speech recognition method, such as... Figure 11 As shown, it can be divided into three modules: data augmentation, voiceprint recognition system (voiceprint recognition network), and SAE model structure (speech recognition model).
[0105] For the data augmentation module, in order to make the SAE model (speech recognition model) highly robust, in addition to applying the prepared training data, data augmentation processing will also be performed on the training data, specifically including the following methods: variable speed without pitch, adding noise, and adding reverb.
[0106] Since video data contains a large number of audio recordings with varying speeds, a speed-variable, pitch-invariant processing method can be chosen to handle the changes in speech rate in real-world scenarios. Specifically, for each audio recording, the speed is varied to 0.9, 1.0, and 1.1 times the original speed, while maintaining the pitch unchanged. This speed-variable, pitch-invariant processing simulates the changes in speaker speed in real-world scenarios, thereby improving the model's robustness.
[0107] The MUSAN dataset, an open-source dataset, was used to add noise to the original input speech signal. The MUSAN dataset contains three types of noise: speech, music, and noise. Specifically, for each training speech segment, one segment of each type of noise was randomly selected for noise addition. For speech noise, the added signal-to-noise ratio (SNR) was randomly selected between 0dB and 15dB; for music noise, the SNR was randomly selected between 5dB and 15dB; and for noise noise, the SNR was randomly selected between 10dB and 20dB. This dataset covers a wide range of noise types and employs noise addition techniques covering different SNRs, thus effectively simulating complex scenarios in user photo album videos and improving the model's accuracy.
[0108] We used the simulated_rirs reverberation data from the open-source RIRS dataset to add reverberation to the speech signal. Specifically, the method involved randomly selecting one reverberation data point and applying it to the input speech. This dataset covers a large amount of reverberation data, and the added reverberation effectively simulates the reverberation scenarios of real users, improving the model's accuracy.
[0109] Voiceprint recognition systems can integrate voiceprint information into the SAE system framework, thereby making the model's estimation of the age attribute of speech more accurate. For example... Figure 11 As shown, for each speech, the r-vector that can represent the voiceprint features is extracted using the voiceprint recognition system and then fused into the SAE model.
[0110] like Figure 2 As shown, the network structure of the neural network-based voiceprint recognition system consists of a frame-level deep feature learning layer, a feature aggregation layer, and a segment-level embedding learning layer. The model is trained using a speaker discriminative training criterion, and speaker embedding vectors are extracted at the segment-level layer. This application uses a ResNet-based network structure; therefore, the extracted speaker embedding vector is simply referred to as the r-vector.
[0111] like Figure 3 As shown, a 34-layer ResNet network structure was used, with an input speech duration of 2 seconds, a frame length of 25 milliseconds, and a frame shift of 10 milliseconds. 40-dimensional FABNK (Filter Bank) features were extracted, resulting in a speech FBANK feature dimension of 40x200x1. Specifically, the first part is a Conv2d layer, ensuring that the time and frequency dimensions remain unchanged while increasing the number of input feature channels. The second part consists of 32 stacked Conv2d layers (ResNetBlock-1, ResNetBlock-2, ResNetBlock-3, and ResNetBlock-4), which convolve the time and frequency dimensions to extract high-dimensional speech features. The third part consists of StatsPooling and Flatten layers, used to extract the sentence-level mean and variance, converting the frame-level feature representation into a sentence-level feature representation. The fourth part consists of fully connected neural networks Dense1 and Dense2, where Dense1 reduces the sentence-level feature representation to 256 dimensions, which is the vector representation r-vector of the voiceprint features. Dense2 maps the r-vector to the final number of speakers N and uses the softmax function to obtain the posterior probability representation of speaker classification.
[0112] like Figure 11 As shown, since the SAE model needs to fuse the voiceprint features of speech, this application divides the SAE network structure into two modules: Self-Attention1 and Self-Attention2. Self-Attention1 extracts the age attribute information (called a-vector) in speech, and then adds it to the r-vector representing the voiceprint. Self-Attention2 is then used to learn the features after fusing a-vector and r-vector.
[0113] Exemplary, in some embodiments, Figure 12 This is a schematic diagram of the structure of the speech recognition model proposed in the embodiments of this application, as shown below. Figure 12 As shown, in the SAE network structure based on Self-Attention, Self-Attention1 and Self-Attention2 networks are completely identical, both consisting of N stacked Self-Attention layers (N is an integer greater than 0). Each self-attention layer has two sub-layers. The first is a multi-head attention mechanism (MHA), and the second is a simple feed-forward network (FFN). We use residual connections (RC) for each sub-layer and then perform layer normalization (LN). To maintain consistency with the dimension of the voiceprint vector r-vector, all sub-layers in the model produce outputs with a dimension of 256.
[0114] Multi-Head Attention (MHA) can be described as mapping a query (Q) and a set of key-value pairs to an output, where the query (Q), key (K), value (V), and output are all vectors. The output is calculated by a weighted sum of the values, where the weight assigned to each value is calculated using Q and the corresponding value in K. The core idea of MHA is to project Q, K, and V through h different linear transformations and then concatenate the different attention results. The Scaled Dot-Product Attention (SDPA) mechanism is implemented as follows: a set of queries is packaged into matrix Q, and the keys and values are also packaged into matrices K and V, respectively.
[0115] The FFN layer in the Self-Attention layer is essentially a two-layer fully connected neural network. The activation function of the first layer is generally a rectified linear unit (ReLU), and the second layer is a linear activation function. If the output of the MHA layer is represented as Z, then the FFN can represent the above formula (2).
[0116] Based on the speech recognition method proposed in this application, for the input FBANK features, on the one hand, it passes through the Self-Attention1 network to extract the a-vector vector (e.g., with a dimension of 256) representing age information, and on the other hand, it passes through the voiceprint recognition network to extract the r-vector vector (e.g., with a dimension of 256) representing voiceprint information. Then, the a-vector and r-vector are added together to achieve the purpose of fusing voiceprint features. The added vector is then passed through the Self-Attention2 network and finally mapped to the output node (recognition result, such as age category including children, middle-aged, and elderly) using FC.
[0117] As can be seen, compared with traditional machine learning (non-deep learning) and the current mainstream deep learning-based SAE method, the speech recognition method proposed in this application can overcome the complex situation of background speech diversity and low signal-to-noise ratio in the complex scenario of user album videos. At the same time, it makes full use of the voiceprint information in the speech. In addition, in terms of model structure, it uses SAN network layer to model the speech temporal information.
[0118] On the one hand, in this application, at the front end of the scheme, data augmentation processing is performed, mainly by using methods such as variable speed without changing pitch, adding noise, and adding reverberation to increase the robustness of the model.
[0119] On the other hand, this application innovatively integrates the voiceprint feature vector r-vector into the a-vector age feature vector at the architectural level. For user videos in real-world scenarios, voiceprint features can supplement the SAE model with information, thereby improving the accuracy of the model's age estimation. Specifically, firstly, a voiceprint model with a ResNet network structure is trained using voiceprint data. Then, the voiceprint model is used to extract r-vectors from the speech to represent voiceprint information. Finally, the r-vectors are added to the a-vectors in the SAE model framework of this scheme, allowing the model to simultaneously learn the voiceprint features in the speech. In practical applications, the SAE model that integrates voiceprint features can improve the accuracy of age estimation.
[0120] Furthermore, this application innovatively applies a self-attention network structure to SAE technology at the model level. This enables the network structure to better model the temporal information of speech and fully utilize the contextual information in the speech, thereby improving the model's recognition accuracy. This is because applying the self-attention network structure to SAE technology leverages the strong ability of self-attention networks to model the temporal information of speech, allowing the model to better learn the contextual information in the speech sequence, thus improving the model's accuracy in estimating the age of the speaker in the user's video.
[0121] It should be noted that in the embodiments of this application, a self-attention network structure is applied to SAE technology, enabling the model to better learn the contextual information in the speech sequence. Alternatively, a gated recurrent unit (GRU) can be used to model temporal speech signals.
[0122] It should be noted that in the embodiments of this application, the a-vector and r-vector can be fused by addition, which is a relatively simple method. Alternatively, dot product or concatenation can be used to fuse the two vectors.
[0123] This application provides a speech recognition method. A terminal device extracts voiceprint features from the speech to be recognized to determine the first voiceprint information corresponding to the speech. The speech to be recognized and the first voiceprint information are input into a speech recognition model to obtain the recognition result corresponding to the speech. The speech recognition model is obtained by training a self-attention network with enhanced speech data. In other words, in this application, the terminal device integrates voiceprint features during speech recognition, improving the accuracy of speech recognition. Furthermore, the speech recognition model used by the terminal device is trained with enhanced speech data, effectively improving its robustness and enabling it to be better applied to complex audio scenarios. The speech recognition model can also fully utilize the contextual information of the speech through a self-attention network, further improving the accuracy of speech recognition. Therefore, in this application, the combined use of voiceprint features and a speech recognition model can accurately identify the speaker's age information, greatly improving speech recognition accuracy and thus enhancing the intelligence of the terminal device.
[0124] Based on the above embodiments, in another embodiment of this application... Figure 13 This is a schematic diagram of the composition structure of the terminal device proposed in the embodiments of this application. Figure 1 ,like Figure 13 As shown, the terminal device 10 proposed in this application embodiment may include a determining unit 111 and an obtaining unit 112.
[0125] The determining unit 111 is used to extract voiceprint features from the speech to be recognized and determine the first voiceprint information corresponding to the speech to be recognized.
[0126] The acquisition unit 112 is used to input the speech to be recognized and the first voiceprint information into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data.
[0127] In the embodiments of this application, further, Figure 14 This is a schematic diagram of the composition structure of the terminal device proposed in the embodiments of this application. Figure 2 ,like Figure 14 As shown, the terminal device 10 proposed in this application embodiment may further include a processor 121, a memory 122 storing instructions executable by the processor 121, and further, the terminal device 10 may also include a communication interface 123 and a bus 124 for connecting the processor 121, the memory 122 and the communication interface 123.
[0128] In the embodiments of this application, the processor 121 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other types, and this application embodiment does not specifically limit this. The terminal device 10 may also include a memory 122, which can be connected to the processor 121. The memory 122 is used to store executable program code, which includes computer operation instructions. The memory 122 may include high-speed RAM memory and may also include non-volatile memory, such as at least two disk drives.
[0129] In embodiments of this application, bus 124 is used to connect communication interface 123, processor 121, and memory 122, as well as the mutual communication between these devices.
[0130] In embodiments of this application, memory 122 is used to store instructions and data.
[0131] Furthermore, in the embodiments of this application, the processor 121 is used to extract voiceprint features from the speech to be recognized, determine the first voiceprint information corresponding to the speech to be recognized, and input the speech to be recognized and the first voiceprint information into a speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data.
[0132] In practical applications, the aforementioned memory 122 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 121.
[0133] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0134] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] This application provides a terminal device that extracts voiceprint features from the speech to be recognized, determining the first voiceprint information corresponding to the speech; inputting the speech and the first voiceprint information into a speech recognition model to obtain the recognition result corresponding to the speech; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data. In other words, in this application embodiment, the terminal device integrates voiceprint features during speech recognition, improving the accuracy of speech recognition. Simultaneously, the speech recognition model used by the terminal device for speech recognition processing is trained with enhanced speech data, effectively improving the robustness of the speech recognition model and enabling it to be better applied to complex audio scenarios. Furthermore, the speech recognition model can fully utilize the contextual information of the speech through a self-attention network, further improving the accuracy of speech recognition. Therefore, in this application embodiment, the combined use of voiceprint features and a speech recognition model can accurately identify the speaker's age information, greatly improving speech recognition accuracy and thus enhancing the intelligence of the terminal device.
[0136] This application provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the speech recognition method described above.
[0137] Specifically, the program instructions corresponding to a speech recognition method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to a speech recognition method in the storage media are read or executed by an electronic device, the following steps are included:
[0138] Voiceprint features are extracted from the speech to be recognized to determine the first voiceprint information corresponding to the speech to be recognized;
[0139] The speech to be recognized and the first voiceprint information are input into the speech recognition model to obtain the recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training a self-attention network with enhanced speech data.
[0140] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0141] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the schematic and / or block diagrams can be implemented by computer program instructions, and combinations of blocks in the schematic and / or block diagrams can be implemented. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the schematic and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0144] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A speech recognition method, characterized in that, The method includes: Acquire the album video and extract the audio from the album video to obtain the speech to be recognized; Voiceprint features are extracted from the speech to be identified to determine the first voiceprint information corresponding to the speech to be identified; The speech recognition model extracts attribute features from the speech to be recognized using a first self-attention network to obtain first attribute information corresponding to the speech. This first attribute information is used to determine the age attribute of the speech. The speech recognition model is obtained by training the self-attention network with enhanced speech data, which is obtained through speed adjustment, noise addition, and reverberation processing. The speed adjustment is used to simulate the speaker's speech rate changes in real-world scenarios. The noise addition is used to simulate scenes in video formats. The reverberation processing is used to simulate real-world reverberation scenarios. The first attribute information and the first voiceprint information are fused according to a preset fusion strategy to obtain the fused feature information corresponding to the speech to be recognized; wherein, the preset fusion strategy includes additive fusion, or dot product fusion, or concatenation fusion. The recognition result corresponding to the speech to be recognized is determined by the second self-attention network of the speech recognition model and the fused feature information, wherein the recognition result corresponding to the speech to be recognized includes age information; The step of extracting voiceprint features from the speech to be identified and determining the first voiceprint information corresponding to the speech to be identified includes: The voiceprint recognition network is used to perform voiceprint recognition processing on the speech to be recognized to determine the first voiceprint information; the voiceprint recognition network includes a frame-level feature learning layer, a feature aggregation layer, and a sentence-level feature learning layer.
2. The method according to claim 1, characterized in that, The speech recognition model includes a fully connected layer. The determination of the recognition result corresponding to the speech to be recognized using the second self-attention network of the speech recognition model and the fused feature information includes: The fused feature information is learned through the second self-attention network to obtain the feature information corresponding to the speech to be recognized; The fully connected layer is used to classify the feature information corresponding to the speech to be recognized, and the recognition result corresponding to the speech to be recognized is determined.
3. The method according to claim 1 or 2, characterized in that, Both the first self-attention network and the second self-attention network include multiple self-attention layers; wherein, the self-attention layer includes a multi-head self-attention mechanism network and a feedforward network.
4. The method according to claim 1, characterized in that, The initial recognition model is trained using the enhanced speech data to obtain the speech recognition model; the initial recognition model includes a first self-attention network, a second self-attention network, and a fully connected layer; training the initial recognition model using the enhanced speech data to obtain the speech recognition model includes: A voiceprint recognition network is used to perform voiceprint recognition processing on the voice data in the voice dataset to determine the second voiceprint information corresponding to the enhanced voice data. The enhanced speech data is subjected to attribute feature extraction by the first self-attention network to obtain the second attribute information corresponding to the enhanced speech data; The enhanced speech data is determined by learning the fusion features through the second self-attention network; wherein the fusion features are obtained by fusing the second attribute information and the second voiceprint information. The enhanced speech data is classified by the fully connected layer to determine the training result corresponding to the enhanced speech data. The loss function is calculated based on the training results, and the initial recognition model is corrected according to the loss function to obtain the speech recognition model.
5. The method according to claim 1 or 4, characterized in that, The voiceprint recognition network is a ResNet network structure.
6. A terminal device, characterized in that, The terminal device includes: a determining unit and an acquiring unit. The determining unit is used to acquire album videos and extract audio from the album videos to obtain the speech to be identified; and to use a voiceprint recognition network to perform voiceprint recognition processing on the speech to be identified to determine the first voiceprint information; the voiceprint recognition network includes a frame-level feature learning layer, a feature aggregation layer, and a sentence-level feature learning layer. The acquisition unit is configured to extract attribute features from the speech to be recognized using a first self-attention network of a speech recognition model to obtain first attribute information corresponding to the speech to be recognized. The first attribute information is used to determine the age attribute corresponding to the speech to be recognized. The speech recognition model is obtained by training the self-attention network with enhanced speech data, which is obtained through speed-changing processing, noise-adding processing, and reverberation processing. The speed-changing processing is used to simulate the speaker's speech rate changes in real-world scenarios. The noise-adding processing is used to simulate scenarios in video types. The reverberation processing is used to simulate real reverberation scenarios. The first attribute information and the first voiceprint information are fused according to a preset fusion strategy to obtain fused feature information corresponding to the speech to be recognized. The preset fusion strategy includes additive fusion, or dot-multiplication fusion, or concatenation fusion. The recognition result corresponding to the speech to be recognized is determined using a second self-attention network of the speech recognition model and the fused feature information. The recognition result corresponding to the speech to be recognized includes age information.
7. A terminal device, characterized in that, The terminal device includes a processor and a memory storing processor-executable instructions, wherein when the instructions are executed by the processor, the method described in any one of claims 1-5 is implemented.
8. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Age identification method, device and terminal equipment
CN109817222A
Age recognition method, device and equipment and computer readable storage medium
CN111312286A
Speech emotion recognition method and system based on three-dimensional depth feature fusion
CN114566189A
Speech recognition method, device and equipment and computer readable storage medium
CN120220656A