Effective speech recognition method and device
By training an effective speech recognition model and filtering out unwanted speech fragments, the problem of poor recognition performance in audio data with low signal-to-noise ratio was solved, achieving high-precision speech recognition in noisy environments.
Patent Information
- Application Number
- CN202510008493.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-01-03
AI Technical Summary
Existing technologies perform poorly in audio data with low signal-to-noise ratios, and are prone to misidentifying speech segments, especially when there is a lot of background noise.
An effective speech recognition model is employed. This model is trained by minimizing the difference between the effective predicted speech and the speech label, reducing the distance between the sample audio features and the noisy audio features, and maximizing the distance between the sample audio features and the pure noise features. Simultaneously, by extracting speakerprint features to filter out spam audio segments, the model's recognition accuracy in noisy environments is improved.
In scenarios with low signal-to-noise ratio and high background noise, it improves the accuracy and noise resistance of speech recognition, ensuring the quality of effective speech data.
Smart Images

Figure CN119763618B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and in particular to an effective speech recognition method and apparatus. Background Technology
[0002] In the field of audio signal processing technology, VAD (Voice Activity Detection) technology is often used to identify speech segments in audio data, and to perform speech recognition, semantic recognition and other processing on the identified speech segments according to different audio data processing needs.
[0003] Currently, most methods extract audio features from audio data using neural networks and then distinguish speech segments from non-speech segments based on these features. However, this method performs poorly in audio data with low signal-to-noise ratios. Summary of the Invention
[0004] This invention provides an effective speech recognition method and apparatus to address the deficiencies in the prior art.
[0005] This invention provides an effective speech recognition method, comprising the following steps:
[0006] Identify the audio data to be recognized;
[0007] Based on an effective speech recognition model, audio features of the audio data to be recognized are extracted, and the audio features of the audio data to be recognized are applied to determine effective speech data from the audio data to be recognized.
[0008] The effective speech recognition model is trained with the following objectives: minimizing the difference between the effective predicted speech and the effective speech label, minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data.
[0009] According to an effective speech recognition method provided by the present invention, the training steps of the effective speech recognition model include:
[0010] Based on the initial model, the audio features of the sample audio data, the audio features of the noisy sample audio data, and the audio features of the pure noise data are extracted.
[0011] Based on the initial model, the effective predicted speech is determined by applying the audio features of the sample audio data;
[0012] Based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, the training loss value is determined.
[0013] Based on the training loss value, the parameters of the initial model are updated to obtain the effective speech recognition model.
[0014] According to an effective speech recognition method provided by the present invention, the step of determining a training loss value based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data includes:
[0015] The classification loss value is determined based on the difference between the effective predicted speech and the effective speech label;
[0016] The intra-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data.
[0017] The inter-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the pure noise data.
[0018] Based on preset weights, the classification loss value, the intra-class loss value, and the inter-class loss value are weighted and summed to obtain the training loss value.
[0019] According to an effective speech recognition method provided by the present invention, after determining effective speech data from the audio data to be recognized, the method further includes:
[0020] Extract the voiceprint features of each segment from the valid speech data;
[0021] Based on the similarity between the voiceprint features of each segment and the voiceprint features of noise spam, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam, spam segments are filtered out from the effective speech data.
[0022] The voiceprint features of the noise garbage sound are determined based on the voiceprint features corresponding to multiple sample noise garbage sounds, and the voiceprint features of the distorted signal garbage sound are determined based on the voiceprint features corresponding to multiple sample distorted signal garbage sounds.
[0023] According to an effective speech recognition method provided by the present invention, the step of filtering out spam speech segments from the effective speech data based on the similarity between the voiceprint features of each segment and the voiceprint features of noisy spam speech, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam speech, includes:
[0024] If the similarity between the voiceprint features of any segment and the voiceprint features of the noise spam is greater than or equal to a threshold, and / or the similarity between the voiceprint features of any segment and the voiceprint features of the distorted signal spam is greater than or equal to the threshold, then any segment is considered as the spam segment, and the spam segment is filtered out from the valid speech data.
[0025] According to an effective speech recognition method provided by the present invention, the step of treating any segment as the spam segment further includes:
[0026] The noise spam voiceprint features and / or the distorted signal spam voiceprint features are updated based on the voiceprint features of any of the segments.
[0027] According to an effective speech recognition method provided by the present invention, the audio data to be recognized is acquired through multiple methods;
[0028] The step of determining valid speech data from the audio data to be identified further includes:
[0029] Based on the duration of the effective speech data corresponding to the audio data to be identified under each acquisition method, the target speech data is determined from each effective speech data.
[0030] The present invention also provides an effective speech recognition device, comprising the following modules:
[0031] The determining unit is used to determine the audio data to be identified.
[0032] The recognition unit is used to extract audio features of the audio data to be recognized based on an effective speech recognition model, and to apply the audio features of the audio data to be recognized to determine effective speech data from the audio data to be recognized.
[0033] The effective speech recognition model is trained with the following objectives: minimizing the difference between the effective predicted speech and the effective speech label, minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data.
[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the effective speech recognition methods described above.
[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the effective speech recognition method as described above.
[0036] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the effective speech recognition methods described above.
[0037] The effective speech recognition method and apparatus provided by this invention improve the accuracy of the effective speech recognition model in recognizing effective speech by minimizing the difference between effective predicted speech and effective speech labels, improve the noise resistance of the effective speech recognition model by minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and improve the noise discrimination ability of the effective speech recognition model by maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data. As a result, the effective speech recognition model can accurately perform effective speech recognition on the audio data to be recognized in scenarios with low speech signal-to-noise ratio and high background noise, thereby improving the accuracy of effective speech recognition. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the effective speech recognition method provided by the present invention.
[0040] Figure 2 This is a flowchart illustrating the effective speech recognition model training method provided by the present invention.
[0041] Figure 3 This is a schematic flowchart of the noise filtering method provided by the present invention.
[0042] Figure 4 This is a schematic diagram of the structure of the effective speech recognition device provided by the present invention.
[0043] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0045] Audio data for the same spoken content is typically collected through multiple methods. From these, a high-quality audio data set with the longest effective duration is selected for subsequent application in tasks such as speech recognition and semantic recognition. Furthermore, because the various acquisition methods used for these audio data sets differ, the background noise levels can vary significantly. For example, some audio data may contain almost only noise, while others, although with relatively low background noise, suffer from severe speech quality distortion, making the speaker's content difficult to understand. Still others, while containing audible audio, require selection based on their longest effective duration to ensure the acquisition of as much speech content as possible.
[0046] If we rely on manual screening to obtain the optimal audio data from the above audio data, it will not only require a lot of human resources, but also the speed of manual screening is limited, resulting in the inability to delete non-optimal audio data in a timely manner and wasting a lot of storage space.
[0047] Currently, most solutions involve extracting audio features from audio data using neural networks and then distinguishing between speech segments and non-speech segments based on these features. However, neural networks are prone to misidentifying speech and non-speech segments in audio data when faced with high background noise and low signal-to-noise ratios. For example, they may incorrectly identify noisy segments as speech or noisy speech segments as non-speech. Here, the neural network can be an endpoint detection network built using a Deep Neural Network (DNN). This endpoint detection network connects the DNN to a fully connected layer. By inputting audio features (such as filter bank features) into the DNN, which then transforms the audio features before inputting them into the fully connected layer, effective speech recognition results at the frame level are obtained, such as the probability of each frame being a speech or a non-speech segment.
[0048] Based on this, the present invention provides an effective speech recognition method to improve the performance of speech recognition in scenarios with high background noise and low signal-to-noise ratio. Specifically, Figure 1This is a flowchart illustrating the effective speech recognition method provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110 and 120.
[0049] Step 110: Determine the audio data to be recognized.
[0050] Here, the audio data to be recognized refers to the audio data that needs to be effectively recognized. The audio data to be recognized can be pure audio data or audio data with noise. The audio data to be recognized can be pre-collected audio data. This embodiment of the invention does not limit the language type of the audio data to be recognized; for example, the audio data to be recognized can be Chinese audio data, English audio data, etc. At the same time, this embodiment of the invention does not limit the length of the audio data to be recognized; for example, the audio data to be recognized can be a sentence or a paragraph, etc.
[0051] It is understood that the audio data to be identified can be obtained in real time through a sound pickup device, or it can be downloaded from the Internet, depending on actual needs. This embodiment of the invention does not make specific limitations in this regard.
[0052] Step 120: Based on the effective speech recognition model, extract the audio features of the audio data to be recognized, and apply the audio features of the audio data to be recognized to determine the effective speech data from the audio data to be recognized.
[0053] The effective speech recognition model aims to minimize the difference between the effective predicted speech and the effective speech label, minimize the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and maximize the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data.
[0054] Specifically, the audio features of the audio data to be identified can be used to describe the characteristics of the audio signal to be identified, and these audio features can distinguish different audio data. The audio features of the audio data to be identified can be any type of audio feature, such as audio data spectrum, filterbank features, bottleneck features, PLP (Perceptual Linear Prediction) features, Mel features, log-Mel features, etc. For example, the audio features of the aforementioned audio data to be identified can be represented in the form of feature vectors.
[0055] After extracting audio features from the audio data to be recognized, the effective speech recognition model determines effective speech data based on these features. Effective speech data can be understood as speech segments in the audio data that clearly express the speaker's speech content. Optionally, the effective speech recognition model can output the probability that each frame segment in the audio data to be recognized is a speech, and concatenate the frame segments with probabilities greater than a threshold to obtain the corresponding effective speech data.
[0056] Considering that traditional methods are prone to misidentifying speech and nospeech in audio data when faced with high background noise and low speech signal-to-noise ratio, this embodiment of the invention sets three training objectives during the training phase of the effective speech recognition model: ① minimizing the difference between effective predicted speech and effective speech labels; ② minimizing the distance between the audio features of sample audio data and the audio features of sample audio data after adding noise; ③ maximizing the distance between the audio features of sample audio data and the audio features of pure noise data.
[0057] In training objective ①, the effective predicted speech is the effective speech predicted by the effective speech recognition model from the sample audio data, and the effective speech label is the real effective speech corresponding to the sample audio data. The difference between the effective predicted speech and the effective speech label is used to characterize the accuracy of the effective speech recognition model in performing effective speech recognition based on the features of the audio data itself (i.e., audio features). The smaller the difference between the two, the more accurately the effective speech recognition model can recognize speech from the sample audio data, that is, the stronger the effective speech recognition model's effective speech recognition ability. In other words, training objective ① aims to enable the effective speech recognition model to perform effective speech recognition more accurately.
[0058] In training objective ②, the noisy sample audio data refers to the audio data obtained after adding noise to the sample audio data. For example, one or more types of pure noise can be randomly added to the sample audio data according to business needs to obtain noisy sample audio data. The distance between the audio features of the sample audio data and the audio features of the noisy sample audio data is used to characterize the effective speech recognition capability of the effective speech recognition model in noisy scenarios. The smaller the distance between the two, the more accurately the effective speech recognition model can perform effective speech recognition in noisy scenarios. That is, training objective ② considers that in practical applications, audio data is often accompanied by various noises, such as environmental noise and equipment noise. Therefore, training objective ② aims to minimize the distance between the audio features before and after adding noise, so that the effective speech recognition model can still maintain a high recognition accuracy for the noisy sample audio data. In other words, it improves the noise resistance of the effective speech recognition model, enabling the effective speech recognition model to better adapt to noisy data and improve the accuracy of effective speech recognition under noisy conditions.
[0059] In training objective ③, the distance between the audio features of the sample audio data and the audio features of the pure noise data is used to characterize the ability of the effective speech recognition model to distinguish between noise and speech data. The larger the distance, the more accurately the effective speech recognition model can distinguish between noise and speech data in noisy scenarios. In other words, training objective ③ aims to maximize the distance between the audio features of the sample audio data and the audio features of the pure noise data, so that the effective speech recognition model can accurately distinguish between speech and noise signals, thereby avoiding interference from noise signals on the effective speech recognition results, i.e., improving the noise discrimination ability of the effective speech recognition model. Here, the sample audio data can be understood as pure speech data.
[0060] Therefore, the effective speech recognition method provided by the embodiments of the present invention improves the accuracy of the effective speech recognition model in recognizing effective speech by minimizing the difference between effective predicted speech and effective speech labels, improves the noise resistance of the effective speech recognition model by minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and improves the noise discrimination ability of the effective speech recognition model by maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data. Thus, the effective speech recognition model can accurately perform effective speech recognition on the audio data to be recognized in scenarios with low speech signal-to-noise ratio and high background noise, thereby improving the accuracy of effective speech recognition.
[0061] Based on the above embodiments, Figure 2 This is a flowchart illustrating the effective speech recognition model training method provided by the present invention, as shown below. Figure 2 As shown, the training steps for an effective speech recognition model include:
[0062] Step 210: Based on the initial model, extract the audio features of the sample audio data, the audio features of the sample audio data after adding noise, and the audio features of the pure noise data.
[0063] Specifically, the initial model can be built based on a pre-trained network and a DNN network. The pre-trained network can include CPC neural networks, wav2vec2.0 neural networks, etc. Taking a CPC network as an example, the pre-trained network can use non-linear coding layers (such as 5-layer CNN + ReLU) to downsample and encode the input sample audio data, the noisy sample audio data, and the pure noise data to obtain latent variables. Then, an autoregressive model (such as a single-layer GRU RNN) is used to encode the latent variables of historical moments to obtain contextual features, i.e., the corresponding audio features.
[0064] The pre-trained network can be trained based on historical audio data and corresponding audio feature labels. During training, the pre-trained network extracts corresponding historical audio features from the historical audio data, determines the contrastive learning loss based on the historical audio features and audio feature labels, and updates the network parameters of the pre-trained network based on the contrastive learning loss. The contrastive learning loss can be determined using InfoNCE (Information Noise Contrastive Estimation), or other methods can be used; this embodiment of the invention does not specifically limit the method used.
[0065] After obtaining the pre-trained network, the pre-trained network is connected to the DNN network to obtain the above initial model. Then, the sample audio data, the sample audio data with added noise, and the pure noise data are respectively input into the initial model. The pre-trained network of the initial model extracts the audio features of the sample audio data, the audio features of the sample audio data with added noise, and the audio features of the pure noise data.
[0066] Step 220: Based on the initial model, apply the audio features of the sample audio data to determine the effective predicted speech.
[0067] Specifically, after extracting the audio features from the sample audio data in step 210, the audio features are input into the DNN network of the initial model. The DNN network then determines the valid speech in the sample audio data, thus obtaining the valid predicted speech. Optionally, the DNN network can output the probability that each frame segment in the sample audio data is valid speech, and concatenate the frame segments with probabilities greater than a threshold to obtain the valid predicted speech.
[0068] Step 230: Based on the difference between effective predicted speech and effective speech labels, the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, determine the training loss value.
[0069] Step 240: Based on the training loss value, update the parameters of the initial model to obtain an effective speech recognition model.
[0070] Specifically, the difference between the effective predicted speech and the effective speech label is used to characterize the accuracy of the effective speech recognition model in recognizing speech based on the features of the audio data itself (i.e., audio features). The smaller the difference between the two, the more accurately the effective speech recognition model can recognize speech from the sample audio data, that is, the stronger the effective speech recognition capability of the effective speech recognition model.
[0071] The distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise is used to characterize the effective speech recognition ability of the effective speech recognition model in noisy scenarios. The smaller the distance between the two, the more accurately the effective speech recognition model can perform effective speech recognition in noisy scenarios. In other words, minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise can improve the noise resistance of the effective speech recognition model, enabling the effective speech recognition model to better adapt to noisy data and improve the accuracy of effective speech recognition under noisy conditions.
[0072] The distance between the audio features of the sample audio data and the audio features of the pure noise data is used to characterize the ability of an effective speech recognition model to distinguish between noise data and speech data. The larger the distance between the two, the more accurately the effective speech recognition model can distinguish between noise data and speech data in noisy scenarios. In other words, maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data enables the effective speech recognition model to accurately distinguish between speech signals and noise signals, thereby avoiding the interference of noise signals on the effective speech recognition results, which improves the noise discrimination ability of the effective speech recognition model.
[0073] This invention combines the differences between effective predicted speech and effective speech labels, the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and the distance between the audio features of the sample audio data and the audio features of the pure noise data to determine the training loss value. Thus, after updating the parameters of the initial model based on the training loss value, the resulting effective speech recognition model can accurately perform effective speech recognition on the audio data to be recognized in scenarios with low speech signal-to-noise ratio and high background noise, thereby improving the accuracy of effective speech recognition.
[0074] It is understandable that when updating the parameters of the initial model based on the training loss value, the convergence condition can be met when the training loss value stabilizes (e.g., is less than a threshold) or the maximum number of iterations is reached, thus obtaining an effective speech recognition model.
[0075] Based on any of the above embodiments, the training loss value is determined based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, including:
[0076] The classification loss value is determined based on the difference between the effective predicted speech and the effective speech label;
[0077] The intra-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data.
[0078] The inter-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the pure noise data.
[0079] Based on preset weights, the classification loss value, intra-class loss value, and inter-class loss value are weighted and summed to obtain the training loss value.
[0080] Specifically, the difference between the effective predicted speech and the effective speech label is used to characterize the accuracy of the effective speech recognition model in recognizing speech based on the features of the audio data itself (i.e., audio features). The smaller the difference between the two, the smaller the corresponding classification loss value, indicating that the effective speech recognition model is more accurate in recognizing speech from the sample audio data, that is, the stronger the effective speech recognition capability of the effective speech recognition model. Optionally, the classification loss value can be determined using the Binary Cross Entropy Loss (BCELoss) function.
[0081] The distance between the audio features of the sample audio data and the audio features of the noisy sample audio data is used to characterize the effective speech recognition capability of the effective speech recognition model in noisy scenarios. The smaller the distance between the two, the smaller the corresponding intra-class loss value, indicating that the effective speech recognition model can perform effective speech recognition more accurately in noisy scenarios. Optionally, the intra-class loss value can be determined using the Mean Squared Error Loss (MSELoss) function.
[0082] The distance between the audio features of the sample audio data and the audio features of the pure noise data is used to characterize the ability of an effective speech recognition model to distinguish between noise data and speech data. The larger the distance, the smaller the inter-class loss value, indicating that the effective speech recognition model can more accurately distinguish between noise data and speech data in noisy scenarios. Optionally, the inter-class loss value can be determined using the triplet loss function.
[0083] During training, the classification loss value, intra-class loss value, and inter-class loss value are weighted and summed based on preset weights to obtain the training loss value. This training loss value takes into account the model's effective speech recognition ability, noise resistance, and noise discrimination ability. As a result, the trained effective speech recognition model can accurately perform speech recognition on the audio data to be recognized in scenarios with low speech signal-to-noise ratio and high background noise, thereby improving the accuracy of effective speech recognition.
[0084] Alternatively, the training loss value can be determined based on the following formula:
[0085]
[0086] in, This represents the training loss value. Represents the classification loss value. Represents the intra-class loss value. Represents the inter-class loss value. for Corresponding weights for Corresponding weights and It can be set according to the actual situation, such as... It is 0.02. It is 0.01.
[0087] While the aforementioned effective speech recognition models can accurately recognize audio data in scenarios with low signal-to-noise ratios and high background noise, thus improving the accuracy of effective speech recognition, these models may still misidentify untrained data (such as purely noisy data or audio data with severe speech distortion that makes the speaker's content incomprehensible). Therefore, this embodiment of the invention, after determining effective speech data from the audio data to be recognized, further identifies and filters out non-functional audio segments (including purely noisy segments and severely distorted segments) within the effective speech data, thereby obtaining higher-quality effective speech data.
[0088] in, Figure 3 This is a flowchart illustrating the noise filtering method provided by the present invention, as shown below. Figure 3 As shown, the method for filtering out junk noise includes the following steps:
[0089] Step 310: After determining the valid speech data from the audio data to be recognized, extract the voiceprint features of each segment in the valid speech data;
[0090] Step 320: Based on the similarity between the voiceprint features of each segment and the voiceprint features of the noise spam sound, and the similarity between the voiceprint features of each segment and the voiceprint features of the distorted signal spam sound, filter out spam sound segments from the effective speech data.
[0091] The voiceprint features of noisy garbage sounds are determined based on the voiceprint features corresponding to multiple samples of noisy garbage sounds, while the voiceprint features of distorted signal garbage sounds are determined based on the voiceprint features corresponding to multiple samples of distorted signal garbage sounds.
[0092] Specifically, the voiceprint features of each segment in valid speech data refer to the speech features corresponding to the sound in each segment. Voiceprint features can include features such as timbre, duration, and pitch. Based on features such as timbre, duration, and pitch, the quality of the corresponding valid speech can be determined, and thus it can be determined whether the valid speech is distorted.
[0093] Noise-generated speech voiceprint features refer to the speech characteristics corresponding to sounds generated by noise. The voiceprint features corresponding to sounds generated by noise are usually characterized by unstable frequencies, amplitude variations, and irregular harmonic structures.
[0094] To determine the voiceprint features of spam audio, multiple sample spam audio recordings can be collected, and voiceprint features can be extracted from them. These voiceprint features can include timbre, duration, pitch, and spectral characteristics. The voiceprint features of spam audio can be determined based on the following formula:
[0095]
[0096] in, This indicates the voiceprint characteristics of noise and garbage sounds. Indicates the first Voiceprint features of individual samples of noise and spam sounds This represents the total number of noise samples.
[0097] Furthermore, the voiceprint features of distorted signal spam audio refer to the speech features corresponding to the sound produced due to signal distortion. The voiceprint features corresponding to distorted sound typically manifest as sound quality degradation, pitch variation, and spectral distortion. Similar to the voiceprint features of noisy spam audio, to determine the voiceprint features of distorted signal spam audio, multiple sample distorted signal spam audios can be collected, and voiceprint features can be extracted from them. These features can also include sound quality, duration, pitch, and spectral characteristics. The voiceprint features of distorted signal spam audio can be determined based on the following formula:
[0098]
[0099] in, This indicates the voiceprint characteristics of distorted signals and spam audio. Indicates the first Voiceprint features of distorted garbage audio samples This represents the total number of junk sounds in the distorted sample signal.
[0100] Furthermore, the similarity between the voiceprint features of each segment and the voiceprint features of the noise spam is used to characterize the matching degree between the voiceprint features of each segment and the voiceprint features of the noise spam. The higher the similarity, the higher the matching degree between the voiceprint features of the corresponding segment and the voiceprint features of the noise spam, and thus the higher the probability that the corresponding segment is noise spam.
[0101] Similarly, the similarity between the voiceprint features of each segment and the voiceprint features of the distorted signal spam is used to characterize the matching degree between the voiceprint features of each segment and the voiceprint features of the distorted signal spam. The higher the similarity, the higher the matching degree between the voiceprint features of the corresponding segment and the voiceprint features of the distorted signal spam, and thus the higher the probability that the corresponding segment is the distorted signal spam. The aforementioned similarity can be measured using cosine similarity or Euclidean distance; this embodiment of the invention does not specifically limit the method.
[0102] Based on this, segments with high similarity to the voiceprint features of noisy spam audio, and / or segments with high similarity to the voiceprint features of distorted signal spam audio can be filtered out as spam audio segments, thereby further ensuring the quality of effective speech data.
[0103] Furthermore, the filtering of spam audio clips here is done using voiceprint features, which eliminates the need for complex training processes and saves on training costs.
[0104] Based on any of the above embodiments, filtering out spam audio segments from valid speech data based on the similarity between the voiceprint features of each segment and the voiceprint features of noisy spam audio, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam audio, includes:
[0105] If the similarity between the voiceprint features of any segment and the voiceprint features of the noise spam is greater than or equal to a threshold, and / or the similarity between the voiceprint features of any segment and the voiceprint features of the distorted signal spam is greater than or equal to a threshold, then any segment is considered a spam segment and spam segments are filtered out from the valid speech data.
[0106] Specifically, if the similarity between the voiceprint features of any segment and the voiceprint features of noise spam is greater than or equal to a threshold, it indicates that the corresponding segment is highly likely to be a noise spam segment, that is, the segment can be regarded as a noise spam segment.
[0107] If the similarity between the voiceprint features of any segment and the voiceprint features of distorted signal spam is greater than or equal to the threshold, it indicates that the corresponding segment has a high probability of being a distorted signal spam segment, that is, the segment can be regarded as a distorted signal spam segment.
[0108] If the similarity between the voiceprint features of any segment and the voiceprint features of noise spam is greater than or equal to the threshold, and the similarity between the voiceprint features of any segment and the voiceprint features of distorted signal spam is greater than or equal to the threshold, then it indicates that the corresponding segment contains both noise spam and distorted signal spam, meaning that the corresponding segment can also be considered as a spam segment.
[0109] After identifying the corresponding segment as a spam segment, the spam segment is filtered out from the valid speech data, thereby obtaining high-quality valid speech data.
[0110] Based on any of the above embodiments, if any segment is designated as a spam segment, the following steps are also included:
[0111] Update the voiceprint features of noisy spam sounds and / or update the voiceprint features of distorted signal spam sounds based on the voiceprint features of any segment.
[0112] Specifically, if any segment is determined to be a spam segment, the noise spam voiceprint features and / or the distortion signal spam voiceprint features can be updated based on that segment, thereby dynamically updating the noise spam voiceprint features and the distortion signal spam voiceprint features and ensuring the reliability of the spam voiceprint features and the distortion signal spam voiceprint features.
[0113] For example, if any segment is determined to be a noisy spam segment, the noisy spam voiceprint features are updated based on the voiceprint features of that segment. For instance, the segment can be used as a sample noisy spam, and the formula for solving the noisy spam voiceprint features described above can be updated using this sample noisy spam to obtain the updated noisy spam voiceprint features. Similarly, if any segment is determined to be a distorted signal spam, the distorted signal spam voiceprint features are updated based on the voiceprint features of that segment. For instance, the segment can be used as a sample distorted signal spam, and the formula for solving the distorted signal spam voiceprint features described above can be updated using this sample distorted signal spam to obtain the updated distorted signal spam voiceprint features.
[0114] Based on any of the above embodiments, the audio data to be identified is acquired through various methods;
[0115] After determining the valid speech data from the audio data to be recognized, the process also includes:
[0116] Based on the duration of the effective speech data corresponding to the audio data to be identified under each acquisition method, the target speech data is determined from each effective speech data.
[0117] Specifically, the audio data to be recognized can be acquired in various ways, such as through microphone recording, telephone recording, network audio recording, and mobile phone recording. Considering that different acquisition methods may introduce varying degrees of noise and distortion, this embodiment of the invention, after performing effective speech recognition on the audio data to be recognized under each acquisition method using the methods described above, calculates the duration of the corresponding effective speech data for each acquisition method. A longer duration indicates higher quality effective speech data under the corresponding acquisition method. To select the optimal speech data from multiple effective speech data, the effective speech data corresponding to the longest duration can be used as the target speech data. Based on this high-quality target speech data, tasks such as speech recognition and semantic understanding can be effectively performed.
[0118] Based on any of the above embodiments, the present invention also provides an effective speech recognition method, the method comprising:
[0119] The audio data to be identified is determined, and based on an effective speech recognition model, audio features of the audio data to be identified are extracted. Then, the audio features of the audio data to be identified are applied to determine the effective speech data from the audio data to be identified.
[0120] The effective speech recognition model is trained based on training samples and corresponding effective speech labels. The training samples include sample audio data, noisy sample audio data, and pure noise data, with a minimum effective speech duration of 500 hours. The noisy sample audio data can be obtained by randomly adding -5 to 15 dB of noise to the sample audio data.
[0121] After obtaining the training samples, they can be segmented, for example, randomly divided into segments ranging from 5 seconds to 20 seconds. Next, the segmented training samples are input in batches into a pre-trained network (this pre-trained network can be a CPC neural network, a wav2vec2.0 neural network, etc.). The pre-trained network extracts the audio features of the segmented training samples, and then inputs these audio features into a DNN network. The DNN network performs effective speech prediction based on these audio features, resulting in effective predicted speech. The pre-trained network and the DNN network constitute the initial model.
[0122] The classification loss value is determined based on the difference between the effective predicted speech and the effective speech label. The intra-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data. The inter-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the pure noise data. The classification loss value, intra-class loss value, and inter-class loss value are weighted and summed according to preset weights to obtain the training loss value. Based on the training loss value, the parameters of the initial model are updated to obtain an effective speech recognition model.
[0123] In addition, after determining the valid speech data from the audio data to be identified, the voiceprint features of each segment in the valid speech data are extracted; if the similarity between the voiceprint features of any segment and the voiceprint features of noise spam is greater than or equal to a threshold, and / or the similarity between the voiceprint features of any segment and the voiceprint features of distorted signal spam is greater than or equal to a threshold, then any segment is regarded as spam segment and spam segment is filtered out from the valid speech data.
[0124] Specifically, the voiceprint features of the noisy spam audio are determined based on the voiceprint features corresponding to multiple samples of noisy spam audio, while the voiceprint features of the distorted signal spam audio are determined based on the voiceprint features corresponding to multiple samples of distorted signal spam audio. Furthermore, if any segment is taken as a spam audio segment, the voiceprint features of the noisy spam audio and / or the voiceprint features of the distorted signal spam audio are updated based on the voiceprint features of that segment.
[0125] Finally, based on the duration of the effective speech data corresponding to the audio data to be identified under each acquisition method, the optimal target speech data is selected from each effective speech data.
[0126] The effective speech recognition device provided by the present invention is described below. The effective speech recognition device described below can be referred to in correspondence with the effective speech recognition method described above.
[0127] Based on any of the above embodiments Figure 4 This is a schematic diagram of the structure of the effective speech recognition device provided by the present invention, as shown below. Figure 4 As shown, the device includes:
[0128] Determining unit 410 is used to determine the audio data to be recognized;
[0129] The recognition unit 420 is used to extract audio features of the audio data to be recognized based on an effective speech recognition model, and to apply the audio features of the audio data to be recognized to determine effective speech data from the audio data to be recognized.
[0130] The effective speech recognition model aims to minimize the difference between the effective predicted speech and the effective speech label, minimize the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and maximize the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data.
[0131] Based on any of the above embodiments, the training steps of an effective speech recognition model include:
[0132] Based on the initial model, audio features of sample audio data, audio features of sample audio data after adding noise, and audio features of pure noise data are extracted.
[0133] Based on the initial model, the audio features of the sample audio data are applied to determine the effective predicted speech;
[0134] The training loss value is determined based on the difference between effective predicted speech and effective speech labels, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data.
[0135] Based on the training loss value, the parameters of the initial model are updated to obtain an effective speech recognition model.
[0136] Based on any of the above embodiments, the training loss value is determined based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, including:
[0137] The classification loss value is determined based on the difference between the effective predicted speech and the effective speech label;
[0138] The intra-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data.
[0139] The inter-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the pure noise data.
[0140] Based on preset weights, the classification loss value, intra-class loss value, and inter-class loss value are weighted and summed to obtain the training loss value.
[0141] Based on any of the above embodiments, after determining valid speech data from the audio data to be identified, the method further includes:
[0142] Extract the voiceprint features of each segment from the valid speech data;
[0143] Based on the similarity between the voiceprint features of each segment and the voiceprint features of noise spam, as well as the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam, spam segments are filtered out from the effective speech data.
[0144] The voiceprint features of noisy garbage sounds are determined based on the voiceprint features corresponding to multiple samples of noisy garbage sounds, while the voiceprint features of distorted signal garbage sounds are determined based on the voiceprint features corresponding to multiple samples of distorted signal garbage sounds.
[0145] Based on any of the above embodiments, filtering out spam audio segments from valid speech data based on the similarity between the voiceprint features of each segment and the voiceprint features of noisy spam audio, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam audio, includes:
[0146] If the similarity between the voiceprint features of any segment and the voiceprint features of the noise spam is greater than or equal to a threshold, and / or the similarity between the voiceprint features of any segment and the voiceprint features of the distorted signal spam is greater than or equal to a threshold, then any segment is considered a spam segment and spam segments are filtered out from the valid speech data.
[0147] Based on any of the above embodiments, if any segment is designated as a spam segment, the following steps are also included:
[0148] Update the voiceprint features of noisy spam sounds and / or update the voiceprint features of distorted signal spam sounds based on the voiceprint features of any segment.
[0149] Based on any of the above embodiments, the audio data to be identified is acquired through various methods;
[0150] After determining the valid speech data from the audio data to be recognized, the process also includes:
[0151] Based on the duration of the effective speech data corresponding to the audio data to be identified under each acquisition method, the target speech data is determined from each effective speech data.
[0152] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute an effective speech recognition method. This method includes: determining audio data to be recognized; extracting audio features from the audio data to be recognized based on an effective speech recognition model, and applying the audio features to determine effective speech data from the audio data to be recognized; the effective speech recognition model is trained with the following objectives: minimizing the difference between effective predicted speech and effective speech labels, minimizing the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model performing effective speech recognition on the sample audio data.
[0153] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the effective speech recognition method provided by the above methods. The method includes: determining audio data to be recognized; extracting audio features of the audio data to be recognized based on an effective speech recognition model, and applying the audio features of the audio data to be recognized to determine effective speech data from the audio data to be recognized; the effective speech recognition model is trained with the following objectives: minimizing the difference between effective predicted speech and effective speech labels, minimizing the distance between the audio features of the sample audio data and the audio features of the sample audio data after adding noise, and maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data; the effective predicted speech is obtained by the effective speech recognition model performing effective speech recognition on the sample audio data.
[0155] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an effective speech recognition method provided by the methods described above. The method includes: determining audio data to be recognized; extracting audio features from the audio data to be recognized based on an effective speech recognition model, and applying the audio features to determine effective speech data from the audio data to be recognized; wherein the effective speech recognition model is trained with the following objectives: minimizing the difference between effective predicted speech and effective speech labels, minimizing the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and maximizing the distance between the audio features of the sample audio data and the audio features of the pure noise data; and wherein the effective predicted speech is obtained by the effective speech recognition model performing effective speech recognition on the sample audio data.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An effective speech recognition method, characterized in that, include: Identify the audio data to be recognized; Based on an effective speech recognition model, audio features of the audio data to be recognized are extracted, and the audio features of the audio data to be recognized are applied to determine effective speech data from the audio data to be recognized. The effective speech recognition model aims to minimize the difference between the effective predicted speech and the effective speech label, minimize the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and maximize the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data. The training loss value of the effective speech recognition model is determined based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data.
2. The effective speech recognition method according to claim 1, characterized in that, The training steps for the effective speech recognition model include: Based on the initial model, the audio features of the sample audio data, the audio features of the noisy sample audio data, and the audio features of the pure noise data are extracted. Based on the initial model, the effective predicted speech is determined by applying the audio features of the sample audio data; Based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, the training loss value is determined. Based on the training loss value, the parameters of the initial model are updated to obtain the effective speech recognition model.
3. The effective speech recognition method according to claim 2, characterized in that, The training loss value is determined based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data, including: The classification loss value is determined based on the difference between the effective predicted speech and the effective speech label; The intra-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data. The inter-class loss value is determined based on the distance between the audio features of the sample audio data and the audio features of the pure noise data. Based on preset weights, the classification loss value, the intra-class loss value, and the inter-class loss value are weighted and summed to obtain the training loss value.
4. The effective speech recognition method according to any one of claims 1 to 3, characterized in that, The step of determining valid speech data from the audio data to be identified further includes: Extract the voiceprint features of each segment from the valid speech data; Based on the similarity between the voiceprint features of each segment and the voiceprint features of noise spam, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam, spam segments are filtered out from the effective speech data. The voiceprint features of the noise garbage sound are determined based on the voiceprint features corresponding to multiple sample noise garbage sounds, and the voiceprint features of the distorted signal garbage sound are determined based on the voiceprint features corresponding to multiple sample distorted signal garbage sounds.
5. The effective speech recognition method according to claim 4, characterized in that, The method of filtering out spam audio segments from the valid speech data based on the similarity between the voiceprint features of each segment and the voiceprint features of noisy spam audio, and the similarity between the voiceprint features of each segment and the voiceprint features of distorted signal spam audio, includes: If the similarity between the voiceprint features of any segment and the voiceprint features of the noise spam is greater than or equal to a threshold, and / or the similarity between the voiceprint features of any segment and the voiceprint features of the distorted signal spam is greater than or equal to the threshold, then any segment is considered as the spam segment, and the spam segment is filtered out from the valid speech data.
6. The effective speech recognition method according to claim 5, characterized in that, The step of designating any one of the segments as the garbage audio segment further includes: The noise spam voiceprint features and / or the distorted signal spam voiceprint features are updated based on the voiceprint features of any of the segments.
7. The effective speech recognition method according to any one of claims 1 to 3, characterized in that, The audio data to be identified was acquired through multiple methods; The step of determining valid speech data from the audio data to be identified further includes: Based on the duration of the effective speech data corresponding to the audio data to be identified under each acquisition method, the target speech data is determined from each effective speech data.
8. An effective speech recognition device, characterized in that, include: The determining unit is used to determine the audio data to be identified. The recognition unit is used to extract audio features of the audio data to be recognized based on an effective speech recognition model, and to apply the audio features of the audio data to be recognized to determine effective speech data from the audio data to be recognized. The effective speech recognition model aims to minimize the difference between the effective predicted speech and the effective speech label, minimize the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and maximize the distance between the audio features of the sample audio data and the audio features of the pure noise data. The effective predicted speech is obtained by the effective speech recognition model through effective speech recognition of the sample audio data. The training loss value of the effective speech recognition model is determined based on the difference between the effective predicted speech and the effective speech label, the distance between the audio features of the sample audio data and the audio features of the noisy sample audio data, and the distance between the audio features of the sample audio data and the audio features of the pure noise data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the effective speech recognition method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the effective speech recognition method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the effective speech recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
End-to-end voice enhancement method based on generation of countermeasure network
CN110390950A
Speech enhancement method of hopping connection deep neural network
CN111192598A