Voice data processing method and device, storage medium and electronic equipment

By denoising and scoring the original speech data, and combining this with noise reduction metrics to select high-quality speech data samples, the problem of poor performance of speech models in existing technologies is solved, and the robustness and selection accuracy of speech models are improved.

CN120932622APending Publication Date: 2025-11-11BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410585353.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies alter the data distribution during voice data screening through preprocessing, resulting in poor voice model performance. Furthermore, existing screening criteria are singular, leading to low screening accuracy.

Method used

By denoising the original speech data, noise reduction indicators and scoring results are determined. High-quality speech data samples are selected by combining preset screening conditions. Noise data is then denoised to select high-quality noise data for training the speech model.

Benefits of technology

It improves the accuracy of voice data selection, enhances the robustness of the trained voice model, and improves call quality and voice recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932622A_ABST
    Figure CN120932622A_ABST
Patent Text Reader

Abstract

The invention provides a voice data processing method and device, a storage medium and electronic equipment, and the method comprises the steps: carrying out the noise reduction of original voice data; according to the voice data after noise reduction, a first noise elimination index is determined, and the first noise elimination index is used for representing the elimination amount of the original voice data after noise reduction; the original voice data is scored, a scoring result is determined, and the scoring result is used for representing the cleanliness degree of the original voice data; and in response to a first noise elimination index and the scoring result satisfying a preset screening condition, determining the original voice data as sample data for training a model. The original voice data is screened by combining the scoring model and the noise elimination index, and the voice model is trained according to the original voice data meeting the screening condition, so that the accuracy of screening the voice data is improved, and the robustness of the trained voice model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, storage medium and electronic device for processing voice data. Background Technology

[0002] In existing technologies, when filtering speech data, it is necessary to first preprocess and score the original speech data, and then filter the preprocessed speech data based on the scoring results to train a speech model. Because preprocessing alters the data distribution of the original speech data, and the scoring model's filtering effect is poor, this results in relatively poor performance of the speech model. Summary of the Invention

[0003] In view of this, the present disclosure provides a method, apparatus, storage medium and electronic device for processing voice data, which can effectively improve the accuracy of voice data screening.

[0004] According to a first aspect of the present disclosure, a method for processing voice data is provided, the method comprising:

[0005] Noise reduction processing is performed on the raw speech data;

[0006] Based on the denoised speech data, a first noise reduction index is determined, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data;

[0007] The original speech data is scored to determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data;

[0008] If the first noise cancellation index and the scoring result meet the preset screening conditions, the original speech data is determined as sample data for training the model.

[0009] In one embodiment, before performing noise reduction processing on the original speech data, the method further includes:

[0010] The original voice data is subjected to DC removal processing to obtain DC-removed voice data.

[0011] In one embodiment, scoring the original speech data and determining the scoring result includes:

[0012] Based on the original voice data after DC removal, the original voice data is scored, and the scoring result is determined.

[0013] In one embodiment, after performing DC-de-DC processing on the original speech data to obtain DC-de-DC speech data, the method further includes:

[0014] Calculate the amplitude spectrum of the original speech data after DC removal;

[0015] Based on the amplitude spectrum, determine the energy value of the original speech data after DC removal;

[0016] The energy values ​​of the original voice data after DC removal are standardized.

[0017] In one embodiment, the noise reduction processing of the original speech data includes:

[0018] Obtain the original speech data and the noise reduction model;

[0019] The original speech data is denoised using the denoising model.

[0020] In one embodiment, determining the first noise cancellation index based on the denoised speech data includes:

[0021] Based on the original speech data, determine the first energy value of the original speech data;

[0022] Based on the denoised speech data, determine the second energy value of the denoised speech data;

[0023] The first noise cancellation index is determined based on the first energy value and the second energy value.

[0024] In one embodiment, the response to the first noise cancellation index and the scoring result satisfying preset screening conditions includes:

[0025] Based on preset weighting coefficients, calculate the weighted average of the first noise cancellation index and the scoring result;

[0026] If the weighted average value is greater than or equal to the first threshold, it indicates that the first noise reduction index and the scoring result meet the preset screening conditions.

[0027] In one embodiment, the method further includes:

[0028] Obtain the raw noise data;

[0029] The original noise data is subjected to noise reduction processing;

[0030] Based on the original noise data after noise reduction, a second noise reduction index is determined, wherein the second noise reduction index represents the amount of noise reduction after noise reduction of the original noise data;

[0031] If the second noise cancellation index is greater than or equal to the second threshold, the original noise data is determined as sample data for training the model.

[0032] In one embodiment, determining the second noise cancellation index based on the original noise data after noise reduction includes:

[0033] Based on the original noise data, determine the third energy value of the original noise data;

[0034] Based on the original noise data after noise reduction, determine the fourth energy value of the original noise data after noise reduction;

[0035] The second noise cancellation index is determined based on the third energy value and the fourth energy value.

[0036] According to a second aspect of the present disclosure, a voice data processing apparatus is provided, the apparatus comprising:

[0037] The first noise reduction unit is used to process the original speech data for noise reduction.

[0038] The first determining unit is used to determine a first noise reduction index based on the denoised speech data, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data.

[0039] The second determining unit is used to score the original speech data and determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data;

[0040] The third determining unit, in response to the first noise cancellation index and the scoring result satisfying the preset screening conditions, determines the original speech data as sample data for training the model.

[0041] In one embodiment, the device further includes:

[0042] The DC removal unit is used to perform DC removal processing on the original voice data to obtain DC-removed voice data.

[0043] In one embodiment, the second determining unit is specifically used for:

[0044] Based on the original voice data after DC removal, the original voice data is scored, and the scoring result is determined.

[0045] In one embodiment, the device further includes:

[0046] The acquisition unit is used to acquire the amplitude spectrum of the original speech data after DC removal;

[0047] The fourth determining unit is used to determine the energy value of the original voice data after DC removal based on the amplitude spectrum;

[0048] A unified unit is used to unify the energy value of the original voice data after DC removal.

[0049] In one embodiment, the first noise reduction unit is specifically used for:

[0050] Obtain the original speech data and the noise reduction model;

[0051] The original speech data is denoised using the denoising model.

[0052] In one embodiment, the first determining unit is specifically used for:

[0053] Based on the original speech data, determine the first energy value of the original speech data;

[0054] Based on the denoised speech data, determine the second energy value of the denoised speech data;

[0055] The first noise cancellation index is determined based on the first energy value and the second energy value.

[0056] In one embodiment, the third determining unit is specifically used for:

[0057] Based on preset weighting coefficients, calculate the weighted average of the first noise cancellation index and the scoring result;

[0058] If the weighted average value is greater than or equal to the first threshold, it indicates that the first noise reduction index and the scoring result meet the preset screening conditions.

[0059] In one embodiment, the device further includes:

[0060] The acquisition unit is used to acquire raw noise data;

[0061] The second noise reduction unit is used to perform noise reduction processing on the original noise data;

[0062] The fifth determining unit is used to determine a second noise elimination index based on the original noise data after noise reduction, wherein the second noise elimination index represents the amount of noise reduction after noise reduction of the original noise data;

[0063] The sixth determining unit, in response to the second noise cancellation index being greater than or equal to the second threshold, determines the original noise data as sample data for training the model.

[0064] In one embodiment, the fifth determining unit is specifically used for:

[0065] Based on the original noise data, determine the third energy value of the original noise data;

[0066] Based on the original noise data after noise reduction, determine the fourth energy value of the original noise data after noise reduction;

[0067] The second noise cancellation index is determined based on the third energy value and the fourth energy value.

[0068] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the methods described in the first aspect above.

[0069] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising:

[0070] processor;

[0071] Memory used to store processor-executable instructions;

[0072] The processor implements the steps of any of the methods described in the first aspect above by running the steps.

[0073] According to a fifth aspect of the present disclosure, a computer program product is provided that, when executed by a processor, implements the steps of any of the methods described in the first aspect above.

[0074] The technical solutions provided in this disclosure may have the following beneficial effects:

[0075] By combining a scoring model and a noise reduction metric to filter the raw speech data, and then training the speech model based on the raw speech data that meets the filtering criteria, the accuracy of the filtered speech data is improved, thereby enhancing the robustness of the trained speech model.

[0076] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0077] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0078] Figure 1 This disclosure is a flowchart illustrating a method for processing voice data according to an exemplary embodiment;

[0079] Figure 2 This disclosure is a flowchart illustrating a method for processing raw speech data according to an exemplary embodiment;

[0080] Figure 3 This disclosure is a flowchart illustrating a method for processing raw noise data according to an exemplary embodiment;

[0081] Figure 4 This disclosure is a block diagram of a voice data processing apparatus according to an exemplary embodiment;

[0082] Figure 5 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0083] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0084] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0085] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0086] In the field of deep learning, the training data of a model directly affects its performance and generalization ability after training. Therefore, to improve model performance, it is usually necessary to screen the training data. For the screening process of speech data, existing technologies typically require preprocessing and scoring the raw speech data, and then screening the preprocessed speech data based on the scoring results.

[0087] In the field of speech enhancement technology, speech data can be selected by setting various quality standards. For example, a uniform signal-to-noise ratio (SNR) can be set; if the SNR of speech data meets the requirements, the speech data is considered to be of high quality and can be used to train a speech model. Besides SNR, quality standards can also include speech clarity, speech features, and other selection criteria to filter out low-quality speech. However, this method uses relatively simple selection criteria, resulting in low accuracy in selecting speech data and consequently poor performance of the trained speech model.

[0088] Furthermore, automated tools and algorithms can be used to label and clean speech data. Automated tools can automatically identify errors, noise, or anomalies in speech data. However, these tools usually require preprocessing of the speech data before processing it, which alters the original distribution of the speech data. As a result, the speech model cannot learn the relevant knowledge during training, leading to poor performance.

[0089] Furthermore, multimodal data integration can be used to filter speech data. For example, combining other sensor data (e.g., image data) with speech data allows for multimodal information filtering from multiple angles, resulting in high-quality speech data. However, this approach requires adding other modal data relative to the speech data, leading to relatively low filtering efficiency.

[0090] Furthermore, in order for the speech model to better handle noise, distortion, or other interference, it is necessary to train the speech model using noisy data. The quality of the noisy data directly affects the performance of the speech model; therefore, the noisy data needs to be filtered before training the model. However, currently, it is usually only possible to filter speech data, not noisy data.

[0091] To address the aforementioned problems, this disclosure provides a method for processing voice data, referring to... Figure 1 This disclosure is a flowchart illustrating a method for processing voice data according to an exemplary embodiment, which will be further described below.

[0092] S101. Perform noise reduction processing on the original speech data.

[0093] In this embodiment of the disclosure, various methods can be used to denoise the original speech data and eliminate noise in the original speech data. For example, methods such as linear filters, spectral subtraction, and Wiener filtering can be used to denoise the original speech data, or machine learning models can be used to denoise the original speech data.

[0094] It should be noted that before using the noise reduction model to process the original speech data, it is necessary to remove DC and adjust the energy value of the original speech data. The specific processing will be explained in detail later and will not be repeated here.

[0095] S102. Based on the denoised speech data, determine a first noise reduction index, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data.

[0096] In this embodiment of the disclosure, after the denoising model performs denoising processing on the original speech data, denoised speech data is obtained. Then, by comparing the amount of noise reduction between the original speech data before denoising and the denoised speech data, a first noise reduction index can be determined. Based on the magnitude of the first noise reduction index, the noise level in the original speech data can be preliminarily determined. The larger the first noise reduction index, the greater the noise in the original speech data, the lower the quality of the original speech data, and the less suitable it is for training a speech model.

[0097] S103. The original speech data is scored to determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data.

[0098] In this embodiment, existing scoring models or tools can be used to score the original speech data, such as using the MOS (Mean Opinion Score) value to evaluate the quality of the original speech. In the field of speech enhancement, MOS is usually used to measure people's subjective evaluation of speech quality. MOS is typically a score between 1 and 5, representing different quality levels from "poor" to "excellent". Different scores correspond to different quality levels of speech data, as shown in Table 1:

[0099] Table 1

[0100]

[0101] For example, when the scoring model rates the raw speech data between 4.0 and 5.0, it means that the raw speech data can be heard clearly by the user and has low latency.

[0102] Based on this, a scoring model can be trained using a large amount of manually labeled speech data. This model can then score and evaluate the raw speech data based on user perception. The scores output by the scoring model are closer to human subjective listening experience and can approximate the cleanliness of the raw speech data. Therefore, the scoring results can also preliminarily represent the cleanliness of the raw speech data; the higher the score, the cleaner and higher the quality of the raw speech data.

[0103] It should be noted that before scoring the raw voice data, the raw voice data also needs to be de-DC processed.

[0104] S104. In response to the first noise cancellation index and the scoring result satisfying the preset screening conditions, the original speech data is determined as sample data for training the model.

[0105] In this embodiment of the disclosure, if the first noise cancellation index and the scoring result meet the preset screening conditions, the original speech data can be determined as sample data for training the speech model. By learning from the original speech data, the speech model can accurately identify information in other speech data, thereby improving the performance of the speech model.

[0106] The trained speech model can be applied to various applications and systems. For example, in the field of communication and voice calls, it can improve call quality, clarity, and clarity. It can also be used in speech recognition systems to help identify and utilize high-quality speech data, improving speech recognition accuracy. Furthermore, it can be applied to audio analysis and processing to extract clean speech signals, aiding in voice feature extraction, emotion analysis, and other speech processing operations. Further, the trained speech model can be applied to devices such as smart speakers and voice assistants. This technology can improve the quality of the device's speech recognition and response, enabling it to more accurately understand user commands and provide services to customers.

[0107] By combining a scoring model and a noise reduction metric to filter the raw speech data, and then training the speech model based on the raw speech data that meets the filtering criteria, the accuracy of the filtered speech data is improved, thereby enhancing the robustness of the trained speech model.

[0108] In this embodiment of the disclosure, before performing noise reduction processing on the original speech data, the method further includes: performing DC removal processing on the original speech data to obtain DC-removed speech data.

[0109] Before denoising the raw speech data, it is necessary to remove the DC component. The DC component in speech data refers to the DC part or DC offset in the speech signal. It refers to the constant component of the signal, that is, the part that does not change over time. In speech signals, the DC component usually manifests as a baseline offset, that is, the amount of offset of the signal on the time axis, causing the overall signal to deviate from the zero baseline, affecting signal quality and processing results. Therefore, in speech signal processing, it is usually necessary to remove the DC component.

[0110] In some embodiments, the DC component can be removed by calculating the mean of the original speech data and then subtracting that mean from each sample of the original speech data. This reduces the DC component of the entire original speech data to zero. Alternatively, a high-pass filter can be used to remove the DC component from the original speech data.

[0111] By performing DC removal on the raw speech data, the quality of the raw speech data can be improved, making it easier to process and analyze, and thus increasing the efficiency of subsequent speech data processing.

[0112] In this embodiment of the disclosure, scoring the original speech data and determining the scoring result includes: scoring the original speech data based on the original speech data after DC removal, and determining the scoring result.

[0113] After removing DC from the original speech data, the DC-removed original speech data is input into the scoring model. The scoring model scores the DC-removed original speech data to obtain the scoring result of the original speech data.

[0114] In this embodiment of the disclosure, after performing DC removal processing on the original speech data to obtain DC-removed speech data, the method further includes: obtaining the amplitude spectrum of the DC-removed original speech data; determining the energy value of the DC-removed original speech data based on the amplitude spectrum; and unifying the energy value of the DC-removed original speech data.

[0115] After removing DC from the original speech data, it is also necessary to adjust the audio energy of the original speech data after removing DC to unify the audio energy of all the original speech data, so as to compare the noise reduction indicators before and after noise reduction.

[0116] In some embodiments, the amplitude spectrum of the original speech data after DC removal can be obtained using STFT (Short-Time Fourier Transform). STFT is a spectral analysis method used to transform a signal from the time domain to the frequency domain and allows for spectral analysis of the signal over time. STFT uses a window function to segment the signal locally and performs a Fourier transform on the signal within each window. The method for obtaining the amplitude spectrum of the original speech data after DC removal is shown in formula (1):

[0117]

[0118] Where X[m,w] is the spectral component of the original speech data after DC removal at frequency w over time interval m, x[n] is the discrete-time sequence of the original speech data after DC removal, w[nm] is the window function, which refers to segmenting the signal locally in the time domain, m represents the window's step size, and e -jwn It is the complex exponential term of the Fourier transform, used to convert a signal from the time domain to the frequency domain.

[0119] By performing a short-time Fourier transform on the original speech data after DC removal, the amplitude spectrum of the original speech data after DC removal can be obtained. Based on the amplitude information of the original speech data after DC removal in the amplitude spectrum, the RMS (Root Mean Square) energy of the original speech data after DC removal can be determined. The RMS value (i.e., the energy value mentioned above) can be calculated using the method shown in formula (2):

[0120]

[0121] Where RMS is the RMS value of the original voice data after DC removal, N is the number of spectrum bands, and X[i] is the amplitude value on the i-th spectrum band.

[0122] After obtaining the RMS value of the original speech data after DC removal, the RMS values ​​of the original speech data after DC removal are unified. For example, the RMS values ​​of all the original speech data after DC removal are adjusted to -25dB. Since the comfortable range of human hearing is -15dB to -35dB, the RMS values ​​of the original speech data after DC removal can be adjusted to this range. The specific adjustment method is shown in formula (3):

[0123]

[0124] Among them, RMS scalar It is the target coefficient value, target dB This represents the target RMS value, which can range from -15dB to -35dB.

[0125] After obtaining the target coefficient value, multiplying the target coefficient value by the original voice data after DC removal will achieve the unification of the energy value of the original voice data after DC removal.

[0126] After unifying the energy value of the original speech data after DC removal, the accuracy and efficiency of noise reduction can be improved. Furthermore, after noise reduction of the original speech data after DC removal, the energy change of the speech data before and after noise reduction can be directly and clearly obtained based on the energy value of the denoised speech data.

[0127] In this embodiment of the disclosure, the noise reduction processing of the original speech data includes: acquiring the original speech data and a noise reduction model; and performing noise reduction processing on the original speech data using the noise reduction model.

[0128] In this embodiment of the disclosure, a noise reduction model can be used to process the original speech data to eliminate noise. The noise reduction model can be any pre-trained model, such as the RNNoise noise reduction model or a sound source separation model.

[0129] By denoising the original speech data, noise interference can be eliminated. Based on the changes in the original speech data before and after denoising, the magnitude of noise interference in the original speech data can be determined, so as to filter out the original speech data with large noise interference and improve the accuracy of filtering the original speech data.

[0130] In this embodiment of the disclosure, determining the first noise cancellation index based on the denoised speech data includes: determining a first energy value of the original speech data based on the original speech data; determining a second energy value of the denoised speech data based on the denoised speech data; and determining the first noise cancellation index based on the first energy value and the second energy value.

[0131] In some embodiments, the RMS value (i.e., the first energy value mentioned above) of the original speech data can be calculated based on the amplitude spectrum, and then the RMS value of the denoised speech data (i.e., the second energy value mentioned above) can be calculated. The difference between the two can be used to obtain the first noise reduction index.

[0132] It should be noted that if the RMS energy value of the original speech data has been standardized before noise reduction, the first energy value can be directly set to the standardized default value when calculating the first noise reduction index, without having to calculate it again, thus improving the processing efficiency of speech data.

[0133] The noise level in the original speech data can be directly determined based on the first noise cancellation index, which is then used to filter the original speech data.

[0134] In this embodiment of the disclosure, the step of responding to the first noise cancellation index and the scoring result satisfying the preset screening conditions includes: calculating the weighted average of the first noise cancellation index and the scoring result based on the preset weighting coefficient; and responding to the weighted average being greater than or equal to a first threshold, indicating that the first noise cancellation index and the scoring result satisfy the preset screening conditions.

[0135] The preset screening conditions refer to the quality of the original speech data meeting the requirements for training the speech model. To determine whether the first noise reduction index and the scoring result meet the preset screening conditions, the relationship between the two and the preset threshold can be compared. If the threshold judgment requirements are met, the original speech data is considered to meet the preset screening conditions.

[0136] A specific threshold determination method could be as follows: First, calculate the weighted average between the first noise reduction index and the scoring result based on preset weighting coefficients. Then, determine whether the weighted average is greater than or equal to the first threshold. If the weighted average is greater than or equal to the first threshold, it indicates that the original speech data meets the preset screening conditions and can be used to train a speech model. For example, if the first noise reduction index is -0.2, the scoring result is 4.7, the preset weighting coefficients are 0.3 for the first noise reduction index and 0.7 for the scoring result, and the first threshold is 4, then it is only necessary to determine whether -0.2*0.3 + 4.7*0.7 = 3.23 is greater than or equal to 4. If 3.23 is less than 4, it indicates that the original speech data does not meet the preset screening conditions. Using this original speech data to train a speech model would result in poor model performance.

[0137] By evaluating the raw speech data in two dimensions—noise cancellation metrics and scoring results—the accuracy of the evaluation of the raw speech data is improved, thereby enhancing the accuracy of the selection of raw speech data.

[0138] In this embodiment of the disclosure, the method further includes: acquiring original noise data; performing noise reduction processing on the original noise data; determining a second noise elimination index based on the noise-reduced original noise data, wherein the second noise elimination index represents the amount of noise reduction after noise reduction of the original noise data; and determining the original noise data as sample data for training the model in response to the second noise elimination index being greater than or equal to a second threshold.

[0139] When training a speech model, it's not enough for the model to learn the speech information from the original speech data; it also needs to be able to recognize noise. Therefore, noisy data is also required for training. However, if the noisy data contains babble noise (human-like noise), the trained speech model may identify the user's voice as noise, leading to distortion of the user's speech. Babble noise refers to background noise from multiple speakers speaking simultaneously or from the simultaneous presence of many sounds. It typically contains multiple sounds of different frequencies and volumes, simulating the background sounds humans experience in complex environments. This type of noise frequently occurs in real-life environments such as conference rooms, cafes, and public transportation.

[0140] Therefore, before using noisy data to train a speech model, it is necessary to filter the noisy data to ensure that there is no babble noise in the filtered data, so as to improve the performance of the trained speech model.

[0141] In some embodiments, the original noise data can first be denoised using techniques such as denoising models. Then, a second noise reduction index is determined based on the energy change between the original noise data before and after denoising. The denoising model can eliminate all noise in the original noise data except for human voice. Therefore, if there is no bubble noise in the noise data, the energy of the denoised noise data will be relatively low, and the energy change between the two will be significant. Therefore, if the second noise reduction index is greater than or equal to a second threshold, it indicates that there is no bubble noise or very little bubble noise in the original noise data, which can be used to train a speech model.

[0142] Correspondingly, if the original noise data contains babble noise, the energy of the noise data after denoising will be smaller than the energy of the noise data before denoising, and the second noise reduction index will also be smaller. In this case, the original noise data will not be suitable for training the speech model and needs to be screened out.

[0143] By denoising the original noise data and comparing the energy changes before and after denoising, the magnitude of the bubble noise in the original noise data can be determined. The original noise data with no or little bubble noise can be selected as sample data for training the speech model, resulting in a high-quality noise dataset, thereby improving the performance of the trained speech model.

[0144] In this embodiment of the disclosure, determining the second noise elimination index based on the original noise data after noise reduction includes: determining a third energy value of the original noise data based on the original noise data; determining a fourth energy value of the original noise data after noise reduction based on the original noise data; and determining the second noise elimination index based on the third energy value and the fourth energy value.

[0145] In some embodiments, the amplitude spectrum of the original noise data can be obtained first through STFT analysis, and then the RMS value of the original noise (i.e., the third energy value mentioned above) can be calculated based on the amplitude spectrum of the original noise data. Then, the RMS value of the original noise data after noise reduction processing (i.e., the fourth energy value mentioned above) can be calculated. Finally, the difference between the third energy value and the fourth energy value is obtained to obtain the second noise reduction index.

[0146] In some embodiments, reference Figure 2The flowchart shown illustrates the process of filtering raw audio data, as detailed below:

[0147] First, step 201 is executed to acquire the original speech data and perform DC removal processing on the original speech data to remove the DC component. Then, the DC-removed speech data is processed through two processing flows. One of the processing flows executes step 202, which scores the speech data using a scoring model to obtain a score indicating the cleanliness of the speech data. The other processing flow executes step 203, which uses a short-time Fourier transform to obtain the amplitude spectrum of the speech data. After obtaining the amplitude spectrum, step 204 is executed to adjust the root mean square value of the speech data, so that the root mean square value of all speech data is unified before noise reduction. Then, step 205 is executed to perform noise reduction processing on the speech data using a noise reduction model. Step 206 calculates the first noise reduction amount of the speech data before and after noise reduction. Finally, step 207 comprehensively judges the scoring result and the first noise reduction amount, and determines whether the comprehensive value of the two is greater than or equal to a first threshold. If the comprehensive value of the two is greater than or equal to the first threshold, it means that the quality of the original speech data meets the requirements and can be used to train the speech model. The speech data is then stored in the database. If the comprehensive value of the two is less than the first threshold, it means that the quality of the original speech data does not meet the requirements and the original speech data is directly discarded.

[0148] In some embodiments, reference Figure 3 The flowchart shown illustrates how to filter raw noisy data, as detailed below:

[0149] First, step 301 is executed to obtain the original noise data. Then, step 302 is executed to obtain the amplitude spectrum of the original noise data through short-time Fourier transform. Next, step 303 is executed to denoise the original noise data using a denoising model to obtain denoised noise data. Then, step 304 is executed to determine the second noise reduction amount based on the energy changes of the noise data before and after denoising. Finally, step 305 is executed to determine whether the noise reduction amount is greater than or equal to the second threshold. If the second noise reduction amount is greater than or equal to the second threshold, it means that the quality of the original noise data meets the requirements and can be used to train the speech model. The original noise data is then stored in the database. If the second noise reduction amount is less than the second threshold, it means that the quality of the original noise data does not meet the requirements and the original noise data is discarded directly.

[0150] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should know that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps may be performed in other orders or simultaneously.

[0151] Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by this disclosure.

[0152] Corresponding to the aforementioned application function implementation method embodiments, this disclosure also provides embodiments of application function implementation apparatus and corresponding terminals.

[0153] Reference Figure 4 A block diagram of a voice data processing apparatus according to an exemplary embodiment shows that the apparatus may include:

[0154] The first noise reduction unit 401 is used to perform noise reduction processing on the original speech data;

[0155] The first determining unit 402 is used to determine a first noise reduction index based on the noise-reduced speech data, wherein the first noise reduction index is used to represent the amount of noise reduction after the original speech data is denoised.

[0156] The second determining unit 403 is used to score the original speech data and determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data;

[0157] The third determining unit 404, in response to the first noise cancellation index and the scoring result satisfying the preset screening conditions, determines the original speech data as sample data for training the model.

[0158] In this embodiment of the disclosure, the apparatus further includes:

[0159] The DC removal unit is used to perform DC removal processing on the original voice data to obtain DC-removed voice data.

[0160] In this embodiment of the disclosure, the second determining unit 403 is specifically used for:

[0161] Based on the original voice data after DC removal, the original voice data is scored, and the scoring result is determined.

[0162] In this embodiment of the disclosure, the apparatus further includes:

[0163] The acquisition unit is used to acquire the amplitude spectrum of the original speech data after DC removal;

[0164] The fourth determining unit is used to determine the energy value of the original voice data after DC removal based on the amplitude spectrum;

[0165] A unified unit is used to unify the energy value of the original voice data after DC removal.

[0166] In this embodiment of the disclosure, the first noise reduction unit 401 is specifically used for:

[0167] Obtain the original speech data and the noise reduction model;

[0168] The original speech data is denoised using the denoising model.

[0169] In this embodiment of the disclosure, the first determining unit 402 is specifically used for:

[0170] Based on the original speech data, determine the first energy value of the original speech data;

[0171] Based on the denoised speech data, determine the second energy value of the denoised speech data;

[0172] The first noise cancellation index is determined based on the first energy value and the second energy value.

[0173] In this embodiment of the disclosure, the third determining unit 404 is specifically used for:

[0174] Based on preset weighting coefficients, calculate the weighted average of the first noise cancellation index and the scoring result;

[0175] If the weighted average value is greater than or equal to the first threshold, it indicates that the first noise reduction index and the scoring result meet the preset screening conditions.

[0176] In this embodiment of the disclosure, the apparatus further includes:

[0177] The acquisition unit is used to acquire raw noise data;

[0178] The second noise reduction unit is used to perform noise reduction processing on the original noise data;

[0179] The fifth determining unit is used to determine a second noise elimination index based on the original noise data after noise reduction, wherein the second noise elimination index represents the amount of noise reduction after noise reduction of the original noise data;

[0180] The sixth determining unit, in response to the second noise cancellation index being greater than or equal to the second threshold, determines the original noise data as sample data for training the model.

[0181] In this embodiment of the disclosure, the fifth determining unit is specifically used for:

[0182] Based on the original noise data, determine the third energy value of the original noise data;

[0183] Based on the original noise data after noise reduction, determine the fourth energy value of the original noise data after noise reduction;

[0184] The second noise cancellation index is determined based on the third energy value and the fourth energy value.

[0185] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0186] Accordingly, in one aspect, embodiments of this disclosure provide an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to:

[0187] Noise reduction processing is performed on the raw speech data;

[0188] Based on the denoised speech data, a first noise reduction index is determined, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data;

[0189] The original speech data is scored to determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data;

[0190] If the first noise cancellation index and the scoring result meet the preset screening conditions, the original speech data is determined as sample data for training the model.

[0191] Figure 5 This is a schematic diagram illustrating the structure of an electronic device 500 according to an exemplary embodiment. For example, device 500 can be a user device, specifically a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, wearable device such as smartwatch, smart glasses, smart bracelet, smart running shoes, etc.

[0192] Reference Figure 5 The device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.

[0193] Processing component 502 typically controls the overall operation of device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.

[0194] Memory 504 is configured to store various types of data to support the operation of device 500. Examples of this data include instructions for any application or method operating on device 500, contact data, phonebook data, messages, pictures, videos, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0195] Power supply component 506 provides power to various components of device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 500.

[0196] Multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the device 500 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0197] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.

[0198] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0199] Sensor assembly 514 includes one or more sensors for providing status assessments of various aspects of device 500. For example, sensor assembly 514 can detect the on / off state of device 500, the relative positioning of components such as the aforementioned display and keypad of device 500, changes in the position of device 500 or a component of device 500, the presence or absence of user contact with device 500, the orientation or acceleration / deceleration of device 500, and temperature changes of device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0200] Communication component 516 is configured to facilitate wired or wireless communication between device 500 and other devices. Device 500 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR, or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the aforementioned communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0201] In an exemplary embodiment, device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0202] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, such as a memory 504 including instructions, which, when executed by a processor 520 of device 500, enables device 500 to perform a method for sending information, the method including:

[0203] Noise reduction processing is performed on the raw speech data;

[0204] Based on the denoised speech data, a first noise reduction index is determined, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data;

[0205] The original speech data is scored to determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data;

[0206] If the first noise cancellation index and the scoring result meet the preset screening conditions, then the original speech data is determined as sample data for training the model.

[0207] The non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0208] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0209] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for processing voice data, characterized in that, The method includes: Noise reduction processing is performed on the raw speech data; Based on the denoised speech data, a first noise reduction index is determined, wherein the first noise reduction index is used to represent the amount of noise reduction after denoising the original speech data; The original speech data is scored to determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data; If the first noise cancellation index and the scoring result meet the preset screening conditions, the original speech data is determined as sample data for training the model.

2. The method according to claim 1, characterized in that, Before performing noise reduction processing on the original speech data, the method further includes: The original voice data is subjected to DC removal processing to obtain DC-removed voice data.

3. The method according to claim 2, characterized in that, The step of scoring the original speech data and determining the scoring result includes: Based on the original voice data after DC removal, the original voice data is scored, and the scoring result is determined.

4. The method according to claim 2, characterized in that, After performing DC removal processing on the original speech data to obtain DC-removed speech data, the method further includes: Calculate the amplitude spectrum of the original speech data after DC removal; Based on the amplitude spectrum, determine the energy value of the original speech data after DC removal; The energy values ​​of the original voice data after DC removal are standardized.

5. The method according to claim 1, characterized in that, The noise reduction process for the original speech data includes: Obtain the original speech data and the noise reduction model; The original speech data is denoised using the denoising model.

6. The method according to claim 1, characterized in that, The determination of the first noise reduction index based on the denoised speech data includes: Based on the original speech data, determine the first energy value of the original speech data; Based on the denoised speech data, determine the second energy value of the denoised speech data; The first noise cancellation index is determined based on the first energy value and the second energy value.

7. The method according to claim 1, characterized in that, The response to the first noise cancellation index and the scoring result satisfying the preset screening conditions includes: Based on preset weighting coefficients, calculate the weighted average of the first noise cancellation index and the scoring result; If the weighted average value is greater than or equal to the first threshold, it indicates that the first noise reduction index and the scoring result meet the preset screening conditions.

8. The method according to claim 1, characterized in that, The method further includes: Obtain the raw noise data; The original noise data is subjected to noise reduction processing; Based on the original noise data after noise reduction, a second noise reduction index is determined, wherein the second noise reduction index represents the amount of noise reduction after noise reduction of the original noise data; If the second noise cancellation index is greater than or equal to the second threshold, the original noise data is determined as sample data for training the model.

9. The method according to claim 8, characterized in that, The determination of the second noise elimination index based on the original noise data after noise reduction includes: Based on the original noise data, determine the third energy value of the original noise data; Based on the original noise data after noise reduction, determine the fourth energy value of the original noise data after noise reduction; The second noise cancellation index is determined based on the third energy value and the fourth energy value.

10. A voice data processing apparatus, characterized in that, include: The first noise reduction unit is used to process the original speech data for noise reduction. The first determining unit is used to determine a first noise reduction index based on the denoised speech data, wherein the first noise reduction index is used to represent the amount of noise reduction after the original speech data is denoised. The second determining unit is used to score the original speech data and determine the scoring result, wherein the scoring result is used to represent the cleanliness of the original speech data; The third determining unit, in response to the first noise cancellation index and the scoring result satisfying the preset screening conditions, determines the original speech data as sample data for training the model.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 9.

12. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor implements the steps of any one of claims 1 to 9 by running the process.

13. A computer program product, characterized in that, When the computer program product is executed by a processor, it implements the steps of the method described in any one of claims 1 to 9.