Sound data processing method, electronic device, storage medium and computer program product
By utilizing feedback optimization techniques of voiceprint filters and correlators during voice interaction, the problem of strong dependence on physical hardware in existing technologies has been solved, achieving better background noise filtering and improved voice interaction quality.
Patent Information
- Application Number
- CN202511210294.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies rely heavily on physical hardware when filtering background noise during voice interaction. They are particularly ineffective at identifying and eliminating diverse background noise on small devices with limited space, thus affecting the quality of voice interaction.
By acquiring the target user's voice data, the initial voice data is labeled using a pre-trained voiceprint recognition model and then input into a voiceprint filter carrying the target user's voiceprint features for filtering. The feedback parameters are optimized by combining a correlator, and the voiceprint filter is dynamically adjusted to improve its performance.
It effectively reduces reliance on physical hardware, improves background noise filtering during voice interaction, and enhances the quality of voice interaction.
Smart Images

Figure CN120977309A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a sound data processing method, electronic device, storage medium, and computer program product. Background Technology
[0002] Voice interaction technology has penetrated into various aspects of daily life, such as voice calls between users and voice interaction between users and terminals. In these scenarios, a common problem is that background noise in the user's environment can easily affect the quality of voice interaction. Since these background noises are diverse and difficult to control, such as noisy voices in public places, the roar of vehicles, the sound of keyboard typing in office environments, and the noise generated by home appliances (such as air conditioners, fans, and washing machines), how to effectively identify and eliminate these background noises is one of the technical problems that must be solved to improve the quality of voice interaction.
[0003] Currently, most noise reduction methods rely on the physical hardware layout of the terminal, such as configuring multiple microphones (e.g., dual-MIC arrays or triple-MIC arrays) and utilizing their spatial differences to achieve noise identification and cancellation. This method has certain limitations; for example, in some small devices with limited space, it may not be possible to configure multiple microphones. Summary of the Invention
[0004] This application provides a sound data processing method, electronic device, and computer program product to reduce the reliance on physical hardware in filtering background noise during voice interaction in the prior art, effectively identify and eliminate background noise during voice interaction, and improve the quality of voice interaction.
[0005] A first aspect of this application provides a sound data processing method, the method comprising:
[0006] Acquire initial audio data, including at least the target user's voice data;
[0007] The target user's voice data and background sound data in the initial sound data are labeled to obtain labeled initial sound data;
[0008] The initial sound data is input into a pre-constructed voiceprint filter carrying the voiceprint features of the target user, so that the voiceprint filter performs filtering processing on the initial sound data based on the voiceprint features of the target user to obtain filtered sound data.
[0009] input the labeled initial sound data and the filtered sound data into a correlator, so that the correlator performs correlation analysis on the filtered sound data and the labeled initial sound data to obtain a feedback parameter for evaluating performance of the voiceprint filter;
[0010] output the feedback parameter to the voiceprint filter, optimize the voiceprint filter according to the feedback parameter, and after returning to perform the step of obtaining initial sound data including sound data of at least a target user, filter the initial sound data using the optimized voiceprint filter.
[0011] In some embodiments, the feedback parameter includes a harmonic-to-noise ratio;
[0012] The optimization of the voiceprint filter according to the feedback parameter includes:
[0013] According to the threshold interval in which the harmonic-to-noise ratio is located, an optimization strategy of the voiceprint filter is determined;
[0014] The voiceprint filter is optimized according to the optimization strategy.
[0015] In some embodiments, the optimization strategy of the voiceprint filter is determined according to the threshold interval in which the harmonic-to-noise ratio is located, including:
[0016] If the harmonic-to-noise ratio is in a low threshold interval, the optimization strategy of the voiceprint filter includes at least one of adjusting a cutoff frequency of the voiceprint filter, increasing a filter order of the voiceprint filter, selecting an adaptive filter algorithm for optimization, and optimizing a voiceprint feature extraction process of the voiceprint filter;
[0017] If the harmonic-to-noise ratio is in a medium threshold interval, the optimization strategy of the voiceprint filter includes at least one of adjusting a cutoff frequency of the voiceprint filter, adjusting a filter strength of the voiceprint filter, and adding normalization processing to a voiceprint feature extracted by the voiceprint filter;
[0018] If the harmonic-to-noise ratio is in a high threshold interval, the optimization strategy of the voiceprint filter includes at least one of adjusting a cutoff frequency of the voiceprint filter, reducing a filter order of the voiceprint filter, and maintaining a current performance of the voiceprint filter.
[0019] In some embodiments, after the labeled initial sound data and the filtered sound data are input into the correlator, the method further includes:
[0020] Based on the correlation analysis performed by the correlator on the filtered sound data and the labeled initial sound data, an analysis result is obtained;
[0021] processing the filtered sound data according to the analysis result to obtain target sound data;
[0022] extracting a voiceprint feature of the target user in the target sound data;
[0023] outputting the extracted voiceprint feature to the voiceprint filter, optimizing the voiceprint filter according to the voiceprint feature, and after returning to perform the step of obtaining initial sound data including sound data of a target user, filtering the initial sound data using the optimized voiceprint filter.
[0024] In some embodiments, the processing the filtered sound data according to the analysis result to obtain target sound data includes:
[0025] rejecting the background sound data contained in the filtered sound data and / or adding sound data of the target user to the filtered sound data to obtain target sound data.
[0026] In some embodiments, further comprising:
[0027] obtaining sample sound data of the target user under a preset environmental condition;
[0028] extracting a voiceprint feature of the target user contained in the preprocessed sample sound data, and training a voiceprint recognition model according to the voiceprint feature of the target user;
[0029] obtaining a voiceprint filter based on the trained voiceprint recognition model, and determining a similarity threshold, the similarity threshold being used to determine whether the initial sound data contains background sound data.
[0030] In some embodiments, the inputting the initial sound data into a voiceprint filter pre-constructed to carry a voiceprint feature of the target user to enable the voiceprint filter to filter the initial sound data based on the voiceprint feature of the target user to obtain filtered sound data includes:
[0031] extracting a voiceprint feature contained in the initial sound data;
[0032] matching the extracted voiceprint feature with a voiceprint feature of the target user in the voiceprint filter to obtain a similarity between the extracted voiceprint feature and the voiceprint feature of the target user;
[0033] when the similarity is less than the similarity threshold, rejecting sound data having the extracted voiceprint feature from the initial sound data to obtain filtered sound data that retains sound data of the target user.
[0034] The second aspect of the embodiment of the present application provides an electronic device, comprising:
[0035] a central processing unit, a memory and an input and output interface;
[0036] The memory is a temporary storage memory or a persistent storage memory.
[0037] The central processing unit is configured to communicate with the memory and execute the instruction operation in the memory to execute the voice data processing method in any specific implementation manner of the first aspect of the embodiment of the present application.
[0038] The third aspect of the embodiment of the present application provides a computer readable storage medium, comprising instructions, when the instructions run on a computer, make the computer execute the voice data processing method in any specific implementation manner of the first aspect of the embodiment of the present application.
[0039] The fourth aspect of the embodiment of the present application provides a computer program product comprising instructions, when the computer program product runs on a computer, makes the computer execute the voice data processing method in any specific implementation manner of the first aspect of the embodiment of the present application.
[0040] From the above technical solutions, the embodiment of the first aspect of the present application has the following advantages:
[0041] The voice data processing method provided by the embodiment of the first aspect of the present application can effectively reduce the dependence on physical hardware when filtering background noise in the voice interaction process. In the voice interaction process, the correlator continuously obtains feedback parameters, and the voiceprint filter is optimized based on the feedback parameters, which can effectively improve the performance of the voiceprint filter, so that the background sound data can be better filtered in the subsequent voice interaction process, and the quality of voice interaction of the user is improved. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 The first flowchart of the voice data processing method provided by the embodiment of the present application is shown.
[0043] Figure 2 The second flowchart of the voice data processing method provided by the embodiment of the present application is shown.
[0044] Figure 3 The third flowchart of the voice data processing method provided by the embodiment of the present application is shown.
[0045] Figure 4 The fourth flowchart of the voice data processing method provided by the embodiment of the present application is shown.
[0046] Figure 5A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0047] The terms "first", "second", "third", and so on, if any, in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product, or apparatus that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatuses.
[0048] In daily life, in scenarios such as voice calls between users or voice interactions between users and terminals, a common problem is that background noise in the environment in which the user is located can easily affect the quality of voice interaction. For example, when users interact through fixed telephone calls, mobile cellular network calls (such as calls relying on 2G, 3G, 4G, and 5G technologies), Internet voice calls (including social software calls (such as WeChat voice calls, QQ voice calls, and the like), conference tool calls (such as Tencent Conference, DingTalk voice calls), and dedicated network calls (such as walkie-talkie calls), various background noises in the environment around the user are often captured by the microphone and transmitted to the other party of the interaction, seriously affecting the quality of voice interaction. Moreover, these background noises are diverse and difficult to control. Therefore, how to accurately identify and eliminate these background noises has become a key technical challenge to improve user experience and ensure the quality of voice interaction. At present, most noise reduction technologies rely on the design of physical hardware, for example, by configuring multiple microphones (such as a dual-microphone array or a three-microphone array), using their differences in spatial position to identify and eliminate noise. However, this method has certain limitations, especially on small devices with limited space, which may not be able to configure multiple microphones at the same time, thereby limiting the performance of noise reduction.
[0049] Based on this, the embodiments of the present application provide a sound data processing method, an electronic device, a storage medium, and a computer program product, which can effectively reduce the dependence on physical hardware when filtering background noise in the voice interaction process. In the voice interaction process, the correlator continuously obtains feedback parameters, and the voiceprint filter is optimized based on the feedback parameters, which can effectively improve the performance of the voiceprint filter, so that the background sound data can be better filtered in the subsequent voice interaction process, and the quality of voice interaction of the user can be improved.
[0050] Referring to Figure 1 An embodiment of the sound data processing method provided in the present application includes the following steps 101 to 107:
[0051] 101. Obtain initial sound data including at least sound data of a target user.
[0052] Generally, the initial sound data includes sound data of the target user and background sound data, and the background sound data includes sound data of other users (non-target users) and environmental sound (including animal sounds and other noises produced by various devices or existing in nature).
[0053] When obtaining the initial sound data including at least sound data of the target user in the voice interaction process of the target user, the sound data acquisition device built-in the user terminal can be used for acquisition, such as a microphone or other devices with sound pickup function, etc. In actual application, the sound data acquisition device of the user terminal peripheral (for example, earphones with microphone or sound pickup function, including wireless earphones and wired earphones) can also be used for acquisition of the initial sound data of the target user. It can be understood that the sound data acquisition device of the peripheral is in communication connection with the user terminal (for example, Bluetooth connection between the wireless earphones and the user terminal, wired connection between the wired earphones and the user terminal, etc.), and the sound data acquisition device of the peripheral transmits the initial voice data of the target user to the user terminal after collecting the initial voice data of the target user.
[0054] The user terminal (including the user terminal of the target user and the user terminal of the calling user in communication with the target user) involved in the embodiments of the present application can include but is not limited to mobile phones, tablet computers, notebook computers, desktop computers, smart voice interaction devices, interphones, etc.
[0055] 102. Label the sound data of the target user and the background sound data in the initial sound data respectively to obtain the labeled initial sound data.
[0056] When labeling the voice data of the target user and the background sound data in the initial voice data, a pre-trained voiceprint recognition model can be used, where the pre-trained voiceprint recognition model includes but is not limited to any one of a pre-trained Gaussian Mixture Model (GMM), a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), and a Long Short-Term Memory Network (LSTM).
[0057] In the application, when obtaining the pre-trained voiceprint recognition model, the following steps can be included: obtaining sample voice data of the target user in a preset environment condition (for example, an environment with a noise intensity of 30 dB or below or 20 dB or below); pre-processing the sample voice data (the pre-processing includes at least one of initial noise reduction, echo cancellation, mute segment cutting, frame windowing, normalization, endpoint detection, and data format unification processing, etc., to improve the quality of the sample voice data and facilitate subsequent processing); extracting parameters reflecting the voiceprint features of the target user from the pre-processed sample voice data, including but not limited to at least one of Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coefficients (LPC), Bark-Frequency Cepstral Coefficients (BFCC), Gammatone-Frequency Cepstral Coefficients (GFCC), formant frequency, pitch (fundamental frequency), and timbre features, etc., which are relatively stable for the target user and are not easily affected by time and environmental changes, and can achieve accurate identification and verification of the user's identity; then inputting these voiceprint feature parameters into the voiceprint recognition model to be trained, and obtaining the pre-trained voiceprint recognition model after the training is completed, at this time, the pre-trained voiceprint recognition model includes the voiceprint features of the target user.
[0058] It can be understood that the number of sample voice data used to train the voiceprint recognition model can be determined according to actual needs, which is not limited here.
[0059] In the application, after obtaining the pre-trained voiceprint recognition model, the initial sound data can be labeled, and the initial sound data can also be labeled by including at least one of preprocessing of the initial sound data, such as preprocessing noise reduction, echo cancellation, mute segment cutting, frame windowing, normalization, endpoint detection, and data format unification processing, to improve the quality of the initial sound data and facilitate subsequent processing. Then, the preprocessed initial sound data is input into the pre-trained voiceprint recognition model, and the pre-trained voiceprint recognition model is used to automatically label the voice data and background sound data of the target user in the initial sound data.
[0060] 103. The initial sound data is input into the pre-constructed voiceprint filter carrying the voiceprint features of the target user, so that the voiceprint filter filters the initial sound data based on the voiceprint features of the target user to obtain filtered sound data.
[0061] It should be noted that after obtaining the initial sound data including at least the voice data of the target user in step 101, the initial sound data can be directly input into the voiceprint filter for filtering processing, or the initial sound data can be preprocessed (the preprocessing of the initial sound data includes at least one of preprocessing noise reduction (spectral subtraction, wiener filtering, wavelet transform noise reduction, and noise reduction based on deep learning), echo cancellation, mute segment cutting, frame windowing, normalization, endpoint detection (locating the starting point and ending point of the voice data of the target user to define a more accurate range for subsequent feature extraction), and data format unification processing) to improve the quality of the initial sound data and facilitate subsequent processing, and then the preprocessed initial sound data is input into the voiceprint filter for filtering processing.
[0062] In this step, if the initial sound data is directly input into the voiceprint filter, the voiceprint filter can also combine a corresponding noise estimation algorithm (such as spectral subtraction and wiener filtering) to estimate the spectral characteristics of the noise from the mute segment in the signal of the initial sound data, and then subtract the noise component from the noisy speech to restore a clearer signal.
[0063] It can be understood that the voiceprint filter mentioned in the embodiments of the present application is pre-constructed by the pre-trained voiceprint recognition model, and the voiceprint filter has the voiceprint features of the target user, and the voiceprint filter also has a voiceprint feature extraction function, which can extract all voiceprint features contained in the initial sound data input thereto.
[0064] In this step, all voiceprint features contained in the initial sound data can be extracted first. Then, the extracted voiceprint features are matched with the voiceprint features of the target user stored in the voiceprint filter. Based on the matching results, the background sound data in the initial sound data is filtered, and the corresponding filtered sound data is obtained after the filtering is completed.
[0065] 104. Input the labeled initial sound data and filtered sound data into the correlator so that the correlator can perform correlation analysis on the filtered sound data and the labeled initial sound data to obtain feedback parameters for evaluating the performance of the acoustic filter.
[0066] By analyzing the filtered sound data and the initial sound data, the correlator can obtain some feedback parameters for evaluating the performance of the voiceprint filter. These feedback parameters include, but are not limited to, the voiceprint filter's false rejection rate (i.e., the ratio in which the voiceprint filter incorrectly identifies the voiceprint features of the target user as those of a non-target user), false acceptance rate (i.e., the ratio in which the voiceprint filter incorrectly identifies the voiceprint features of a non-target user as those of the target user), and harmonic noise ratio (HNR, the ratio between the energy of the harmonic components and the energy of the noise components in the sound signal carrying the filtered sound data).
[0067] Among them, the false rejection rate and false acceptance rate are used to quantify the performance of the voiceprint filter. The lower the false rejection rate and false acceptance rate, the better the performance of the voiceprint filter system and the higher the accuracy of voiceprint recognition. The harmonic noise ratio is used to measure the effect of background noise data removal in the filtered sound data. The higher the harmonic noise ratio, the better the effect of background noise data removal in the filtered sound data. Conversely, the lower the harmonic noise ratio, the worse the effect of background noise data removal in the filtered sound data. For cases with relatively low harmonic noise, the voiceprint filter should be adjusted and optimized to achieve better noise reduction effect.
[0068] 105. Output the feedback parameters to the voiceprint filter, optimize the voiceprint filter according to the feedback parameters, and after returning to the step of obtaining initial sound data including at least the target user's voice data, use the optimized voiceprint filter to filter the initial sound data.
[0069] By analyzing the correlation between the filtered sound data and the initial sound data using a correlator, the corresponding feedback parameters are obtained. These feedback parameters are then output to the voiceprint filter, allowing the filter to dynamically adjust the relevant parameters. This enables better matching and comparison of the target user's voiceprint features and better filtering out background noise data from the initial sound data.
[0070] In this embodiment, step 105 optimizes the acoustic filter based on the feedback parameters, specifically including:
[0071] 1051. Determine the optimization strategy for the acoustic text filter based on the threshold range of the harmonic noise ratio.
[0072] 1052. Optimize the acoustic filter according to the optimization strategy.
[0073] To better optimize filter performance, this embodiment divides the harmonic noise ratio (HNR) threshold into intervals. Different optimization strategies are set for different threshold intervals of the HNR, thereby obtaining a better-performing acoustic signature filter. Specifically, the threshold intervals can be divided into low threshold intervals (e.g., HNR < 10 dB), medium threshold intervals (e.g., 10 dB ≤ HNR < 20 dB), and high threshold intervals (e.g., HNR ≥ 20 dB). This is just an example and is not a limitation.
[0074] When determining the optimization strategy for the acoustic text filter based on the threshold range of the harmonic noise ratio, it may also include:
[0075] If the harmonic noise ratio is in the low threshold range, the optimization strategy for the acoustic text filter includes at least one of the following: adjusting the cutoff frequency of the acoustic text filter, increasing the filter order of the acoustic text filter, selecting an adaptive filtering algorithm for optimization, and optimizing the acoustic text feature extraction process of the acoustic text filter.
[0076] If the harmonic noise ratio is in the middle threshold range, the optimization strategy for the acoustic text filter includes at least one of the following: adjusting the cutoff frequency of the acoustic text filter, adjusting the filtering strength of the acoustic text filter, and adding normalization processing to the acoustic text features extracted by the acoustic text filter.
[0077] If the harmonic noise ratio is in the high threshold range, the optimization strategy for the acoustic strip filter includes at least one of adjusting the cutoff frequency of the acoustic strip filter, reducing the filter order of the acoustic strip filter, and maintaining the current performance of the acoustic strip filter.
[0078] When the harmonic noise ratio is in the low threshold range (e.g., HNR < 10dB), it indicates that the noise component accounts for a large proportion, and the noise reduction capability of the filter needs to be strengthened. At this time, the cutoff frequency of the acoustic filter can be reduced to retain the frequency band where harmonics are concentrated and filter out the frequency band dominated by noise (such as high-frequency noise and low-frequency noise). For example, in the signal of filtered audio data with HNR in the low threshold range (contaminated by high-frequency noise), the harmonics are concentrated below 3kHz. At this time, the cutoff frequency of the acoustic filter can be reduced from 4kHz to 3kHz, thereby effectively filtering out high-frequency noise while retaining most of the harmonics.
[0079] When the harmonic noise ratio is in the low threshold range, the filter order of the acoustic filter can be increased to enhance the attenuation characteristics of the acoustic filter and thus suppress noise more effectively. For example, increasing the filter order of the acoustic filter from the 2nd order to the 4th order makes the attenuation speed of the acoustic filter in the stopband faster, thereby effectively filtering out unwanted high-frequency signals and improving the clarity and fidelity of the signal.
[0080] When the harmonic noise ratio is in the low threshold range, an adaptive filtering algorithm can also be introduced to automatically adjust the parameters of the acoustic filter according to the characteristics of the input signal, so as to better track the changes in noise and improve the noise reduction effect. For example, the parameters of the acoustic filter can be adjusted by using the Least Mean Squares (LMS) adaptive filtering algorithm and the Recursive Least Squares (RLS) adaptive filtering algorithm.
[0081] When the harmonic noise ratio is in the low threshold range, the voiceprint feature extraction process of the voiceprint filter is also optimized; for example, when extracting the Mel frequency cepstral coefficients, the pre-emphasis coefficients can be increased to highlight the high-frequency part of the sound signal, and the extracted features can be smoothed to reduce the influence of noise.
[0082] When the harmonic noise ratio is in the low threshold range, the following voiceprint feature extraction processes can be optimized. For example, in the preprocessing stage, sound activity can be detected based on spectral entropy, subband energy distribution, and deep learning to locate the true speech segments from noisy audio (excluding pure noise segments). In the preprocessing noise reduction stage, at least one of spectral subtraction, Wiener filtering, wavelet transform noise reduction, and deep learning-based noise reduction can be used to reduce the noise floor. The detected speech segments can be further refined to locate the start and end points of the target user's voice, defining a more precise range for subsequent feature extraction. In the feature extraction stage, at least one of Mel-frequency cepstral coefficients, linear prediction coefficients, and deep learning-based feature extraction methods can be used to extract the corresponding voiceprint features. After extracting the original features, further processing can be performed to reduce feature fluctuations caused by noise, specifically including at least one of normalization, dimensionality reduction, linear discriminant analysis, and principal component analysis. Finally, first-order and second-order differences can be used to enhance the characterization of speech dynamics, thereby improving the robustness of the voiceprint filter for voiceprint recognition.
[0083] It is understandable that when the harmonic noise ratio is in the low threshold range, at least one of the above optimization strategies can be used to optimize the acoustic filter, and no limitation is made here.
[0084] When the harmonic noise ratio (HNR) is in the middle threshold range (e.g., 10dB ≤ HNR < 20dB), it indicates that noise components still exist, but harmonic components have become somewhat dominant. In this case, it is necessary to preserve the voiceprint characteristics of the speech signal as much as possible while reducing noise. Depending on the actual situation, the cutoff frequency of the filter can be fine-tuned. For example, if important voiceprint characteristics are found in the high-frequency part of the speech signal, the cutoff frequency of the voiceprint filter can be appropriately increased, but care should be taken to avoid introducing too much high-frequency noise.
[0085] When the harmonic noise ratio is in the middle threshold range, the filtering intensity of the voiceprint filter can be appropriately reduced to avoid over-filtering and loss of voiceprint features. For example, the attenuation coefficient of the filter can be adjusted from -30dB to -20dB.
[0086] When the harmonic noise ratio is in the middle threshold range, the extracted voiceprint features can also be normalized to reduce the differences between different speech signals, including but not limited to using mean normalization or variance normalization, so that the features have better stability in different environments.
[0087] It is understandable that when the harmonic noise ratio is in the middle threshold range, at least one of the above optimization strategies can be used to optimize the acoustic filter, and no limitation is made here.
[0088] When the harmonic noise ratio (HNR) is in the high threshold range (e.g., HNR ≥ 20 dB), it indicates that there is less noise. In this case, the interference of the filter on the user's voice signal should be minimized to preserve the complete voiceprint characteristics. This can be achieved by appropriately increasing the cutoff frequency of the voiceprint filter to increase its passband width, allowing more components of the user's voice signal to pass through. Alternatively, the filter order of the voiceprint filter can be reduced to decrease its complexity and avoid over-processing of the speech signal. For example, reducing the filter order from 4th to 2nd order will also reduce the processing time and computational resource requirements.
[0089] It should be noted that when the harmonic noise ratio is in the high threshold range, the voiceprint features extracted by the voiceprint filter may already be quite accurate. Therefore, the current parameters and performance of the voiceprint filter can be kept unchanged to avoid over-processing that leads to the loss of harmonic details.
[0090] It is understandable that when the harmonic noise ratio is in the high threshold range, at least one of the above optimization strategies can be used to optimize the acoustic filter or not, and no limitation is made here.
[0091] In this embodiment, feedback parameters are continuously obtained through the correlator, and the voiceprint filter is optimized based on these feedback parameters. This can effectively improve the performance of the voiceprint filter, thereby better filtering background noise data in subsequent voice interaction processes and improving the quality of user voice interaction.
[0092] Please see Figure 2 This application provides another embodiment of a sound data processing method, comprising the following steps:
[0093] 201. Obtain initial audio data that includes at least the target user's audio data.
[0094] 202. Label the target user's voice data and background sound data in the initial sound data to obtain the labeled initial sound data.
[0095] 203. Input the initial sound data into a pre-constructed voiceprint filter carrying the voiceprint features of the target user, so that the voiceprint filter can filter the initial sound data based on the voiceprint features of the target user to obtain filtered sound data.
[0096] 204. Input the labeled initial sound data and filtered sound data into the correlator, and perform correlation analysis on the filtered sound data and labeled initial sound data based on the correlator to obtain the analysis results.
[0097] Due to the inherent performance limitations of the voiceprint filter, it may misfilter or miss filtering during the process of filtering background noise data in the initial sound data. That is, the voiceprint filter may not filter out some background noise data in the initial sound data, or it may misfilter out some target user's voice data in the initial sound data. Therefore, in order to address the potential errors of the voiceprint filter, it is necessary to perform a correlation analysis on the filtered sound data obtained by the voiceprint filter using a correlator to determine whether there is still some background noise data in the filtered sound data and whether some target user's voice data is missing, thereby ensuring the accuracy of data processing.
[0098] Based on correlator analysis, there are three possible analysis results: the first is that the filtered audio data contains only some background noise data (missed filtering); the second is that the filtered audio data is missing only some target user audio data (false filtering); and the third is that the filtered audio data contains both some background noise data and some target user audio data (both missed filtering and false filtering exist).
[0099] In applications, correlation analysis includes: the correlator first aligns the signal of the filtered sound data and the signal of the labeled initial sound data, then multiplies the signal after alignment, and finally integrates (for continuous signals) or sums (for discrete signals) the result of the signal multiplication to obtain the correlation value between the two signals and outputs it.
[0100] 205. Process the filtered sound data based on the analysis results to obtain the target sound data.
[0101] If the first analysis result mentioned above is that the filtered sound data only contains some background sound data, then the background sound data contained in the filtered sound data needs to be removed, and the target sound data is obtained after the removal.
[0102] If the second analysis result mentioned above is that the filtered audio data is missing only part of the target user's audio data, then the missing part of the target user's audio data needs to be added to the filtered audio data to obtain the target audio data.
[0103] For the third analysis result mentioned above, that is, the filtered sound data contains both background sound data and missing target user sound data, it is necessary to remove the background sound data contained in the filtered sound data and add the missing target user sound data to the filtered sound data to obtain the target sound data.
[0104] 206. Extract the voiceprint features of the target user from the target audio data.
[0105] In this embodiment, after obtaining the target voice data, the voiceprint features of the target user can be extracted from the target voice data using a model with voiceprint feature extraction capabilities. The model with voiceprint feature extraction capabilities includes, but is not limited to, any one of the following voiceprint recognition models: Gaussian mixture model, deep neural network model, convolutional neural network model, recurrent neural network model, long short-term memory network model, etc. The specific language type of the target user includes, but is not limited to, any one of the following: Chinese, English, French, Italian, Mongolian, etc.
[0106] 207. Output the extracted voiceprint features to the voiceprint filter, optimize the voiceprint filter based on the voiceprint features, and after returning to execute the step of obtaining initial sound data that includes at least the target user's voice data, use the optimized voiceprint filter to filter the initial sound data.
[0107] In this embodiment, after the extracted voiceprint features are output to the voiceprint filter, the voiceprint features of the target user stored in the voiceprint filter are compared with the voiceprint features extracted from the target sound data in this instance. This optimizes and saves the voiceprint features stored in the voiceprint filter, so that after returning to the step of obtaining initial sound data that includes at least the sound data of the target user, the optimized voiceprint filter can be used to filter the initial sound data, thereby achieving better recognition and noise reduction effects.
[0108] It is understood that in this embodiment, steps 201 to 203 are the same as steps 101 to 103 in the previous embodiment. Therefore, the relevant content of steps 201 to 203 refers to the relevant description in the previous embodiment, and will not be repeated here.
[0109] Please see Figure 3 This application provides another embodiment of a sound data processing method, which further includes the following steps related to pre-constructing a voiceprint filter:
[0110] 301. Obtain sample sound data of the target user under preset environmental conditions.
[0111] In applications, preset environmental conditions include environments with noise levels within a certain range, such as environments below 30dB or 20dB. This is just an example and is not a limitation.
[0112] When acquiring sample audio data of a target user, it can be obtained through a built-in audio data acquisition device in the user terminal, such as a microphone or other device with sound pickup function. In practical applications, it can also be acquired through an external audio data acquisition device of the user terminal (e.g., headphones with microphone or sound pickup function, including wireless headphones and wired headphones). There is no limitation here.
[0113] It is understandable that the number of sample audio data obtained in this step can be determined according to actual needs, and is not limited here.
[0114] 302. Extract the voiceprint features of the target user contained in the preprocessed sample audio data, and train the voiceprint recognition model based on the voiceprint features of the target user.
[0115] The sample audio data is preprocessed (preprocessing includes at least one of the following: initial noise reduction, echo cancellation, silence segment clipping, frame windowing, normalization, endpoint detection, and data format unification, to improve the quality of the sample audio data and facilitate subsequent processing). Parameters that reflect the voiceprint characteristics of the target user are extracted from the preprocessed sample audio data. These voiceprint characteristic parameters are then input into the voiceprint recognition model to be trained, and a pre-trained voiceprint recognition model is obtained after training.
[0116] 303. Obtain a voiceprint filter based on the trained voiceprint recognition model and determine the similarity threshold.
[0117] The similarity threshold is used to determine whether the initial sound data contains background noise data.
[0118] It is understandable that the voiceprint recognition model trained here can be the same model as the pre-trained voiceprint recognition model mentioned above. In some cases, they can also be two different voiceprint recognition models. This is not a limitation here.
[0119] Using a trained voiceprint recognition model, the voiceprint embedding vector of the target user is extracted from the sample sound data. This vector can represent the voiceprint features of the target user. A voiceprint filter is constructed based on the extracted voiceprint embedding vector of the target user. At the same time, the voiceprint embedding vector of the new sample sound data (including the voiceprint embedding vector of the target user and the voiceprint embedding vector of the non-target user) and the extracted voiceprint embedding vector of the target user are input, and the similarity between the two is calculated. The voiceprint embedding vector of the target user is used as a positive sample, and the voiceprint embedding vector of the non-target user is used as a negative sample. The voiceprint filter is trained to distinguish the voiceprint of the target user and the voiceprint of the non-target user. The similarity threshold is determined and saved according to the corresponding training data, and finally a voiceprint filter with filtering function is obtained.
[0120] It should be noted that a preliminary voiceprint filter can be obtained through steps 301 to 303. In this embodiment, after step 303, the audio data processing method steps 101 to 105 or steps 201 to 210 of the aforementioned embodiment are also included. For details, please refer to the relevant descriptions of the aforementioned embodiments, which will not be repeated here.
[0121] Please see Figure 4 This application provides another embodiment of a sound data processing method, comprising:
[0122] 401. Obtain initial audio data that includes at least the target user's audio data.
[0123] 402. Label the target user's voice data and background sound data in the initial sound data to obtain the labeled initial sound data.
[0124] 403. Extract the voiceprint features contained in the initial sound data.
[0125] As can be seen from the relevant content of the foregoing embodiments, the voiceprint filter mentioned in the embodiments of this application has the voiceprint features of the target user, and the voiceprint filter also has the voiceprint feature extraction function, which can extract all the voiceprint features contained in the initial sound data input therein (including the voiceprint features of the target user and the voiceprint features of non-target users in the initial sound data).
[0126] 404. Match the extracted voiceprint features with the voiceprint features of the target user in the voiceprint filter to obtain the similarity between the extracted voiceprint features and the voiceprint features of the target user.
[0127] As described in the foregoing embodiments, this voiceprint filter can determine, based on a similarity threshold, which voiceprint features of the target user and which voiceprint features of non-target users are contained in the input initial sound data. Specifically, the determination method compares the numerical relationship between the similarity between the voiceprint features extracted from the initial sound data and the voiceprint features of the target user stored in the voiceprint filter, and the similarity threshold, thereby determining the sound data corresponding to the voiceprint features that need to be filtered.
[0128] 405. When the similarity is less than the similarity threshold, the sound data with the extracted voiceprint features is removed from the initial sound data to obtain filtered sound data that retains the sound data of the target user.
[0129] If the similarity between the voiceprint features extracted by the voiceprint filter from the initial sound data and the voiceprint features of the target user stored in the voiceprint filter is less than the similarity threshold, then all sound data with the extracted voiceprint features will be removed from the initial sound data. If the similarity between the voiceprint features extracted by the voiceprint filter from the initial sound data and the voiceprint features of the target user stored in the voiceprint filter is greater than or equal to the similarity threshold, then all sound data with the extracted voiceprint features will be retained, and finally, filtered sound data with background noise data filtered out and only the target user's sound data retained will be obtained.
[0130] It is understood that, in this embodiment, the initial sound data input to the voiceprint filter can be preprocessed initial sound data, or the initial sound data can be preprocessed by the voiceprint filter after being input to the voiceprint filter, and then the voiceprint features can be extracted from the preprocessed initial sound data. There is no limitation here.
[0131] 406. Input the labeled initial sound data and filtered sound data into the correlator so that the correlator can perform correlation analysis on the filtered sound data and the labeled initial sound data to obtain feedback parameters for evaluating the performance of the acoustic filter.
[0132] 407. Output the feedback parameters to the voiceprint filter, optimize the voiceprint filter according to the feedback parameters, and after returning to the step of obtaining initial sound data including at least the target user's voice data, use the optimized voiceprint filter to filter the initial sound data.
[0133] It is understood that this embodiment may or may not include the relevant content of steps 201 to 203 of the aforementioned embodiments, and this is not limited here. Steps 401 to 402 of this embodiment are consistent with steps 101 to 102 of the aforementioned embodiments, and steps 406 to 407 of this embodiment are consistent with steps 104 to 105 of the aforementioned embodiments. Therefore, the relevant content of steps 401 to 402 and 406 to 407 refers to the relevant description of the aforementioned embodiments, and will not be repeated here.
[0134] Based on the description of the foregoing embodiments, it can be seen that the sound data processing method provided by this application can not only effectively identify noise data from non-target users, but also effectively filter it out. Compared with traditional methods, this solution does not rely on the hardware design of the terminal device, which can reduce the expenditure of hardware design costs. It is also applicable to dealing with multi-source interference. Furthermore, in the process of sound data processing, the combination of the corresponding processing of the acoustic filter and correlator can not only avoid over-filtering of user voice data, but also avoid weakening the important low-frequency components of user voice data, which could lead to sound blurring or distortion.
[0135] It is understood that any parts not detailed in this embodiment can be found in the relevant descriptions in the other embodiments described above, and will not be repeated here.
[0136] It is understood that in the various method embodiments of this application, the order of the steps does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0137] Please see Figure 5 The electronic device 500 provided in this application embodiment may include one or more central processing units (CPUs) 501 and a memory 505, wherein the memory 505 stores one or more applications or data.
[0138] The memory 505 can be volatile or persistent storage. The program stored in the memory 505 can include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the central processing unit 501 can be configured to communicate with the memory 505 and execute the series of instruction operations stored in the memory 505 on the electronic device 500.
[0139] Electronic device 500 may also include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input / output interfaces 504, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0140] The central processing unit 501 can perform the operations in any specific method embodiment of the first aspect described above, which will not be described in detail here.
[0141] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.
[0142] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.
[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0144] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0145] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0146] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a server or terminal device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0148] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for processing sound data, characterized in that, The method includes: Acquire initial audio data, including at least the target user's voice data; The target user's voice data and background sound data in the initial sound data are labeled to obtain labeled initial sound data; The initial sound data is input into a pre-constructed voiceprint filter carrying the voiceprint features of the target user, so that the voiceprint filter performs filtering processing on the initial sound data based on the voiceprint features of the target user to obtain filtered sound data. The labeled initial sound data and the filtered sound data are input into a correlator so that the correlator performs correlation analysis on the filtered sound data and the labeled initial sound data to obtain feedback parameters for evaluating the performance of the acoustic filter. The feedback parameters are output to the voiceprint filter, the voiceprint filter is optimized according to the feedback parameters, and after returning to the step of obtaining initial sound data including at least the target user's voice data, the optimized voiceprint filter is used to filter the initial sound data.
2. The sound data processing method according to claim 1, characterized in that, The feedback parameters include the harmonic noise ratio; The optimization of the acoustic filter based on the feedback parameters includes: The optimization strategy for the acoustic texture filter is determined based on the threshold range of the harmonic noise ratio. The acoustic filter is optimized according to the optimization strategy described above.
3. The sound data processing method according to claim 2, characterized in that, The step of determining the optimization strategy for the acoustic signature filter based on the threshold range of the harmonic noise ratio includes: If the harmonic noise ratio is in the low threshold range, the optimization strategy for the acoustic print filter is determined to include at least one of the following: adjusting the cutoff frequency of the acoustic print filter, increasing the filter order of the acoustic print filter, selecting an adaptive filtering algorithm for optimization, and optimizing the acoustic print feature extraction process of the acoustic print filter. If the harmonic noise ratio is in the middle threshold range, then the optimization strategy for the acoustic text filter is determined to include at least one of the following: adjusting the cutoff frequency of the acoustic text filter, adjusting the filtering strength of the acoustic text filter, and adding normalization processing to the acoustic text features extracted by the acoustic text filter. If the harmonic noise ratio is in the high threshold range, the optimization strategy for the acoustic pattern filter is determined to include at least one of adjusting the cutoff frequency of the acoustic pattern filter, reducing the filter order of the acoustic pattern filter, and maintaining the current performance of the acoustic pattern filter.
4. The sound data processing method according to claim 1, characterized in that, After inputting the labeled initial sound data and the filtered sound data into the correlator, the method further includes: Based on the correlator, a correlation analysis is performed on the filtered sound data and the labeled initial sound data to obtain the analysis results; The filtered sound data is processed based on the analysis results to obtain the target sound data; Extract the voiceprint features of the target user from the target sound data; The extracted voiceprint features are output to the voiceprint filter. The voiceprint filter is optimized based on the voiceprint features. After returning to the step of obtaining initial sound data that includes at least the target user's voice data, the optimized voiceprint filter is used to filter the initial sound data.
5. The sound data processing method according to claim 4, characterized in that, The step of processing the filtered sound data according to the analysis results to obtain the target sound data includes: The background noise data contained in the filtered sound data is removed and / or the target user's sound data is added to the filtered sound data to obtain the target sound data.
6. The sound data processing method according to any one of claims 1 to 5, characterized in that, Also includes: Acquire sample audio data of the target user under preset environmental conditions; Extract the voiceprint features of the target user contained in the preprocessed sample audio data, and train a voiceprint recognition model based on the voiceprint features of the target user; Based on the trained voiceprint recognition model, a voiceprint filter is obtained, and a similarity threshold is determined. The similarity threshold is used to determine whether the initial sound data contains background sound data.
7. The sound data processing method according to claim 6, characterized in that, The step of inputting the initial sound data into a pre-constructed voiceprint filter carrying the voiceprint features of the target user, so that the voiceprint filter performs filtering processing on the initial sound data based on the voiceprint features of the target user to obtain filtered sound data, includes: Extract the voiceprint features contained in the initial sound data; The extracted voiceprint features are matched with the voiceprint features of the target user in the voiceprint filter to obtain the similarity between the extracted voiceprint features and the voiceprint features of the target user. When the similarity is less than the similarity threshold, the sound data with the extracted voiceprint features is removed from the initial sound data to obtain filtered sound data that retains the sound data of the target user.
8. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the sound data processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on a computer, cause the computer to perform the sound data processing method as described in any one of claims 1 to 7.
10. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the sound data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Animal sound automatic labeling method and system based on space-time adaptive threshold value
CN122050396A
Animal call automatic labeling method and system based on spatio-temporal adaptive threshold
CN122050396B