Speech enhancement method, device and equipment based on environmental noise adaptation
By adopting an environmental noise-adaptive speech enhancement method, and combining a microphone array and a scheme library, dynamic perception and in-sampling noise reduction are achieved. This solves the problem of noise spectrum and human voice aliasing in complex acoustic environments in traditional speech enhancement techniques, improves speech quality and recognition accuracy, and saves computing resources.
Patent Information
- Application Number
- CN202510474340.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Traditional speech enhancement technologies suffer from noise spectrum mixing with human voice in signals collected in complex acoustic environments, resulting in poor back-end processing and affecting speech recognition accuracy.
An environmental noise-adaptive speech enhancement method is adopted. Audio is collected in real time through a microphone array, sound features are extracted, and a noise reduction scheme is determined by combining spatial distribution features and a scheme library. The microphone array is controlled to collect audio and perform speech enhancement, realizing a closed-loop architecture of dynamic perception, in-process noise reduction and back-end enhancement.
It improves the quality of signals input to back-end processing in complex acoustic environments, reduces noise spectrum aliasing with human voice, enhances speech quality, improves the reliability of speech recognition, and saves computing resources.
Smart Images

Figure CN120340515B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a speech enhancement method, apparatus, and device based on environmental noise adaptation. Background Technology
[0002] Speech recognition accuracy is the cornerstone of reliability in smart home and in-vehicle interaction scenarios, while speech quality directly affects recognition performance. Traditional speech enhancement technology adopts a separate architecture of "fixed acquisition + back-end processing". However, in the fixed acquisition mode, the signal acquired in complex acoustic environments has noise spectrum and human voice highly mixed, resulting in poor signal quality at the input back-end. The back-end processing algorithm is forced to extract speech from the damaged signal, resulting in poor effect of back-end processing in improving speech quality. Summary of the Invention
[0003] Based on this, it is necessary to address the technical problem that the existing technology's "fixed acquisition + back-end processing" split architecture is not effective in improving speech quality in complex acoustic environments. Therefore, a speech enhancement method, device, and equipment based on environmental noise adaptation are proposed.
[0004] Firstly, a speech enhancement method based on environmental noise adaptation is provided, the method comprising:
[0005] Based on the first time interval, the audio collected in real time by the microphone array in the target microphone is obtained as the first audio.
[0006] Based on the first audio, sound features are extracted to obtain the first feature;
[0007] Based on the solution library, and according to the spatial distribution characteristics of the microphone array and the first feature, a noise reduction solution is determined.
[0008] According to the noise reduction scheme, the microphone array is controlled to collect audio as the second audio.
[0009] Speech enhancement is performed based on the second audio to obtain the target audio.
[0010] Secondly, a speech enhancement device based on environmental noise adaptation is provided, the device comprising:
[0011] The first acquisition module is used to acquire audio collected in real time by the microphone array in the target microphone based on a first time interval, as the first audio.
[0012] The feature extraction module is used to extract sound features based on the first audio to obtain the first feature;
[0013] The scheme determination module is used to determine the noise reduction scheme based on the scheme library, the spatial distribution characteristics of the microphone array and the first feature;
[0014] The second acquisition module is used to control the microphone array to acquire audio according to the noise reduction scheme, as the second audio.
[0015] The processing module is used to perform speech enhancement based on the second audio to obtain the target audio.
[0016] Thirdly, a computer device is provided, the target device comprising: a target microphone and a device body, the device body being communicatively connected to the target microphone, the device body comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the aforementioned speech enhancement method based on environmental noise adaptation.
[0017] This application discloses a speech enhancement method, apparatus, and device based on adaptive environmental noise. It extracts sound features (first features) from a first audio recording and then combines these with the spatial distribution features of the microphone array and a scheme library to determine an in-sampling noise reduction scheme. This method accurately determines the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio) according to this in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio to obtain the target audio. Overall, this speech enhancement method based on adaptive environmental noise achieves a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement." Compared to the traditional separate architecture of "fixed acquisition + back-end processing," this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improving the signal quality input to the back-end processing and thus enhancing the speech quality improvement effect of the back-end processing. Furthermore, by activating the in-sampling noise reduction scheme based on a first time interval, frequent determination of the in-sampling noise reduction scheme is not required, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computational resources. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] in:
[0020] Figure 1 This is an application environment diagram of a speech enhancement method based on environmental noise adaptation in one embodiment;
[0021] Figure 2 This is a flowchart of a speech enhancement method based on environmental noise adaptation in one embodiment;
[0022] Figure 3 This is a structural block diagram of a speech enhancement device based on environmental noise adaptation in one embodiment;
[0023] Figure 4 This is a structural block diagram of the target device in one embodiment. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] The speech enhancement method based on environmental noise adaptation provided in this invention can be applied to, for example... Figure 1 In the application environment, the application environment includes a target device, which comprises a target microphone 1 and a device body 2. The target microphone 1 and the device body 2 are connected via wired or wireless communication technology. The target microphone 1 is equipped with a microphone array. The microphone array includes at least two microphones, which work collaboratively through software. Each microphone in the microphone array faces a different direction to collect sound from sound sources in different directions and form audio.
[0026] The target device is configured to implement the environmental noise-adaptive speech enhancement method of this application. The device body 2 of the target device is used for: acquiring audio collected in real-time by a microphone array in a target microphone based on a first time interval, as a first audio; extracting sound features based on the first audio to obtain a first feature; determining a noise reduction scheme based on a scheme library, according to the spatial distribution features corresponding to the microphone array and the first feature; controlling the microphone array to collect audio according to the noise reduction scheme, as a second audio; and performing speech enhancement based on the second audio to obtain the target audio.
[0027] This application extracts sound features (first features) from a first audio recording and then combines these features with the spatial distribution characteristics of the microphone array and a scheme library to determine an in-sampling noise reduction scheme. This method can accurately determine the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio) according to this in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio to obtain the target audio. Overall, this speech enhancement method based on environmental noise adaptation achieves a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement." Compared to the traditional separate architecture of "fixed acquisition + back-end processing," this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improving the signal quality input to the back-end processing and thus enhancing the speech quality improvement effect of the back-end processing. Furthermore, starting the in-sampling noise reduction scheme based on a first time interval eliminates the need for frequent determination of the in-sampling noise reduction scheme, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computational resources.
[0028] The application environment may further include: a client and a server. The client communicates with the main body 2 of the target device. The server communicates with the client.
[0029] Optionally, users can interact with the target device through a client.
[0030] Optionally, users can input control data through the client, and the client will upload the control data to the server.
[0031] Optionally, the user obtains initial data from the server through the client, modifies the initial data on the client to obtain modified data, and controls the target device based on the modified data.
[0032] The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0033] The present invention will now be described in detail through specific embodiments.
[0034] Please see Figure 2 As shown, Figure 2 A flowchart illustrating an environmental noise-adaptive speech enhancement method provided in an embodiment of the present invention includes the following steps:
[0035] S1: Based on the first time interval, acquire the audio collected in real time by the microphone array in the target microphone, and use it as the first audio;
[0036] The first time interval is a time period. The first time interval can be data set by default in the application program, data entered by the user in advance, or data actively generated by the application program file according to preset conditions.
[0037] Specifically, at the beginning of each time period of the first time interval, the audio collected in real time by the microphone array in the target microphone is acquired, and the acquired audio is used as the first audio.
[0038] Optionally, the first time interval is determined by scene pattern recognition to adapt to the non-stationarity of environmental noise.
[0039] Scene modes include, but are not limited to: quiet environment mode, noisy environment mode, and meeting scene mode.
[0040] The Quiet Environment mode is primarily suitable for indoor environments with minimal background noise, such as quiet offices or studies. In this mode, audio processing may focus on high-fidelity reproduction of sound details, such as accurately presenting the timbre of instruments and the layering of music. For speech processing, the emphasis may be on the naturalness of the speech, avoiding excessive noise reduction to prevent affecting the original timbre.
[0041] Noisy Environment Mode is used in scenarios with significant background noise, such as construction sites, busy streets, or noisy factory workshops. In this mode, noise reduction is enhanced, employing more sophisticated algorithms to identify and remove various types of noise, such as the low-frequency roar of industrial equipment and the high-frequency honking of cars on the street. For speech processing, techniques such as beamforming may be used to focus on the target speech source, while simultaneously enhancing speech intensity and clarity to ensure accurate speech recognition even in noisy environments.
[0042] The meeting scenario mode, designed for conference room environments, focuses on handling multi-person voice interactions. It employs multi-microphone array technology to differentiate the direction of different speakers' voices, performing sound separation and enhancement. Echo cancellation may also be included, as echoes can occur due to sound reflections in conference rooms. Furthermore, voice optimization ensures that each speaker's speech is clear and intelligible, maintaining good voice quality during long-distance calls (such as video conferencing).
[0043] S2: Extract sound features based on the first audio to obtain the first feature;
[0044] Specifically, a pre-trained sound feature extraction model is used to extract sound features from the first audio, and the extracted sound features are used as the first feature.
[0045] The pre-trained speech feature extraction model is a model pre-trained on a neural network (e.g., convolutional neural network, recurrent neural network) using the publicly available LibriNoise dataset, which contains 100,000 pairs of noisy and clean speech. LibriNoise is an open-source speech dataset primarily used for speech separation and noise robustness research. This dataset is derived from the LibriSpeech dataset and simulates real-world noise interference by adding various types of noise to the original speech.
[0046] S3: Based on the solution library, determine the noise reduction solution in the sampling process according to the spatial distribution characteristics of the microphone array and the first feature;
[0047] The scheme library is a pre-defined database storing noise reduction schemes. It describes the association data between scheme identifiers and noise reduction schemes. A scheme identifier can be data that uniquely identifies a noise reduction scheme, such as the scheme name or scheme ID. For example, the scheme library includes the following parameter mappings: Scheme ID-001 corresponds to a beamforming main lobe angle of 30° and a high-pass filter cutoff frequency of 200Hz; Scheme ID-002 corresponds to enabling the blind source separation algorithm and a noise suppression threshold of -15dB.
[0048] Specifically, classification and prediction are performed based on the spatial distribution characteristics of the microphone array and the first feature. The vector element with the largest value is extracted from the prediction result. The scheme identifier corresponding to the extracted vector element is determined as the target scheme identifier. Noise reduction schemes are selected from the scheme library based on the determined target scheme identifier. The selected noise reduction scheme is used as the noise reduction scheme in the sampling.
[0049] Optionally, the spatial distribution features and the first feature are concatenated and input into a pre-trained first scheme model for classification and prediction to obtain a target scheme identifier. The target scheme identifier is then used to select a denoising scheme from the scheme library, and the selected denoising scheme is used as the denoising scheme in the sampling.
[0050] The first pre-trained model is a pre-trained multi-class classification model. The model structure and training method of the first model can be selected from existing technologies, which will not be elaborated here.
[0051] In-speech noise reduction schemes utilize control data from the microphone array's audio acquisition process. The purpose of these schemes is to reduce noise during audio acquisition. Optionally, the in-speech noise reduction scheme may include one or more of the following: hardware control parameters, signal processing strategies, and environmental adaptation information.
[0052] Hardware control parameters are the physical adjustment data of the microphone array. These parameters include beamforming parameters and microphone on / off strategies. Beamforming parameters specifically include: gain weights for each microphone (e.g., main lobe enhancement, side lobe suppression), signal delays for each microphone calculated based on the sound source direction (for phase alignment), beam patterns (e.g., selection of omnidirectional, cardioid, supercardioid, etc.), and main lobe coverage angle (e.g., narrow beam focusing on a single sound source, wide beam covering multi-person conversations). Microphone on / off strategies specifically include: dynamically enabling / disabling some microphones based on the sound source location (reducing noise interference), and simultaneously forming multiple beams to track different sound sources (e.g., multi-speaker separation in conference scenarios).
[0053] Signal processing strategies are the core data for algorithm-level noise reduction, specifically including: noise suppression parameters, speech enhancement parameters, and spatial filtering parameters. Noise suppression parameters include filter types (such as the cutoff frequencies of high-pass / low-pass / band-pass filters in adaptive LMS / NLMS algorithms) and spectral attenuation thresholds based on noise estimation (used for spectral subtraction to suppress steady-state noise). Speech enhancement parameters cover Wiener filter gain (dynamically adjusting frequency domain gain based on signal-to-noise ratio) and deep learning model selection (such as real-time speech enhancement models like RNNoise and DEMUCS). Spatial filtering parameters involve mixture matrix estimation for blind source separation (BSS) (such as Independent Component Analysis (ICA)) and sound source localization methods (such as calculating the sound source direction angle based on TDOA time difference or SRP-PHAT phase transform).
[0054] Environmental adaptation information is configuration data for dynamic scene adjustments, specifically including: spatial distribution characteristics and dynamic scene adaptation. Spatial distribution characteristics: encompass the geometry of the microphone array (e.g., linear / circular / spherical layout coordinates) and acoustic environment parameters (e.g., room reverberation time RT60, background noise spectral characteristics). Dynamic scene adaptation: includes noise type identification (classifying steady-state / transient, broadband / narrowband noise to switch processing strategies) and sound source tracking strategies (e.g., Kalman filtering to predict moving sound source trajectories, multi-target sound source priority allocation).
[0055] S4: According to the noise reduction scheme, control the microphone array to collect audio as the second audio;
[0056] Specifically, according to the in-sampling noise reduction scheme, the array control data of the microphone array is updated, and the microphone array acquires audio based on the updated array control data, using the acquired audio as the second audio. The in-sampling noise reduction scheme enables the microphone array to perform noise reduction during audio acquisition based on the updated array control data, thereby improving the signal quality of the second audio.
[0057] S5: Perform speech enhancement based on the second audio to obtain the target audio.
[0058] Specifically, speech enhancement is performed based on the second audio to achieve one or more of the following operations: noise reduction, dereverberation, interference removal, distortion correction, and adaptation to downstream tasks. The second audio with enhanced speech is then used as the target audio.
[0059] Noise Reduction: Suppresses background noise (such as wind noise, keyboard noise, and crowd noise). De-reverberation: Reduces speech blurring caused by room reflections. Interference Removal: Eliminates non-target sound sources (such as non-target speakers in multi-person conversations). Distortion Correction: Compensates for speech impairments caused by equipment or transmission (such as compression distortion and stuttering). Downstream Task Adaptation: Provides cleaner input signals for tasks such as Automatic Speech Recognition (ASR) and Voiceprint Recognition.
[0060] Optionally, the second audio is subjected to channel separation to obtain separated channel data; echo cancellation is performed on the separated channel data to obtain clear channel output data; sound localization is performed on the clear channel output data to obtain audio data with spatial localization; sound field simulation is performed on the audio data with spatial localization to obtain simulated audio field data; fine-grained enhancement is performed on the simulated audio field data to obtain enhanced audio data; dynamic range adjustment and sound balance are performed on the enhanced audio data to output the target audio.
[0061] This embodiment extracts sound features (first features) from the first audio recording and then combines them with the spatial distribution features of the microphone array and a scheme library to determine the in-sampling noise reduction scheme. This method can accurately determine the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio recording) according to the in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio recording to obtain the target audio. Overall, this speech enhancement method based on environmental noise adaptation realizes a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement". Compared with the traditional separate architecture of "fixed acquisition + back-end processing", this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improve the signal quality input to the back-end processing, and help improve the effect of back-end processing in improving speech quality. In addition, starting the in-sampling noise reduction scheme based on the first time interval eliminates the need for frequent determination of the in-sampling noise reduction scheme, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computing resources.
[0062] In one embodiment, the step of determining the noise reduction scheme based on the scheme library, according to the spatial distribution characteristics corresponding to the microphone array and the first feature, includes:
[0063] S31: Based on a preset time window and taking the end time of the first audio as the end time of the time window, obtain the sound features corresponding to the audio collected by the microphone array as a feature packet;
[0064] Specifically, the end time of the first audio is taken as the end time of the time window, and the sound features corresponding to the audio collected by the microphone array within the preset time window are obtained, and all the obtained sound features are taken as a feature packet.
[0065] A feature packet is a data sequence. The sound features in the feature packet are sorted from earliest to latest according to the acquisition time of the audio corresponding to the sound feature.
[0066] S32: Based on the feature package, predict the trend of future sound features to obtain the feature trend prediction result;
[0067] Specifically, the feature package is input into a pre-trained trend prediction model to predict the trend of future sound features, and the predicted data is used as the feature trend prediction result.
[0068] The pre-trained trend prediction model is a model pre-trained based on an LSTM (Long Short-Term Memory) time-series prediction network. It is understood that the pre-trained trend prediction model can also be trained based on other models / neural networks capable of trend prediction; this is not a limitation here.
[0069] The characteristic trend prediction results describe the trend of noise spectrum changes and the probability of human voice distribution within a future time window.
[0070] S33: The feature trend prediction result, the spatial distribution feature, and the first feature are input into the pre-trained first scheme model for classification and prediction to obtain the target scheme identifier;
[0071] Specifically, the feature trend prediction result, the spatial distribution feature, and the first feature are concatenated and input into a pre-trained first scheme model for classification and prediction to obtain the target scheme identifier.
[0072] S34: Determine the noise reduction scheme based on the target scheme identifier and the scheme library.
[0073] Specifically, the noise reduction scheme with the same scheme identifier as the target scheme identifier in the scheme library is selected as the noise reduction scheme in the sampling.
[0074] This embodiment first acquires sound features to form a feature package by using a preset time window and the end time of the first audio recording as the end time of the time window. This fully utilizes the sound features of recently acquired audio, making the acquired feature package highly timely and representative. Based on this feature package, future sound feature trends are predicted to obtain feature trend prediction results. This helps to understand potential changes in the sound environment in advance (such as predicting transient noise in scenarios like airports and subways), thus enabling more proactive development of in-sampling noise reduction schemes. Then, the feature trend prediction results, spatial distribution features, and the first feature are input into a pre-trained first scheme model for classification and prediction to obtain a target scheme identifier. Utilizing the model's learning ability, a suitable scheme identifier can be accurately determined based on various relevant factors. Finally, based on the target scheme identifier and the scheme library, an in-sampling noise reduction scheme is determined. This accurately selects the in-sampling noise reduction scheme that best matches the current sound environment and sound features from the existing scheme library, thereby effectively improving the noise reduction effect during the process of acquiring audio from a microphone array in complex and ever-changing environments. It can also adapt to changes in environmental noise, improving the quality of the final target audio.
[0075] In one embodiment, before the step of acquiring the audio collected in real time by the microphone array in the target microphone based on a first time interval as the first audio, the method further includes:
[0076] S11: Obtain the current scene mode based on the second time interval and use it as the target mode;
[0077] Optionally, at the start of the second time interval, the user-inputted scene mode is obtained as the current scene mode, and this current scene mode is used as the target mode.
[0078] Optionally, at the start of the second time interval, the scene mode can be automatically determined by comprehensively analyzing the audio collected by the microphone array. For example, in addition to matching the spatial sound intensity distribution of the microphone array with a historical noise database, specific frequency components in the collected audio can also be analyzed. If the high-frequency components are high and have a certain regularity, it may indicate a scene with a lot of metal collision sounds, such as in a machine shop; if the low-frequency components are prominent and there is a continuous low-frequency hum, it may indicate a scene near large motors or transformers. At the same time, the sound type in the audio signal can be identified. For example, if a high proportion of natural sounds such as birdsong and wind are identified, it may be an outdoor natural scene; if a lot of human voices are identified, such as keyboard typing, it may be an office scene. In addition, the rate of change of sound energy distribution in different directions can be analyzed. If the sound energy changes rapidly and frequently in a certain direction, it may be a traffic scene, because changes in the direction of vehicle travel will cause the sound energy to change rapidly in different directions. By comprehensively judging these multi-dimensional audio features, the scene mode can be automatically determined.
[0079] Understandably, when the target device is idle, the scene mode can be automatically determined by comprehensively analyzing the audio collected by the microphone array. This determined scene mode is then stored, and at the beginning of the second time interval, it is retrieved as the target mode. First, utilizing idle time for analysis fully leverages the device's resources, avoiding additional computational burden during peak usage and improving the overall operating efficiency of the target device. Second, determining the scene mode through comprehensive analysis of the collected audio accurately reflects the sound characteristics of different scenes, resulting in high accuracy. Storing this mode and retrieving it at the beginning of the second time interval reduces the time required for real-time scene mode determination, enabling rapid acquisition of suitable scene modes and improving response speed. Furthermore, the stored scene mode serves as empirical data that can be directly used in subsequent operations, avoiding repeated analysis of the same or similar audio data, saving computational resources, and contributing to improved adaptability and stability of this application in different scenarios.
[0080] S12: Determine the time interval using a lookup table method based on the target pattern, and use it as the initial time interval;
[0081] Specifically, the time interval is determined by a lookup table method based on the target pattern, and the time interval obtained from the lookup table is used as the initial time interval.
[0082] S13: Obtain the acquisition error corresponding to the target mode;
[0083] The acquisition error corresponding to the target mode refers to the error generated when the microphone array acquires audio based on the default array control data in the target mode. Optionally, the acquisition error includes errors in spectral aliasing, input signal-to-noise ratio (SNR), temporal fluctuation index (TFI), spatial consistency, and historical matching deviation. Spectral aliasing is the proportion of noise energy overlapping with human voice in the frequency domain, reflecting the degree of noise contamination of the speech signal's frequency band. Input SNR is the ratio of the energy of the speech segment to the energy of the noise segment in the acquired signal, reflecting the quality of the input signal. TFI is the number of abrupt changes in signal energy per unit time, characterizing the non-stationarity of noise (such as sudden noise). Spatial consistency error is the spatial directivity deviation of the signals in each channel of the microphone array, reflecting the degree of array beamforming failure. Historical matching deviation is the difference between the current time interval and the historical optimal interval, used to evaluate scene adaptability. The specific calculation formulas for spectral aliasing, input SNR, TFI, spatial consistency, and historical matching deviation can be selected from existing technologies and will not be elaborated here.
[0084] Specifically, the acquisition error corresponding to the target mode can be obtained from the preset storage space.
[0085] It is understood that the acquisition error corresponding to the target mode can be determined when the target device is idle or based on the user-input start signal, and the determined acquisition error corresponding to the target mode can be stored. When needed, the stored acquisition error corresponding to the target mode can be retrieved.
[0086] S14: Based on the acquisition error, the initial time interval is calibrated to obtain the first time interval.
[0087] Optionally, based on the acquisition error, the initial time interval is calibrated using a preset calibration formula, and the calibrated initial time interval is used as the first time interval.
[0088] The preset calibration formula can be obtained by fitting historical data, and no limitation is made here.
[0089] The specific operations of calibration include: keeping it unchanged, extending it, and shortening it.
[0090] In another embodiment of this application, before the step of acquiring the audio collected in real time by the microphone array in the target microphone based on the first time interval as the first audio, the method further includes: acquiring the current scene mode based on the second time interval as the target mode; and determining the time interval based on the target mode using a lookup table method as the first time interval.
[0091] This embodiment first obtains the target mode based on the second time interval, which helps to accurately locate the current environmental scene type and provides a basic framework for subsequent operations. Next, a lookup table method is used to determine the initial time interval. This method can quickly match the corresponding preset time interval based on the scene mode, improving the system's response speed. Then, the acquisition error corresponding to the target mode is obtained, which allows the system to clearly identify possible deviations during the acquisition process in the current scene mode. Finally, the initial time interval is calibrated based on the acquisition error to obtain the first time interval, which can effectively compensate for various errors during the acquisition process, making the timeliness of re-determining the noise reduction scheme based on the first audio acquired based on the first time interval more accurate. This series of operations improves the adaptability of the speech enhancement method in different environmental scenes and the timeliness of re-determining the noise reduction scheme, thereby optimizing the speech enhancement effect.
[0092] In one embodiment, after the step of obtaining the current scene mode as the target mode according to the second time interval, the method further includes:
[0093] S15: Obtain the adjustment data corresponding to the target mode;
[0094] Adjusting the data refers to adjusting the parameters in the noise reduction scheme.
[0095] It is understood that when the target device is idle, the adjustment data corresponding to the target mode can be determined and stored. When the application is needed, the stored adjustment data corresponding to the target mode can be retrieved.
[0096] First, a detailed analysis of the environmental noise characteristics under the target mode is required, including the noise's frequency distribution, intensity, and periodicity. This can be achieved by collecting a large number of noise samples in the scene corresponding to the target mode and using techniques such as spectrum analysis. Next, the characteristics of the speech signal under the target mode are analyzed, including the frequency range and common speech intensities, to clarify the speech features that need to be preserved in this scenario. Then, based on existing noise reduction scheme models, the impact of different parameter adjustments on the current noise and speech signal processing results is simulated and continuously compared and optimized. Successful parameter adjustment experiences in similar scenarios from historical data can also be referenced, combined with the subtle differences in the current scenario for further adjustments. Finally, considering the overall goals of speech enhancement, such as reducing noise while preserving speech clarity and intelligibility to the greatest extent possible, the adjustment data for the parameters in the noise reduction scheme are determined.
[0097] S16: Update the values of the parameters in the scheme library according to the adjustment data.
[0098] Specifically, based on the adjustment data, the values of the parameters in the scheme library are updated, and the update operations include: replacement or numerical adjustment.
[0099] This embodiment first acquires the adjustment data corresponding to the target mode. This operation enables the system to obtain key information that matches the current specific scenario mode. This adjustment data exists specifically for the particular needs or characteristics of this scenario mode. Next, the parameter values in the solution library are updated based on this adjustment data. This update operation allows the solution library to better adapt to the speech enhancement needs of the target mode. It optimizes the configuration of parameters in the solution library, making them more suitable for the environmental noise characteristics, speech propagation characteristics, and other factors of the current scenario. This improves the accuracy, effectiveness, and adaptability of the entire speech enhancement method in this scenario, helps to more accurately adaptively process environmental noise, and thus improves the overall effect of speech enhancement.
[0100] In one embodiment, the microphone array includes: a main microphone and multiple secondary microphones;
[0101] The step of extracting sound features based on the first audio to obtain the first feature includes:
[0102] S21: Extract human voice features and noise features from the main audio in the first audio, and use them as the first human voice feature and the first noise feature, wherein the main audio is the audio of the speaker collected by the main microphone;
[0103] Specifically, based on the human voice feature extraction model, human voice features are extracted from the main audio in the first audio and used as the first human voice feature; based on the noise feature extraction model, noise features are extracted from the main audio in the first audio and used as the first noise feature.
[0104] Human voice feature extraction models are trained using convolutional neural networks (such as ResNet-34). Noise feature extraction models are also trained using convolutional neural networks (such as ResNet-34). ResNet-34 (Residual Network-34) is a deep residual network.
[0105] Understandably, the main audio includes the speaker's audio, while the secondary audio includes the surrounding environment.
[0106] Understandably, before extracting the main audio features of the first audio, the blind source separation (BSS) algorithm is used to preprocess the main audio of the first audio, and the speaker's speech and environmental noise are separated by independent component analysis (ICA) before features are extracted separately.
[0107] S22: Extract human voice features and noise features from the secondary audio in the first audio, and use them as the second human voice feature and the second noise feature, wherein the secondary audio is the audio of the surrounding environment collected by the secondary microphone;
[0108] Specifically, based on the human voice feature extraction model, human voice features are extracted from the sub-audio frequencies of the first audio and used as the second human voice features; based on the noise feature extraction model, noise features are extracted from the sub-audio frequencies of the first audio and used as the second noise features.
[0109] Understandably, the secondary audio mainly includes the audio of the surrounding environment, and secondarily includes the audio of people.
[0110] Understandably, before extracting the sub-audio features of the first audio, the blind source separation (BSS) algorithm is used to preprocess the sub-audio of the first audio, and independent component analysis (ICA) is used to separate human speech from environmental noise before extracting features separately.
[0111] S23: Concatenate the first human voice feature and the main voice marker to obtain the first data; concatenate the first noise feature and the main noise marker to obtain the second data; concatenate the second human voice feature and the secondary human voice marker to obtain the third data; and concatenate the first noise feature and the secondary noise marker to obtain the fourth data.
[0112] S24: Using a target marker, the first data, the second data, the third data, and the fourth data are concatenated to obtain the first feature.
[0113] The primary voice marker emphasizes the weight of the first voice feature when processing all voice features in the subsequent processing of the first feature (its function is to preserve and enhance the audio corresponding to the first voice feature). The primary noise marker weakens the weight of the first noise feature when processing all noise features in step S24 (its function is to weaken the audio corresponding to the first noise feature). The secondary voice marker weakens the weight of the second voice feature when processing all voice features in step S24 (its function is to weaken the audio corresponding to the second voice feature). The secondary noise marker emphasizes the weight of the second noise feature when processing all noise features in step S24 (its function is to strengthen the weakened audio corresponding to the second noise feature).
[0114] Specifically, the data is concatenated based on the order of the first data, the target marker, the second data, the target marker, the third data, the target marker, and the fourth data, and the concatenated data is used as the first feature.
[0115] The target marker is a string of characters used to accurately distinguish the first data, the second data, the third data, and the fourth data from the first feature during subsequent processing, thereby improving the accuracy of subsequent processing of the first feature.
[0116] This embodiment first extracts human voice features and noise features from the main audio and secondary audio sources respectively. This method of extracting audio sources separately can comprehensively obtain relevant features of the speaker and the surrounding environment. By concatenating different human voice features and noise features with their respective tags, and then further concatenating them with the target tag to obtain the first feature, this concatenation method cleverly utilizes the weight adjustment effect of different tags. The main voice tag emphasizes the weight of the first human voice feature, the main noise tag weakens the weight of the first noise feature, the secondary human voice tag weakens the weight of the second human voice feature, and the secondary noise tag emphasizes the weight of the second noise feature. This helps to treat human voice and noise features from different sources more reasonably during the processing. Overall, this technology can more accurately distinguish the speaker's voice from environmental noise, thereby better enhancing the target speech, suppressing noise, and improving speech quality and intelligibility in subsequent speech processing.
[0117] In one embodiment, the step of performing speech enhancement based on the second audio to obtain the target audio includes:
[0118] S51: Obtain specified sound features;
[0119] The specified voice features are the pre-acquired and stored voice features of a specified person.
[0120] Specifically, it retrieves the specified sound features from the preset storage space.
[0121] S52: Perform speech enhancement on the human audio corresponding to the specified sound features based on the second audio to obtain the target audio.
[0122] Optionally, firstly, extract the human audio corresponding to the specified sound features from the second audio and use it as the audio to be enhanced. Then, perform speech enhancement on the audio to be enhanced (one or more of the following operations: noise reduction, dereverberation, interference removal, distortion correction, and adaptation to downstream tasks). Use the speech-enhanced audio to be enhanced as the target audio.
[0123] The second audio stream is segmented into frames to facilitate subsequent analysis. For each frame, its acoustic features are calculated, such as spectral features and Mel-frequency cepstral coefficients (MFCCs). Next, the calculated acoustic features of each frame are compared with specified acoustic features in a acoustic feature database, using algorithms such as Euclidean distance and cosine similarity to measure the degree of matching. When the matching degree between the acoustic features of a frame and the specified acoustic features reaches a certain threshold, this frame is marked as belonging to the audio portion of the person corresponding to the specified acoustic features. Finally, all the marked audio frames are reassembled in their original order to obtain the audio of the person corresponding to the specified acoustic features, i.e., the audio to be enhanced.
[0124] Optionally, firstly, locate the human audio corresponding to the specified sound features from the second audio, perform speech enhancement on the located audio (one or more operations among noise reduction, dereverberation, interference removal, distortion correction, and adaptation to downstream tasks), and use the second audio with completed speech enhancement as the target audio.
[0125] The first step involves a detailed analysis of the specified sound features to identify their unique acoustic characteristics, such as specific pitch ranges, timbre characteristics, and pronunciation habits. The second audio file is then preprocessed, potentially including noise reduction, to minimize interference and improve localization. Next, acoustic feature analysis techniques, such as speech signal processing algorithms, are used to scan the audio signal, starting from the beginning. Parameters like frequency and amplitude are analyzed to find segments matching the specified sound features. A sliding window approach can be used to progressively analyze different time periods. When a segment of the audio signal exhibits characteristics matching the specified sound features, such as a specific frequency pattern or pitch variation, the location of the person's audio likely corresponding to the specified sound features is preliminarily identified. The analysis is then expanded to further verify the location against surrounding audio samples, ensuring accuracy and ultimately completing the localization of the person's audio from the second audio file.
[0126] From the perspective of improving voice quality, this embodiment can effectively enhance the voice of a designated person in complex audio environments, making the voice clearer and more intelligible, and reducing the impact of ambient noise or other interference on the voice. This helps improve the accuracy of information transmission in applications such as voice communication and voice recognition. From the perspective of personalized processing, enhancing the audio of a person specifically corresponding to a designated voice feature demonstrates strong targeting, and can meet the special needs of different users for the voice of a specific person in different scenarios. For example, it can highlight the speaker's voice in a multi-person conference scenario, or enhance the voice of a specific speaker in a noisy environment to facilitate subsequent voice analysis, transcription, and other operations.
[0127] In one embodiment, after the step of extracting sound features based on the first audio to obtain the first feature, the method further includes:
[0128] S61: Determine a post-harvest noise reduction scheme based on the first feature;
[0129] Specifically, the first feature is input into the pre-trained second scheme model for classification prediction. The vector element with the largest value is extracted from the predicted vector. The category corresponding to the extracted vector element is used as the hit label. The scheme corresponding to the hit label is used as the post-sampling denoising scheme.
[0130] The second pre-trained model is a pre-trained multi-class classification model. The model structure and training method of the second model can be selected from existing technologies, which will not be elaborated here.
[0131] Post-acquisition noise reduction is a method for reducing noise in audio captured by a microphone array. The purpose of post-acquisition noise reduction is to reduce noise in the acquired audio.
[0132] Optional post-collection noise reduction solutions include: noise type identification results and noise reduction algorithms, such as white noise, pink noise, or other specific types of environmental noise.
[0133] The step of performing speech enhancement on the human audio corresponding to the specified sound features based on the second audio to obtain the target audio includes:
[0134] S521: Based on the post-acquisition noise reduction scheme, perform speech enhancement on the human audio corresponding to the specified sound features according to the second audio to obtain the target audio.
[0135] Specifically, firstly, based on the noise type identification results in the post-acquisition noise reduction scheme, a suitable noise reduction algorithm is determined. For white noise, a filter-based noise reduction algorithm, such as a Wiener filter, might be chosen; for colored noise, a statistical model-based noise reduction method might be used. Then, the parameters of the noise reduction algorithm are adjusted based on the noise intensity estimate. For cases with high noise intensity, the filter attenuation coefficient is increased or the weights in the statistical model are adjusted. Next, the second audio is processed by frequency band segmentation. For frequency bands determined to be primarily noise, appropriate noise reduction methods are used for suppression; for frequency bands containing human audio corresponding to specified sound features, while performing noise reduction, speech enhancement is emphasized (one or more operations among noise reduction, dereverberation, interference removal, distortion correction, and adaptation to downstream tasks). The frequency bands containing human audio corresponding to specified sound features, after noise reduction and speech enhancement processing, are then recombined to obtain the target audio.
[0136] This embodiment determines the post-acquisition noise reduction scheme based on the first feature, so it can perform customized processing of the acquired noise according to different sound features, which is beneficial to improve the noise reduction effect of the acquired noise. It specifically enhances the audio of the person corresponding to the specified sound features, which shows strong targeting and can meet the special needs of different users for the voice of a specific person in different scenarios. For example, it can highlight the speaker's voice in a multi-person conference scenario, or enhance the voice of a specific speaker in a noisy environment to facilitate subsequent speech analysis, transcription and other operations.
[0137] In one embodiment, the step of performing speech enhancement on the second audio corresponding to the specified sound features based on the post-acquisition noise reduction scheme to obtain the target audio includes:
[0138] S5211: Obtain the vocal cord pathology speech feature compensation parameters corresponding to the specified sound feature;
[0139] The vocal cord pathology speech feature compensation parameters corresponding to the specified voice features are pre-acquired and stored data.
[0140] Specifically, the vocal cord pathological speech feature compensation parameters corresponding to the specified sound features are obtained from the preset storage space.
[0141] The vocal cord pathology speech feature compensation parameters include one or more of the following: frequency-related data, fundamental frequency-related data, time-domain feature data, and comprehensive adjustment values related to speech quality.
[0142] Frequency-related data includes energy compensation values for specific frequencies under different pathological conditions. For example, vocal cord paralysis may lead to energy loss in certain low-frequency components, and specific energy boost values for these low-frequency components will be provided. Frequency-related data also includes adjustment values for formant frequencies and bandwidth. Vocal cord lesions may alter the characteristics of the formant, and compensation parameters will include the values that the formant frequencies should be adjusted for and the amount of bandwidth change under different pathological conditions.
[0143] Fundamental frequency-related data includes fundamental frequency offset compensation values. For example, vocal cord nodules may slightly raise or lower the fundamental frequency, and the compensation parameters will provide corresponding adjustment values to correct the fundamental frequency and bring it closer to its normal state.
[0144] Temporal feature data includes adjusted values for the start and end times of speech. In some vocal cord pathologies, the start and end of speech may be unclear; compensation parameters will include correction data for the start and end times of speech.
[0145] Comprehensive adjustment values related to speech quality, such as quantitative adjustment values for clarity and intelligibility. These values are obtained based on subjective and objective evaluations of a large number of pathological speech samples and are used to improve the overall quality of vocal cord pathology speech.
[0146] This process involves analyzing specific sound features, which may include extracting parameters such as spectral characteristics, fundamental frequency, and formants. These parameters are then matched with data in a vocal cord pathology speech feature database. For example, if the extracted fundamental frequency is unstable and falls within a specific range, and the frequency and bandwidth of the formants exhibit a similar pattern to those in the vocal cord nodule pathology speech feature database, then the corresponding vocal cord pathology type can be determined. Once the vocal cord pathology type is identified, corresponding vocal cord pathology speech feature compensation parameters can be obtained from the vocal cord pathology speech feature database. These parameters, derived from the analysis of a large number of speech samples with that pathology type, are used for subsequent speech repair.
[0147] The vocal cord pathology speech feature compensation parameters are generated by collecting speech samples from patients with different pathological types (such as vocal cord polyps and paralysis) and analyzing the spectral offset and fundamental frequency fluctuation range of the speech samples.
[0148] S5212: Based on the post-acquisition noise reduction scheme, perform speech enhancement on the human audio corresponding to the specified sound features according to the second audio to obtain the third audio;
[0149] Specifically, based on the post-acquisition noise reduction scheme, speech enhancement is performed on the second audio corresponding to the specified sound features of the person, and the speech-enhanced second audio is used as the third audio.
[0150] S5213: Using a beamforming algorithm, suppress sound sources from directions other than the speaker in the main audio of the third audio, perform spectral analysis on the third audio, identify frequency components that may lose or reduce quality, and obtain identification results;
[0151] Specifically, when using beamforming algorithms, the layout of the microphone array (if multiple microphones are used) must first be determined or a virtual microphone array simulated. The direction of arrival (DOA) of the sound is calculated by measuring the time and amplitude differences of the sound signals received by each microphone. For the main audio signal in the third audio stream, the weight vector of the beamforming algorithm is adjusted according to its DOA to maximize the gain in the direction of the main audio signal while suppressing sound sources from other directions. During spectral analysis, a Fast Fourier Transform (FFT) is performed on the third audio stream to convert the time-domain audio signal into a frequency-domain signal. Then, the energy, phase, and other parameters of each frequency component are analyzed and compared with the original, high-quality reference audio or a normal frequency range determined based on prior knowledge. If the energy of a frequency component is below the normal range or the phase is abnormal, it is marked as a frequency component that may suffer from loss or degraded quality, ultimately yielding the identification result.
[0152] The time delay difference of the microphone array is calculated using the generalized cross-correlation (GCC-PHAT) algorithm to estimate the sound source direction angle; a spatial filter is generated based on the minimum variance distortionless response (MVDR) algorithm to enhance the speaker's direction signal and suppress sidelobe noise.
[0153] S5214: Based on the recognition result and the vocal cord pathology speech feature compensation parameters, a linear interpolation method is used to repair the third audio to obtain the target audio.
[0154] Specifically, the frequency components requiring repair and their degree of damage are determined based on the identification results. Then, vocal cord pathology speech feature compensation parameters are used. These parameters contain adjustment data for frequency components under vocal cord pathology conditions; for example, specific frequency energy compensation values may be required in cases of vocal cord polyps. Using linear interpolation, for each damaged frequency component, the value of the repaired frequency component is calculated based on its relationship with adjacent normal frequency components and the vocal cord pathology speech feature compensation parameters. For example, if the energy of a certain frequency component needs to be increased, the increase magnitude is determined through linear interpolation to make it as close as possible to a normal audio state while conforming to the vocal cord pathology speech feature compensation parameters. After performing this operation on all frequency components requiring repair in the third audio, the target audio is obtained.
[0155] This embodiment first obtains the vocal cord pathology speech feature compensation parameters corresponding to specified voice features, enabling targeted processing of speech that may have vocal cord pathology issues. This helps improve speech quality, making the speech of people with vocal cord pathology clearer and more intelligible. A third audio is obtained through post-acquisition noise reduction, effectively removing noise interference and enhancing the designated speaker's voice. Then, beamforming is applied to focus on the main audio and suppress sound sources from directions other than the speaker, further improving the clarity of the designated speaker's voice. Simultaneously, spectral analysis identifies frequency components that may be lost or degraded, providing an accurate basis for subsequent repair operations. Finally, based on the identification results and the vocal cord pathology speech feature compensation parameters, linear interpolation is used to repair the third audio to obtain the target audio. This improves overall audio quality, especially for speech with vocal cord pathology issues, and effectively highlights the audio of the person corresponding to the specified voice features even in multi-person scenarios.
[0156] Please see Figure 3 As shown, in one embodiment, a speech enhancement device based on environmental noise adaptation is provided, the device comprising:
[0157] The first acquisition module 901 is used to acquire audio collected in real time by the microphone array in the target microphone based on a first time interval, and use it as the first audio.
[0158] Feature extraction module 902 is used to extract sound features based on the first audio to obtain a first feature;
[0159] The scheme determination module 903 is used to determine the noise reduction scheme based on the scheme library, according to the spatial distribution characteristics of the microphone array and the first feature;
[0160] The second acquisition module 904 is used to control the microphone array to acquire audio according to the acquisition noise reduction scheme, as the second audio.
[0161] Processing module 905 is used to perform speech enhancement based on the second audio to obtain the target audio.
[0162] This embodiment extracts sound features (first features) from the first audio recording and then combines them with the spatial distribution features of the microphone array and a scheme library to determine the in-sampling noise reduction scheme. This method can accurately determine the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio recording) according to the in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio recording to obtain the target audio. Overall, this speech enhancement method based on environmental noise adaptation realizes a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement". Compared with the traditional separate architecture of "fixed acquisition + back-end processing", this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improve the signal quality input to the back-end processing, and help improve the effect of back-end processing in improving speech quality. In addition, starting the in-sampling noise reduction scheme based on the first time interval eliminates the need for frequent determination of the in-sampling noise reduction scheme, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computing resources.
[0163] In one embodiment, the step of determining the noise reduction scheme based on the scheme library in the scheme determination module 903, according to the spatial distribution characteristics corresponding to the microphone array and the first feature, includes:
[0164] Based on a preset time window and taking the end time of the first audio as the end time of the time window, the sound features corresponding to the audio collected by the microphone array are obtained as feature packets.
[0165] Based on the feature package, the trend of future sound features is predicted to obtain the feature trend prediction result;
[0166] The feature trend prediction result, the spatial distribution feature, and the first feature are input into a pre-trained first scheme model for classification and prediction to obtain the target scheme identifier;
[0167] The noise reduction scheme is determined based on the target scheme identifier and the scheme library.
[0168] In one embodiment, the apparatus further includes: an interval determination module, the interval determination module being used to:
[0169] The current scene pattern is obtained based on the second time interval and used as the target pattern;
[0170] The time interval is determined using a lookup table method based on the target pattern and used as the initial time interval.
[0171] Obtain the acquisition error corresponding to the target mode;
[0172] Based on the acquisition error, the initial time interval is calibrated to obtain the first time interval.
[0173] In one embodiment, after the step of obtaining the current scene mode as the target mode according to the second time interval in the interval determination module, the method further includes:
[0174] Obtain the adjustment data corresponding to the target mode;
[0175] Based on the adjustment data, the values of the parameters in the scheme library are updated.
[0176] In one embodiment, the microphone array includes: a main microphone and multiple secondary microphones;
[0177] The step of extracting sound features based on the first audio and obtaining the first feature in the feature extraction module 902 includes:
[0178] Based on the main audio in the first audio, human voice features and noise features are extracted as the first human voice feature and the first noise feature, wherein the main audio is the audio of the speaker collected by the main microphone;
[0179] Based on the secondary audio in the first audio, human voice features and noise features are extracted as the second human voice feature and the second noise feature, wherein the secondary audio is the audio of the surrounding environment collected by the secondary microphone;
[0180] The first human voice feature and the main voice marker are concatenated to obtain the first data; the first noise feature and the main noise marker are concatenated to obtain the second data; the second human voice feature and the secondary human voice marker are concatenated to obtain the third data; and the first noise feature and the secondary noise marker are concatenated to obtain the fourth data.
[0181] The first feature is obtained by concatenating the first data, the second data, the third data, and the fourth data using a target marker.
[0182] In one embodiment, the step of performing speech enhancement based on the second audio to obtain the target audio in the processing module 905 includes:
[0183] Obtain the specified sound features;
[0184] Based on the second audio, speech enhancement is performed on the human audio corresponding to the specified voice features to obtain the target audio.
[0185] In one embodiment, the apparatus further includes: a scheme generation module, the scheme generation module being configured to: determine a post-sampling noise reduction scheme based on the first feature;
[0186] The step in the processing module 905 of performing speech enhancement on the human audio corresponding to the specified sound features based on the second audio to obtain the target audio includes:
[0187] Based on the post-acquisition noise reduction scheme, the target audio is obtained by performing speech enhancement on the human audio corresponding to the specified sound features according to the second audio.
[0188] In one embodiment, the step in the processing module 905 of performing speech enhancement on the human audio corresponding to the specified sound features based on the post-acquisition noise reduction scheme to obtain the target audio includes:
[0189] Obtain the vocal cord pathology speech feature compensation parameters corresponding to the specified sound features;
[0190] Based on the post-acquisition noise reduction scheme, the third audio is obtained by enhancing the speech of the human audio corresponding to the specified sound features according to the second audio.
[0191] A beamforming algorithm is used to suppress sound sources from directions other than the speaker in the main audio of the third audio. Spectral analysis is performed on the third audio to identify frequency components that may lose or reduce quality, and the identification results are obtained.
[0192] Based on the recognition results and the vocal cord pathology speech feature compensation parameters, a linear interpolation method is used to repair the third audio to obtain the target audio.
[0193] Please see Figure 4 As shown, in one embodiment, a target device is proposed, comprising: a target microphone 1 and a device body 2, the device body 2 being communicatively connected to the target microphone 1. The device body 2 includes a memory 21, a processor 22, and a computer program stored in the memory 21 and executable on the processor 22. When the processor 22 executes the computer program, it performs the following steps:
[0194] Based on the first time interval, the audio collected in real time by the microphone array in the target microphone is obtained as the first audio.
[0195] Based on the first audio, sound features are extracted to obtain the first feature;
[0196] Based on the solution library, and according to the spatial distribution characteristics of the microphone array and the first feature, a noise reduction solution is determined.
[0197] According to the noise reduction scheme, the microphone array is controlled to collect audio as the second audio.
[0198] Speech enhancement is performed based on the second audio to obtain the target audio.
[0199] This embodiment extracts sound features (first features) from the first audio recording and then combines them with the spatial distribution features of the microphone array and a scheme library to determine the in-sampling noise reduction scheme. This method can accurately determine the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio recording) according to the in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio recording to obtain the target audio. Overall, this speech enhancement method based on environmental noise adaptation realizes a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement". Compared with the traditional separate architecture of "fixed acquisition + back-end processing", this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improve the signal quality input to the back-end processing, and help improve the effect of back-end processing in improving speech quality. In addition, starting the in-sampling noise reduction scheme based on the first time interval eliminates the need for frequent determination of the in-sampling noise reduction scheme, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computing resources.
[0200] In one embodiment, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, performs the following steps:
[0201] Based on the first time interval, the audio collected in real time by the microphone array in the target microphone is obtained as the first audio.
[0202] Based on the first audio, sound features are extracted to obtain the first feature;
[0203] Based on the solution library, and according to the spatial distribution characteristics of the microphone array and the first feature, a noise reduction solution is determined.
[0204] According to the noise reduction scheme, the microphone array is controlled to collect audio as the second audio.
[0205] Speech enhancement is performed based on the second audio to obtain the target audio.
[0206] This embodiment extracts sound features (first features) from the first audio recording and then combines them with the spatial distribution features of the microphone array and a scheme library to determine the in-sampling noise reduction scheme. This method can accurately determine the in-sampling noise reduction scheme based on the actual sound conditions and the spatial layout of the microphones. Then, the microphone array is controlled to acquire audio (second audio recording) according to the in-sampling noise reduction scheme. This step ensures that the acquired audio is effective audio under environmental noise conditions. Finally, speech enhancement is performed on the second audio recording to obtain the target audio. Overall, this speech enhancement method based on environmental noise adaptation realizes a closed-loop architecture of "dynamic perception → in-sampling noise reduction → front-end suppression → back-end enhancement". Compared with the traditional separate architecture of "fixed acquisition + back-end processing", this application can more effectively avoid acquiring signals with highly mixed noise spectra and human voices in complex acoustic environments, improve the signal quality input to the back-end processing, and help improve the effect of back-end processing in improving speech quality. In addition, starting the in-sampling noise reduction scheme based on the first time interval eliminates the need for frequent determination of the in-sampling noise reduction scheme, avoiding fluctuations in audio acquisition that may be caused by frequent switching of noise reduction schemes, improving the stability of audio acquisition, and saving computing resources.
[0207] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0208] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0209] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0210] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech enhancement method based on environmental noise adaptation, characterized in that, The method includes: Based on the first time interval, the audio collected in real time by the microphone array in the target microphone is obtained as the first audio. Based on the first audio, sound features are extracted to obtain the first feature; Based on the solution library, and according to the spatial distribution characteristics of the microphone array and the first feature, a noise reduction solution is determined during sampling. The noise reduction solution includes one or more data from all data corresponding to hardware control parameters, signal processing strategies, and environmental adaptation information. According to the noise reduction scheme, the microphone array is controlled to collect audio as the second audio. Speech enhancement is performed based on the second audio to obtain the target audio; Before the step of acquiring the audio collected in real time by the microphone array in the target microphone based on the first time interval as the first audio, the method further includes: The current scene pattern is obtained based on the second time interval and used as the target pattern; The time interval is determined using a lookup table method based on the target pattern and used as the initial time interval. Obtain the acquisition error corresponding to the target mode, wherein the acquisition error corresponding to the target mode is the error generated by the microphone array acquiring audio based on the default array control data in the target mode; Based on the acquisition error, the initial time interval is calibrated to obtain the first time interval; The step of obtaining the current scene mode as the target mode according to the second time interval includes: when the device executing the speech enhancement method based on environmental noise adaptation is idle, determining the scene mode by comprehensively analyzing the audio collected by the microphone array, storing the determined scene mode, and obtaining the stored scene mode as the target mode at the beginning of the time period of the second time interval. The microphone array includes: a main microphone and multiple secondary microphones; The step of extracting sound features based on the first audio to obtain the first feature includes: Based on the main audio in the first audio, human voice features and noise features are extracted as the first human voice feature and the first noise feature, wherein the main audio is the audio of the speaker collected by the main microphone; Based on the secondary audio in the first audio, human voice features and noise features are extracted as the second human voice feature and the second noise feature, wherein the secondary audio is the audio of the surrounding environment collected by the secondary microphone; The first human voice feature and the main voice marker are concatenated to obtain the first data; the first noise feature and the main noise marker are concatenated to obtain the second data; the second human voice feature and the secondary human voice marker are concatenated to obtain the third data; and the first noise feature and the secondary noise marker are concatenated to obtain the fourth data. The first feature is obtained by concatenating the first data, the second data, the third data, and the fourth data using a target marker.
2. The speech enhancement method based on adaptive environmental noise according to claim 1, characterized in that, The step of determining the noise reduction scheme based on the scheme library, according to the spatial distribution characteristics of the microphone array and the first feature, includes: Based on a preset time window and taking the end time of the first audio as the end time of the time window, the sound features corresponding to the audio collected by the microphone array are obtained as feature packets. Based on the feature package, the trend of future sound features is predicted to obtain the feature trend prediction result; The feature trend prediction result, the spatial distribution feature, and the first feature are input into a pre-trained first scheme model for classification and prediction to obtain the target scheme identifier; The noise reduction scheme is determined based on the target scheme identifier and the scheme library.
3. The speech enhancement method based on adaptive environmental noise according to claim 1, characterized in that, After the step of obtaining the current scene mode as the target mode according to the second time interval, the method further includes: Obtain the adjustment data corresponding to the target mode; Based on the adjustment data, the values of the parameters in the scheme library are updated.
4. The speech enhancement method based on adaptive environmental noise according to claim 1, characterized in that, The step of performing speech enhancement based on the second audio to obtain the target audio includes: Obtain the specified sound features; Based on the second audio, speech enhancement is performed on the human audio corresponding to the specified voice features to obtain the target audio.
5. The speech enhancement method based on adaptive environmental noise according to claim 4, characterized in that, After the step of extracting sound features based on the first audio to obtain the first feature, the method further includes: Based on the first feature, a post-harvest noise reduction scheme is determined; The step of performing speech enhancement on the human audio corresponding to the specified sound features based on the second audio to obtain the target audio includes: Based on the post-acquisition noise reduction scheme, the target audio is obtained by performing speech enhancement on the human audio corresponding to the specified sound features according to the second audio.
6. The speech enhancement method based on adaptive environmental noise according to claim 5, characterized in that, The step of performing speech enhancement on the second audio based on the post-acquisition noise reduction scheme and the human audio corresponding to the specified sound features to obtain the target audio includes: Obtain the vocal cord pathology speech feature compensation parameters corresponding to the specified sound features; Based on the post-acquisition noise reduction scheme, the third audio is obtained by enhancing the speech of the human audio corresponding to the specified sound features according to the second audio. A beamforming algorithm is used to suppress sound sources from directions other than the speaker in the main audio of the third audio. Spectral analysis is performed on the third audio to identify frequency components that may lose or reduce quality, and the identification results are obtained. Based on the recognition results and the vocal cord pathology speech feature compensation parameters, a linear interpolation method is used to repair the third audio to obtain the target audio.
7. A speech enhancement device based on adaptive environmental noise, characterized in that, The device includes: The first acquisition module is used to acquire audio collected in real time by the microphone array in the target microphone based on a first time interval, as the first audio. The feature extraction module is used to extract sound features based on the first audio to obtain the first feature; The scheme determination module, based on the scheme library, determines the noise reduction scheme in the sampling process according to the spatial distribution characteristics corresponding to the microphone array and the first feature. The noise reduction scheme in the sampling process includes one or more data from all data corresponding to hardware control parameters, signal processing strategies and environmental adaptation information. The second acquisition module is used to control the microphone array to acquire audio according to the noise reduction scheme, as the second audio. The processing module is used to perform speech enhancement based on the second audio to obtain the target audio; Before the step of acquiring the audio collected in real time by the microphone array in the target microphone based on the first time interval as the first audio, the method further includes: The current scene pattern is obtained based on the second time interval and used as the target pattern; The time interval is determined using a lookup table method based on the target pattern and used as the initial time interval. Obtain the acquisition error corresponding to the target mode, wherein the acquisition error corresponding to the target mode is the error generated by the microphone array acquiring audio based on the default array control data in the target mode; Based on the acquisition error, the initial time interval is calibrated to obtain the first time interval; The step of obtaining the current scene mode as the target mode according to the second time interval includes: when the device executing the speech enhancement method based on environmental noise adaptation is idle, determining the scene mode by comprehensively analyzing the audio collected by the microphone array, storing the determined scene mode, and obtaining the stored scene mode as the target mode at the beginning of the time period of the second time interval. The microphone array includes: a main microphone and multiple secondary microphones; When extracting sound features based on the first audio to obtain the first feature: Based on the main audio in the first audio, human voice features and noise features are extracted as the first human voice feature and the first noise feature, wherein the main audio is the audio of the speaker collected by the main microphone; Based on the secondary audio in the first audio, human voice features and noise features are extracted as the second human voice feature and the second noise feature, wherein the secondary audio is the audio of the surrounding environment collected by the secondary microphone; The first human voice feature and the main voice marker are concatenated to obtain the first data; the first noise feature and the main noise marker are concatenated to obtain the second data; the second human voice feature and the secondary human voice marker are concatenated to obtain the third data; and the first noise feature and the secondary noise marker are concatenated to obtain the fourth data. The first feature is obtained by concatenating the first data, the second data, the third data, and the fourth data using a target marker.
8. A target device, characterized in that, The target device includes a target microphone and a device body, the device body being communicatively connected to the target microphone, the device body including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the speech enhancement method based on environmental noise adaptation as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Noise elimination method, device and equipment, and readable storage medium
CN111627456A
Audio noise reduction method for VR equipment, electronic equipment and storage medium
CN114255779A
Intelligent glasses audio processing method and system
CN119296564A