An audio processing method, an electronic device, and a storage medium
By using technical means such as multi-channel filtering, dynamic noise suppression and adaptive filtering in Bluetooth headphones, interfering signals in complex environments are identified and suppressed in real time, and the output parameters of audio signals are dynamically adjusted according to environmental conditions, the problem of unstable audio quality in the existing technology is solved, and more efficient audio processing effects are achieved.
Patent Information
- Application Number
- CN202510436452.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The prior art cannot effectively identify and suppress multiple interference sources in complex environments, resulting in unstable audio quality and lack of dynamic adjustment mechanisms for audio signal output parameters.
Multi-channel filtering, dynamic noise suppression, short-time Fourier transform based on dynamic time window length, spectrum analysis and adaptive filtering are used to detect and suppress interfering signals in real time, and dynamically adjust the output parameters of the audio signal according to environmental conditions.
It improves the audio processing effect in Bluetooth headphones, enhances the clarity and stability of the audio signal, and ensures that the audio output can maintain the best sound quality performance in different environments.
Smart Images

Figure CN119946506B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and particularly to an audio processing method, an electronic device, and a storage medium. Background Art
[0002] With the popularization of wireless earphones, especially Bluetooth earphones, users have higher and higher requirements for sound quality. Bluetooth earphones have become an indispensable part of people's daily lives due to their convenient wireless connection and compact design. However, Bluetooth earphones are prone to interference from external environmental noises during use. Especially in a noisy environment, the audio output quality of the earphones is often greatly reduced, which has an adverse impact on the user's auditory experience.
[0003] Although existing audio processing technologies can reduce noise interference to a certain extent, their processing effects are limited. Current audio noise suppression technologies mainly rely on a single noise cancellation algorithm, and most systems adopt a fixed noise suppression method, which cannot be dynamically adjusted according to the real-time changes of the environment. In addition, the accuracy of existing technologies in spectrum analysis is insufficient, and they cannot effectively identify and dynamically suppress multiple interference sources in a complex environment, resulting in the audio quality not reaching the best effect.
[0004] The main problem faced by existing technologies is the inability to perform adaptive optimization for complex environmental conditions. Especially in the case of coexistence of multiple interference sources, traditional audio processing methods often cannot identify and process these interference signals in real time and accurately, resulting in unstable audio output quality. In addition, existing technologies also lack a dynamic adjustment mechanism for audio signal output parameters and cannot automatically optimize audio quality according to different environmental conditions. Therefore, there is an urgent need for a technology that can adaptively adjust audio signal processing strategies in a dynamic environment. Summary of the Invention
[0005] In view of this, the purpose of the embodiments of the present invention is to provide an audio processing method, an electronic device, and a storage medium to solve the above technical problems.
[0006] To achieve the above purpose, in the first aspect, the present invention provides an audio processing method applicable to Bluetooth earphones, and the method includes the following steps:
[0007] Obtain an audio signal, and perform preprocessing on the obtained audio signal by using a multi-channel filtering algorithm and a dynamic noise suppression algorithm to obtain a preprocessed audio signal;
[0008] Use short-time Fourier transform based on a dynamic time window length to perform spectrum decomposition on the preprocessed audio signal to obtain the spectrum characteristics and power distribution of the preprocessed audio signal in multiple frequency bands;
[0009] Detect external interference sources in the preprocessed audio signal in real time, identify interference signals based on the spectral characteristics and power distribution of multiple frequency bands, and use an adaptive filtering algorithm to suppress the interference signals to obtain an audio signal after interference suppression;
[0010] Dynamically adjust multiple output parameters of the audio signal after interference suppression according to real-time change data of environmental conditions to obtain an optimized audio signal;
[0011] Obtain a final audio output signal according to the optimized audio signal, and output the final audio output signal to the speaker of the Bluetooth headset for playback.
[0012] In a second aspect, the present invention provides an electronic device, which includes: a processor; and a memory for storing instructions executable by the processor;
[0013] Wherein, the processor is configured to execute the instructions to implement the audio processing method described in the first aspect.
[0014] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the audio processing method described in the first aspect.
[0015] The above technical solutions have the following beneficial effects:
[0016] By adopting technical means such as multi-channel filtering, dynamic noise suppression, short-time Fourier transform based on dynamic time window length, spectral analysis, and adaptive filtering, the present invention can improve the audio processing effect in Bluetooth headsets. Through the combined application of multi-channel filtering and dynamic noise suppression algorithms, the interference of external noise is effectively reduced, and the clarity of the audio signal is improved. The short-time Fourier transform based on dynamic time window length can accurately analyze the spectral characteristics and power distribution of the audio signal in different frequency bands, so as to more accurately identify interference sources and reduce the influence of noise on the audio signal. The adaptive filtering algorithm can dynamically adjust the suppression intensity of the interference signal to achieve targeted interference suppression and further optimize the audio signal quality. By obtaining real-time environmental condition change data and dynamically adjusting multiple output parameters of the audio signal according to these data, the present invention ensures that the audio output can maintain the best sound quality performance in different environments. Therefore, the present invention improves the audio processing ability of Bluetooth headsets in complex environments and provides a more stable, clear, and real auditory experience. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 is a flowchart of the audio processing method according to an embodiment of the present invention;
[0019] Figure 2 is a detailed flowchart of step S10 of the audio processing method according to an embodiment of the present invention;
[0020] Figure 3 is a detailed flowchart of step S20 of the audio processing method according to an embodiment of the present invention;
[0021] Figure 4 is a detailed flowchart of step S30 of the audio processing method according to an embodiment of the present invention;
[0022] Figure 5 is a detailed flowchart of step S40 of the audio processing method according to an embodiment of the present invention;
[0023] Figure 6 is a detailed flowchart of step S50 of the audio processing method according to an embodiment of the present invention;
[0024] Figure 7 is a functional block diagram of a computer-readable storage medium according to an embodiment of the present invention;
[0025] Figure 8 is a functional block diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0027] The method of the embodiment of the present invention realizes the optimized processing of the audio signal in the Bluetooth headset through technical means such as multi-channel filtering, dynamic noise suppression, spectrum analysis, and adaptive filtering, and can provide clearer and more real audio output in a complex environment, improving the user's wearing experience.
[0028] Embodiment 1
[0029] Figure 1is a flowchart of the audio processing method according to an embodiment of the present invention. As Figure 1 shown, the audio processing method according to an embodiment of the present invention includes the following steps:
[0030] S10: Obtain an audio signal, and preprocess the obtained audio signal by using a multi-channel filtering algorithm and a dynamic noise suppression algorithm to obtain a preprocessed audio signal.
[0031] In this embodiment, first, an audio signal in the external environment is collected by a microphone array in a Bluetooth headset. These audio signals include the user's voice, background noise, traffic noise, crowd noise, etc. in the environment. For the collected original audio signal, a multi-channel filtering algorithm is used for preliminary filtering. This filtering is performed by jointly analyzing the signals of multiple microphone channels to remove some common low-frequency noises, such as wind noise or air-conditioning noise. Then, environmental noise data is obtained, and environmental noise estimation is performed based on the noise data to obtain a relatively clean audio signal. Finally, a dynamic noise suppression algorithm is used to suppress the remaining noise in the audio signal to remove the remaining background noise and obtain a preprocessed audio signal.
[0032] S20: Perform spectral decomposition on the preprocessed audio signal by using a short-time Fourier transform based on a dynamic time window length to obtain spectral characteristics and power distribution of the preprocessed audio signal in multiple frequency bands.
[0033] In this embodiment, the preprocessed audio signal dynamically determines an appropriate time window length according to its frequency characteristics. The dynamic adjustment of the time window length is based on the spectral characteristics of the audio signal. For example, when the frequency range of the audio signal is wide, the time window length is long to ensure a detailed analysis of the spectrum; while when the frequency is concentrated, the time window is short to more accurately capture instantaneous changes. Then, an algorithm based on the short-time Fourier transform is used to perform spectral decomposition on the audio signal. This step converts the audio signal into a frequency-domain representation, extracts spectral characteristics of multiple frequency bands, and calculates the power distribution of each frequency band. Through this process, the spectral characteristics and power distribution information of each frequency band can be obtained, further providing a basis for the identification and suppression of interference sources.
[0034] S30: Real-time detect an external interference source in the preprocessed audio signal, identify an interference signal based on the spectral characteristics and power distribution of the multiple frequency bands, and use an adaptive filtering algorithm to suppress the interference signal to obtain an audio signal after interference suppression.
[0035] In step S30, based on the spectral characteristics and power distribution of the acquired audio signal, the system will detect external interference sources in the audio signal in real time. For example, when the system identifies a sudden change in the power distribution of a specific frequency band in the spectrum, the system will analyze the frequency characteristics of these interference signals through spectral analysis and compare them with a known noise library to identify the specific interference source. For example, it is caused by external noise sources such as traffic noise or crowd conversations. After identification, an adaptive filtering algorithm is used to dynamically suppress the identified interference signal. The adaptive filtering algorithm adjusts the filter parameters according to the real-time noise characteristics, effectively suppressing the interference signal and maintaining the integrity of the required audio signal. Finally, an audio signal after interference suppression is obtained.
[0036] S40: Dynamically adjust multiple output parameters of the audio signal after interference suppression according to the real-time change data of environmental conditions and the user's multimodal status data to obtain an optimized audio signal.
[0037] Specifically, the real-time change data of environmental conditions is obtained through sensors built into the Bluetooth headset or data exchange with devices such as smartphones, covering environmental noise level, environmental temperature, environmental humidity, environmental air pressure, etc. The user's multimodal status data includes information such as the user's motion state, user location, Bluetooth connection state, and the usage mode of the headset. The system generates context awareness parameters based on these real-time change data of environmental conditions and the user's multimodal status data, and combines the processing requirements of the audio signal to adjust various parameters of the audio output, including volume, frequency response, gain, etc. For example, in a situation with high environmental noise, the system automatically increases the volume and moderately adjusts the spectrum to ensure that the audio output is clear and not masked by external noise. In a quiet environment, the volume is appropriately reduced and the low-frequency response is optimized to make the sound quality softer.
[0038] S50: Obtain the final audio output signal according to the optimized audio signal, and output the final audio output signal to the speaker of the Bluetooth headset for playback.
[0039] Specifically, the optimized audio signal undergoes multi-modal perception collaborative processing to complete the final signal optimization. The specific processing flow includes: physiological feature enhancement processing, acoustic environment adaptation processing, and multi-source signal fusion output.
[0040] Physiological feature enhancement processing refers to obtaining the user's physiological vibration characteristics through a biosensor, extracting enhancement factors related to the speech signal, and performing frequency band selective enhancement on the optimized audio signal to improve the energy focus and phase consistency of the speech frequency band.
[0041] Acoustic environment adaptation processing refers to dynamically analyzing the changes in the acoustic characteristics of the ear canal based on the internal environment sensing data of the earphone, adaptively suppressing the resonance distortion frequency bands caused by physical state changes (such as wearing tightness, environmental air pressure fluctuations), and eliminating acoustic resonance interference.
[0042] Multi-source signal fusion output refers to mixing the enhanced signal and the original optimized signal according to a preset weight ratio, ensuring that the amplitude-frequency characteristics of the output signal meet the acoustic standards, and finally driving the speaker to achieve high-fidelity playback.
[0043] Through the above technical solutions, the embodiments of the present invention provide a Bluetooth headset audio processing method capable of adaptively optimizing the audio processing process in complex environments, which not only improves the quality of audio signals but also ensures the adaptability and stability of audio output in various environments.
[0044] Figure 2 It is a specific flowchart of step S10 of the audio processing method of the embodiments of the present invention. As Figure 2 shown, step S10 specifically includes the following steps:
[0045] S11: Collect audio signals in the external environment to obtain the original audio signals.
[0046] In this embodiment, the Bluetooth headset is equipped with multiple microphones for collecting audio signals in the surrounding environment. These microphones are located at different positions of the earphone to form a microphone array for simultaneously collecting sound signals from different directions. Through the array microphone design, the spatial coverage of sound collection can be improved and single-directional noise can be reduced. The audio signals collected by the microphones include the user's speech and environmental noise (such as traffic noise, wind noise, crowd noise, etc.). For example, when the user uses the Bluetooth headset on the street, the microphones will simultaneously capture the user's speech and the surrounding traffic noise. The collected original audio signals will be converted into digital signals and sent to the subsequent processing module.
[0047] S12: Use a non-uniform sub-band decomposition method based on the human ear auditory masking effect to divide the original audio signal into several sub-band signals that conform to the characteristics of the human ear critical band, enhance the sub-band signals where the fundamental frequency harmonics are located, perform noise suppression processing on the other sub-band signals outside the fundamental frequency harmonics, and recombine the enhanced sub-band signals where the fundamental frequency harmonics are located with the other sub-band signals after noise suppression processing to obtain a preliminarily filtered audio signal.
[0048] Step S12 can specifically include the following steps:
[0049] S121: Based on the human ear auditory masking effect, determine the critical band distribution of the original audio signal, and perform non-uniform sub-band decomposition on the original audio signal according to the critical band distribution to obtain multiple sub-band signals.
[0050] Specifically, according to the physiological characteristics of the human ear auditory system, decompose the original audio signal into non-uniform sub-band signals that match auditory perception. In specific implementation, a filter bank with critical band matching is used for sub-band division. The frequency band boundaries of the filter bank are set according to the critical band model defined in the international standard ISO 532-1:2017, covering the auditory sensitive frequency band from 20 Hz to 16 kHz. The bandwidth of each sub-band expands approximately logarithmically as the center frequency increases. The bandwidth in the low-frequency band (e.g., 20 - 100 Hz) is set to 10 - 30 Hz, and the bandwidth in the high-frequency band (e.g., above 4 kHz) expands to 200 - 500 Hz. The filter bank is implemented using a time-frequency analysis structure with frequency selection characteristics (e.g., gamma-tone filter, equivalent rectangular bandwidth filter bank), and its impulse response is consistent with the frequency resolution characteristics of the human ear basilar membrane, ensuring that the signals after sub-band decomposition conform to the auditory masking effect in the time-frequency domain. The decomposed sub-band signals do not overlap in the frequency domain, and a preset transition band is reserved between adjacent sub-bands to prevent spectral leakage.
[0051] S122: Detect the target sub-band signal where the fundamental frequency harmonics are located among the multiple sub-band signals, and perform harmonic enhancement processing on the target sub-band signal to enhance the signal energy of the fundamental frequency harmonic components.
[0052] Specifically, detect and enhance the fundamental frequency harmonic components of each decomposed sub-band signal. First, extract the voice fundamental frequency through an improved autocorrelation method: perform frame processing on each sub-band signal (frame length 20 ms, frame shift 10 ms), calculate the autocorrelation function of each frame signal, and search for the maximum peak within the fundamental frequency range of 50 - 400 Hz to determine the fundamental frequency candidate value; subsequently, use the harmonic product spectrum algorithm to verify the continuity of the fundamental frequency harmonics, that is, detect whether there are obvious harmonic peaks at integer multiples of the fundamental frequency (2 、3 …n ). If stable harmonic structures are detected in more than 5 consecutive frames, determine that the sub-band is the target sub-band where the fundamental frequency harmonics are located. Apply dynamic gain control to the confirmed target sub-band signal: dynamically adjust the gain coefficient according to the proportion of the harmonic energy in the sub-band to the total energy (the calculation formula is η = E harmonic / E total ). When η > 70%, use a 4 dB gain boost, otherwise maintain linear amplification. The target sub-band signal after gain processing retains the original phase information to avoid waveform distortion. Eharmonic represents the energy of the harmonic components within the target sub - band, that is, the total energy of the frequency components that are integer multiples of the fundamental frequency ( ). E total represents the total energy of all frequency components within the target sub - band, including the energy of harmonics, non - harmonics (such as noise), and other non - periodic components.
[0053] S123: Perform noise suppression processing on sub - band signals other than the target sub - band signal.
[0054] Specifically, perform noise suppression on the remaining sub - bands that are not marked as the target sub - band. First, estimate the noise floor power of each sub - band through the minimum value statistical method, that is, continuously track the minimum power value of the sub - band signal within a 1 - second time window and use it as the noise power estimate value P noise . Subsequently, calculate the instantaneous signal - to - noise ratio of the sub - band ( SNR = 10log 10 ( P current / P noise ), where P current is the instantaneous power value of the current sub - band signal. Apply dynamic attenuation to the sub - bands with a signal - to - noise ratio lower than a preset threshold (such as 6 dB). The attenuation coefficient α is calculated according to the following formula: , and the maximum attenuation amount does not exceed 12 dB. To suppress the remaining "musical noise" after the suppression process, only the amplitude spectrum of the sub - band signal is adjusted during the attenuation process, while keeping the original phase information unchanged. The finally output non - harmonic sub - band signal effectively suppresses the broadband noise component while retaining the naturalness of the ambient sound.
[0055] S124: Recombine the target sub - band signal after harmonic enhancement processing and the other sub - band signals after noise suppression processing to obtain a preliminarily filtered audio signal.
[0056] Specifically, synthesize and reconstruct the enhanced target sub - band signal and the suppressed non - harmonic sub - band signal. In specific implementation, a synthesis filter bank matching the decomposition filter bank is used for frequency - domain to time - domain conversion. During the synthesis process, the phase response of each sub - band signal is pre - calibrated to ensure the alignment accuracy of the time - domain waveform. The reconstructed time - domain signal is processed by a pre - emphasis filter (transfer function is ), to compensate for the high - frequency attenuation during signal transmission. To eliminate the inter - frame discontinuity introduced by frame processing, the overlap - add method is used for signal smoothing, and the length of the overlap region is set to 50% of the frame length (i.e., 10 ms). The finally output preliminarily filtered signal reduces the ambient noise interference while retaining the speech clarity, providing an optimized input for subsequent processing.
[0057] S13: Obtain environmental noise data, and perform environmental noise estimation on the filtered audio signal according to the environmental noise data to obtain environmental noise estimation information.
[0058] Specifically, some environmental noise sensors are built into the Bluetooth headset, or real-time data of the environmental noise level is obtained by interacting with external devices such as smartphones. For example, the system can work in cooperation with the GPS sensor and acceleration sensor of the smartphone to obtain the noise estimation value of the current environment. If the Bluetooth headset or an external device working in cooperation with the Bluetooth headset detects that the user is in a noisy environment (such as a busy street), the weight of the noise estimation will be increased. According to the collected noise data, the system uses a noise estimation algorithm (such as the least mean square error estimation method or Kalman filtering) to perform noise estimation on the filtered audio signal. In this way, the system can intelligently adjust the accuracy and range of the noise estimation according to different changes in the environment.
[0059] S14: Filter the noise of the filtered audio signal according to the environmental noise estimation information using a dynamic noise suppression algorithm to obtain a preprocessed audio signal.
[0060] Specifically, the dynamic noise suppression algorithm includes noise suppression based on spectral subtraction, Wiener filtering, adaptive filtering, etc. Through the dynamic noise suppression algorithm, the system can identify and remove the noise components in the audio signal. When using spectral subtraction, the system estimates the spectrum of the noise according to the environmental noise data and the spectral characteristics of the audio signal, and then subtracts these estimated noise spectra from the audio signal to effectively eliminate the background noise. Wiener filtering is a method of noise suppression based on the power spectral density estimation of the signal and the noise. The Wiener filter calculates the ratio of the power spectrum of the signal to the power spectrum of the noise to obtain an optimal filter, which can maximize the signal-to-noise ratio and suppress the noise. Adaptive filtering is a technique for dynamically adjusting the filter parameters, and its algorithms include the Least Mean Squares (LMS) and Recursive Least Squares (RLS) algorithms. The adaptive filter can automatically optimize the filter parameters according to the changes in the signal, so it is particularly suitable for processing noise in a changing environment. At this time, the process of noise filtering will be adjusted in real time according to the environmental changes. If the environmental noise level changes greatly, the system will correspondingly adjust the parameters of the dynamic noise suppression algorithm to obtain the best effect. The formula of the dynamic noise suppression algorithm can be expressed as:
[0061] ;
[0062] Where is the spectrum of the audio signal after noise filtering, X is the spectrum of the filtered audio signal, is the noise spectrum estimated based on the noise.
[0063] This process can be implemented in a variety of alternative ways. For example, a neural network-based noise suppression algorithm can be used, which can more accurately identify complex noise patterns and suppress them. In this way, dynamic noise suppression can achieve more flexible and accurate noise filtering effects under different environmental conditions.
[0064] Through the above steps, the system first collects the audio signal of the external environment and performs processing such as multi-channel filtering, noise estimation, and dynamic noise suppression. Finally, a preprocessed audio signal with noise suppression is obtained. This process can effectively improve the sound quality performance of the Bluetooth headset, especially in scenarios with complex noise and rapid environmental changes, providing a clearer and more stable audio experience for users.
[0065] Figure 3 is the specific flowchart of step S20 in the embodiment of the present invention, as Figure 3 shown, step S20 includes the following steps:
[0066] S21: Dynamically determine the time window length according to the frequency characteristics of the preprocessed audio signal.
[0067] The traditional Short-Time Fourier Transform (STFT) uses a time window with a fixed length for decomposition. However, in practical applications, the frequency characteristics of the audio signal change over time. Therefore, in this embodiment, a strategy of dynamically adjusting the time window length is adopted to adapt to the changes of different frequency components in the audio signal. Specifically, the system first analyzes the spectral characteristics of the audio signal, including frequency components, frequency changes, and spectral width, etc. If the frequency range of the signal is wide, the time window length can be appropriately increased to ensure that the details in the spectrum can be captured; for a signal with a narrow frequency range, the system shortens the time window length to reduce the computational amount and improve the time-domain resolution. In this way, the time window length can be adaptively adjusted according to the frequency characteristics of the audio signal, thereby optimizing the effect of spectral analysis.
[0068] In an alternative implementation, the time window length is dynamically adjusted based on the spectral entropy of the signal. The spectral entropy measures the complexity of the signal spectrum. When the spectrum of the signal is more complex, the entropy value is higher, and the system will select a longer time window; when the spectrum is relatively flat, the entropy value is lower, and the system uses a shorter time window.
[0069] S22: Based on the length of the time window, perform a short-time Fourier transform on the preprocessed audio signal to obtain the spectral representation of the audio signal.
[0070] The short-time Fourier transform is a method for converting a time-domain signal into a frequency-domain representation. It divides the signal into multiple small time periods and performs a Fourier transform on each period to obtain the time-frequency information of the signal. Specifically, the system uses a dynamically adjusted time window to frame the audio signal, and then performs a Fourier transform on each frame of the signal to convert it from the time domain to the frequency domain, obtaining the spectral representation.
[0071] In this embodiment, the length of the time window determines the trade-off between the frequency resolution and the time-domain resolution of the short-time Fourier transform. For example, if the window length is long, it has better frequency resolution in the spectrum, but poor time-domain resolution; while a shorter window length can improve the time-domain resolution but sacrifices the frequency resolution. By dynamically adjusting the time window length, the system can flexibly handle this trade-off in different environments.
[0072] S23: Decompose the spectral representation, extract the spectral features of multiple frequency bands, and calculate the power distribution of each frequency band to obtain the spectral features and power distribution of the preprocessed audio signal in multiple frequency bands.
[0073] In step S23, the system decomposes the spectral representation obtained by the short-time Fourier transform, extracts the spectral features of multiple frequency bands, and calculates the power distribution of each frequency band. The goal of this process is to extract the features of multiple frequency bands from the spectrum for use in subsequent interference source identification and suppression processes. Specifically, the system first divides the spectrum into multiple frequency bands, for example, divides the spectrum into low-frequency, medium-frequency, and high-frequency regions, corresponding to different audio features (such as human voices, music, background noise, etc.).
[0074] The power distribution of each frequency band can be calculated by the following formula:
[0075] ;
[0076] where P(f k ) is the power distribution of frequency band f k , X(t, f k ) is the spectral value of this frequency band f k at time t, and T is the total number of time windows. By calculating the power distribution of each frequency band, the system can obtain the energy distribution of each frequency band in the audio signal, providing a basis for the detection and suppression of interference signals.
[0077] In addition, to improve the accuracy of spectral feature extraction, the system can also adopt other feature extraction methods, such as Mel-Frequency Cepstral Coefficients (MFCC) extraction, to further process the spectral features of the audio signal. This method is applicable to speech signal analysis and can provide more robust audio features in the case of high ambient noise.
[0078] Through the above steps, the system can dynamically adjust the time window length according to the frequency characteristics of the preprocessed audio signal, thereby implementing the short-time Fourier transform and extracting the spectral features and power distribution of multiple frequency bands from the spectrum.
[0079] Figure 4 It is the specific flowchart of step S30 in the embodiment of the present invention. As Figure 4 shown, step S30 includes the following steps:
[0080] S31: Real-time detect the preprocessed audio signal, and identify the interference sources in the audio signal based on the spectral features and power distribution of the preprocessed audio signal in multiple frequency bands.
[0081] Specifically, the system first uses the spectral decomposition and power distribution calculation methods mentioned in the previous steps to obtain the spectral features of each frequency band. On this basis, the system compares with the known models of interference signals to identify the interference sources in the audio signal in real time. For example, the system can judge whether the current audio signal contains these noise sources by comparing the spectral characteristics of known traffic noise, crowd noise or wind noise.
[0082] To achieve this goal, the system can adopt a pattern matching algorithm based on spectral features, such as Dynamic Time Warping (DTW) or spectral feature matching method, to compare the real-time audio signal with the pre-stored noise library to quickly detect the appearance of interference signals. If an interference source is detected, the system will mark the frequency band where the interference signal is located in the spectrum and provide the necessary data for the subsequent steps.
[0083] S32: Perform spectral analysis on the interference signals corresponding to the identified interference sources to obtain the frequency characteristics of the interference signals.
[0084] Specifically, the system performs more refined spectral analysis on these signals according to the frequency band range of the interference signals detected in real time, and extracts their frequency features, including information such as the main frequency, bandwidth, and frequency distribution. For example, if traffic noise is identified as an interference source, the system will analyze that the frequency range of traffic noise usually concentrates in the lower frequency band, about dozens of hertz to hundreds of hertz, while wind noise appears in a higher frequency range.
[0085] The implementation of spectrum analysis can be achieved through the Fast Fourier Transform (FFT) or other frequency-domain analysis methods to perform high-precision analysis on the identified interference signals and obtain the frequency characteristics of the interference signals. In this way, the system can provide accurate frequency feature data for subsequent noise suppression algorithms.
[0086] S33: Use an adaptive filtering algorithm to dynamically suppress the interference signal according to the frequency characteristics of the interference signal, and obtain an audio signal after interference suppression.
[0087] Specifically, the adaptive filtering algorithm can adjust the parameters of the filter in real time according to the frequency characteristics of the signal, thereby effectively suppressing the influence of the interference signal. Taking the Least Mean Squares (LMS) algorithm as an example, the system will dynamically adjust the weights of the filter based on the frequency characteristics of the identified interference signal to minimize the error between the output signal and the desired signal. In specific implementation, the filter will adjust its amplitude response according to the frequency characteristics of the noise, weaken the energy in the interference frequency band, thereby realizing the dynamic suppression of the interference signal.
[0088] In addition, to improve the dynamic adaptability, the system can select other adaptive filtering algorithms, such as the recursive least squares algorithm or the deep learning-based adaptive filter. These methods can provide higher robustness, especially in an environment where the interference sources are complex and changing rapidly.
[0089] Through the above steps, the system can detect the interference sources in the audio signal in real time, accurately identify their frequency characteristics, and then use the adaptive filtering algorithm to dynamically suppress these interference signals, thereby effectively improving the quality of the audio signal. This process can achieve precise noise suppression in various complex environments, especially in an environment where multiple interference sources coexist. Compared with the traditional static noise cancellation method, the embodiment of the present invention can respond to environmental changes in real time through an adaptive dynamic filtering mechanism, enabling the Bluetooth headset to always maintain high-quality audio output and enhancing the user's auditory experience in various scenarios.
[0090] Figure 5 is the specific flowchart of step S40 in the embodiment of the present invention. As Figure 5 shown, step S40 includes the following steps:
[0091] S41: Obtain real-time change data of environmental conditions and user multimodal status data.
[0092] Specifically, the system obtains real-time change data of environmental conditions, which can be obtained through the sensors of the Bluetooth headset itself or through data exchange with other devices (such as smartphones, smartwatches, smart glasses, smart bracelets, etc.). Specifically, the Bluetooth headset can be equipped with environmental noise sensors, temperature and humidity sensors, acceleration sensors, barometric pressure sensors, etc., for real-time monitoring of changes in the current environment. Or, the Bluetooth headset can obtain the real-time change data of the above environmental conditions through wireless communication with smart devices or wearable devices such as smartphones, smartwatches, and smart glasses. Through these sensors, the system can obtain the intensity of environmental noise, the geographical location of the user, the user's motion state (such as walking, stationary, riding in a vehicle, etc.), and other environmental factors.
[0093] For example, when the user walks from a quiet indoor environment to a noisy outdoor environment, the headset can detect the change in the user's motion state through the acceleration sensor and simultaneously detect the change in the noise level through the environmental noise sensor. These data provide the basis for subsequent steps to ensure that audio processing can be automatically adjusted according to environmental changes.
[0094] S42: Generate context-aware parameters according to the real-time change data of the environmental conditions and the user's multimodal state data.
[0095] Specifically, the context-aware parameters include at least one of the following: volume gain parameter, used to dynamically adjust the volume according to the environmental noise level; adaptive noise reduction parameter, which can optimize the noise reduction level according to the surrounding noise situation; spectrum optimization parameter, which improves the sound quality by enhancing or attenuating specific frequencies; voice enhancement parameter, which improves the voice clarity during a call; sound effect mode parameter, which can match the best sound effect according to the user's activity scenario; sidetone parameter, which enables the user to hear their own voice during a call to enhance the call experience; echo suppression parameter, used to reduce the impact of environmental echo on audio quality; low-latency mode parameter, which ensures audio synchronization during games or video playback; inter-device audio synchronization parameter, which optimizes audio switching and synchronization effects when multiple devices are connected; voice pickup mode parameter, which optimizes the pickup direction of the microphone to improve voice recognition and call quality.
[0096] Specifically, these context-aware parameters will adjust the processing of the audio signal to ensure that the headset outputs the best sound quality under different environmental conditions. For example, when the environmental noise is strong, the system needs to increase the gain of the audio signal to ensure the clarity of the audio; while in a quiet environment, it is necessary to reduce the gain of the audio signal and optimize the spectrum to reduce unnecessary noise.
[0097] In this embodiment, fuzzy logic or machine learning algorithms can be used to generate context-aware parameters according to the real-time change data of the environmental conditions and the user's multimodal state data.
[0098] In the previous data collection stage, the sensor data provided by the Bluetooth headset and its associated devices (such as smartphones, smartwatches, smart glasses, and smart bracelets) can be used to detect real-time changes in the environment, including ambient noise level, temperature, humidity, air pressure, as well as user multimodal status data, including user motion state, user location, Bluetooth connection state, and the usage mode of the headset, etc. Among them, the ambient noise level can be detected by an ambient noise sensor to judge the noise level of the current environment; temperature and humidity are measured by a temperature and humidity sensor to sense the climate change of the external environment; an air pressure sensor is used to detect air pressure changes to identify whether the user is in a high-altitude area or a confined space. The user's motion state can be judged by an acceleration sensor, such as walking, running, taking a vehicle, or stationary. A positioning system or base station positioning can be used to determine the geographical location of the user, and further combined with indoor positioning technologies such as Wi-Fi fingerprint, Bluetooth beacon, and near-field communication tag, specific scenarios (such as the airport check-in area, subway platform) can be identified through signal strength fingerprint matching. In addition, the Bluetooth connection state can be used to judge whether the headset is currently connected to multiple devices and the audio synchronization situation between devices. The Bluetooth connection state can be obtained in real time by querying the device connection information and signal parameters in the Bluetooth protocol stack; the usage mode of the headset (such as call, music, game, etc.) determines the priority of audio processing and the direction of parameter adjustment. The usage mode of the headset can be detected based on the analysis of the current audio stream characteristics or user interaction instructions.
[0099] In the data analysis and feature extraction stage, the collected raw data is analyzed to extract multiple target features, including noise intensity change features, user motion mode features, environment type features, call status features, and audio mode demand features. For example, the noise intensity change features can be used to judge the noise level of the environment by detecting the trend of decibel changes. The motion mode features can identify whether the user is in a walking state, running state, cycling state, taking a vehicle state, or stationary state, while the environment type features can be further divided into scenarios such as quiet indoor, noisy outdoor, in-vehicle, or subway. In addition, this step also analyzes whether the user is in a call mode and the current audio mode demand features, such as call, music, game, or video playback. To achieve accurate classification, machine learning algorithms or rule-based logical analysis can be used. For example, when the noise sensor detects that the ambient noise exceeds 70 dB, it can be judged that the user is in a noisy environment; if the acceleration sensor shows that the user is in a stable motion state, it is determined that the user is walking or running; and when the positioning system identifies that the user is moving at high speed, it is determined that the user is in a vehicle. Through the comprehensive analysis of these features, the user's current environment and the user's current multimodal state or situation can be accurately identified.
[0100] The characteristics of noise intensity change are further subdivided into: low-noise environment, medium-noise environment, and high-noise environment; the characteristics of user movement patterns are further subdivided into: stationary, walking, running, cycling, and riding in a vehicle; the call status characteristics include: call mode and non-call mode; the audio mode demand characteristics include: call mode, music mode, game mode, and video mode; the environmental type characteristics are further subdivided into: quiet indoor environment, including offices, studies, meeting rooms, libraries, etc.; ordinary indoor environment, including homes, cafes, restaurants, shopping malls, etc.; noisy indoor environment, including gyms, playgrounds, exhibition halls, etc.; open outdoor environment, including parks, squares, campuses, etc.; traffic flow environment, including streets, sidewalks, cycling environments, etc.; enclosed traffic environment, including inside a vehicle, buses, subways, airplanes, high-speed trains, etc.; special environment, including airports, stations, subway stations, concerts, etc. The subdivision and classification of environmental type characteristics are achieved through the fusion of multi-modal sensor data (location information, noise, air pressure, wireless signal strength, motion state, etc.) and machine learning classification algorithms (random forest or support vector machine) for accurate classification.
[0101] In the stage of generating corresponding context-aware parameters, one or more context-aware parameters can be obtained through fuzzy logic, look-up table method, or machine learning algorithms based on multiple target features obtained in the data analysis and feature extraction stage.
[0102] Fuzzy Logic is an inference method suitable for uncertain environments. It allows the system to make reasonable decisions even when there are fuzzy relationships between input features. In the process of generating context-aware parameters, fuzzy logic can be used to process multiple target features, such as noise level, motion state, and environmental type, to generate the most suitable audio parameters. For example, when the noise sensor detects that the environmental noise is 65 dB, which is between quiet and noisy, the system does not directly set "low noise reduction" or "strong noise reduction", but dynamically adjusts the noise reduction level to "medium noise reduction" according to the fuzzy membership degree of "noise level". Similarly, if the acceleration sensor shows that the user's motion state is between walking and stationary (such as slow walking), the system does not simply classify it as walking or stationary, but appropriately adjusts the sound effect mode based on fuzzy reasoning to retain some environmental sounds while maintaining the clarity of the sound quality. Fuzzy logic can ensure a smooth transition of parameter adjustment, avoid abrupt audio experiences, and provide more natural adaptability in various environmental changes.
[0103] The lookup table method (Lookup Table, LUT) is an efficient way to generate parameters and is suitable for scenarios that require quick matching of audio optimization strategies. In this method, the system pre-establishes a mapping table between context features and audio parameters. Whenever the input target feature or context feature changes, the system directly looks up the corresponding optimization parameters based on the input features without the need for complex calculations. For example, if the system detects that the noise intensity is 75 dB, the user is in a walking state, and the current mode is the call mode, the system can quickly match the set of parameters "high noise reduction, voice enhancement, sidetone enhancement" by looking up the table. When the user enters a quiet indoor environment, the system can look up the set of parameters "low noise reduction, standard volume, voice enhancement optimization". The advantage of the lookup table method is low computational complexity and fast response speed, making it suitable for devices with limited computing resources such as Bluetooth headsets.
[0104] Machine Learning (ML) can be used to intelligently generate context-aware parameters, especially suitable for adaptive audio optimization in complex environments. Compared with the lookup table method of rule mapping and fuzzy logic, the machine learning method can learn the optimal audio parameters in specific situations from a large amount of user data and automatically adjust them. Machine learning methods include Decision Tree, Neural Network, Reinforcement Learning, etc. For example, the system can use a supervised learning model (such as a random forest or a deep neural network) to train audio optimization schemes in different environments and predict the optimal volume gain, noise reduction level, and sound effect mode based on the current input features. If reinforcement learning is adopted, the system can also continuously optimize according to user feedback. For example, if the user often manually adjusts the volume in a certain environment, the system will learn this preference and automatically apply the same volume gain in similar situations. The machine learning method can continuously optimize parameter matching as data accumulates, making the audio processing of the headset more intelligent. By combining edge computing or cloud training, the computational burden on the Bluetooth headset device side can be reduced.
[0105] S43: Dynamically adjust multiple output parameters of the audio signal according to the context-aware parameters to obtain the adjusted multiple output parameters.
[0106] Specifically, these output parameters include the volume, frequency response, gain, time delay, etc. of the audio. For example, when the system detects that the user is in a noisy environment (such as on the street), the system will automatically increase the volume and enhance the low-frequency response according to the volume gain parameter used to dynamically adjust the volume size, the adaptive noise reduction parameter used to optimize the noise reduction level, and the spectral optimization parameter used to enhance or attenuate specific frequencies, to ensure that the user can clearly hear the audio signal. When the user is in a quiet environment (such as in the office), the system will appropriately reduce the volume according to the volume gain parameter and the spectral optimization parameter, optimize the high-frequency response, and make the sound quality softer and more comfortable. In addition, the system can also adjust the frequency response of the audio and use a dynamic equalizer to optimize the gain of different frequency bands. For example, in a high-noise environment, the low-frequency response will be increased to counteract the environmental noise.
[0107] In a noisy environment, the system will automatically increase the gain of the audio signal, especially in the low-frequency range, to ensure that voices or other important audio components can still be clearly audible against the noise background. In a quiet environment, the low-frequency gain is reduced to make the sound quality smoother and more comfortable.
[0108] In a high-noise environment, the system can moderately increase the time delay to have enough time for audio processing steps such as noise suppression and echo cancellation. By increasing the time delay, the audio processing system can better filter and suppress unnecessary noise in the environment, thereby improving the clarity and intelligibility of the audio signal. In a quiet environment, the system will reduce the time delay to ensure the immediacy of the audio signal. Especially in real-time voice communication or application scenarios that require immediate response, reducing the time delay helps to ensure the fluency and interactivity of the audio.
[0109] S44: Generate an optimized audio signal according to the adjusted multiple output parameters and output it to the speaker of the Bluetooth headset.
[0110] Specifically, the optimized audio signal will be compensated according to the context awareness parameters to ensure the clarity, balance, and comfort of the audio. For example, the audio signal after gain adjustment and spectral optimization will be sent to the speaker for output, and the user can hear the best sound quality experience in different environments.
[0111] Through the above steps, the embodiments of the present invention can automatically adjust the audio output parameters of the Bluetooth headset according to the real-time changes in environmental conditions, thereby effectively coping with noise interference and user requirements in different environments. For example, in a noisy environment, the system automatically increases the volume and optimizes the frequency response, enabling the user to still clearly hear the audio content; in a quiet environment, a softer and more comfortable sound quality is provided. This dynamic adjustment mechanism ensures that the Bluetooth headset can provide the best auditory experience under various complex environmental conditions. The embodiments of the present invention not only improve the adaptability of audio processing but also enable the Bluetooth headset to maintain high sound quality stability under different environmental conditions. Compared with traditional static audio processing methods, the present invention has stronger environmental adaptability and can effectively improve the user experience.
[0112] Figure 6 is the specific flowchart of step S50 of the embodiments of the present invention, as Figure 6 shown, step S50 includes the following steps:
[0113] S51: Obtain the temporal bone vibration signal of the user through the bone conduction sensor built in the Bluetooth headset, and extract the harmonic features of the preset speech frequency band from the time-frequency spectrum of the temporal bone vibration signal.
[0114] In this embodiment, the Bluetooth headset is built with a bone conduction sensor for obtaining the vibration signal of the temporal bone of the wearer. This vibration signal can be collected through the accelerometer of the bone conduction sensor. By performing time-frequency spectrum analysis on the vibration signal of the temporal bone, specific speech frequency bands can be extracted. Specifically, in step S511, wavelet transform or Hilbert transform is used to perform spectrum analysis on the temporal bone vibration signal obtained through the bone conduction sensor. These methods decompose the temporal bone vibration signal into multiple time-frequency intervals, revealing the change characteristics of the temporal bone vibration signal at different times and frequencies. In step S512, a time-frequency spectrum is generated based on these time-frequency intervals. The time-frequency spectrum is a two-dimensional representation of the distribution of the temporal bone vibration signal in time and frequency. It integrates the spectrum information of the temporal bone vibration signal at different time points to form a comprehensive time-frequency diagram. In step S513, according to the preset speech frequency band (in the range of 300 Hz to 3400 Hz), the harmonic features within the corresponding frequency band are extracted from the time-frequency spectrum. These harmonic features can accurately reflect the frequency components of the speech signal, thereby effectively distinguishing the speech signal from other noise or interference signals.
[0115] The mathematical formulas for extracting harmonic features include but are not limited to the following: ;
[0116] where X(t,f) is the time-frequency spectrum obtained by wavelet transform or Hilbert transform, H(f) is the harmonic feature of the speech frequency band, f0 is the center frequency of the set speech frequency band, and δ(f f0) is a narrowband filter.
[0117] The variable t represents the time index, indicating a specific time point in the time-frequency spectrum X(t, f), that is, the time frame or discrete time point after short-time analysis of the temporal bone vibration signal.
[0118] The variable T represents the total time range or the number of time frames of the analysis window, indicating the summation from t = 0 to t = T to aggregate the harmonic characteristics within the entire time window.
[0119] In this formula, |X(t, f)| represents the signal intensity at time t and frequency f, while δ(f f0) serves as a narrowband filter to select a specific speech frequency band f0, and then sums the energy of this frequency band over all time frames to extract the harmonic characteristics of this frequency band.
[0120] S52: Convolve and mix the harmonic characteristics of the preset speech frequency band with the optimized audio signal in the frequency domain to generate a phase-synchronized enhanced audio signal.
[0121] In this embodiment, by convolving and mixing the extracted harmonic characteristics of the speech frequency band with the optimized audio signal in the frequency domain, the clarity and audibility of the speech signal are enhanced. This step includes the following steps:
[0122] S521: Obtain the harmonic characteristics of the preset speech frequency band from the audio signal through time-frequency analysis methods.
[0123] First, based on the harmonic characteristics of the preset speech frequency band (the frequency band between 300 Hz and 3400 Hz) extracted in the previous steps. These harmonic characteristics are extracted from the time-frequency analysis of the audio signal through technical means such as wavelet transform or short-time Fourier transform. The harmonic characteristics represent important information such as pitch and frequency components in the speech signal and are an accurate description of the speech part in the audio signal. The harmonic characteristics can be expressed as H(f), where f is the frequency, representing the spectral characteristics within the speech frequency band.
[0124] S522: Perform spectral analysis on the optimized audio signal to obtain the spectrum of the optimized audio signal.
[0125] This step performs spectral analysis on the already optimized audio signal. The optimized audio signal is the audio signal after optimization steps such as noise suppression and dynamic processing, which has eliminated most of the background noise and has high signal quality. For this signal, spectral analysis can be performed using methods such as short-time Fourier transform and fast Fourier transform. The result of spectral analysis is the energy distribution of the audio signal at each frequency point, expressed as S(f), where (f) is the amplitude information of the optimized audio signal at frequency f.
[0126] S523: Perform a convolution operation between the harmonic features of the preset voice frequency band and the spectrum of the optimized audio signal to achieve frequency-domain mixing and generate an enhanced audio signal.
[0127] Perform a convolution operation between H(f) (harmonic features of the voice frequency band) obtained from spectral analysis and S(f) (spectrum of the optimized audio signal). The convolution operation is achieved by weighted superposition of these two frequency-domain signals. The goal is to combine the harmonic features of the voice with the optimized audio signal, thereby enhancing the voice component.
[0128] The convolution operation can be expressed by the following formula: Y(f) = H(f) S(f);
[0129] where Y(f) is the enhanced audio signal, H(f) is the harmonic features of the preset voice frequency band, S(f) is the spectrum of the optimized audio signal, representing the convolution operation in the frequency domain.
[0130] Through the convolution operation, the harmonic feature H(f) and the optimized signal S(f) are weighted and mixed in the frequency domain to ensure that the voice signal is enhanced in the final audio signal.
[0131] S524: Adjust the phase information of the frequency band to ensure the phase alignment of the harmonic features of the preset voice frequency band and the spectrum of the optimized audio signal in the frequency domain, thereby generating a phase-synchronized enhanced audio signal.
[0132] When performing frequency-domain convolution mixing, attention should be paid to the phase synchronization problem of the signal. Phase synchronization means maintaining the phase consistency of each frequency band of the signal during spectrum mixing to ensure that the time-domain waveform of the signal does not distort. To achieve this, during convolution mixing, the phase information of each frequency band can be adjusted to ensure the phase alignment of H(f) and S(f) in the frequency domain, thereby generating the synthesized enhanced audio signal Y(f).
[0133] By ensuring phase synchronization during the frequency-domain convolution process, the finally obtained enhanced audio signal Y(f) can better match the characteristics of the original signal in terms of both frequency and phase, while enhancing the voice component and reducing noise interference.
[0134] S53: Obtain real-time air pressure data of the ear canal closed space based on the air pressure sensor in the inner cavity of the Bluetooth headset, calculate the current ear canal resonance peak offset according to the real-time air pressure data of the ear canal closed space and the acoustic resonance physical model, and determine the frequency band range affected by resonance according to the current ear canal resonance peak offset.
[0135] Specifically, this process can include the following steps:
[0136] S531: Obtain real-time air pressure data of the ear canal closed space.
[0137] The air pressure sensor built in the Bluetooth headset monitors the air pressure change in the ear canal in real time. The air pressure data of the ear canal is the comprehensive result of factors such as ambient noise and external pressure change, which has an impact on the acoustic characteristics of the ear canal. By measuring and collecting the air pressure data in the ear canal, the real-time ear canal air pressure value can be obtained, denoted as P, which is the real-time ear canal air pressure data that changes with time.
[0138] S532: Calculate the offset of the current ear canal resonance peak according to the real-time air pressure data of the ear canal closed space and the acoustic resonance physical model.
[0139] According to the air pressure change of the ear canal, use the acoustic resonance physical model to calculate the current ear canal resonance peak offset. The resonance frequency of the ear canal will shift with the air pressure change. Through the acoustic resonance physical model, the current ear canal resonance peak frequency f res (P) and the relationship between the ear canal air pressure value P can be obtained. The acoustic resonance physical model describes the relationship between the ear canal resonance peak frequency and the air pressure through the following formula: ;
[0140] where, f res (P) is the resonance peak frequency calculated according to the current ear canal air pressure value P; f res0 is the ear canal resonance peak frequency under the standard ear canal air pressure P0; P0 is the standard ear canal air pressure value; P is the currently measured ear canal air pressure value; γ is an empirical constant used to adjust the relationship between the air pressure and the resonance frequency. Through this formula calculation, the current ear canal resonance peak frequency f res (P) can be obtained, and the change of the resonance frequency band is determined according to the offset.
[0141] S533: Determine the frequency band range affected by resonance according to the current ear canal resonance peak offset.
[0142] By calculating the current ear canal resonance peak offset: Δfres = f res (P) f res0 , the offset of the resonance peak relative to the standard ear canal air pressure can be obtained. According to this offset, the frequency band range affected by resonance is determined. Since the influence of ear canal resonance is concentrated in a certain frequency range, a frequency band width Δf can be set to define the frequency band affected by resonance. This frequency band width will be adjusted according to the offset of the resonance peak to ensure that the frequency range affected by the resonance effect is covered.
[0143] Specifically, the frequency band range affected by resonance can be determined as [f res (P) Δf / 2, f res(P) + Δf / 2], the signals within this range will be affected by ear canal resonance and need to be appropriately corrected in subsequent steps.
[0144] S54: Perform dynamic notch compensation on the signals in the enhanced audio signal that are within the frequency band range affected by resonance to obtain a notch correction signal.
[0145] Specifically, based on the offset frequency of the ear canal resonance peaks, select and locate the frequency band range affected by resonance. For the audio signals in these frequency bands, use a dynamic notch filter for compensation to eliminate the deterioration of sound quality caused by resonance. During dynamic notch compensation, adjust the center frequency of the notch filter in real time according to the actual resonance peak frequency of the ear canal to ensure that the compensation process can effectively repair the audio distortion caused by the resonance phenomenon.
[0146] The filter for dynamic notch compensation can be implemented by the following formula: ;
[0147] where W(f) is the frequency response of the notch filter, f res is the center frequency of the frequency band affected by resonance, Δf is the bandwidth of the notch filter, and f is the currently processed frequency.
[0148] S55: Perform weighted fusion of the notch correction signal and the optimized audio signal to obtain the final audio output signal.
[0149] After performing notch compensation, perform weighted fusion of the corrected signal and the optimized audio signal. The purpose of weighted fusion is to balance the contributions of the correction signal and the optimized signal to ensure the quality of the final audio output. Specifically, by adjusting the fusion coefficient, make the correction signal play a greater role within the frequency band affected by resonance, while maintaining the original characteristics of the optimized signal in other frequency bands.
[0150] The mathematical expression for weighted fusion is: S final (f) = α Y(f) + (1 α) X(f);
[0151] where, S final (f) is the final audio output signal, Y(f) is the notch correction signal, X(f) is the optimized audio signal, and α is the fusion coefficient that controls the weights of the two.
[0152] S56: Output the final audio output signal to the Bluetooth headset speaker for playback.
[0153] Finally, the final audio signal after weighted fusion is output to the speaker of the Bluetooth headset for playback. At this time, the user can hear the high-quality audio signal after optimization, enhancement, and correction of the resonance effect, ensuring the best guarantee of audio clarity and naturalness in different environments.
[0154] In this embodiment, the bone conduction sensor is used to directly capture the temporal bone vibration signal, which can extract the harmonic features related to voice production from the physical vibration level, effectively overcome the interference of environmental noise on the voice components, and improve the purity of voice feature extraction in noisy scenarios; by real-time monitoring the air pressure change in the closed space of the ear canal and establishing a dynamic acoustic model, the ear canal resonance frequency shift caused by the change of wearing state can be accurately identified, and the adaptive frequency response compensation for the current ear canal morphology can be realized; by combining the phase synchronization mixing and the dynamic notch technology, while enhancing the voice harmonic energy, the specific resonance frequency band is suppressed, which not only maintains the improvement of voice clarity but also avoids the common timbre distortion problem in traditional noise reduction processing; finally, through the intelligent weighted fusion of multiple signal sources, the enhanced audio output not only has the practical characteristics of prominent human voice but also maintains a natural and balanced auditory experience. This solution systematically solves the technical problems that are difficult to balance in noise suppression, ear canal resonance compensation, and sound quality maintenance of traditional headphones through the organic cooperation of biometric sensing, environmental adaptation modeling, and digital signal processing.
[0155] Embodiment 2
[0156] As Figure 7 shown, the embodiment of the present invention further provides a computer-readable storage medium 600, in which a computer program 610 for executing each step in the method embodiment is stored. When the computer program 610 is executed by a processor, the above-mentioned any audio processing method is realized.
[0157] Figures 1 to 6When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. Of course, there are also other ways of readable storage media, such as quantum memory, graphene memory, and so on. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0158] Embodiment III
[0159] Figure 8 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. An embodiment of the present invention also provides an electronic device. Please refer to Figure 8 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0160] The processor, network interface, and memory can be interconnected through an internal bus. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a bidirectional arrow is used in
[0161] A memory for storing programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming an automated disaster recovery system based on centralized configuration at the logical level. A processor that executes the program stored in the memory and is specifically used to execute Figures 1 to 6 the audio processing method disclosed in the illustrated embodiment.
[0162] As described above Figures 1 to 6 the processing method of the distributed file system disclosed in the illustrated embodiment can be applied to or implemented by the processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0163] Of course, in addition to the software implementation, the electronic device of the present invention does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device. The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0164] Although the present invention provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many execution orders of steps and does not represent the only execution order. When the actual device or terminal product executes, it can be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment).
[0165] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices, and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0167] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0168] Each embodiment in this specification is described in a related manner. For the same and similar parts between the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, electronic device and readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant parts.
[0169] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. An audio processing method, applicable to a Bluetooth headset, characterized in that: The method comprises the following steps: S10: Acquire an audio signal, and preprocess the acquired audio signal using a multi-channel filtering algorithm and a dynamic noise suppression algorithm to obtain a preprocessed audio signal; S20: performing spectrum decomposition on the preprocessed audio signal using a short-time Fourier transform based on a dynamic time window length to obtain spectrum characteristics and power distribution of the preprocessed audio signal in multiple frequency bands; S30: detecting external interference sources in the preprocessed audio signal in real time, identifying interference signals based on the spectrum characteristics and power distribution of the multiple frequency bands, and suppressing the interference signals using an adaptive filtering algorithm to obtain an audio signal after interference suppression; S40: dynamically adjusting multiple output parameters of the audio signal after interference suppression according to real-time change data of environmental conditions and multimodal state data of the user to obtain an optimized audio signal; S50: Obtain a final audio output signal according to the optimized audio signal, and output the final audio output signal to a speaker of the Bluetooth headset for playback.
2. The audio processing method according to claim 1, characterized in that: Step S10 specifically includes: S11: Collecting audio signals in the external environment to obtain original audio signals; S12: using a non-uniform sub-band decomposition method based on the masking effect of human hearing, the original audio signal is divided into a plurality of sub-band signals that meet the critical frequency band characteristics of the human ear, the sub-band signal where the fundamental frequency harmonics are located is enhanced, and other sub-band signals other than the fundamental frequency harmonics are subjected to noise suppression processing, and the enhanced sub-band signal where the fundamental frequency harmonics are located is recombined with other sub-band signals after the noise suppression processing to obtain a preliminarily filtered audio signal; S13: Acquire environmental noise data, and perform environmental noise estimation on the audio signal after preliminary filtering according to the environmental noise data to obtain environmental noise estimation information; S14: According to the environmental noise estimation information, a dynamic noise suppression algorithm is used to filter out the noise of the audio signal after the preliminary filtering to obtain a preprocessed audio signal.
3. The audio processing method according to claim 1, characterized in that: Step S20 specifically includes: S21: dynamically determining the time window length according to the frequency characteristics of the preprocessed audio signal; S22: Based on the time window length, perform short-time Fourier transform on the preprocessed audio signal to obtain a frequency spectrum representation of the audio signal; S23: Decomposing the frequency spectrum representation, extracting frequency spectrum features of multiple frequency bands, and calculating the power distribution of each frequency band, to obtain the frequency spectrum features and power distribution of the preprocessed audio signal in multiple frequency bands.
4. The audio processing method according to claim 1, characterized in that: Step S30 specifically includes: S31: detecting the preprocessed audio signal in real time, and identifying interference sources in the audio signal based on the spectrum characteristics and power distribution of the preprocessed audio signal in multiple frequency bands; S32: Performing spectrum analysis on the interference signal corresponding to the identified interference source to obtain the frequency characteristics of the interference signal; S33: using an adaptive filtering algorithm to dynamically suppress the interference signal according to the frequency characteristics of the interference signal to obtain an audio signal after interference suppression.
5. The audio processing method according to claim 1, characterized in that: Step S40 specifically includes: S41: Acquire real-time change data of environmental conditions and multimodal status data of users; S42: generating context awareness parameters according to the real-time change data of the environmental conditions and the multimodal state data of the user; S43: dynamically adjusting multiple output parameters of the audio signal according to the context perception parameter to obtain multiple adjusted output parameters; S44: Generate an optimized audio signal according to the adjusted multiple output parameters, and output the optimized audio signal to the speaker of the Bluetooth headset.
6. The audio processing method according to claim 1, characterized in that: Step S50 specifically includes: S51: obtaining a user's temporal bone vibration signal through a bone conduction sensor built into the Bluetooth headset, and extracting harmonic features of a preset voice frequency band from a time spectrum of the temporal bone vibration signal; S52: performing convolution mixing on the harmonic features of the preset speech band and the optimized audio signal in the frequency domain to generate a phase-synchronized enhanced audio signal; S53: obtaining real-time air pressure data of the closed space of the ear canal based on the air pressure sensor in the inner cavity of the Bluetooth headset, calculating the current ear canal resonance peak offset according to the real-time air pressure data of the closed space of the ear canal and the acoustic resonance physical model, and determining the frequency band range affected by the resonance according to the current ear canal resonance peak offset; S54: performing dynamic notch compensation on the signal in the enhanced audio signal within the frequency band affected by the resonance to obtain a notch correction signal; S55: performing weighted fusion of the notch correction signal and the optimized audio signal to obtain a final audio output signal; S56: Output the final audio output signal to the speaker of the Bluetooth headset for playback.
7. The audio processing method according to claim 6, characterized in that: Step S51 specifically includes: S511: performing spectrum analysis on the temporal bone vibration signal acquired by the bone conduction sensor using wavelet transform or Hilbert transform, and decomposing the temporal bone vibration signal into multiple time-frequency intervals; S512: Generate a time-frequency spectrum based on the multiple time-frequency intervals, where the time-frequency spectrum is a two-dimensional representation of the distribution of the temporal bone vibration signal in time and frequency; S513: According to the preset speech frequency band, extract harmonic features in the corresponding frequency band from the time-frequency spectrum.
8. The audio processing method according to claim 6, characterized in that: Step S52 specifically includes: S521: Obtaining harmonic features of a preset speech frequency band from the audio signal by using a time-frequency analysis method; S522: performing spectrum analysis on the optimized audio signal to obtain a spectrum of the optimized audio signal; S523: performing a convolution operation between the harmonic characteristics of the preset speech frequency band and the spectrum of the optimized audio signal to achieve frequency domain mixing and generate an enhanced audio signal; S524: Adjusting the phase information of the frequency band to ensure that the harmonic characteristics of the preset speech band and the spectrum of the optimized audio signal are aligned in phase in the frequency domain, thereby generating a phase-synchronized enhanced audio signal.
9. An electronic device, characterized in that: It includes: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the audio processing method as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the audio processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Voice endpoint detection method and device based on non-uniform sub-band separation variances
CN110610724A
Wake-up method and device of MEMS earphone, equipment and storage medium
CN118354237A