A voice enhancement method, system and storage medium of a Bluetooth earphone

By extracting the spectral features and auditory scene cues of the Bluetooth headset audio signal, and combining the acoustic features modulated by the auditory scene cues, a joint feature representation is generated and input into the speech enhancement network. This solves the problem of insufficient speech clarity and intelligibility in traditional audio processing technology, and achieves speech enhancement effect in complex environments.

CN122138092AInactive Publication Date: 2026-06-02SHENZHEN GAOWEI COMM TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN GAOWEI COMM TECH CO LTD
Filing Date
2026-05-08
Publication Date
2026-06-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional audio processing techniques mainly focus on simple signal filtering and gain adjustment, which makes it difficult to effectively improve the voice clarity and intelligibility of Bluetooth headphones in complex environments.

Method used

By acquiring audio signals from a microphone, spectral features, acoustic features, and auditory scene cues are extracted. The acoustic features are modulated in conjunction with the auditory scene cues to generate a joint feature representation, which is then input into a speech enhancement network for processing to enhance the speech data.

Benefits of technology

It significantly improves speech intelligibility in noisy environments, maintains the stability and clarity of speech signals, and enhances the robustness and enhancement effect of speech signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138092A_ABST
    Figure CN122138092A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio processing technology and provides a method, system, and storage medium for voice enhancement in Bluetooth headsets. The voice enhancement method includes: acquiring an audio signal to be enhanced collected by a microphone; extracting spectral features, acoustic features, and auditory scene cues corresponding to the audio signal; modulating the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features; concatenating the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent voice enhancement networks; and inputting the joint feature representation into the voice enhancement network to obtain enhanced voice data. This solution, by combining different audio features, makes the enhanced voice clearer, more natural, and adaptable to varying auditory environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of audio processing, and particularly relates to a method, system and storage medium for voice enhancement of Bluetooth headsets. Background Technology

[0002] With the development of wireless technology, Bluetooth headsets have become widely used in daily life, especially in scenarios such as listening to music, making and receiving calls, and voice recognition. The portability and convenience of Bluetooth headsets make them an indispensable audio device for modern people. However, due to environmental noise, echoes, and other interference factors, users often encounter problems such as poor sound quality and low voice recognition rates when using Bluetooth headsets. These issues not only affect the user experience but also limit the application of Bluetooth headsets in a wider range of scenarios.

[0003] Traditional audio processing techniques primarily focus on simple signal filtering and gain adjustment, which often fails to effectively improve speech clarity and intelligibility. Especially in complex auditory environments, simply relying on signal enhancement methods cannot fully utilize the rich information contained in the audio signal. Therefore, a novel audio processing solution is urgently needed to enhance the intelligibility of speech signals in noisy environments while preserving their characteristics. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a voice enhancement method, system and storage medium for Bluetooth headsets to solve the technical problem that traditional audio processing technologies mainly focus on simple signal filtering and gain adjustment, which often makes it difficult to effectively improve the clarity and intelligibility of speech.

[0005] A first aspect of this invention provides a voice enhancement method for a Bluetooth headset, the voice enhancement method for the Bluetooth headset comprising: Acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features, and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index, and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; Based on the auditory scene cues, the acoustic features are modulated to obtain modulated acoustic embedding features; The modulated acoustic embedding features and the spectral features are concatenated along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The joint feature representation is input into the speech enhancement network to obtain enhanced speech data.

[0006] Furthermore, the step of acquiring the audio signal to be enhanced collected by the microphone and extracting the spectral features, acoustic features, and auditory scene cues corresponding to the audio signal to be enhanced includes: Acquire the audio signal to be enhanced, captured by the microphone; The audio signal to be enhanced is subjected to frame segmentation, windowing, and short-time Fourier transform to obtain a complex spectrum; Calculate the amplitude spectrum corresponding to the complex spectrum; The amplitude spectrum is converted into a Mel spectrum and used as the spectral feature; Based on the Mel spectrum, acoustic features and auditory scene cues are extracted.

[0007] Furthermore, the step of extracting acoustic features and auditory scene cues based on the Mel spectrum includes: Extract the spectral centroid and spectral energy concentration from the Mel spectrum, and extract the zero-crossing rate of the signal to be enhanced; Calculate the peak value of the autocorrelation function between signal frames in the audio signal to be enhanced, and use it as the harmonicity index; Calculate the energy change smoothness of each signal frame in the time direction, and scale the energy change smoothness to obtain the time continuity index; For each signal frame, a set of spectral peak frequencies representing local maxima of energy are extracted from the amplitude spectrum; Obtain a preset candidate fundamental frequency, and for the frequency band corresponding to the signal frame, calculate the similarity between the spectral peak frequency in the frequency band and the harmonic sequence with each candidate fundamental frequency as the fundamental tone. The maximum value among multiple similarities is used as the frequency domain correlation index of that frequency band in the signal frame.

[0008] Furthermore, the step of modulating the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features includes: Based on the auditory scene cues, the probability of each signal frame belonging to multiple predefined auditory component categories is calculated and used as the prior probability of auditory attention; wherein, the multiple predefined auditory component categories include foreground speech components, stable background noise components, and sudden interference components; The acoustic features are linearly transformed and projected onto a dimension with the same number of channels as the Mel spectrum to obtain the initial acoustic embedding features; Based on the prior probability of auditory attention, the dynamic modulation weight is calculated; wherein, the dynamic modulation weight = , This represents the prior probability of auditory attention corresponding to the foreground speech component. This represents the prior probability of auditory attention corresponding to a stable background noise component. This represents the prior probability of auditory attention corresponding to the sudden interference component. , and Represents a learnable scalar parameter; The dynamic modulation weights are constrained by the Sigmoid function to obtain the modulation coefficients; The initial acoustic embedding features are multiplied element-wise with the modulation coefficients to obtain the modulated acoustic embedding features.

[0009] Furthermore, the step of calculating the probability that each signal frame belongs to multiple predefined auditory component categories based on the auditory scene cues, and using this probability as the prior probability of auditory attention, includes: Multiply the time continuity index and the frequency domain correlation index to obtain the initial value corresponding to the foreground speech component; Subtracting 1 from the frequency domain correlation index and multiplying it by the time continuity index yields the initial value corresponding to the stable background noise component. Subtract 1 from the frequency domain correlation index and multiply it by the peak energy ratio of the signal frame to obtain the initial value corresponding to the burst interference component; The initial values ​​corresponding to the foreground speech component, the stable background noise component, and the sudden interference component are normalized by softmax to obtain the prior probability of auditory attention belonging to each category of the signal frame. The auditory attention prior probabilities of all signal frames are downsampled to the same time-frequency resolution as the Mel spectrum by mapping the Mel filter bank.

[0010] Further, the step of inputting the joint feature representation into the speech enhancement network to obtain enhanced speech data includes: The joint feature is represented as the first modality feature; The joint feature representation is downsampled in the time dimension, and the auditory attention prior probabilities corresponding to multiple predefined auditory component categories in multiple signal frames are downsampled in the time dimension. The downsampled joint feature representation is concatenated with the downsampled auditory attention prior probability to obtain the second modality feature; The speech enhancement network performs intramodal encoding on the first modal features to obtain the first intermediate features; The speech enhancement network performs intramodal encoding on the second modality features to obtain the second intermediate features; The speech enhancement network uses the first intermediate feature as the query vector and the second intermediate feature as the key vector and value vector to calculate the first cross-modal attention feature; The speech enhancement network uses the second intermediate feature as the query vector and the first intermediate feature as the key vector and value vector to calculate the second cross-modal attention feature. The speech enhancement network outputs the enhanced speech data based on the first cross-modal attention features and the second cross-modal attention features.

[0011] Further, the step of calculating the first cross-modal attention feature using the first intermediate feature as the query vector and the second intermediate feature as the key vector and value vector includes: The first intermediate feature is projected through the first query linear transformation layer to obtain the first query vector; The second intermediate feature is projected through the first key linear transformation layer and the first value linear transformation layer respectively to obtain the first key vector and the first value vector; wherein the first query vector, the first key vector and the first value vector have the same feature dimension. Calculate the attention weight of the dot product between the first query vector and the first key vector; The first cross-modal attention feature is obtained by weighting and summing the first value vector based on the dot product attention weights.

[0012] Furthermore, the step of the speech enhancement network outputting the enhanced speech data based on the first cross-modal attention features and the second cross-modal attention features includes: The first cross-modal attention feature is projected through the first output linear transformation layer, then weighted and summed with the first intermediate feature, and then normalized by the layer to obtain the first updated feature; wherein the weights of the weighted summation are learnable parameters. The second cross-modal attention feature is projected through the second output linear transformation layer, then weighted and summed with the second intermediate feature, and then normalized by the layer to obtain the second updated feature; wherein, the weights of the weighted summation are learnable parameters; The first updated feature and the second updated feature are respectively used as the first modal feature and the second modal feature input to the next fusion layer; The first updated feature and the second updated feature are processed by a feedforward neural network and residual connections respectively. Based on the final first updated feature and the final second updated feature after processing by L fusion layers, a deep speech feature representation is generated. The decoder converts the deep speech feature representation into the original time-frequency domain corresponding to the enhanced speech data.

[0013] A second aspect of this invention provides a voice enhancement system for a Bluetooth headset, the voice enhancement system for the Bluetooth headset comprising: The acquisition unit is used to acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; A modulation unit is used to modulate the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features; The splicing unit is used to splice the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The processing unit is used to input the joint feature representation into the speech enhancement network to obtain enhanced speech data.

[0014] A third aspect of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the voice enhancement method for Bluetooth headsets described in the first aspect.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the voice enhancement method for Bluetooth headsets described in the first aspect.

[0016] The beneficial effects of this invention compared to existing technologies are as follows: by extracting spectral and acoustic features from the audio signal to be enhanced collected by the microphone, speech signals and environmental noise can be effectively separated. Especially in noisy environments, the modulated acoustic embedding features can significantly improve speech intelligibility. This method combines auditory scene cues, such as the temporal continuity and frequency domain correlation of signal frames, to better adapt to dynamically changing environments. By modulating acoustic features based on these cues, the enhanced speech signal maintains better stability and clarity in variable environments, improving signal robustness. Concatenating the modulated acoustic embedding features and spectral features along the channel dimension generates a joint feature representation that comprehensively considers information from different features, thus providing richer input for subsequent speech enhancement networks. This effective feature fusion method helps deep learning models better learn and recognize complex patterns in speech signals, improving enhancement performance. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic flowchart of a voice enhancement method for a Bluetooth headset provided by the present invention is shown; Figure 2 A schematic diagram of a voice enhancement system for a Bluetooth headset according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of a terminal device provided in an embodiment of the present invention is shown. Detailed Implementation

[0019] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0020] This invention provides a method, system, and storage medium for enhancing voice in Bluetooth headsets, addressing the technical problem that traditional audio processing technologies mainly focus on simple signal filtering and gain adjustment, which often fails to effectively improve voice clarity and intelligibility.

[0021] First, this invention provides a voice enhancement method for Bluetooth headsets. Please refer to [link / reference]. Figure 1 , Figure 1 A schematic flowchart of a voice enhancement method for a Bluetooth headset provided by the present invention is shown. Figure 1 As shown, the voice enhancement method for this Bluetooth headset may include the following steps: Step 101: Acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; Spectral features are fundamental time-frequency representations (such as STFT spectra and Mel spectra), carrying the most comprehensive information on the signal energy distribution of speech and noise. Acoustic features refer to a set of low-dimensional statistical features with clear physical meaning. Acoustic features include, but are not limited to, spectral centroid, spectral energy concentration, harmonicity index, and / or zero-crossing rate.

[0022] The centroid of the spectrum reflects the center of gravity of the spectrum and can distinguish between low-pitched noise and bright speech. The spectral energy concentration describes the degree of concentration of energy distribution; speech is generally more concentrated than stationary noise. The harmonicity index quantifies the harmonic structure of a signal; speech (especially vowels) has strong harmonicity, while many noises have weak harmonicity. The zero-crossing rate refers to the frequency at which the signal crosses zero, and can roughly distinguish between unvoiced sounds (high zero-crossing rate) and voiced sounds / low-frequency noise (low zero-crossing rate). These features characterize the acoustic properties of a signal from different perspectives, require little computation, and have a natural ability to distinguish certain types of noise.

[0023] Auditory scene cues refer to key clues used in scene analysis that simulate the human auditory system. Auditory scene cues include, but are not limited to, the temporal continuity index of signal frames and the frequency domain correlation index of signal frames.

[0024] The temporal continuity index measures the degree of stability or abrupt change of a signal along the time axis. Speech and some continuous noise exhibit temporal continuity, while transient impulse noise does not. The frequency domain correlation index measures the degree of coordination between changes in different frequency bands. Speech frequency bands, excited by the same sound source, often exhibit some correlation, while the correlation patterns of some broadband noises may differ. These clues do not directly describe what the signal is, but rather how the signal is organized, providing a higher-level basis for subsequently distinguishing speech from noise and determining the type of noise.

[0025] Specifically, step 101 includes steps 1011 to 1015: Step 1011: Acquire the audio signal to be enhanced by the microphone; Step 1012: Perform frame segmentation, windowing, and short-time Fourier transform on the audio signal to be enhanced to obtain a complex spectrum; Framing breaks down a continuous audio stream into short time intervals (e.g., 20ms-40ms per frame). Windowing (e.g., Hamming windowing) reduces signal truncation (spectral leakage) caused by framing, allowing the signal at both ends of the frame to smoothly transition to zero. Short-time Fourier transform converts the time-domain signal of each frame into a complex spectrum. This spectrum contains the amplitude and phase information of each frequency component and is the foundation for all subsequent feature calculations.

[0026] Step 1013: Calculate the amplitude spectrum corresponding to the complex spectrum; For tasks such as speech enhancement, the energy distribution (i.e., amplitude) of the signal is usually more critical than phase information (especially in deep learning models, where phase is often ignored or processed separately in the early stages).

[0027] Taking the amplitude spectrum, or the magnitude of the complex spectrum, reflects the energy strength of the signal at different times and frequencies, and is the core input for subsequent extraction of Mel spectrum and various features.

[0028] Step 1014: Convert the amplitude spectrum into a Mel spectrum and use it as the spectral feature; The Mel scale is a non-linear frequency scale based on the characteristics of human hearing. The human ear is sensitive to low-frequency differences (on the Hz scale) but insensitive to high-frequency differences. The Mel scale maps linear frequencies to a scale that approximates human auditory perception.

[0029] Mel spectrum is achieved by applying a set of Mel filter banks (triangular filters, dense at low frequencies and sparse at high frequencies) to the amplitude spectrum. This process is equivalent to reducing and smoothing the spectrum, and focusing on the frequency region that the human ear is sensitive to.

[0030] Mel spectra, as spectral features, are closer to speech perception, helping models learn patterns important to human hearing. Their dimensionality is typically much lower than the original STFT spectrum, reducing the computational cost and parameter requirements of subsequent neural networks. They are insensitive to small frequency shifts, enhancing feature stability.

[0031] Step 1015: Based on the Mel spectrum, extract acoustic features and auditory scene cues.

[0032] Acoustic features (spectral centroid, energy concentration, etc.) and auditory scene cues (temporal continuity, frequency domain correlation) are not calculated directly from the original waveform or linear spectrum, but are based on Mel spectrum calculation.

[0033] All features share the same data source preprocessed using an auditory model (Mel scale), ensuring inherent consistency among features. Features computed at the Mel scale inherently incorporate the perceptual weights of the human ear. For example, the spectral centroid computed at the Mel scale reflects perceived brightness more accurately than that computed at the Hertzian scale. The Mel spectrum has a lower dimensionality, making it less computationally intensive to compute statistical features (such as centroid and energy concentration) and correlation indices, thus making it more suitable for embedded devices such as Bluetooth headsets.

[0034] In the embodiments corresponding to steps 1011 to 1015, feature information that helps enhance the speech signal is extracted through the processes of framing, windowing, short-time Fourier transform, amplitude spectrum calculation, and Mel spectrum conversion. These steps not only improve the accuracy of signal analysis but also make the extracted features more consistent with the characteristics of human hearing, laying a solid foundation for subsequent speech enhancement processing.

[0035] Specifically, step 1015 includes steps A1 to A6: Step A1: Extract the spectral centroid and spectral energy concentration from the Mel spectrum, and extract the zero-crossing rate of the signal to be enhanced; The Mel frequency centroid is the Mel frequency calculated by weighting the energy of the Mel frequency bands, and it reflects the perceived brightness of the sound.

[0036] Mel spectrum energy concentration refers to the variance or kurtosis of the energy in a Mel spectrum, reflecting whether the energy is concentrated in a narrow band or dispersed. Calculated as follows: ; in, express, Indicates time, Indicates the Mel filter bank index. This represents the total number of Mel filter banks. This represents a variable that accumulates from the first Mel band to the kth Mel band.

[0037] Step A2: Calculate the peak value of the autocorrelation function between signal frames in the audio signal to be enhanced, and use it as the harmonicity index; The autocorrelation function is used to measure the similarity of a signal to its delayed copy. For periodic or quasi-periodic signals (such as voiced speech), the autocorrelation function will show a significant peak at a delay equal to the period (pitch period).

[0038] The height or significance of the autocorrelation function peak (such as the ratio of peak value to sidelobes) directly reflects the strength of the signal's periodicity, i.e., its harmonicity. The more pronounced the peak value, the closer the signal is to a pure periodic signal, and the clearer the harmonic structure.

[0039] Harmonicity Index The calculation method is as follows: First, perform the inverse transform on the amplitude spectrum in the frequency domain: .

[0040] Then calculate the harmonic index. .in, Represents the spectrum of complex numbers. Indicates frequency, Indicates the fundamental frequency. This indicates that the autocorrelation function has zero lag ( The value when =0), Represents a constant term. and Indicates the fundamental frequency range, max This indicates the autocorrelation function in the fundamental frequency range [ , Peak values ​​within the range of ].

[0041] Step A3: Calculate the energy change smoothness of each signal frame in the time direction, and scale the energy change smoothness to obtain the time continuity index; The energy of each signal frame here refers to the change in the total energy or Mel spectrum energy of that frame over time.

[0042] Smoothness can be measured by calculating the stability of the energy ratio between adjacent frames or by calculating the magnitude of the first / second difference of the short-time energy curve. The smoother the change, the higher the smoothness, and the larger the exponential value.

[0043] The energy changes of speech and steady-state noise are relatively gradual, while the energy of transient impulse noise (such as knocking sounds) changes drastically. This index can effectively capture this temporal characteristic of the scene.

[0044] Energy change smoothness The calculation method is as follows: ; in, This represents the amplitude spectrum corresponding to time t and frequency f. This represents the amplitude spectrum corresponding to time t-1 and frequency f. Represents a constant.

[0045] Time continuity index Calculated as follows: ; in, This is the scaling factor.

[0046] Step A4: For each signal frame, extract a set of spectral peak frequencies representing local maxima of energy from the amplitude spectrum; For each frame, a set of spectral peak frequencies characterizing local maxima of energy are found from the original amplitude spectrum (rather than the Mel spectrum).

[0047] Spectral peaks are prominent components in the frequency spectrum, typically corresponding to the fundamental frequency and harmonics of the sound source (vocal cord vibration), or strong resonance peaks. Extracting spectral peaks is the first step in extracting key frequency clues from a dense frequency spectrum.

[0048] Step A5: Obtain the preset fundamental frequency candidate frequency, and for the frequency band corresponding to the signal frame, calculate the similarity between the spectral peak frequency in the frequency band and the harmonic sequence with each fundamental frequency candidate frequency as the fundamental tone. Set a preset candidate baseband frequency (e.g., 80Hz-400Hz, corresponding to the baseband frequency of human speech) and generate a series of candidate baseband frequencies F0.

[0049] For each candidate fundamental frequency F0_i, generate a sequence of its integer multiples of frequency: F0_i, 2F0_i, 3F0_i, ... (i.e., the theoretical pure harmonic frequency positions).

[0050] Within a certain frequency band of the current signal frame, calculate the degree of matching (similarity) between the actual extracted spectral peak frequencies and this theoretical harmonic sequence. Similarity calculation can be achieved by: counting the number of spectral peaks falling in the vicinity of the theoretical harmonic frequency, or calculating the percentage difference between the weighted sums of spectral peak energy or theoretical position weights.

[0051] This step is essentially a fundamental frequency detection and harmonic matching process. It assesses the extent to which the spectral structure of the current frequency band conforms to a harmonic model generated by a periodic sound source (i.e., speech).

[0052] Step A6: Take the maximum value among multiple similarities as the frequency domain correlation index of the frequency band in the signal frame.

[0053] The maximum similarity among multiple candidate fundamental frequencies is used as the frequency domain relevance index for that frequency band in the current frame. This maximum value reflects the degree of agreement between the spectral structure and the optimally matched harmonic sequence within the current frequency band. A higher value indicates that the energy of that frequency band is more likely to originate from a single, periodic sound source (i.e., the target speech).

[0054] Correlation does not refer to statistical correlation between frequency bands, but rather to the correlation between frequency domain spectral peaks and an ideal harmonic template. It measures the degree of harmonic structuring of the spectrum, making it an extremely powerful indicator of speech presence. The spectral peak distribution of noise (especially non-steady-state noise) is typically random or non-harmonic, thus this index effectively distinguishes speech from many types of noise.

[0055] The index can be calculated separately in the low-frequency band, mid-frequency band, and high-frequency band, thereby obtaining more refined, frequency-band-related scene analysis capabilities.

[0056] In the embodiments corresponding to steps A1 to A6, rich information is provided for the speech enhancement algorithm by extracting various acoustic features (such as spectral centroid, zero-crossing rate, and energy concentration), calculating harmonic and time continuity indices, extracting spectral peak frequencies, and calculating frequency domain correlation indices. These features help improve the clarity and intelligibility of speech signals, and also provide important basis for subsequent signal processing and enhancement.

[0057] Step 102: Based on the auditory scene cues, modulate the acoustic features to obtain modulated acoustic embedding features; This is not a simple feature concatenation, but a conditional processing or attention mechanism. When the temporal continuity index is high, the weight of the harmonicity index is strengthened, because continuous harmonic signals are likely speech or steady-state noise. When the frequency domain correlation index shows a specific pattern, the feature values ​​of the spectral centroid and energy concentration are adjusted to better distinguish between relevant speech bands and irrelevant noise bands. This makes traditional, static acoustic features adaptive. Based on cues from the current auditory scene, the acoustic feature dimensions that are most discriminative in the current scene are highlighted, while dimensions that may cause confusion are suppressed. This is equivalent to performing feature optimization using prior knowledge of human hearing before the data enters the neural network.

[0058] Specifically, step 102 includes steps 1021 to 1025: Step 1021: Based on the auditory scene cues, calculate the probability that each signal frame belongs to multiple predefined auditory component categories, and use it as the prior probability of auditory attention; wherein, the multiple predefined auditory component categories include foreground speech components, stable background noise components, and sudden interference components; Based on auditory scene cues, the probability of each signal frame (which may be refined to each Mel band (t, m)) belonging to one of the three categories—foreground speech, stable background noise, and sudden interference—is calculated as the prior probability of auditory attention.

[0059] The complex auditory scene is decomposed into three auditory components with distinct acoustic characteristics. The categories of auditory components include, but are not limited to, foreground speech components, stable background noise components, and sudden interference components.

[0060] Foreground speech components refer to the target speaker's voice, which typically exhibits harmonicity and a degree of continuity. Stable background noise components include sounds such as fan noises and background human voices, which have relatively stable energy distribution. Sudden interference components include keyboard sounds, door slamming sounds, and instantaneous collision sounds, which are characterized by abrupt changes in time.

[0061] Based on the time continuity index and the frequency domain correlation index, three probability values ​​are output, and their sum is 1 (or normalized by Softmax). This forms a pixel-level (time-frequency point level) soft segmentation map of the auditory scene.

[0062] Specifically, step 1021 includes steps B1 to B5: Step B1: Multiply the temporal continuity index and the frequency domain correlation index to obtain the initial value corresponding to the foreground speech component; Initial value of the prospect = time continuity index × frequency domain correlation index.

[0063] Pure, continuous foreground speech typically possesses both high temporal continuity (energy changes smoothly over time) and high-frequency domain correlation (spectral structure exhibits harmonic characteristics).

[0064] Using multiplication implies that two conditions must be met simultaneously. A low exponent (such as poor temporal continuity of sudden interference or low frequency domain correlation of non-harmonic noise) will lead to a decrease in the initial value of the foreground speech. This aligns with the requirement that speech should possess structural acoustic priors in both the time and frequency domains.

[0065] Step B2: Subtract 1 from the frequency domain correlation index and multiply it by the time continuity index to obtain the initial value corresponding to the stable background noise component; Initial value of stable background noise = (1 - frequency domain correlation index) × time continuity index.

[0066] Stable background noise (such as fan or air conditioner noise) is characterized by low-frequency domain correlation (usually without a clear harmonic structure) and high temporal continuity (energy remains stable over a long period of time).

[0067] The (1-frequency domain correlation index) measures harmonicity. Multiplying this value by the time continuity index means looking for signal components that are stable in time but lack harmonic structure, which is characteristic of many steady-state noises.

[0068] Step B3: Subtract 1 from the frequency domain correlation index and multiply it by the peak energy ratio of the signal frame to obtain the initial value corresponding to the burst interference component; Initial value of sudden interference = (1 - frequency domain correlation index) × energy peak ratio of signal frame.

[0069] The peak energy ratio of a signal frame refers to the ratio of the instantaneous energy of that frame to the recent average energy, or the ratio of the maximum sample value within the frame to the RMS energy. A high ratio indicates a sudden burst of energy.

[0070] Sudden disturbances (such as collision sounds or knocking sounds) are characterized by low-frequency domain correlation (usually non-harmonic broadband impulses) and high energy peak ratio (energy increases sharply in a short period of time).

[0071] This formula seeks out non-harmonic components with abrupt energy changes. (1 - frequency domain correlation index) filters out speech, while the peak energy ratio captures its transient characteristics.

[0072] Step B4: Perform softmax normalization on the initial values ​​corresponding to the foreground speech component, the stable background noise component, and the sudden interference component to obtain the prior probability of auditory attention belonging to each category of the signal frame. The three initial values ​​are input into the Softmax function for normalization to obtain the probability that the signal frame belongs to one of the three categories.

[0073] The Softmax function ensures that the sum of the three probability values ​​is 1, which conforms to the definition of probability. It transforms three initial confidence levels into an interpretable, mutually exclusive probability distribution.

[0074] For each signal frame t, three scalar probability values ​​are obtained: P_fore(t), P_bg(t), and P_burst(t). At this point, frequency bands have not yet been distinguished, and the probabilities are global probabilities at the frame level.

[0075] Step B5: Map the auditory attention prior probabilities of all signal frames through a Mel filter bank and downsample them to the same time-frequency resolution as the Mel spectrum to obtain the auditory attention prior probabilities.

[0076] The auditory attention prior probabilities of all signal frames are mapped through a Mel filter bank and downsampled to the same time-frequency resolution as the Mel spectrum.

[0077] Steps B1 to B4 calculate the probability ((T,1) dimension) of each signal frame belonging to each category. However, subsequent modulation requires the probability ((T,M) dimension, with the same resolution as the Mel spectrum) for each Mel band.

[0078] Mel filter bank mapping is the core of resolution conversion. Mel filter bank mapping copies the frame-level probability vector [P_fore(t), P_bg(t), P_burst(t)] of each time frame M times, forming a coarse (T, M) probability map.

[0079] Considering the characteristics of the Mel filter bank (each Mel band covers a continuous linear frequency range), the frame-level probabilities can be weighted and averaged or interpolated based on the mapping relationship between the original STFT bands and the Mel bands to generate a smoother Mel-scale probability map. This is equivalent to using the Mel filter bank as a smoother or mapper.

[0080] In the embodiments corresponding to steps B1 to B5, the probability of each signal frame belonging to a different auditory component category can be effectively calculated based on auditory scene cues. The logic of these steps relies on in-depth analysis of time and frequency domain features, helping the system identify foreground speech, stable background noise, and sudden interference. This process not only considers the physical characteristics of audio signals but also incorporates an understanding of human hearing, enabling the final auditory attention prior probability to better reflect the human ear's attention allocation to different sound components.

[0081] Step 1022: Project the acoustic features onto a dimension with the same number of channels as the Mel spectrum through a linear transformation to obtain the initial acoustic embedding features; Acoustic features (such as scalars or low-dimensional vectors like spectral centroids and harmonic exponents) are projected onto a dimension with the same number of Mel spectral channels through a linear transformation (fully connected layer).

[0082] Each time frame of the Mel spectrum has M Mel bands (channels). In order for acoustic features to be stitched together with spectral features along the channel dimension, they must be extended to the same spatial dimension (T,M).

[0083] The linear transformation diffuses the global acoustic properties of each time frame (such as the overall strong harmonicity of this frame) across all M frequency bands, forming a two-dimensional initial acoustic embedding feature. Each frequency band contains global acoustic property information.

[0084] Step 1023: Calculate the dynamic modulation weights based on the prior probability of auditory attention; wherein, the dynamic modulation weights = , This represents the prior probability of auditory attention corresponding to the foreground speech component. This represents the prior probability of auditory attention corresponding to a stable background noise component. This represents the prior probability of auditory attention corresponding to the sudden interference component. , and Represents a learnable scalar parameter; A learnable linear combination formula is used to fuse the prior probabilities of the three auditory components into a dynamic modulation weight.

[0085] Dynamic modulation weight = .

[0086] This serves as a positive stimulus. When the probability of foreground speech is high, the modulation weight is increased to enhance the representational power of the acoustic features at that time-frequency point, since this is likely the speech that needs to be preserved and enhanced.

[0087] This is for negative suppression. When the probability of background noise or sudden interference is high, the modulation weight is reduced to suppress the acoustic characteristics of these regions, as the acoustic properties of these regions may be misleading (e.g., sudden interference may also have high harmonic instants).

[0088] For global bias. A learnable baseline value to ensure the modulation weights have a reasonable initial range.

[0089] Learnable parameters , and It is not fixed, but rather automatically learned from the data during model training. This allows the modulation strategy to adapt to specific datasets, headphone hardware, or user preferences, thereby finding the optimal enhancement-inhibition balance. For example, , and It can be set to 1.42, 0.53, and 0.2.

[0090] This weighted map can be understood as a confidence map of acoustic features based on the auditory scene. Regions with high weights indicate that the acoustic features at that location are highly reliable and have a positive effect on speech enhancement; regions with low weights indicate that the acoustic features at that location may be contaminated by noise and need to be weakened.

[0091] Step 1024: The dynamic modulation weights are constrained by the Sigmoid function to obtain the modulation coefficients; The dynamic modulation weights are constrained using the Sigmoid function to obtain the modulation coefficients. The Sigmoid function compresses the input values ​​to the (0,1) interval. This ensures that the modulation coefficients are a gain value between 0 and 1, preventing excessively large or small modulation weights from causing training instability or excessive feature distortion.

[0092] The coefficient can be viewed as a soft switch or attention gate. A value close to 1 indicates that the acoustic feature is completely allowed to pass through, while a value close to 0 indicates that the feature is almost completely blocked.

[0093] Step 1025: Multiply the initial acoustic embedding features element-wise with the modulation coefficients to obtain the modulated acoustic embedding features.

[0094] At each time-frequency point (t, m), the value of the initial acoustic embedding feature is scaled using the corresponding modulation coefficient.

[0095] Modulated acoustic embedding features are generated. These features are no longer the original, global acoustic properties, but rather scene-aware weighted features. In speech-dominant regions, the acoustic features are amplified; in noise-dominant regions, the acoustic features are attenuated. This is equivalent to a soft masking or attention focusing at the feature level.

[0096] In the embodiments corresponding to steps 1021 to 1025, acoustic features are modulated based on auditory scene cues to obtain modulated acoustic embedding features. This process not only emphasizes the importance of foreground speech but also ensures the flexibility and stability of the processing through dynamic modulation weights and sigmoid constraints, providing an effective feature foundation for the final speech enhancement.

[0097] Step 103: Concatenate the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The modulated acoustic embedding features are concatenated with the original spectral features along the channel dimension. In deep learning, this is equivalent to adding a new set of feature descriptors to each time-frequency (TF-bin). The spectral features provide rich low-level signal information, while the modulated acoustic features provide physically meaningful statistical information refined from high-level scene cues. The network simultaneously obtains fine-grained spectral details and coarse-grained, modulated acoustic properties. This combination helps the network learn the nonlinear mapping relationship between speech and noise under complex conditions more quickly and accurately, potentially improving the model's convergence speed and final performance.

[0098] Step 104: Input the joint feature representation into the speech enhancement network to obtain enhanced speech data.

[0099] Operation: Input the joint feature representation into a speech enhancement network (such as DNN, CNN, LSTM or a combination thereof), and the network outputs enhanced speech data (usually a time-frequency mask or clean speech spectrum).

[0100] Technical implications: This step is the execution phase of the method. The core task of the network is to learn and separate the target speech from the joint feature representation.

[0101] Since the input features have already been preprocessed and modulated, the network does not need to learn from scratch how to utilize basic acoustic properties and scene cues, and can focus more on learning complex separation functions. This is expected to reduce the requirements for network size and training data volume, and improve the generalization ability (i.e., reliability) in unknown noise scenarios.

[0102] Specifically, step 104 includes steps 1041 to 1048: Step 1041: Represent the joint feature as the first modality feature; The joint feature representation is directly used as the first mode feature. The first mode feature carries high temporal resolution original information, including spectral details and modulated acoustic features. It is the most comprehensive and detailed representation of the signal.

[0103] Step 1042: Downsample the joint feature representation in the time dimension, and downsample the auditory attention prior probabilities corresponding to multiple predefined auditory component categories in multiple signal frames in the time dimension; Step 1043: Concatenate the downsampled joint feature representation with the downsampled auditory attention prior probability to obtain the second modality feature; The joint feature representation is downsampled along the temporal dimension (e.g., using convolution or pooling with a stride > 1). The auditory attention prior probability is also downsampled along the temporal dimension. The downsampled features are then concatenated to obtain the second modality feature.

[0104] The previously calculated soft segmentation map of the auditory scene (i.e., which time-frequency points might be speech, noise, or interference) is input as explicit guiding information into the deeper layers of the network. This is equivalent to providing the network with high-level, rule-based scene understanding as side information. Therefore, the second modality features are feature representations with low temporal resolution but rich in macro-context and explicit scene prior knowledge.

[0105] The speech enhancement network employs a two-stream encoder network architecture. Each modality passes through an independent encoder stack, and each stack contains L identical coding layers. Each coding layer contains an intra-modal attention module and a cross-modal attention module.

[0106] For the first modality of features, standard multi-head self-attention (MSA) or convolutional attention modules (CBAM) are used. For the second modality of features, due to their shorter sequences, recurrent neural networks (such as GRU / LSTM) or long sequence Transformers are more suitable for capturing long-term dependencies.

[0107] Step 1044: The speech enhancement network performs intramodal encoding on the first modality features to obtain the first intermediate features; Step 1045: The speech enhancement network performs intra-modal encoding on the second modality features to obtain the second intermediate features; The speech enhancement network encodes the first modality features within the first modality to obtain the first intermediate features.

[0108] The speech enhancement network encodes the second modality features within the second modality to obtain the second intermediate features.

[0109] Intramodal encoding consists of several layers of convolutional, recurrent neural network, or Transformer encoder layers. Its purpose is to refine and abstract features within their respective modalities.

[0110] The first modal encoder focuses on learning local spectral patterns and detailed features (such as phoneme boundaries and subtle harmonic structures) at high resolution.

[0111] The second modal encoder focuses on learning the global scene structure and long-term dependencies at low resolution (such as the overall distribution of speech segments / silence segments and the macroscopic type of noise).

[0112] Step 1046: The speech enhancement network uses the first intermediate feature as the query vector and the second intermediate feature as the key vector and value vector to calculate the first cross-modal attention feature; Attention is calculated using the first intermediate feature (high-resolution details) as the query vector and the second intermediate feature (low-resolution scene) as the key and value vectors.

[0113] This can be understood as querying the scene for details. High-resolution detailed features actively query and retrieve relevant global contextual information from low-resolution scene features. For example, a vague speech fragment (details) can query the scene prior: "Based on long-term scene analysis, is it likely that this time period and frequency band I am in is speech? What are the overall characteristics of the surrounding noise?" Thus, scene knowledge is used to clarify and correct the details.

[0114] Specifically, step 1046 includes steps C1 to C4: Step C1: Project the first intermediate feature through the first query linear transformation layer to obtain the first query vector; The first intermediate feature (high-resolution detail feature) is projected through a first query linear transformation layer (i.e., a fully connected layer or a 1x1 convolution).

[0115] The purpose of this linear transformation is to learn how to formulate effective questions from detailed features. It maps detailed features to a new feature space where the feature representations are optimized for matching and querying against scene features.

[0116] Step C2: Project the second intermediate feature through the first key linear transformation layer and the first value linear transformation layer respectively to obtain the first key vector and the first value vector; wherein, the first query vector, the first key vector and the first value vector have the same feature dimension; The second intermediate feature (low-resolution scene feature) is projected through the first key linear transformation layer and the first value linear transformation layer, respectively.

[0117] The key transformation layer learns how to represent scene features using indexes or keywords in order to match queries.

[0118] The value transformation layer learns how to represent the information content carried by scene features, which will be extracted and aggregated after matching.

[0119] The projected Q1, K1, and V1 have the same feature dimension. This is a prerequisite for the attention mechanism to function correctly, ensuring that the dot product operation is mathematically valid.

[0120] Step C3: Calculate the attention weight of the dot product between the first query vector and the first key vector; Calculate the attention weight of the dot product between the first query vector (Q1) and the first key vector (K1).

[0121] Attention weights = Softmax((Q1 * K1^T) / sqrt(d_k)), where d_k is the dimension of the key vector and sqrt(d_k) is the scaling factor.

[0122] The dot product (Q1 * K1^T) calculates the similarity between the query vector and all key vectors. The higher the dot product value, the more relevant the query is to the context information represented by the key.

[0123] Softmax normalization transforms the similarity scores into a probability distribution, ensuring that the sum of all weights equals 1. This guarantees that subsequent weighted summation focuses on the most relevant components and produces a stable output.

[0124] The generated attention weight matrix is ​​an alignment matrix or relevance map. Each row corresponds to a high-resolution detail location (query), and each column corresponds to a low-resolution scene location (key). The weight values ​​indicate how much information each detail location should draw from each scene location.

[0125] Step C4: Based on the dot product attention weights, perform a weighted summation on the first value vector to obtain the first cross-modal attention feature.

[0126] Based on the calculated dot product attention weights, the first value vector (V1) is weighted and summed to obtain the first cross-modal attention feature. First cross-modal attention feature = attention weight * V1.

[0127] For each high-resolution detail query location, information is extracted and fused from all scene values ​​based on its similarity to all scene keys (attention weight). Scene locations with high similarity contribute more to the final output in terms of their value vectors.

[0128] The first cross-modal attention feature is the result of modulating the first intermediate feature (details) with information from the second intermediate feature (scene). It preserves the original high-resolution detail structure, but the details at each location are integrated with the most relevant contextual information from the global scene prior. For example, a speech point contaminated by noise can be enhanced and cleaned by focusing on information from the scene prior that this location is a high-probability speech area.

[0129] In the embodiments corresponding to steps C1 to C4, a cross-modal attention mechanism is implemented through linear transformation and dot product operations, enabling the speech enhancement network to effectively integrate information from different modalities, thereby improving the quality of the enhanced speech signal. In this way, the speech enhancement network can better focus on features relevant to the current task, thus improving performance.

[0130] Step 1047: The speech enhancement network uses the second intermediate feature as the query vector and the first intermediate feature as the key vector and value vector to calculate the second cross-modal attention feature; Attention is calculated using the second intermediate feature (low-resolution scene) as the query vector and the first intermediate feature (high-resolution detail) as the key and value vectors.

[0131] This can be understood as scene focusing on details. Low-resolution scene features actively seek out and focus on local evidence in high-resolution detail features that is crucial for scene judgment. For example, scene priors may initially identify an area as noise, but by focusing on details, weak harmonic structures are discovered, thereby refining and correcting the scene judgment.

[0132] This bidirectional, symmetrical attention mechanism enables bidirectional, iterative information fusion between the detail stream and the scene stream. It is not a simple splicing or addition, but rather allows information from two different perspectives and resolutions to engage in deep, targeted dialogue and complementarity.

[0133] Step 1048: The speech enhancement network outputs the enhanced speech data based on the first cross-modal attention features and the second cross-modal attention features.

[0134] The network decoder (including upsampling layers, convolutional layers, etc.) needs to fuse these two fully interactive and complementary attention features (e.g., re-concatenate or add them) and finally map them back to the time-frequency mask or clean speech spectrum, and then obtain the enhanced time-domain speech signal through inverse transformation.

[0135] In the embodiments corresponding to steps 1041 to 1048, by fusing and encoding features from different modalities and utilizing a cross-modal attention mechanism, the aim is to improve the performance of the speech enhancement network, thereby achieving more effective speech data enhancement.

[0136] Specifically, step 1048 includes steps D1 to D5: Step D1: Project the first cross-modal attention feature through the first output linear transformation layer, then perform a weighted sum with the first intermediate feature, and then perform layer normalization to obtain the first updated feature; wherein, the weights of the weighted sum are learnable parameters; First Update Feature = LayerNorm(α1*First Intermediate Feature + β1*(First Output Linear Transformation Layer(First Cross-Modal Attention Feature))).

[0137] Linear transformation projects cross-modal attention features into the same space as the original intermediate features, potentially performing dimensionality adjustments or feature refinement.

[0138] Weighted summation is a learnable residual connection. α1 and β1 are learnable scalar or vector parameters, allowing the network to automatically learn how much original detail to retain and how much modulation information from the scene to incorporate. This is more flexible than a residual connection that is fixed at 1.

[0139] This step adaptively fuses the contextual information extracted from the scene with the original detail features to generate an updated, informed version of the detail features.

[0140] Step D2: Project the second cross-modal attention feature through the second output linear transformation layer, then perform a weighted sum with the second intermediate feature, and then perform layer normalization to obtain the second updated feature; wherein, the weights of the weighted sum are learnable parameters; Symmetrically, the same operation is performed on the second modality. Evidence information focused from details is adaptively fused with the original scene prior features to generate a refined, updated version of the scene features.

[0141] Step D3: Input the first updated feature and the second updated feature as the first modal feature and the second modal feature of the next fusion layer, respectively; Instead of performing a single cross-modal interaction, the network iterates multiple times through L identical fusion layers. At each layer, features from both modalities are updated based on the output of the previous layer, followed by another bidirectional cross-modal attention interaction. This allows for multi-round, deep dialogue and collaboration between details and scene information, resulting in more comprehensive and accurate information fusion. L is a hyperparameter, which can be 2, 3, 4, etc., representing the network's depth and fusion capability.

[0142] Step D4: Perform feedforward neural network processing and residual connection on the first updated feature and the second updated feature respectively, and generate a deep speech feature representation based on the final first updated feature and the final second updated feature after processing through L fusion layers; Within each fusion layer, the updated first and second updated features are processed by a feedforward neural network (a small fully connected network), and residual connections are also used.

[0143] Feedforward networks introduce additional nonlinearity and transformability into the features of each modality, enhancing the model's expressive power. Following cross-modal interactions, each modality is allowed to undergo further processing and abstraction within its own feature space.

[0144] The cross-modal attention -> weighted summation / layer normalization -> feedforward network / residual here constitutes the structure of a Transformer encoder layer, but is applied to a dual-modal cross-attention scenario.

[0145] After processing through L such fusion layers, the final first updated feature and the final second updated feature are combined in some form (such as concatenation, addition, or taking one) to generate a deep speech feature representation.

[0146] This deep feature representation is the essence of L rounds of detail-scene deep interaction and fusion. It contains high-resolution speech details and fully integrates global auditory scene understanding, and theoretically has extremely strong robustness to noise.

[0147] Step D5: Based on the decoder, the deep speech feature representation is converted into the enhanced speech data corresponding to the original time-frequency domain.

[0148] The decoder is a symmetric neural network structure (such as transposed convolution, upsampling layers, etc.) that is responsible for upsampling the fused deep features to the original time-frequency resolution (i.e., the size of the Mel spectrum or STFT spectrum).

[0149] The decoder output is a time-frequency mask (such as an ideal ratio mask, IRM) or an enhanced complex spectrum. This is combined with the phase (or estimated phase) of the original noisy speech, and the enhanced time-domain speech signal is obtained through an inverse short-time Fourier transform.

[0150] In the embodiments corresponding to steps D1 to D5, the speech enhancement network can effectively utilize cross-modal attention features to generate high-quality enhanced speech data, thereby improving the user's listening experience. Each step aims to improve the expressive power of the features and the performance of the model, ensuring that the final output speech data has better clarity and intelligibility.

[0151] In the embodiments corresponding to steps 101 to 104, by extracting spectral and acoustic features from the audio signal to be enhanced acquired by the microphone, the speech signal and environmental noise can be effectively separated. Especially in noisy environments, the modulated acoustic embedding features can significantly improve speech intelligibility. This method combines auditory scene cues, such as the temporal continuity and frequency domain correlation of signal frames, to better adapt to dynamically changing environments. By modulating the acoustic features based on these cues, the enhanced speech signal maintains better stability and clarity in variable environments, improving signal robustness. Concatenating the modulated acoustic embedding features and spectral features along the channel dimension, the resulting joint feature representation can comprehensively consider information from different features, thus providing richer input for subsequent speech enhancement networks. This effective feature fusion method helps deep learning models better learn and recognize complex patterns in speech signals, improving enhancement results.

[0152] like Figure 2 This invention provides a voice enhancement system for Bluetooth headsets; please refer to [link / reference]. Figure 2 , Figure 2 A schematic diagram of a voice enhancement system for a Bluetooth headset provided by the present invention is shown, such as... Figure 2 The voice enhancement system of a Bluetooth headset shown includes: The acquisition unit 21 is used to acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; Modulation unit 22 is used to modulate the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features; The splicing unit 23 is used to splice the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; Processing unit 24 is used to input the joint feature representation into the speech enhancement network to obtain enhanced speech data.

[0153] This invention provides a voice enhancement system for Bluetooth headsets. By extracting spectral and acoustic features from the audio signal to be enhanced, acquired by a microphone, it effectively separates the speech signal from environmental noise. Especially in noisy environments, the modulated acoustic embedding features significantly improve speech intelligibility. This method incorporates auditory scene cues, such as the temporal continuity and frequency domain correlation of signal frames, enabling better adaptation to dynamically changing environments. By modulating acoustic features based on these cues, the enhanced speech signal maintains better stability and clarity in variable environments, improving signal robustness. Concatenating the modulated acoustic embedding features and spectral features along the channel dimension generates a joint feature representation that comprehensively considers information from different features, thus providing richer input for subsequent speech enhancement networks. This effective feature fusion method helps deep learning models better learn and recognize complex patterns in speech signals, improving enhancement performance.

[0154] Figure 3 This is a schematic diagram of a terminal device provided in an embodiment of the present invention. Figure 3 As shown, a terminal device 3 in this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a voice enhancement program for a Bluetooth headset. When the processor 30 executes the computer program 32, it implements the steps in the various embodiments of the voice enhancement method for a Bluetooth headset described above, for example... Figure 1 Steps 101 to 104 are shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each unit in the above-described device embodiments, for example... Figure 2 The function of the unit shown.

[0155] For example, the computer program 32 can be divided into one or more units, which are stored in the memory 31 and executed by the processor 30 to complete the present invention. The one or more units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 32 in the terminal device 3. For example, the specific functions of each unit of the computer program 32 can be divided as follows: The acquisition unit is used to acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; A modulation unit is used to modulate the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features; The splicing unit is used to splice the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The processing unit is used to input the joint feature representation into the speech enhancement network to obtain enhanced speech data.

[0156] The terminal device includes, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of a terminal device 3 and does not constitute a limitation on a terminal device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0157] The processor 30 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0158] The memory 31 can be an internal storage unit of the terminal device 3, such as a hard disk or memory of the terminal device 3. The memory 31 can also be an external storage device of the terminal device 3, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device 3. Furthermore, the memory 31 can include both internal and external storage units of the terminal device 3. The memory 31 is used to store the computer program and other programs and data required by the roaming control device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0159] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0160] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0161] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0162] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0163] This invention provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.

[0164] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0165] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0166] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0167] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0168] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units.

[0169] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0170] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0171] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."

[0172] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0173] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0174] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for enhancing voice in a Bluetooth headset, characterized in that, The voice enhancement method for the Bluetooth headset includes: Acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features, and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index, and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; Based on the auditory scene cues, the acoustic features are modulated to obtain modulated acoustic embedding features; The modulated acoustic embedding features and the spectral features are concatenated along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The joint feature representation is input into the speech enhancement network to obtain enhanced speech data.

2. The voice enhancement method for Bluetooth headsets as described in claim 1, characterized in that, The process involves acquiring the audio signal to be enhanced, collected by a microphone, and extracting the spectral features, acoustic features, and [other characteristics] corresponding to the audio signal to be enhanced. The steps of auditory scene cues include: Acquire the audio signal to be enhanced, captured by the microphone; The audio signal to be enhanced is subjected to frame segmentation, windowing, and short-time Fourier transform to obtain a complex spectrum; Calculate the amplitude spectrum corresponding to the complex spectrum; The amplitude spectrum is converted into a Mel spectrum and used as the spectral feature; Based on the Mel spectrum, acoustic features and auditory scene cues are extracted.

3. The voice enhancement method for Bluetooth headsets as described in claim 2, characterized in that, The steps for extracting acoustic features and auditory scene cues based on the Mel spectrum include: Extract the spectral centroid and spectral energy concentration from the Mel spectrum, and extract the zero-crossing rate of the signal to be enhanced; Calculate the peak value of the autocorrelation function between signal frames in the audio signal to be enhanced, and use it as the harmonicity index; Calculate the energy change smoothness of each signal frame in the time direction, and scale the energy change smoothness to obtain the time continuity index; For each signal frame, a set of spectral peak frequencies representing local maxima of energy are extracted from the amplitude spectrum; Obtain a preset candidate fundamental frequency, and for the frequency band corresponding to the signal frame, calculate the similarity between the spectral peak frequency in the frequency band and the harmonic sequence with each candidate fundamental frequency as the fundamental tone. The maximum value among multiple similarities is used as the frequency domain correlation index of that frequency band in the signal frame.

4. The voice enhancement method for Bluetooth headsets as described in claim 1, characterized in that, The step of modulating the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features includes: Based on the auditory scene cues, the probability of each signal frame belonging to multiple predefined auditory component categories is calculated and used as the prior probability of auditory attention; wherein, the multiple predefined auditory component categories include foreground speech components, stable background noise components, and sudden interference components; The acoustic features are linearly transformed and projected onto a dimension with the same number of channels as the Mel spectrum to obtain the initial acoustic embedding features; Based on the prior probability of auditory attention, the dynamic modulation weight is calculated; wherein, the dynamic modulation weight = , This represents the prior probability of auditory attention corresponding to the foreground speech component. This represents the prior probability of auditory attention corresponding to a stable background noise component. This represents the prior probability of auditory attention corresponding to the sudden interference component. , and Represents a learnable scalar parameter; The dynamic modulation weights are constrained by the Sigmoid function to obtain the modulation coefficients; The initial acoustic embedding features are multiplied element-wise with the modulation coefficients to obtain the modulated acoustic embedding features.

5. The voice enhancement method for Bluetooth headsets as described in claim 4, characterized in that, The step of calculating the probability that each signal frame belongs to multiple predefined auditory component categories based on the auditory scene cues, and using this probability as the prior probability of auditory attention, includes: Multiply the time continuity index and the frequency domain correlation index to obtain the initial value corresponding to the foreground speech component; Subtracting 1 from the frequency domain correlation index and multiplying it by the time continuity index yields the initial value corresponding to the stable background noise component. Subtract 1 from the frequency domain correlation index and multiply it by the peak energy ratio of the signal frame to obtain the initial value corresponding to the burst interference component; The initial values ​​corresponding to the foreground speech component, the stable background noise component, and the sudden interference component are normalized by softmax to obtain the prior probability of auditory attention belonging to each category of the signal frame. The auditory attention prior probabilities of all signal frames are downsampled to the same time-frequency resolution as the Mel spectrum by mapping the Mel filter bank.

6. The voice enhancement method for Bluetooth headsets as described in claim 1, characterized in that, The step of inputting the joint feature representation into the speech enhancement network to obtain enhanced speech data includes: The joint feature is represented as the first modality feature; The joint feature representation is downsampled in the time dimension, and the auditory attention prior probabilities corresponding to multiple predefined auditory component categories in multiple signal frames are downsampled in the time dimension. The downsampled joint feature representation is concatenated with the downsampled auditory attention prior probability to obtain the second modality feature; The speech enhancement network performs intramodal encoding on the first modal features to obtain the first intermediate features; The speech enhancement network performs intramodal encoding on the second modality features to obtain the second intermediate features; The speech enhancement network uses the first intermediate feature as the query vector and the second intermediate feature as the key vector and value vector to calculate the first cross-modal attention feature; The speech enhancement network uses the second intermediate feature as the query vector and the first intermediate feature as the key vector and value vector to calculate the second cross-modal attention feature. The speech enhancement network outputs the enhanced speech data based on the first cross-modal attention features and the second cross-modal attention features.

7. The voice enhancement method for Bluetooth headsets as described in claim 6, characterized in that, The step of calculating the first cross-modal attention feature using the first intermediate feature as the query vector and the second intermediate feature as the key vector and value vector includes: The first intermediate feature is projected through the first query linear transformation layer to obtain the first query vector; The second intermediate feature is projected through the first key linear transformation layer and the first value linear transformation layer respectively to obtain the first key vector and the first value vector; wherein the first query vector, the first key vector and the first value vector have the same feature dimension. Calculate the attention weight of the dot product between the first query vector and the first key vector; The first cross-modal attention feature is obtained by weighting and summing the first value vector based on the dot product attention weights.

8. The voice enhancement method for Bluetooth headsets as described in claim 6, characterized in that, The steps of the speech enhancement network outputting the enhanced speech data based on the first cross-modal attention features and the second cross-modal attention features include: The first cross-modal attention feature is projected through the first output linear transformation layer, then weighted and summed with the first intermediate feature, and then normalized by the layer to obtain the first updated feature; wherein the weights of the weighted summation are learnable parameters. The second cross-modal attention feature is projected through the second output linear transformation layer, then weighted and summed with the second intermediate feature, and then normalized by the layer to obtain the second updated feature; wherein, the weights of the weighted summation are learnable parameters; The first updated feature and the second updated feature are respectively used as the first modal feature and the second modal feature input to the next fusion layer; The first updated feature and the second updated feature are processed by a feedforward neural network and residual connections respectively. Based on the final first updated feature and the final second updated feature after processing by L fusion layers, a deep speech feature representation is generated. The decoder converts the deep speech feature representation into the original time-frequency domain corresponding to the enhanced speech data.

9. A voice enhancement system for a Bluetooth headset, characterized in that, The voice enhancement system of the Bluetooth headset includes: The acquisition unit is used to acquire the audio signal to be enhanced collected by the microphone, and extract the spectral features, acoustic features and auditory scene cues corresponding to the audio signal to be enhanced; wherein, the acoustic features include the spectral centroid, spectral energy concentration, harmonicity index and / or zero-crossing rate, and the auditory scene cues include the temporal continuity index of the signal frame and the frequency domain correlation index of the signal frame; A modulation unit is used to modulate the acoustic features based on the auditory scene cues to obtain modulated acoustic embedding features; The splicing unit is used to splice the modulated acoustic embedding features and the spectral features along the channel dimension to generate a joint feature representation for subsequent speech enhancement networks; The processing unit is used to input the joint feature representation into the speech enhancement network to obtain enhanced speech data.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the voice enhancement method for Bluetooth headsets as described in any one of claims 1 to 8.