Speech enhancement method and device based on full-band and sub-band fusion, terminal and medium

By decomposing and fusing the full-band and sub-band spectra of the speech signal, information interaction and feature fusion between the full-band and sub-band are realized, improving the speech enhancement effect and solving the problems of insufficient information interaction and inadequate feature fusion in the existing technology.

CN121054024BActive Publication Date: 2026-03-31GUANGZHOU UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing full-band and sub-band fusion methods suffer from insufficient information exchange and inadequate feature fusion, resulting in poor speech enhancement effects, especially when dealing with complex speech signals, making it difficult to preserve details and overall structure.

Method used

By decomposing the full-band spectrum of the noisy speech signal into low-frequency, mid-frequency, and high-frequency sub-bands, enhancing them separately, and then achieving information interaction and fusion through a full-band enhancement module and a spectrum fusion module, a full-band-sub-band fused spectrum is generated, and the target speech enhancement spectrum is finally determined.

Benefits of technology

It effectively solves the problems of insufficient information interaction and inadequate feature fusion in the process of full-band and sub-band fusion, and improves the effect of speech enhancement, especially in the process of complex speech signal processing, preserving details and overall structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121054024B_ABST
    Figure CN121054024B_ABST
Patent Text Reader

Abstract

The application discloses a speech enhancement method and device based on full-band and sub-band fusion, a terminal and a medium. The method determines an initial full-band spectrum of a noisy speech signal; decomposes an amplitude spectrum corresponding to the initial full-band spectrum of the noisy speech into three sub-bands of a low frequency band, a middle frequency band and a high frequency band through a sub-band enhancement module, performs speech enhancement based on each sub-band to determine a preliminary sub-band enhanced amplitude spectrum; determines a preliminary full-band enhanced spectrum from the initial full-band spectrum of the noisy speech through a full-band enhancement module; fuses the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum through a spectrum fusion module to determine a full-band-sub-band fusion spectrum; determines a fusion complex mask based on the full-band-sub-band fusion spectrum; and determines a target speech enhanced spectrum from the fusion complex mask and the initial full-band spectrum of the noisy speech. The method effectively solves the problems of insufficient information interaction and insufficient feature fusion in the full-band and sub-band fusion process in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more particularly to a speech enhancement method, apparatus, terminal, and medium based on full-band and sub-band fusion. Background Technology

[0002] Full-band speech signal processing methods model the contextual relationships of speech signals, which can better handle the temporal information in speech, comprehensively capture the global features of speech signals, and improve the ability to understand and model speech signals. However, full-band methods ignore the time-frequency distribution information of the speech spectrum and do not consider the differences in local patterns in different frequency bands. When dealing with speech signals with special frequency structures or prominent local frequency features, such as mixed music and speech signals or speech with high-frequency harmonic components, full-band methods struggle to fully capture the local structure and frequency components, which adversely affects the noise reduction effect.

[0003] Subband-based speech signal processing methods divide the spectrum into multiple subbands and process each subband separately. These methods can select appropriate processing techniques based on the noise characteristics within each subband's own frequency band. Furthermore, subband-based methods can leverage detailed processing of subband information to uncover subtle features hidden in the time-frequency distribution. However, subband methods may lose full-band spectral correlation and cannot utilize long-distance cross-band dependencies, potentially leading to discontinuities or inconsistencies when reconstructing the overall spectral structure. Therefore, it is difficult to recover clear speech in subbands with low signal-to-noise ratios.

[0004] Given the advantages and disadvantages of full-band and sub-band methods, full-band-sub-band fusion methods have emerged in recent years. By integrating the global information processing capabilities of full-band methods with the local detail analysis advantages of sub-band methods, the problem of balancing detail preservation and overall structure in speech signal processing has been effectively improved. However, current methods still have significant limitations: the interaction between full-band and sub-band features is relatively shallow, global information is difficult to accurately guide sub-band processing, and sub-band details are not effectively fed back into full-band modeling; at the same time, the problem of ignoring the spectral correlation between sub-bands has not been effectively solved, and sub-band errors can spread to the full-band during the full-band-sub-band feature fusion process, thus affecting the overall effect of speech enhancement.

[0005] Therefore, existing technologies still need improvement and development. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a speech enhancement method, device, terminal and medium based on full-band and sub-band fusion, in order to address the above-mentioned defects of the prior art.

[0007] The technical solution adopted by this invention to solve the problem is as follows:

[0008] In a first aspect, embodiments of the present invention provide a speech enhancement method based on full-band and sub-band fusion, wherein the method includes:

[0009] Acquire a noisy speech signal, and determine an initial noisy full-band spectrum of the speech signal based on the noisy speech signal;

[0010] The amplitude spectrum corresponding to the initial noisy full-band spectrum of the speech is decomposed into three sub-bands: low frequency band, mid frequency band and high frequency band by the sub-band enhancement module. Speech enhancement is performed based on each sub-band to determine the preliminary sub-band enhanced amplitude spectrum.

[0011] The initial full-band enhancement spectrum is determined by the full-band enhancement module based on the initial noisy speech full-band spectrum.

[0012] The preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum are fused using a spectrum fusion module to determine the full-band-sub-band fused spectrum;

[0013] The full-band-subband fused spectrum is processed sequentially through a gated loop unit and a linear layer to determine the fused complex value mask;

[0014] The target speech enhancement spectrum is determined based on the fused complex-valued mask and the initial noisy full-band speech spectrum.

[0015] In one implementation, determining a preliminary sub-band enhancement amplitude spectrum based on each of the sub-bands includes:

[0016] Feature extraction is performed on each of the sub-bands to determine the sub-band features corresponding to each of the sub-bands;

[0017] Generate feature masks for each sub-band based on the features of each sub-band;

[0018] Gating mechanism processing and feature fusion are performed on the sub-band features corresponding to each sub-band and the sub-band feature masks corresponding to each sub-band other than the sub-band to determine the sub-band enhancement features;

[0019] The frequency bands of each sub-band enhancement feature are combined to determine the preliminary sub-band enhancement amplitude spectrum.

[0020] In one implementation, a preliminary full-band enhancement spectrum is determined by a full-band enhancement module based on the initial noisy speech full-band spectrum, including:

[0021] Based on the multi-scale fusion encoder, multi-scale features are determined according to the initial noisy full-band spectrum of the speech;

[0022] A bidirectional time-frequency feature extraction module is used to extract features from the multi-scale features to determine the time-frequency domain features;

[0023] The preliminary full-band enhancement spectrum is determined based on the time-frequency domain characteristics.

[0024] In one implementation, determining multi-scale features based on a multi-scale fusion encoder according to the initial noisy full-band spectrum of the speech includes:

[0025] The real and imaginary parts of the amplitude spectrum and complex spectrum corresponding to the initial noisy full-band spectrum of the speech are concatenated to form a three-dimensional hybrid feature;

[0026] A multi-scale fusion encoder is used to extract features from the three-dimensional hybrid features to determine the multi-scale features.

[0027] In one implementation method, determining the preliminary full-band enhanced spectrum based on the time-domain-frequency domain characteristics includes:

[0028] An amplitude spectrum decoder is used to reconstruct the amplitude spectrum features based on the time-domain-frequency domain features and the multi-scale features, thereby determining the amplitude spectrum features.

[0029] The time-domain-frequency domain features are reconstructed using a complex spectrum decoder to determine the complex spectrum features;

[0030] The preliminary full-band enhancement spectrum is determined based on the amplitude spectrum characteristics and the complex spectrum characteristics.

[0031] In one implementation, an amplitude spectrum decoder is used to reconstruct and determine the amplitude spectrum features based on time-domain-frequency domain features and the multi-scale features, including:

[0032] The amplitude spectrum mask is determined by the masking module based on the time-frequency domain characteristics;

[0033] The multi-scale features are masked based on the amplitude spectrum mask, and the masked multi-scale features are determined.

[0034] An amplitude spectrum decoder is used to reconstruct the amplitude spectrum features based on the masked multi-scale features to determine the amplitude spectrum features.

[0035] In one implementation method, the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum are fused by a spectrum fusion module to determine the full-band-sub-band fused spectrum, including:

[0036] Linear transformation and nonlinear activation processing are performed on the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to determine the fused real-valued mask;

[0037] The preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum are weighted according to the fused real-value mask to determine the full-band-sub-band fused spectrum.

[0038] Secondly, embodiments of the present invention also provide a speech enhancement device based on full-band and sub-band fusion, wherein the speech enhancement device based on full-band and sub-band fusion includes:

[0039] The preliminary spectrum extraction module is used to acquire the noisy speech signal and determine the initial noisy full-band spectrum of the speech signal based on the noisy speech signal.

[0040] The preliminary sub-band enhancement module is used to decompose the amplitude spectrum corresponding to the initial noisy full-band spectrum of speech into three sub-bands: low-frequency band, mid-frequency band, and high-frequency band, and to perform speech enhancement based on each sub-band to determine the preliminary sub-band enhancement amplitude spectrum.

[0041] A preliminary full-band enhancement module is used to determine a preliminary full-band enhancement spectrum based on the initial noisy speech full-band spectrum.

[0042] The spectrum fusion module is used to fuse the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to determine the full-band-sub-band fused spectrum.

[0043] The fusion complex value mask determination module is used to process the full-band-sub-band fused spectrum sequentially through a gated loop unit and a linear layer to determine the fusion complex value mask;

[0044] The target spectrum enhancement module is used to determine the target speech enhancement spectrum based on the fused complex value mask and the initial noisy full-band speech spectrum.

[0045] Thirdly, embodiments of the present invention also provide a terminal, the terminal including a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing the speech enhancement method based on full-band and sub-band fusion as described above; the processor is used to execute the programs.

[0046] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions, wherein the instructions are adapted to be loaded and executed by a processor to implement any of the above-described speech enhancement methods based on full-band and sub-band fusion.

[0047] The beneficial effects of this invention are as follows: In this embodiment, an initial noisy full-band spectrum of the speech signal is determined. A sub-band enhancement module decomposes the amplitude spectrum corresponding to the initial noisy full-band spectrum into three sub-bands: low-frequency, mid-frequency, and high-frequency. Speech enhancement is performed based on each sub-band to determine the preliminary sub-band enhancement amplitude spectrum. A full-band enhancement module determines the preliminary full-band enhancement spectrum based on the initial noisy full-band spectrum. A spectrum fusion module fuses the preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum to determine the full-band-sub-band fused spectrum. A fused complex-value mask is determined based on the full-band-sub-band fused spectrum. The target speech enhancement spectrum is determined based on the fused complex-value mask and the initial noisy full-band spectrum. Because this invention sets up two different branches, full-band and sub-band, and achieves information interaction and fusion between different frequency band levels through the spectrum fusion module, the enhancement effect is improved. Therefore, it can effectively solve the problems of insufficient information interaction and inadequate feature fusion in the existing technology during the full-band and sub-band fusion process. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating the speech enhancement method based on full-band and sub-band fusion provided in an embodiment of the present invention.

[0050] Figure 2 This is a schematic diagram of an embodiment of the speech enhancement method based on full-band and sub-band fusion provided by the present invention.

[0051] Figure 3 This is a schematic diagram of an embodiment of the sub-band enhancement module provided in this invention.

[0052] Figure 4 This is a schematic diagram of an embodiment of the multi-scale fusion encoder and decoder provided in this invention.

[0053] Figure 5 This is a schematic diagram of an embodiment of the bidirectional time-frequency feature extraction module provided in this invention.

[0054] Figure 6 This is a schematic diagram of an embodiment of the time-domain bidirectional Conformer provided in this invention.

[0055] Figure 7 This is a schematic diagram of an embodiment of the multi-scale feature fusion extraction module provided in this invention.

[0056] Figure 8 This is a schematic diagram of an embodiment of the spectrum fusion module provided in this invention.

[0057] Figure 9 This is a schematic diagram of the internal modules of the speech enhancement device based on full-band and sub-band fusion provided in an embodiment of the present invention.

[0058] Figure 10 This is a schematic diagram of the terminal provided in an embodiment of the present invention. Detailed Implementation

[0059] This invention discloses a speech enhancement method, apparatus, terminal, and medium based on full-band and sub-band fusion. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.

[0060] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0061] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0062] Given the advantages and disadvantages of full-band and sub-band methods, full-band-sub-band fusion methods have emerged in recent years. By integrating the global information processing capabilities of full-band methods with the local detail analysis advantages of sub-band methods, the problem of balancing detail preservation and overall structure in speech signal processing has been effectively improved. However, current methods still have significant limitations: the interaction between full-band and sub-band features is relatively shallow, global information is difficult to accurately guide sub-band processing, and sub-band details are not effectively fed back into full-band modeling; at the same time, the problem of ignoring the spectral correlation between sub-bands has not been effectively solved, and sub-band errors can spread to the full-band during the full-band-sub-band feature fusion process, thus affecting the overall effect of speech enhancement.

[0063] To address the aforementioned shortcomings of existing technologies, this invention provides a speech enhancement method based on full-band and sub-band fusion. The method determines an initial noisy full-band spectrum of the noisy speech signal; a sub-band enhancement module decomposes the amplitude spectrum corresponding to the initial noisy full-band spectrum into three sub-bands: low-frequency, mid-frequency, and high-frequency. Speech enhancement is performed based on each sub-band to determine a preliminary sub-band enhancement amplitude spectrum; a full-band enhancement module determines a preliminary full-band enhancement spectrum based on the initial noisy full-band spectrum; a spectrum fusion module fuses the preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum to determine a full-band-sub-band fused spectrum; a fused complex-valued mask is determined based on the full-band-sub-band fused spectrum; and a target speech enhancement spectrum is determined based on the fused complex-valued mask and the initial noisy full-band spectrum. Because this invention sets up two different branches, full-band and sub-band, and achieves information interaction and fusion between different frequency band levels through the spectrum fusion module, the enhancement effect is improved. Therefore, it effectively solves the problems of insufficient information interaction and inadequate feature fusion in the full-band and sub-band fusion process in existing technologies.

[0064] Exemplary method:

[0065] like Figure 1 As shown, the method includes:

[0066] Step S100: Obtain the noisy speech signal and determine the initial noisy full-band spectrum based on the noisy speech signal.

[0067] Noisy speech signals are unprocessed full-band speech signals. The initial noisy full-band spectrum is the representation of the noisy speech signal after conversion to the frequency domain, used to represent the frequency components and their amplitudes (i.e., signal strength) contained in the signal. By determining the initial noisy full-band spectrum based on the noisy speech signal, and then modeling based on the initial noisy full-band spectrum, it is possible to better separate the speech signal from the noise signal, accurately reconstructing clean and clear speech in complex acoustic environments, thereby enhancing the full-band speech. Because noisy speech signals have strong overall characteristics, processing based on the initial noisy full-band spectrum can comprehensively capture the global features of the speech signal, demonstrating good processing capabilities for complex speech information. For example, for complex speech signals involving multilingual mixed speech or multiple interwoven sound sources, speech enhancement using the initial noisy full-band spectrum can effectively extract key speech information for subsequent processing.

[0068] In one implementation, determining the initial noisy full-band spectrum of the noisy speech signal includes: performing a short-time Fourier transform on the noisy speech signal to determine the initial noisy full-band spectrum, wherein the initial noisy full-band spectrum is a complex spectrum; and performing a modulo operation on the initial noisy full-band spectrum to determine the amplitude spectrum corresponding to the initial noisy full-band spectrum.

[0069] In simple terms, the Short-Time Fourier Transform (SFT) converts a speech signal into a complex spectrum, determining the complex spectrum corresponding to the initial noisy full-band speech spectrum. Then, a modulo operation is performed on this complex spectrum to obtain the amplitude spectrum. In subsequent processing, the initial noisy full-band speech spectrum includes both its complex spectrum and amplitude spectrum. Compared to the Discrete Fourier Transform (DFT), the SFT preserves the time-domain information of the speech signal. This embodiment uses the SFT to convert the noisy speech signal into an initial noisy full-band speech spectrum for subsequent speech enhancement processing.

[0070] like Figure 1 As shown, the method includes:

[0071] Step S200: The amplitude spectrum corresponding to the initial noisy full-band spectrum of the speech is decomposed into three sub-bands: low-frequency band, mid-frequency band, and high-frequency band by the sub-band enhancement module. Speech enhancement is performed based on each sub-band to determine the preliminary sub-band enhancement amplitude spectrum.

[0072] Considering that local patterns in the spectrogram typically differ across frequency bands—lower frequency bands are more likely to contain high-energy, high-pitched, and sustained sounds, while higher frequency bands often contain low-energy, noise, and rapidly decaying sounds—noisy speech signals tend to ignore the time-frequency distribution information of the speech spectrum. This embodiment uses a sub-band enhancement module to decompose the amplitude spectrum corresponding to the initial noisy full-band speech spectrum into low-frequency, mid-frequency, and high-frequency sub-bands. By processing each sub-band, detailed features in the local speech signal are captured, and noise within each sub-band is selectively removed based on its characteristics, resulting in a preliminary sub-band enhanced amplitude spectrum.

[0073] In one implementation, speech enhancement is performed based on each of the sub-bands to determine a preliminary sub-band enhancement amplitude spectrum, including:

[0074] Step S201: Extract features from each of the sub-bands to determine the sub-band features corresponding to each of the sub-bands;

[0075] Step S202: Generate the feature mask of each sub-band based on the features of each sub-band;

[0076] Step S203: Perform gating mechanism processing and feature fusion on the sub-band features corresponding to each sub-band and the sub-band feature masks corresponding to each sub-band other than the sub-band to determine the sub-band enhancement features;

[0077] Step S204: Perform frequency band merging on each of the sub-band enhancement features to determine the preliminary sub-band enhancement amplitude spectrum.

[0078] Specifically, such as Figure 3 As shown, the amplitude spectrum corresponding to the initial noisy full-band speech spectrum is decomposed into three sub-bands: low-frequency, mid-frequency, and high-frequency. Convolution is used to extract basic features from each of the three sub-bands, yielding sub-band features for each. Six cross-sub-band gating units are constructed, each consisting of a convolutional layer and a sigmoid activation function. Gating signals are generated based on the basic features of different sub-bands, and these gating signals enhance the basic features of each sub-band, resulting in enhanced sub-band features. Specifically, the low-frequency enhanced features are obtained by multiplying the low-frequency basic features with the gating signals generated from the mid-frequency and high-frequency basic features, and then superimposing them. The mid-frequency enhanced features are obtained by multiplying the mid-frequency basic features with the gating signals generated from the low-frequency and high-frequency basic features, and then superimposing them. The high-frequency enhanced features are obtained by multiplying the high-frequency basic features with the gating signals generated from the low-frequency and mid-frequency basic features, and then superimposing them. The output convolutional layer compresses the dimensionality of each sub-band enhanced feature and concatenates them along the frequency dimension to obtain a complete preliminary sub-band enhanced amplitude spectrum, which is used for subsequent full-band-sub-band fusion.

[0079] In one implementation, several sub-bands can be obtained by uniformly dividing the amplitude spectrum corresponding to the full-band spectrum of the initial noisy speech. Alternatively, the amplitude spectrum can be divided non-uniformly (Mel / Bark band) or adaptively (energy threshold) to obtain several sub-bands.

[0080] like Figure 1 As shown, the method includes:

[0081] Step S300: Determine the preliminary full-band enhancement spectrum based on the initial noisy full-band speech spectrum using the full-band enhancement module.

[0082] In simple terms, the initial noisy speech full-band spectrum is a full-band spectrum, covering the entire spectral range of the initial noisy speech. When performing contextual relationship modeling based on the noisy speech full-band spectrum, it is possible to better process the temporal information in the speech, improve the understanding and modeling ability of the speech signal, and thus better separate the speech signal from the noise signal. This allows for accurate reconstruction of clean and clear speech in complex acoustic environments, thereby achieving the goal of speech enhancement. In this embodiment, the full-band enhancement module performs coarse-grained processing and preliminary enhancement based on the initial noisy speech full-band spectrum to obtain the preliminary full-band enhanced spectrum.

[0083] Specifically, such as Figure 2 As shown, the full-band enhancement module includes a multi-scale fusion encoder, a bidirectional time-frequency feature extraction module, a masking module, an amplitude spectrum decoder, and a complex spectrum decoder. In one implementation, the full-band enhancement module determines the initial full-band enhancement spectrum based on the initial noisy speech full-band spectrum, including:

[0084] Step S301: Determine multi-scale features based on the initial noisy full-band spectrum of the multi-scale fusion encoder;

[0085] Step S302: Use a bidirectional time-frequency feature extraction module to extract features from the multi-scale features and determine the time-frequency domain features;

[0086] Step S303: Determine the preliminary full-band enhancement spectrum based on the time-frequency domain characteristics.

[0087] Specifically, the initial noisy full-band spectrum of the speech is input into a multi-scale fusion encoder, and feature extraction is performed through the multi-scale fusion encoder to obtain multi-scale features. For example... Figure 4As shown in (a), a multi-scale fusion encoder can be constructed using multiple (e.g., four) multi-scale fusion feature extractors (MSMFEs). By inputting the initial noisy full-band spectrum of the speech into the MSMFE for feature extraction, multi-scale features are obtained. A time-frequency feature extraction module is then used to perform joint time-domain and frequency-domain modeling based on the multi-scale features, yielding time-domain and frequency-domain features. Speech enhancement is then applied to these time-domain and frequency-domain features to determine the initial full-band enhanced spectrum. For complex speech signals involving multiple languages ​​or multiple interleaved sound sources, extracting the corresponding time-frequency-domain features of the full-band speech signal and performing speech enhancement based on these features effectively extracts key speech information.

[0088] The structure of the bidirectional time-frequency feature extraction module (BiT-F Conformer) is as follows: Figure 5 As shown. In this embodiment, the bidirectional time-frequency feature extraction module is mainly used to extract key time-domain and frequency-domain features. Its structure integrates the deep modeling capabilities of the time and frequency domains, enabling it to capture the dynamic changes and intrinsic dependencies of speech signals in different dimensions, thereby effectively improving feature representation capabilities and speech enhancement effects.

[0089] The bidirectional time-frequency feature extraction module includes a bidirectional time-conformer (Bi-Time-Conformer, such as...) Figure 6 The model employs a frequency-domain Conformer (Freq-Conformer) and a multi-level residual connection structure. Multi-scale features are rearranged and fed into the feature extraction paths in both the time and frequency directions, enabling the separation and modeling of multi-dimensional information. The time-domain bidirectional Conformer uses a bidirectional structure, extracting features from both the forward and reverse directions of the time series, enhancing the model's ability to model long-distance temporal dependencies and helping to maintain the temporal consistency and integrity of speech. The frequency-domain Conformer extracts global features in the frequency dimension, modeling the connections between frequency bands through a self-attention mechanism, thus enhancing the model's understanding of the speech spectral structure.

[0090] The bidirectional time-frequency feature extraction module introduces a dual-path processing structure. Input multi-scale features are directly fed into the Conformer module on one hand, maintaining stable information transmission through residual connections; on the other hand, after a dimension flipping operation, they are fed into another Conformer module to extract features from different perspectives, enhancing the model's global perception capability. The features output from the two paths are concatenated at the end and a linear transformation is applied to output the fusion result, achieving complete information integration.

[0091] In addition, in the bidirectional time-frequency feature extraction module, the processing of time-domain data can also use bidirectional LSTM / Transformer instead of Conformer to handle temporal dependencies; the processing of frequency-domain data can use 2D-CNN instead of self-attention to model correlations through cross-band convolution.

[0092] In one implementation, a multi-scale fusion encoder determines multi-scale features based on the initial noisy full-band spectrum of the speech, including:

[0093] Step S3011: The real and imaginary parts of the amplitude spectrum and complex spectrum corresponding to the initial noisy full-band spectrum of the speech are concatenated into a three-dimensional hybrid feature;

[0094] Step S3012: Use a multi-scale fusion encoder to extract features from the three-dimensional hybrid features and determine the multi-scale features.

[0095] Specifically, the initial noisy full-band spectrum includes an amplitude spectrum and a complex spectrum, where the complex spectrum includes a real part and an imaginary part. Before inputting the initial noisy full-band spectrum into the full-band enhancement module, the amplitude spectrum, the real part, and the imaginary part of the complex spectrum are concatenated into a three-dimensional hybrid feature, providing richer basic features for speech enhancement. The concatenated three-dimensional hybrid is then input into a multi-scale fusion encoder to extract multi-scale features, achieving multi-scale fine-grained modeling of speech features and enhancing the model's feature representation ability under different receptive fields.

[0096] A multi-scale fusion encoder consists of multiple multi-scale fusion feature extractors, such as... Figure 7 As shown, the multi-scale fusion feature extractor consists of multiple independent convolutional units and a Selective Kernel Feature Fusion (SKFF) module. Through multi-level feature extraction and adaptive fusion mechanisms, it enhances the ability to perceive local and global information in speech signals.

[0097] Each convolutional unit comprises a combination of pointwise convolutions and depthwise separable convolutions, used to dynamically adjust the channel structure and efficiently extract local features. Combined with residual connection structures, this helps alleviate the vanishing gradient problem in deep training, improving model stability and convergence speed. The outputs of the convolutional units are normalized and processed with the PReLU activation function to enhance the network's nonlinear modeling capabilities.

[0098] To achieve multi-scale feature coverage, the multi-scale fusion feature extractor design incorporates three depthwise separable convolutional kernels of different sizes: 3×3, 5×5, and 7×7. These convolutional paths focus on local details, balancing regions and overall structure, and global contextual information, respectively, forming a multi-level, multi-perspective model of the input features, effectively improving the model's adaptability to diverse speech scenarios.

[0099] Finally, the features output from each scale path are input into the SKFF module. This module dynamically adjusts the fusion weights of features at each scale based on the current input content through an adaptive weighting mechanism, thereby achieving the optimal information integration strategy. Through this structure, the multi-scale fusion feature extractor extracts multi-scale features using three different sized convolutional kernels. After adaptive fusion by SKFF, it can improve the multi-level perception capability of speech signals and the model's feature representation and discrimination performance under complex speech signals.

[0100] In addition, in multi-scale fusion feature extractors, dilated / deformable convolutions can be used instead of fixed-size kernels to dynamically adjust the receptive field; channel attention can be used instead of SKFF to learn the distribution of feature importance.

[0101] In one implementation, determining the preliminary full-band enhanced spectrum based on the time-frequency domain characteristics includes:

[0102] Step S3031: Use an amplitude spectrum decoder to reconstruct the amplitude spectrum features based on the time-domain-frequency domain features and the multi-scale features;

[0103] Step S3032: Reconstruct the time-domain-frequency-domain features using a complex spectrum decoder to determine the complex spectrum features;

[0104] Step S3033: Determine the preliminary full-band enhancement spectrum based on the amplitude spectrum characteristics and the complex spectrum characteristics.

[0105] In simple terms, this embodiment uses dual decoders to reconstruct the initial full-band enhanced spectrum in parallel based on time-domain and frequency-domain features. Specifically, an amplitude spectrum decoder reconstructs the amplitude spectrum features based on time-domain and frequency-domain features, while a complex spectrum decoder reconstructs the complex spectrum features based on time-domain and frequency-domain features. The complex spectrum contains real and imaginary parts and is a joint time-frequency domain feature; its amplitude and phase together carry the correlation information between the frequency and time domains. The initial full-band enhanced spectrum can be determined based on the amplitude spectrum features and the complex spectrum features.

[0106] like Figure 4 As shown in (b), the amplitude spectrum decoder and the complex spectrum decoder are implemented based on multiple multi-scale fusion feature extractors. The difference lies in the number of channels. The amplitude spectrum decoder has 1 channel, while the complex spectrum decoder has 2 channels, so the complex spectrum decoder can output the real part and the imaginary part.

[0107] In one implementation, an amplitude spectrum decoder is used to reconstruct the amplitude spectrum features based on the time-domain-frequency domain features and the multi-scale features, including:

[0108] The amplitude spectrum mask is determined by the masking module based on the time-frequency domain characteristics;

[0109] The multi-scale features are masked based on the amplitude spectrum mask, and the masked multi-scale features are determined.

[0110] An amplitude spectrum decoder is used to reconstruct the amplitude spectrum features based on the masked multi-scale features to determine the amplitude spectrum features.

[0111] This embodiment determines the amplitude spectrum mask module based on time-frequency domain features using a masking module. It then adaptively weights key features from the time-frequency domain features using the amplitude spectrum mask, and combines this with multi-scale features generated by a multi-scale fusion encoder. An amplitude spectrum decoder is then used to reconstruct and determine the amplitude spectrum features. The masking module is implemented based on a dynamic spectral mask mechanism. This embodiment uses a bidirectional time-frequency feature extraction module for joint time-frequency modeling, combined with adaptive weighting of key features using a dynamic spectral mask, to effectively capture the time-frequency dependencies of speech signals.

[0112] like Figure 1 As shown, the method further includes the following steps:

[0113] Step S400: The preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum are fused using the spectrum fusion module to determine the full-band-sub-band fused spectrum.

[0114] Specifically, the spectrum fusion module is one module in the frequency band fusion model. In this embodiment, the noisy speech signal is first enhanced to obtain a preliminary full-band enhanced spectrum. Then, the spectrum fusion module fuses the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum, achieving information interaction between different frequency band levels. This fully leverages the synergistic effect between full-band and sub-band features, thereby enhancing the expressive power of the speech signal and the model's discriminative ability. This effectively solves the problem that insufficient information interaction between the full-band and sub-band leads to a lack of effective complementarity during the fusion process, resulting in the loss of some details or the advantages of the global structure.

[0115] In one implementation, the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum are fused by a spectrum fusion module to determine the full-band-sub-band fused spectrum, including:

[0116] Linear transformation and nonlinear activation processing are performed on the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to determine the fused real-valued mask;

[0117] The preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum are weighted according to the fused real-value mask to determine the full-band-sub-band fused spectrum.

[0118] Specifically, the spectrum fusion module enables deep fusion between full-band and sub-band features, improving the overall performance of a single-channel speech enhancement system. For example... Figure 8As shown, the steps for fusing the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum using the spectrum fusion module include:

[0119] Linear transformation and nonlinear activation are applied to the initial full-band enhancement spectrum and the initial sub-band enhancement amplitude spectrum to obtain a fused real-valued mask. In this embodiment, the initial full-band enhancement spectrum and the initial sub-band enhancement amplitude spectrum are processed by linear transformation and nonlinear activation operations to generate an initial hidden representation, which is used to construct the feature basis for the fusion process. By introducing an average pooling operation, information exchange between sub-band features is realized, thereby capturing the internal correlation between frequency bands and improving the model's ability to perceive speech structure. Then, the hidden features are transformed through further linear layers and nonlinear activation functions to extract global context information and generate a fused real-valued mask. The fused real-valued mask is obtained through the Sigmoid function and is used to model the importance of features in different frequency bands.

[0120] Based on the fused real-valued mask, an element-wise weighted operation is performed on the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to achieve adaptive fusion of full-band and sub-band information. This fusion strategy can dynamically adjust the information flow path according to the input content, thereby enhancing the spectrum fusion module's ability to express and understand complex frequency patterns in speech signals and outputting the final enhanced feature result.

[0121] In addition to weighting each tuple in the signal by generating a fused real-value mask, the preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum can also be fused by feature splicing and cross-band convolution, thus avoiding explicit mask generation.

[0122] like Figure 1 As shown, the method further includes the following steps:

[0123] Step S500: The full-band-subband fused spectrum is processed sequentially through a gated loop unit and a linear layer to determine the fused complex value mask.

[0124] Specifically, in this embodiment, determining the fused complex-valued mask based on the full-band-sub-band fused spectrum includes: processing the full-band-sub-band fused spectrum sequentially through a gated recurrent unit (GRU) and a linear layer to determine the real part feature mask and the imaginary part feature mask; and determining the fused complex-valued mask based on the real part feature mask and the imaginary part feature mask. The GRU effectively captures long-term dependencies in the sequence data, and the data from the GRU is then linearly transformed by the linear layer to obtain a suitable dimension and representation. Finally, the output of the linear layer is converted into a mask using an activation function.

[0125] like Figure 1 As shown, the method further includes the following steps:

[0126] Step S600: Determine the target speech enhancement spectrum based on the fused complex value mask and the initial noisy full-band speech spectrum.

[0127] In simple terms, after obtaining the fused complex value mask, the fused complex value mask is multiplied by the original noisy full-band spectrum of the speech. The noise in the original noisy full-band spectrum of the speech is removed by the fused complex value mask, thus obtaining the enhanced spectrum of the target speech.

[0128] Based on the above embodiments, the present invention also provides a speech enhancement device based on full-band and sub-band fusion, such as... Figure 9 As shown, the device includes:

[0129] The preliminary spectrum extraction module 01 is used to acquire the noisy speech signal and determine the initial noisy full-band spectrum of the speech signal based on the noisy speech signal.

[0130] The preliminary sub-band enhancement module 02 is used to decompose the amplitude spectrum corresponding to the initial noisy full-band spectrum of the speech into three sub-bands: low frequency band, mid frequency band and high frequency band, and perform speech enhancement based on each sub-band to determine the preliminary sub-band enhancement amplitude spectrum.

[0131] The preliminary full-band enhancement module 03 is used to determine the preliminary full-band enhancement spectrum based on the initial noisy speech full-band spectrum through the full-band enhancement module;

[0132] The spectrum fusion module 04 is used to fuse the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum through the spectrum fusion module to determine the full-band-sub-band fused spectrum;

[0133] The fusion complex value mask determination module 05 is used to process the full-band-sub-band fused spectrum sequentially through a gated loop unit and a linear layer to determine the fusion complex value mask;

[0134] The target spectrum enhancement module 06 is used to determine the target speech enhancement spectrum based on the fused complex value mask and the initial noisy full-band speech spectrum.

[0135] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 10As shown, the terminal includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice enhancement method based on full-band and sub-band fusion. The display screen can be an LCD screen or an e-ink screen.

[0136] Those skilled in the art will understand that Figure 10 The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0137] In one implementation, the terminal's memory stores one or more programs, and these programs are configured to be executed by one or more processors, and the programs contain instructions for performing a speech enhancement method based on full-band and sub-band fusion.

[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0139] In summary, this invention discloses a speech enhancement method, apparatus, terminal, and medium based on full-band and sub-band fusion. The method determines an initial noisy full-band spectrum of the noisy speech signal; decomposes the amplitude spectrum corresponding to the initial noisy full-band spectrum into three sub-bands—low-frequency, mid-frequency, and high-frequency—using a sub-band enhancement module; performs speech enhancement based on each sub-band to determine a preliminary sub-band enhancement amplitude spectrum; determines a preliminary full-band enhancement spectrum using a full-band enhancement module based on the initial noisy full-band spectrum; fuses the preliminary full-band enhancement spectrum and the preliminary sub-band enhancement amplitude spectrum using a spectrum fusion module to determine a full-band-sub-band fused spectrum; determines a fused complex-valued mask based on the full-band-sub-band fused spectrum; and determines the target speech enhancement spectrum based on the fused complex-valued mask and the initial noisy full-band spectrum. Because this invention sets up two different branches—full-band and sub-band—and achieves information interaction and fusion between different frequency band levels through the spectrum fusion module, it improves the enhancement effect. Therefore, it effectively solves the problems of insufficient information interaction and inadequate feature fusion in the existing technology during full-band and sub-band fusion.

[0140] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A speech enhancement method based on full-band and sub-band fusion, characterized in that, The method comprises: acquiring a noisy speech signal, determining an initial noisy speech full-band spectrum according to the noisy speech signal; decomposing an amplitude spectrum corresponding to the initial noisy speech full-band spectrum into three sub-bands of a low frequency band, a medium frequency band and a high frequency band through a sub-band enhancement module, performing speech enhancement based on each of the sub-bands, and determining a preliminary sub-band enhanced amplitude spectrum; determining a preliminary full-band enhanced spectrum according to the initial noisy speech full-band spectrum through a full-band enhancement module; fusing the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum through a spectrum fusion module to determine a full-band-sub-band fused spectrum; processing the full-band-sub-band fused spectrum through a gated recurrent unit and a linear layer in sequence to determine a fused complex mask; determining a target speech enhanced spectrum according to the fused complex mask and the initial noisy speech full-band spectrum; determining a preliminary full-band enhanced spectrum according to the initial noisy speech full-band spectrum through a full-band enhancement module, comprising: determining a multi-scale feature according to the initial noisy speech full-band spectrum based on a multi-scale fusion encoder; extracting a time-frequency feature using a bidirectional time-frequency feature extraction module to determine a time-frequency feature; determining the preliminary full-band enhanced spectrum according to the time-frequency feature.

2. The speech enhancement method based on full-band and sub-band fusion according to claim 1, characterized in that, performing speech enhancement based on each of the sub-bands to determine a preliminary sub-band enhanced amplitude spectrum, comprising: extracting a feature of each of the sub-bands to determine a sub-band feature corresponding to each of the sub-bands; generating a sub-band feature mask corresponding to each of the sub-bands according to each sub-band feature; respectively performing gated mechanism processing and feature fusion on the sub-band feature corresponding to each of the sub-bands and the sub-band feature mask corresponding to each of the sub-bands except the sub-band to determine a sub-band enhanced feature; performing frequency band merging on each of the sub-band enhanced features to determine the preliminary sub-band enhanced amplitude spectrum.

3. The speech enhancement method based on full-band and sub-band fusion according to claim 1, characterized in that, determining a multi-scale feature according to the initial noisy speech full-band spectrum based on a multi-scale fusion encoder, comprising: concatenating the real part and the imaginary part of the amplitude spectrum and the complex spectrum corresponding to the initial noisy speech full-band spectrum into a three-dimensional hybrid feature; extracting a feature of the three-dimensional hybrid feature using a multi-scale fusion encoder to determine the multi-scale feature.

4. The speech enhancement method based on full-band and sub-band fusion according to claim 1, characterized in that, determining the preliminary full-band enhanced spectrum according to the time-frequency feature, comprising: reconstructing the time-frequency feature and the multi-scale feature using an amplitude spectrum decoder to determine an amplitude spectrum feature; reconstructing the time-frequency feature using a complex spectrum decoder to determine a complex spectrum feature; determining the preliminary full-band enhanced spectrum according to the amplitude spectrum feature and the complex spectrum feature.

5. The speech enhancement method based on full-band and sub-band fusion according to claim 4, characterized in that, reconstructing the time-frequency feature and the multi-scale feature using an amplitude spectrum decoder to determine an amplitude spectrum feature, comprising: determining an amplitude spectrum mask according to the time-frequency feature using a mask module; masking the multi-scale feature based on the amplitude spectrum mask to determine a masked multi-scale feature; reconstructing the masked multi-scale feature using an amplitude spectrum decoder to determine an amplitude spectrum feature.

6. The speech enhancement method based on full-band and sub-band fusion according to claim 1, characterized in that, fusing the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum through a spectrum fusion module to determine a full-band-sub-band fused spectrum, comprising: linearly transforming and nonlinearly activating the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to determine a fused real-valued mask; weighting the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum according to the fused real-valued mask to determine the full-band-sub-band fused spectrum.

7. A speech enhancement device based on full-band and sub-band fusion, characterized by, The apparatus comprises: a preliminary spectrum extraction module configured to obtain a noisy speech signal and determine an initial noisy speech full-band spectrum based on the noisy speech signal; a preliminary sub-band enhancement module configured to decompose an amplitude spectrum corresponding to the initial noisy speech full-band spectrum into three sub-bands of a low frequency band, a middle frequency band, and a high frequency band by a sub-band enhancement module, perform speech enhancement based on each of the sub-bands to determine a preliminary sub-band enhanced amplitude spectrum; a preliminary full-band enhancement module configured to determine a preliminary full-band enhanced spectrum based on the initial noisy speech full-band spectrum by a full-band enhancement module; a spectrum fusion module configured to fuse the preliminary full-band enhanced spectrum and the preliminary sub-band enhanced amplitude spectrum to determine a full-band-sub-band fused spectrum by a spectrum fusion module; a fused complex-valued mask determination module configured to determine a fused complex-valued mask by sequentially processing the full-band-sub-band fused spectrum by a gated recurrent unit and a linear layer; a target spectrum enhancement module configured to determine a target speech enhanced spectrum based on the fused complex-valued mask and the initial noisy speech full-band spectrum; determining a preliminary full-band enhanced spectrum based on the initial noisy speech full-band spectrum by a full-band enhancement module comprises: determining a multi-scale feature based on the initial noisy speech full-band spectrum by a multi-scale fusion encoder; performing feature extraction on the multi-scale feature to determine a time-frequency feature by a bidirectional time-frequency feature extraction module; and determining the preliminary full-band enhanced spectrum based on the time-frequency feature.

8. A terminal, characterized by comprising: The terminal comprises a memory and one or more processors; the memory stores one or more programs; the programs contain instructions for executing the full-band and sub-band fusion based speech enhancement method according to any one of claims 1-6; and the processor is configured to execute the programs.

9. A computer-readable storage medium storing a plurality of instructions thereon, characterized in that, The instructions are loaded and executed by the processor to implement the steps of the full-band and sub-band fusion based speech enhancement method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multichannel voice signal enhancement method

    CN119049497A