A speech separation post-filtering method and system based on sub-band envelope features
By employing a post-speech filtering method based on subband envelope features, and utilizing an m-octave bandpass filter bank and envelope masking coefficients, the problem of insufficient utilization of frequency band envelope features and auditory mechanisms in existing technologies is solved, thereby improving the quality of speech separation.
Patent Information
- Application Number
- CN202510145356.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing speech separation techniques do not fully utilize the envelope features of the estimated source signal and mixed signal in different frequency bands, and do not follow the human auditory perception mechanism, resulting in a degradation in speech separation quality.
A speech separation and filtering method based on subband envelope features is adopted. Subband decomposition is performed by designing an m-octave bandpass filter bank, calculating the envelope masking coefficient and suppressing interference components, and finally reconstructing and enhancing the speech signal.
It effectively reduces crosstalk components in speech signals and improves speech separation performance in complex acoustic environments.
Smart Images

Figure CN119964593B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of speech separation, and particularly relates to a speech separation post-filtering method and system based on sub-band envelope features. BACKGROUND
[0002] The goal of speech separation is to extract individual speech sources from mixed speech signals. This is very useful in many applications, such as speech recognition in noisy environments, or separating individual speakers' speech in a teleconference. However, this is a very challenging task because multiple speakers' speech often interferes with each other.
[0003] Traditional speech separation techniques mainly rely on signal processing methods, such as beamforming techniques based on spatial characteristics, or independent component analysis techniques based on statistical characteristics. However, these methods usually require precise prior knowledge of the characteristics of the speech signals, such as the positions of the speakers or the statistical characteristics of the speech signals. In practical applications, these prior knowledge is often unavailable or inaccurate, so the performance of these methods is usually limited.
[0004] In recent years, deep learning techniques have made significant progress in the field of speech separation. These methods usually use neural network models to directly predict individual speech sources from mixed audio signals, without the need for precise prior knowledge of the characteristics of the speech signals. However, these methods usually require a large amount of labeled data for training, and the enhanced performance in complex acoustic environments (such as strong reverberation, multiple noise interference, and severe speech interference) still needs to be improved.
[0005] After obtaining the preliminary separation results of the above two types of algorithms, the estimated target sound source signals can be further enhanced through post-filtering techniques. This technology uses prior information such as the acoustic characteristics of the original mixed signal to improve the quality of the separation results.
[0006] Existing post-filtering methods mainly adopt two strategies: one is to directly perform masking processing on the full frequency band or the separated frequency band, and the other is to adjust the zeros and poles based on specific coding frameworks such as LPC. However, these methods have two main limitations: first, they do not fully utilize the envelope characteristics of the estimated sound source signals and the mixed signals in different frequency bands to achieve accurate masking; second, they do not follow the human auditory perception mechanism, i.e., masking calculation based on sub-band time-domain envelope features. These limitations lead to the processed speech being prone to artificial noise and amplitude modulation quality degradation. SUMMARY
[0007] The present application aims to overcome the defects of the prior art that do not fully utilize the envelope characteristics of the estimated sound source signals and the mixed signals in different frequency bands to achieve accurate masking, or do not follow the human auditory perception mechanism.
[0008] To achieve the above object, the application provides a speech separation post-filtering method based on sub-band envelope characteristics, comprising:
[0009] designing an m-times frequency band-pass filter set, performing sub-band decomposition on the separated estimated sound source signal and the original mixed signal to obtain a frequency band division representation;
[0010] performing Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal to construct an analytic signal, calculating the instantaneous amplitude and removing high-frequency components through low-pass filtering to obtain a corresponding sub-band envelope;
[0011] calculating the ratio of the sub-band envelope of the estimated sound source signal to the original mixed signal as an initial masking value, limiting the initial masking value by setting a lower threshold, and performing nonlinear mapping on the initial masking value through a sine function to obtain an envelope masking coefficient;
[0012] applying the envelope masking coefficient to the corresponding sub-band signal of the estimated sound source to further suppress the interference component;
[0013] performing full-band reconstruction on each sub-band signal subjected to the masking processing to obtain an enhanced target sound source signal.
[0014] As an improvement of the above method, the designing of the m-times frequency band-pass filter set, performing sub-band decomposition on the separated estimated sound source signal and the original mixed signal to obtain a frequency band division representation, comprises:
[0015] obtaining the frequency division of the m-times frequency band-pass filter set, and designing an L-order Butterworth band-pass filter for each frequency band under the m-times frequency division;
[0016] performing zero-phase filtering on the original mixed signal and the estimated sound source signal respectively using the L-order Butterworth band-pass filter to obtain corresponding sub-band signals.
[0017] As an improvement of the above method, the value of m ranges from any rational number between 1 / 6 and 1, and the value of L ranges from a positive integer between 3 and 8.
[0018] As an improvement of the above method, the performing of Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal to construct an analytic signal, calculate the instantaneous amplitude and remove high-frequency components through low-pass filtering to obtain a corresponding sub-band envelope, comprises:
[0019] for each sub-band signal y i [n] of the original mixed signal, i is the sub-band number, the following operations are performed:
[0020] performing discrete Hilbert transform to obtain an orthogonal signal
[0021]
[0022] where DFT and IDFT denote the discrete Fourier transform and inverse transform, respectively; is the frequency index of the DFT; j is the imaginary unit; and sgn is the sign function;
[0023]
[0024] from the quadrature signal and the signal y i [n] to obtain the analytic signal, denoted as
[0025]
[0026] calculate the instantaneous amplitude of the analytic signal, denoted as
[0027]
[0028] where * denotes the conjugate;
[0029] perform low-pass filtering on the signal using a low-pass filter with a cutoff frequency of F l to filter out the high-frequency components and obtain the time-domain envelope of each subband of the original mixed signal;
[0030] for each subband signal x i [n] of the estimated sound source signal, perform the following operations:
[0031] perform a discrete Hilbert transform to obtain the quadrature signal
[0032]
[0033] from the quadrature signal and the signal x i [n] to obtain the analytic signal, denoted as
[0034]
[0035] calculate the instantaneous amplitude of the analytic signal, denoted as
[0036]
[0037] perform low-pass filtering on the signal using a low-pass filter with a cutoff frequency of F l to filter out the high-frequency components and obtain the time-domain envelope of each subband of the estimated sound source signal.
[0038] As an improvement of the above method, the cutoff frequency is any frequency in the range of 5-50 Hz.
[0039] As an improvement of the above method, the ratio of the estimated sound source signal and the sub-band envelope of the mixed signal is calculated as an initial masking value, the initial masking value is limited by setting a lower threshold, and the initial masking value is nonlinearly mapped by a sine function to obtain an envelope masking coefficient, comprising:
[0040] Calculate the initial masking value I' i [n]:
[0041]
[0042] Wherein, represents the i-th sub-band envelope of the estimated sound source signal; represents the i-th sub-band envelope of the mixed signal; I l represents the lower threshold;
[0043] The initial masking value is nonlinearly mapped by a sine function to obtain an envelope masking coefficient I i [n]:
[0044]
[0045] As an improvement of the above method, the lower threshold is set to any real number in the range of 0.001-0.05.
[0046] As an improvement of the above method, the envelope masking coefficient is applied to the corresponding sub-band signal of the estimated sound source to further suppress the interference component, comprising:
[0047] Calculate the i-th sub-band signal of the further suppressed interference component
[0048]
[0049] Wherein, I i [n] represents the envelope masking coefficient; x i [n] represents the m-octave filtered i-th estimated sound source sub-band signal.
[0050] As an improvement of the above method, the sub-band signals after masking processing are reconstructed in full frequency band to obtain an enhanced target sound source signal, comprising:
[0051] Reconstruct all sub-band signals in full frequency band to obtain an enhanced target signal, denoted as x enh [n]:
[0052]
[0053] wherein, represents the signal of the i-th sub-band after the masking processing; B represents the number of sub-bands under m times of frequency range division.
[0054] The application also provides a speech separation post-filtering system based on sub-band envelope characteristics, which is realized based on the above method, and the system comprises:
[0055] A sub-band decomposition module is configured to design an m times frequency range band-pass filter set, perform sub-band decomposition on the separated estimated sound source signal and the original mixed signal, and obtain the frequency band division representation thereof.
[0056] An obtained sub-band envelope module is configured to perform Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal, construct an analytic signal, calculate the instantaneous amplitude thereof, remove high-frequency components through low-pass filtering, and obtain the corresponding sub-band envelope.
[0057] An obtained envelope masking coefficient module is configured to calculate the ratio of the sub-band envelope of the estimated sound source signal to the original mixed signal as an initial masking value, limit the initial masking value by setting a lower threshold, and obtain the envelope masking coefficient by nonlinearly mapping the initial masking value through a sine function.
[0058] An interference component suppression module is configured to apply the envelope masking coefficient to the corresponding sub-band signal of the estimated sound source, and realize further suppression of the interference component.
[0059] A full-band reconstruction module is configured to perform full-band reconstruction on each sub-band signal after the masking processing, and obtain an enhanced target sound source signal.
[0060] Compared with the prior art, the application has the following advantages:
[0061] The method of the application effectively reduces the crosstalk component in the estimated sound source signal through sub-band envelope masking, and effectively improves the performance of the speech separation system in a complex acoustic environment. BRIEF DESCRIPTION OF DRAWINGS
[0062] Figure 1 The speech separation post-filtering method based on sub-band envelope characteristics is shown in the flowchart. DETAILED DESCRIPTION
[0063] The technical solutions of the application will be described in detail below with reference to the accompanying drawings.
[0064] Embodiment 1
[0065] As Figure 1 shown, the application provides a speech separation post-filtering method based on sub-band envelope characteristics, which comprises:
[0066] Step S101: design a m-times frequency band-pass filter set (m is preferably 1 / 3), and perform sub-band decomposition on the separated estimated sound source signal and the original mixed signal to obtain their frequency band division representations.
[0067] Specifically, a frequency division of a m-times frequency band-pass filter set is obtained, where m is any rational number in the range of 1 / 6 to 1, and is preferably 1 / 3; for each frequency band under the m-times frequency division, a L-order Butterworth band-pass filter is designed (L is a positive integer in the range of 3 to 8, and is preferably 4). The filter is used to perform zero-phase filtering on the original mixed signal y[n] and the estimated sound source signal x[n] respectively to obtain corresponding sub-band signals, denoted as y i [n] and x i [n] respectively, where i is the sub-band number.
[0068] Step S102: perform Hilbert transform on each sub-band of the separated estimated sound source signal and the original mixed signal to construct an analytic signal, calculate its instantaneous amplitude and remove high-frequency components through low-pass filtering to obtain the corresponding sub-band envelope.
[0069] Specifically, the sub-bands of the original mixed signal are denoted as y i [n], and the sub-bands of the estimated sound source signal are denoted as x i [n]. The same operation is performed on these sub-bands, and these sub-bands are collectively denoted as s[n], and the following operations are performed one by one: discrete Hilbert transform is performed on the sub-band signal s[n] to obtain the orthogonal signal s h [n];
[0070] The analytic signal is obtained according to the orthogonal signal s h [n] and the signal before transformation s[n], denoted as z[n]: z[n] = s[n] + js h [n]; the instantaneous amplitude of the analytic signal is calculated, denoted as s m [n]: (where s represents conjugate). A low-pass filter with a cutoff frequency of F l is designed, F l is any frequency in the range of 5-50 Hz, and is preferably 10 Hz, and s m [n] is low-pass filtered to filter out high-frequency components. After such processing, the time-domain envelope of each sub-band of the original mixed signal and the estimated sound source signal is obtained, denoted as where i is the sub-band number.
[0071] Step S103: calculate the ratio of the envelope of each sub-band of the estimated sound source signal to the mixed signal as the initial masking value, limit the value with a certain lower threshold, and perform nonlinear mapping on the masking value through a sine function to obtain the masking coefficient.
[0072] Specifically, the i-th sub-band envelope of the estimated sound source signal and the mixed signal is estimated as Calculate the initial envelope mask value I' i [n]
[0073]
[0074] where I l is a lower bound threshold of the mask value in the range of 0.001-0.05, preferably 0.01. The initial mask value calculated is nonlinearly mapped using a sinusoidal function to obtain the mask coefficient
[0075] Step S104: Apply the mask coefficient to the corresponding sub-band signal of the estimated sound source to further suppress the interference component.
[0076] Specifically, the i-th estimated sound source sub-band signal x i using m-octave filtering is multiplied by the corresponding mask coefficient I i to obtain the i-th sub-band signal after further suppression of the interference component
[0077] Step S105: Perform full-band reconstruction on each sub-band signal after masking to obtain the enhanced target sound source signal.
[0078] Specifically, let the i-th sub-band signal after masking be There are B sub-bands under m-octave division. Perform full-band reconstruction on all sub-band signals to obtain the enhanced target signal:
[0079] Embodiment 2
[0080] The application also provides a speech separation post-filtering system based on sub-band envelope features, which is realized based on the above method, and the system comprises:
[0081] A sub-band decomposition module is configured to design an m-octave bandpass filter set, perform sub-band decomposition on the separated estimated sound source signal and the original mixed signal, and obtain the frequency band division representation thereof;
[0082] An obtain sub-band envelope module is configured to perform Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal, construct an analytic signal, calculate the instantaneous amplitude thereof, remove high-frequency components through low-pass filtering, and obtain the corresponding sub-band envelope;
[0083] The envelope masking coefficient acquisition module is configured to calculate a ratio of the estimated sound source signal and the envelope of each subband of the original mixed signal as an initial masking value, limit the initial masking value by a lower threshold, and perform nonlinear mapping on the initial masking value by a sine function to obtain the envelope masking coefficient.
[0084] The interference component suppression module is configured to apply the envelope masking coefficient to the corresponding subband signal of the estimated sound source to further suppress the interference component.
[0085] The full-band reconstruction module is configured to perform full-band reconstruction on the masked subband signals to obtain the enhanced target sound source signal.
[0086] The present application can also provide a computer device, which comprises at least one processor, a memory, at least one network interface and a user interface. The components in the device are coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between the components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus.
[0087] The user interface can include a display, a keyboard or a clicking device, for example, a mouse, a trackball, a touchpad or a touch screen.
[0088] It can be appreciated that the memory in the embodiments disclosed in the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM, PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (Random Access Memory, RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (Dynamic RAM, DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (Synchlink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory described herein is intended to include but not limited to these and any other suitable types of memory.
[0089] In some embodiments, the memory stores elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and an application program.
[0090] Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program includes various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program for implementing the method of the embodiments of the present disclosure can be included in the application program.
[0091] In the above-described embodiments, the processor can also be used to execute the steps of the above method by invoking the programs or instructions stored in the memory, specifically, the programs or instructions stored in the application program.
[0092] execute the steps of the above method.
[0093] The method can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the method can be completed by an integrated logic circuit of hardware in the processor or by an instruction in the form of software. The processor can be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods disclosed above can be implemented or executed by the processor. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed above can be directly embodied as a hardware code executed by the processor or a combination of hardware and software modules in the processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage media is located in the storage memory, and the processor reads information in the storage memory and combines the hardware to complete the steps of the method.
[0094] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions described in the present application or a combination thereof.
[0095] For software implementation, the functions of the present application can be implemented by executing the functional modules (such as processes, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0096] The application further provides a nonvolatile storage medium for storing the computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and all of them should be covered in the scope of the claims of the present application.
Claims
1. A method for post-filtering of speech separation based on sub-band envelope characteristics, comprising: designing an m-times frequency band-pass filter bank to sub-band decompose an estimated sound source signal and an original mixed signal to obtain their frequency band division representations; m is an arbitrary rational number in a range of 1 / 6 to 1; performing Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal to construct an analytic signal, calculating its instantaneous amplitude and removing high frequency components by low-pass filtering to obtain a corresponding sub-band envelope; calculating a ratio of the sub-band envelopes of the estimated sound source signal and the original mixed signal as an initial masking value, limiting the initial masking value by a lower threshold, and performing nonlinear mapping on the initial masking value by a sinusoidal function to obtain an envelope masking coefficient; applying the envelope masking coefficient to a corresponding sub-band signal of the estimated sound source to further suppress interference components; full-band reconstructing each sub-band signal after the masking processing to obtain an enhanced target sound source signal.
2. The method of claim 1, wherein the sub-band envelope feature-based post-filtering method is characterized by, The designing of the m-times frequency band-pass filter bank to sub-band decompose the estimated sound source signal and the original mixed signal to obtain their frequency band division representations comprises: obtaining a frequency division of the m-times frequency band-pass filter bank, and designing an L-order Butterworth band-pass filter for each frequency band under the m-times frequency division; performing zero-phase filtering on the original mixed signal and the estimated sound source signal respectively by using the L-order Butterworth band-pass filter to obtain corresponding sub-band signals; L is a positive integer in a range of 3 to 8. 3.The sub-band envelope feature based speech separation post-filtering method according to claim 1, wherein, The performing of Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal to construct an analytic signal, calculate its instantaneous amplitude and remove high frequency components by low-pass filtering to obtain a corresponding sub-band envelope comprises: For each subband signal y i [n] of the original mixed signal, i is the subband index, and the following operations are performed: performing a discrete hilbert transform to obtain quadrature signals where DFT and IDFT denote the discrete Fourier transform and inverse transform, respectively; is the frequency index of the DFT; j is the imaginary unit; and sgn is the sign function. According to the orthogonal signal and the transformed signal y i [n] to obtain the analytic signal, denoted as The instantaneous amplitude of the analytical signal is calculated, denoted as wherein, * represents conjugate; The low-pass filter with a cut-off frequency of F l is used to perform low-pass filtering on the to filter out high-frequency components and obtain the time-domain envelope of each sub-band of the original mixed signal. For estimating the subband signals x i [n] of the sound source signal, the following operations are performed: performing a discrete Hilbert transform to obtain quadrature signals According to the orthogonal signal and the transformed signal x i [n] to obtain the analytic signal, denoted by The instantaneous amplitude of the analytical signal is calculated, denoted as The low-pass filtered signal is obtained by low-pass filtering the signal with a low-pass filter having a cut-off frequency of F l and removing high-frequency components to obtain a time-domain envelope of each sub-band of the estimated sound source signal. 4. The method of post-filtering for speech separation based on subband envelope characteristics according to claim 3, wherein, the cutoff frequency is an arbitrary frequency in a range of 5-50 Hz.
5. The method of claim 1, wherein the sub-band envelope feature-based post-filtering method is characterized by, The calculating of a ratio of the sub-band envelopes of the estimated sound source signal and the original mixed signal as an initial masking value, limiting the initial masking value by a lower threshold, and performing nonlinear mapping on the initial masking value by a sinusoidal function to obtain an envelope masking coefficient comprises: Calculating the initial masking value I 'i [n]: wherein, represents an estimated i-th subband envelope of the sound source signal; represents an i-th subband envelope of the mixed signal; l represents a set lower bound threshold; The initial masking value is nonlinearly mapped using a sinusoidal function to obtain an envelope masking coefficient I i [n]:
6. The method of speech separation post-filtering based on subband envelope characteristics according to claim 5, wherein, the lower threshold is an arbitrary real number in a range of 0.001-0.
05.
7. The method of post-filtering for speech separation based on subband envelope characteristics according to claim 1, wherein, The applying of the envelope masking coefficient to a corresponding sub-band signal of the estimated sound source to further suppress interference components comprises: The i-th subband signal with further suppression of interference components is calculated where I i [n] denotes the envelope masking coefficient; x i [n] denotes the m-octave filtered i-th estimated sound source subband signal.
8. The method of post-filtering for speech separation based on subband envelope characteristics according to claim 1, wherein, The full-band reconstructing of each sub-band signal after the masking processing to obtain an enhanced target sound source signal comprises: Full-band reconstruction is performed on all sub-band signals to obtain an enhanced target signal, denoted as x enh [n]: wherein, represents the signal of the i-th subband after the masking process; B represents the number of subbands under m octave division.
9. A post-filtering system for speech separation based on sub-band envelope features, implemented based on the method of any of claims 1-8, characterized in that, The system comprises: a sub-band decomposition module configured to design an m-times frequency band-pass filter bank to sub-band decompose an estimated sound source signal and an original mixed signal to obtain their frequency band division representations; m is an arbitrary rational number in a range of 1 / 6 to 1; an obtaining sub-band envelope module configured to perform Hilbert transform on each sub-band of the estimated sound source signal and the original mixed signal to construct an analytic signal, calculate its instantaneous amplitude and remove high frequency components by low-pass filtering to obtain a corresponding sub-band envelope; The envelope masking coefficient obtaining module is configured to calculate a ratio of the envelope of the estimated sound source signal and the envelope of the original mixed signal in each subband as an initial masking value, limit the initial masking value by a lower threshold, and perform nonlinear mapping on the initial masking value by a sine function to obtain the envelope masking coefficient. The interference component suppressing module is configured to apply the envelope masking coefficient to the corresponding subband signal of the estimated sound source to further suppress the interference component. The full-band reconstruction module is configured to perform full-band reconstruction on the masked subband signals to obtain the enhanced target sound source signal.
Citation Information
Patent Citations
Speech separation method based on fuzzy membership function
CN103325381A
Apparatus and method for improving the perceived quality of sound reproduction by combining active noise cancellation and perceptual noise compensation
CN104303227A