Dry and wet sound separation method and related device

By introducing decoupling technology for dry and wet sound states and spatiotemporal information, the problem of poor dry and wet sound separation in existing technologies has been solved, achieving high-precision separation in dynamic audio signals and adapting to different music styles and effects configurations.

CN121789709APending Publication Date: 2026-04-03IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing dry and wet audio separation technologies are prone to distortion or musicality loss in audio scenarios with diverse styles, complexities, and dynamic changes in status, resulting in poor separation performance.

Method used

By introducing dry and wet sound states, employing spatiotemporal information decoupling technology, and combining spectral and spatial information decoupling and separation processing, and utilizing multidimensional features and dynamic adjustment of system parameters, accurate separation of dry and wet sound signals can be achieved.

Benefits of technology

It improves the accuracy and adaptability of dry and wet audio separation, maintaining high-quality separation in dynamically changing audio signals and adapting to different music styles and effects configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789709A_ABST
    Figure CN121789709A_ABST
Patent Text Reader

Abstract

The invention discloses a dry and wet sound separation method and a related device, and relates to the technical field of audio processing, the dry and wet sound state is introduced, and the signal state is determined, so that the separation strategy better fits the dynamic change of the signal. On the basis, when dry and wet sound separation is realized, a spatio-temporal information decoupling technology is adopted, spectrum information which changes rapidly and relatively stable spatial information are respectively processed and cooperatively utilized, interference of spatial aliasing on spectrum separation is avoided, a separation result is verified and optimized through spatial correlation, the problem that stereo spatial information is not fully utilized is solved, and the separation efficiency is improved. Therefore, the spatial information of stereophonic sound is fully mined, and the overall separation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a method and apparatus for separating dry and wet audio. Background Technology

[0002] In the field of audio processing, "dry audio" refers to the original recording signal of vocals, lead instruments, etc., without any effects processing; while "wet audio" refers to the audio signal that has been processed by various effects units to add reverb, delay, compression, and environmental sound effects such as harmony. Simply put, "dry audio" retains the most natural state of the sound, while "wet audio" is the result of adding various post-production techniques to enhance or change its characteristics. In stereo music, "dry audio" and "wet audio" are often presented in a mixed form, and due to the dynamic adjustment of effects unit parameters and the influence of multi-channel spatial superposition effects, the boundary between the two becomes relatively blurred.

[0003] Dry and wet sound separation is particularly important for applications such as professional audio production or audio rendering for in-vehicle surround sound systems, because in these cases, extremely high sound quality and precise detail control are required to achieve the best listening experience.

[0004] Currently, the mainstream audio separation technologies mainly include noise reduction schemes based on spectral subtraction, blind separation schemes based on statistical modeling, and end-to-end separation schemes based on deep learning. However, when applied to audio dry and wet separation scenarios with diverse styles and dynamic changes in state, the current mainstream audio separation technologies are prone to serious distortion, residue, or musical loss in the separation results, resulting in poor separation performance.

[0005] Therefore, how to provide a dry and wet sound separation technology to improve the dry and wet sound separation effect has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] In view of the above problems, this application provides a method and related apparatus for separating wet and dry sound, so as to improve the separation effect of wet and dry sound. The specific solution is as follows:

[0007] The first aspect of this application provides a method for separating wet and dry audio, comprising:

[0008] Acquire stereo audio signals;

[0009] The stereo audio signal is preprocessed to obtain the complex time spectrum of each frame in the stereo audio signal;

[0010] Based on the complex time spectrum of each frame in the stereo audio signal, the dry and wet sound states of each frame in the stereo audio signal are determined.

[0011] By utilizing the dry and wet tones of each frame in the stereo audio signal, the complex time-spectrum of each frame in the stereo audio signal is decoupled and separated from the spectral and spatial information to obtain the dry tone signal and the wet tone signal.

[0012] In one possible implementation, determining the dry / wet tone state of each frame in the stereo audio signal based on the complex time-frequency spectrum of each frame in the stereo audio signal includes:

[0013] For each frame, feature extraction is performed on the complex time-spectrum of that frame to obtain the multidimensional features of that frame;

[0014] Based on the multidimensional features of each frame, the dry and wet sound states of each frame in the stereo audio signal are determined.

[0015] In one possible implementation, determining the dry / wet tone state of each frame in the stereo audio signal based on the multidimensional features of each frame includes:

[0016] For each frame, a reference frame is determined, which is a preset number of frames in the right field of view of the frame.

[0017] Based on the multidimensional features of the frame and the multidimensional features of the reference frame, the dry and wet sound state of the frame is detected to obtain the dry and wet sound state of the frame.

[0018] In one possible implementation, the step of decoupling and separating spectral and spatial information of the complex time-spectrum of each frame in the stereo audio signal by utilizing the dry and wet tones of each frame to obtain dry and wet tones includes:

[0019] For each frame in the stereo audio signal, the complex time spectrum of the frame is decoupled and separated from the spectral information and spatial information using the wet and dry states of the frame to obtain the wet spectrum and dry spectrum of the frame.

[0020] The wet and dry spectra of the frame are transformed in the time domain to obtain the dry and wet signals of the frame.

[0021] In one possible implementation, the step of decoupling and separating spectral and spatial information of the complex time spectrum of the frame using the wet and dry tones of the frame to obtain the wet and dry tones spectrum of the frame includes:

[0022] Obtain the system parameters corresponding to the dry and wet sound states of the frame;

[0023] Based on the system parameters, the complex time spectrum of the frame is decoupled and separated into spectral and spatial information to obtain the wet and dry spectra of the frame.

[0024] In one possible implementation, the step of decoupling and separating spectral and spatial information of the complex time spectrum of the frame according to the system parameters to obtain the wet spectrum and dry spectrum of the frame includes:

[0025] A filter is constructed based on the system parameters; the filter includes a signal filter and a spatial filter.

[0026] The complex time spectrum of the frame is processed using the signal filter to obtain spectral information, and the complex time spectrum of the frame is processed using the spatial filter to obtain spatial information, thereby obtaining the wet spectrum and dry spectrum of the frame.

[0027] In one possible implementation, the method further includes:

[0028] Detect whether the state of the frame has transitioned;

[0029] When a state transition of the frame is detected, the system parameters corresponding to the dry and wet sound states of the frame are dynamically adjusted.

[0030] A second aspect of this application provides a dry / wet sound separation device, comprising:

[0031] Acquisition unit, used to acquire stereo audio signals;

[0032] The preprocessing unit is used to preprocess the stereo audio signal to obtain the complex time spectrum of each frame in the stereo audio signal;

[0033] The dry / wet sound state determination unit is used to determine the dry / wet sound state of each frame in the stereo audio signal based on the complex time spectrum of each frame in the stereo audio signal.

[0034] The separation processing unit is used to decouple and separate the spectral information and spatial information of the complex time spectrum of each frame in the stereo audio signal by utilizing the dry and wet sound states of each frame in the stereo audio signal, so as to obtain the dry sound signal and the wet sound signal.

[0035] In one possible implementation, the dry / wet sound state determination unit includes:

[0036] The feature extraction unit is used to extract features from the complex time-frequency spectrum of each frame to obtain the multidimensional features of the frame.

[0037] The dry / wet sound state determination subunit is used to determine the dry / wet sound state of each frame in the stereo audio signal based on the multidimensional features of each frame.

[0038] In one possible implementation, the dry / wet sound state determining subunit is specifically used for:

[0039] For each frame, a reference frame is determined, which is a preset number of frames in the right field of view of the frame.

[0040] Based on the multidimensional features of the frame and the multidimensional features of the reference frame, the dry and wet sound state of the frame is detected to obtain the dry and wet sound state of the frame.

[0041] In one possible implementation, the separation processing unit includes:

[0042] The separation processing subunit is used to perform decoupling and separation processing of spectral information and spatial information of the complex time spectrum of each frame in the stereo audio signal by utilizing the dry and wet state of the frame, so as to obtain the wet spectrum and dry spectrum of the frame.

[0043] The time-domain conversion unit is used to perform time-domain conversion on the wet and dry spectra of the frame to obtain the dry and wet signals of the frame.

[0044] In one possible implementation, the separation processing subunit is specifically used for:

[0045] Obtain the system parameters corresponding to the dry and wet sound states of the frame;

[0046] Based on the system parameters, the complex time spectrum of the frame is decoupled and separated into spectral and spatial information to obtain the wet and dry spectra of the frame.

[0047] In one possible implementation, the separation processing subunit is specifically used for:

[0048] A filter is constructed based on the system parameters; the filter includes a signal filter and a spatial filter.

[0049] The complex time spectrum of the frame is processed using the signal filter to obtain spectral information, and the complex time spectrum of the frame is processed using the spatial filter to obtain spatial information, thereby obtaining the wet spectrum and dry spectrum of the frame.

[0050] In one possible implementation, the device further includes:

[0051] The system parameter dynamic adjustment unit is used to detect whether the state of the frame has changed; when the state of the frame is changed, the system parameters corresponding to the dry and wet sound states of the frame are dynamically adjusted.

[0052] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the dry and wet sound separation method described in the first aspect or any implementation thereof.

[0053] A fourth aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0054] The memory is used to store computer programs;

[0055] The processor is used to execute the computer program so that the electronic device can implement the dry and wet sound separation method of the first aspect or any implementation thereof.

[0056] The fifth aspect of this application provides a computer-readable storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform the dry and wet sound separation method described in the first aspect or any implementation thereof.

[0057] By employing the aforementioned technical solution, this application provides a method and related apparatus for separating dry and wet tones. This introduces the concept of dry and wet tones, and by determining the signal state, the separation strategy better aligns with dynamic signal changes. Furthermore, in achieving dry and wet tones separation, a spatiotemporal information decoupling technique is used to process and collaboratively utilize rapidly changing spectral information and relatively stable spatial information. This avoids interference from spatial aliasing on spectral separation and optimizes the separation results through spatial correlation verification, addressing the problem of insufficient utilization of stereo spatial information. This fully leverages the spatial information of stereo sound, thereby improving the overall separation effect. Attached Figure Description

[0058] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0059] Figure 1 A schematic flowchart of a method for separating wet and dry sound provided in an embodiment of this application;

[0060] Figure 2 This is a schematic diagram of the structure of a dry and wet sound separation device provided in an embodiment of this application;

[0061] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0062] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0063] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0064] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0065] In the field of audio processing, "dry audio" refers to the original recording signal of vocals, lead instruments, etc., without any effects processing; while "wet audio" refers to the audio signal that has been processed by various effects units to add reverb, delay, compression, and environmental sound effects such as harmony. Simply put, "dry audio" retains the most natural state of the sound, while "wet audio" is the result of adding various post-production techniques to enhance or change its characteristics. In stereo music, "dry audio" and "wet audio" are often presented in a mixed form, and due to the dynamic adjustment of effects unit parameters and the influence of multi-channel spatial superposition effects, the boundary between the two becomes relatively blurred.

[0066] Dry and wet sound separation is particularly important for applications such as professional audio production or audio rendering for in-vehicle surround sound systems, because in these cases, extremely high sound quality and precise detail control are required to achieve the best listening experience.

[0067] In the professional audio production field, with increasingly sophisticated music production and the widespread introduction of "personalized sound effect adjustment" functions by streaming platforms, the demand for high-precision and highly adaptable dry and wet sound separation technology is growing. In the application scenario of audio rendering in in-vehicle surround sound systems, dry and wet sound separation technology helps to accurately reshape the sound field layout or enhance speech clarity (such as highlighting human voices) in the limited space of the car interior, thereby meeting the dual needs of high-fidelity reproduction and personalized listening experience, and improving the driving experience.

[0068] However, due to the extreme diversity of musical styles, from a cappella with predominantly dry vocals to symphonies rich in wet vocal effects, the signal characteristics of different styles vary greatly. Moreover, music itself is highly dynamic, and the energy ratio and spatial properties between dry and wet vocal components can change drastically in a short period of time (such as the transient changes when drumbeats rise and fall), which places extremely high demands on the separation and processing of dry and wet vocals.

[0069] Currently, mainstream audio separation technologies mainly include spectral subtraction-based noise reduction schemes, statistical modeling-based blind separation schemes, and deep learning-based end-to-end separation schemes, among which:

[0070] The core idea of ​​the spectral subtraction-based noise reduction scheme is to treat wet sounds (especially the reverberation component) as a special type of "noise." The target dry sound spectrum is obtained by subtracting the estimated wet sound spectrum from the spectrum of the mixed signal. The key challenge of this scheme lies in accurately estimating the wet sound spectrum. Since wet and dry sounds highly overlap in the mixed signal, direct separation is difficult. Traditional methods typically assume that wet sounds have "statistical stationarity" (e.g., the energy of reverberation changes slowly over a short period), and their spectral energy can be approximated by taking the moving average or minimum value of the mixed signal spectrum. Generally, dry sounds in music (such as vocals and instrumental transients) are short-lived and sudden, and their energy is significantly higher than the subsequent, continuously decaying reverberation tail. After obtaining the estimated wet sound spectrum, this value is subtracted from the mixed signal spectrum in the time-frequency domain, and then the reconstructed dry sound signal is obtained through inverse short-time Fourier transform (ISTFT).

[0071] Blind source separation schemes based on statistical modeling include methods based on principal component analysis (PCA), independent component analysis (ICA), nonnegative matrix factorization (NMF), and late reverberation models (such as WPE). These methods typically assume that dry and wet sounds are statistically independent to some extent, utilizing multi-channel information (left and right channels of stereo) for blind source separation, attempting to find mutually independent or non-overlapping components and classify them as dry and wet sounds respectively. These methods are effective in certain specific music genres or arrangement styles.

[0072] End-to-end separation schemes based on deep learning construct convolutional neural networks (CNNs) or recurrent neural networks (RNNs) and train them with a large number of labeled mixed signal-to-dry / wet samples, enabling the model to learn the feature mapping relationship between dry and wet audio. In practical applications, the stereo mixed signal is input into the trained model, which directly outputs the separated dry and wet audio signals. This type of method models the mixing relationship between dry and wet audio in an end-to-end manner, and can recover a clean dry audio signal to a certain extent.

[0073] A comprehensive analysis of current mainstream audio separation technologies reveals that they are either limited by steady-state assumptions, rely on accurate spatial priors, or lack dynamic state recognition capabilities. When applied to audio dry-wet separation scenarios with diverse styles and dynamic state changes, these technologies are prone to causing severe distortion, residue, or musical loss in the separation results.

[0074] To address the aforementioned problems, this application provides a method for separating wet and dry audio. The wet and dry audio separation method of this application will be described in detail below with reference to the accompanying drawings.

[0075] Reference Figure 1 , Figure 1 This is a flowchart illustrating a method for separating wet and dry sound according to an embodiment of this application, as shown below. Figure 1 As shown in the embodiments of this application, a method for separating wet and dry sound may include the following steps, which are described in detail below.

[0076] S101: Acquire stereo audio signal;

[0077] In this application, the stereo audio signal includes two signals, namely the left channel signal L and the right channel signal R.

[0078] S102: Preprocess the stereo audio signal to obtain the complex time spectrum of each frame in the stereo audio signal;

[0079] In this application, the two signals in the stereo audio signal can be processed into frames according to a preset frame length and frame shift (e.g., 50%), and a Hanning window is applied to each frame obtained by the frame processing to reduce the interference of spectral leakage on subsequent feature extraction. By performing STFT on the windowed signal of each frame, the complex time spectrum X(k, m) of the windowed signal of each frame can be obtained, where k is the frequency point index and m is the frame index.

[0080] S103: Determine the dry / wet tone state of each frame in the stereo audio signal based on the complex time spectrum of each frame in the stereo audio signal;

[0081] Current mainstream audio separation techniques are mostly based on static features or global modeling strategies, failing to fully consider the dynamic evolution characteristics of audio signals over time, especially ignoring the transition process between dry and wet dominant states. For example, when a sound transitions from a pause to a sound, its spectrum, energy, and spatial coherence characteristics all undergo abrupt changes. If the system fails to capture this state change, it will be difficult to perform targeted processing.

[0082] To address this issue, this application introduces dry and wet sound states. By determining the signal state, the separation strategy is made to better align with the dynamic changes of the signal.

[0083] In this application, the dry and wet audio states include a dry audio stable state, a dry audio transitional state, a wet audio stable state, and a wet audio transitional state. The dry audio stable state refers to an audio signal in which the dry audio component dominates and the characteristic parameters (such as amplitude, spectral characteristics, and correlation coefficient of the two stereo signals) fluctuate very little over multiple consecutive frames. The dry audio transitional state is the transitional stage from dry audio to wet audio, in which the characteristic parameters show continuous changes. The wet audio stable state is similar to the dry audio stable state, referring to an audio signal in which the wet audio component dominates and the characteristic parameters (such as amplitude, spectral characteristics, and correlation coefficient of the two stereo signals) fluctuate very little over multiple consecutive frames. The wet audio transitional state is the transitional stage from wet audio to dry audio, in which the characteristic parameters show continuous changes.

[0084] Current mainstream audio separation techniques often require low-latency processing, and most systems model only the current frame and its historical information, failing to utilize observation information from several future frames. However, future information often provides crucial clues when estimating signal states or determining state boundaries. For example, if the signal energy continues to rise in the future, the current frame is very likely in the rising dry tone phase; if the spatial correlation increases in the future, it may enter a stable reverberation phase. Therefore, appropriately introducing "look-ahead" information is of significant value in enhancing the stability and accuracy of estimation.

[0085] Therefore, in one possible implementation, the determination of the dry / wet state of each frame in the stereo audio signal can be made by combining the complex time spectrum of the frame and its future frames.

[0086] S104: Using the dry and wet sound states of each frame in the stereo audio signal, the complex time spectrum of each frame in the stereo audio signal is decoupled and separated into spectral and spatial information to obtain dry sound signal and wet sound signal.

[0087] Current mainstream audio separation techniques often employ a unified modeling strategy for spectral and spatial information, ignoring the differences in their rates of change. In reality, spectral information (such as short-time energy and formants) typically changes rapidly, while spatial information (such as source location and delay structure) changes relatively slowly. Failure to differentiate between them can lead to delays in spectral estimation and fluctuations in spatial modeling, affecting the overall separation quality.

[0088] To address this issue, this application introduces the concept of decoupling spectral and spatial information. The complex time-frequency spectrum of each frame in the stereo audio signal is decoupled and separated to obtain dry and wet signals.

[0089] In one possible implementation, for spectral information, the rapidly changing components of spectral amplitude and phase in the time-frequency domain are extracted, and the changes in spectral information are tracked in real time to ensure accurate capture of the spectral features of dry and wet audio. For spatial information, the focus is on the spatial relative relationships of signals in the time-frequency domain, such as the slowly changing features of phase difference and amplitude ratio between different channels, and the spatial relative relationships are adjusted slowly to avoid audio spatial positioning confusion caused by rapid changes in spatial information.

[0090] It should be noted that the dry audio signal and wet audio signal in this application are in stereo format.

[0091] This application introduces dry and wet tone states, and by determining the signal state, the separation strategy is made more closely aligned with the dynamic changes of the signal. Based on this, when implementing dry and wet tone separation, a spatiotemporal information decoupling technique is employed. This process separately handles and coordinates rapidly changing spectral information with relatively stable spatial information, avoiding interference from spatial aliasing on spectral separation. Furthermore, spatial correlation is used to verify and optimize the separation results, addressing the problem of insufficient utilization of stereo spatial information. This fully leverages the spatial information of stereo sound, improving the overall separation effect.

[0092] In another embodiment of this application, the specific implementation of determining the dry / wet tone state of each frame in the stereo audio signal based on the complex time spectrum of each frame in the stereo audio signal is described. This method may include the following steps:

[0093] S201: For each frame, perform feature extraction on the complex time spectrum of that frame to obtain the multidimensional features of that frame;

[0094] In this application, for each frame, features can be extracted from the complex time-frequency spectrum of the frame to obtain the multidimensional features of the frame; in this application, there are no specific restrictions on the multidimensional features, which can be various combinations of time-domain features, frequency-domain features, and stereo features.

[0095] Time-domain characteristics include the root mean square (RMS) value of the signal, which represents the average energy of the signal. Dry / wet tones tend to be stable in their steady state, while transition states exhibit linear or exponential changes. Frequency-domain characteristics include the flatness of the spectral envelope, which is the ratio of the geometric mean to the arithmetic mean of the spectral amplitudes. Dry tones (such as pure vocals) have a lower spectral envelope flatness because their spectrum is concentrated in a specific frequency band, while wet tones have a higher spectral envelope flatness. Stereo characteristics include the cross-correlation coefficient C. For dry tones (such as vocal recordings with strong two-channel synchronization), C is close to 1, while for wet tones (due to reverberation / echo causing differences in two-channel delay), C decreases. Transition state C shows a continuous and gradual change (e.g., when transitioning from dry to wet tones, C gradually decreases from high to low).

[0096] For example, the multidimensional features include any combination of spectral envelope, amplitude change rate, phase features, cross-correlation spectrum, and cross-correlation coefficient.

[0097] S202: Based on the multidimensional features of each frame, determine the dry and wet sound state of each frame in the stereo audio signal.

[0098] In one possible implementation, determining the dry / wet tone state of each frame in the stereo audio signal based on the multidimensional features of each frame includes:

[0099] S2021: For each frame, determine the reference frame of the frame, wherein the reference frame of the frame is a preset number of frames in the right field of view of the frame;

[0100] In this application, for each frame, a preset number of frames in the right field of view of the frame are determined as reference frames for determining the dry and wet sound state of the frame.

[0101] In this application, the specific value of the preset quantity can be determined through extensive experimental verification based on the characteristics of the audio signal (such as sampling rate, signal change frequency, etc.) to achieve a balance between computational complexity and accuracy improvement. For example, for audio with a sampling rate of 44.1kHz, in a speech signal processing scenario, the next 5-20 frames (corresponding to a time of approximately 100-400ms) can be selected as the right visual field data.

[0102] S2022: Based on the multidimensional features of the frame and the multidimensional features of the reference frame, perform dry and wet sound state detection on the frame to obtain the dry and wet sound state of the frame.

[0103] In this application, a dry / wet sound state detection model can be pre-trained using the multidimensional features of each frame of the sample signal and the corresponding dry / wet sound state labels of each frame. As one possible approach, the multidimensional features of the frame and the multidimensional features of the reference frame can be input into the dry / wet sound state detection model to detect the dry / wet sound state of the frame and obtain the dry / wet sound state of the frame.

[0104] In this application, the introduction of several future frames (i.e., the right field of view) to assist in the state determination of the current frame can compensate for the lack of information in the state determination of a single frame. For example, the continuous characteristics of the wet reverberation effect in several future frames can help determine whether the current frame is in the process of transitioning from a wet reverberation transition state to a stable state. Therefore, for each frame, the introduction of several future frames (i.e., the right field of view) to assist in state determination significantly improves the accuracy of state determination for each frame, providing a more reliable basis for signal processing under different states in the future.

[0105] In another embodiment of this application, the specific implementation method of decoupling and separating spectral information and spatial information of the complex time spectrum of each frame in the stereo audio signal by utilizing the dry and wet sound states of each frame in the stereo audio signal to obtain dry sound signals and wet sound signals is described. This method may include the following steps:

[0106] S301: For each frame in the stereo audio signal, the complex time spectrum of the frame is decoupled and separated from the spectral information and spatial information using the wet and dry sound states of the frame to obtain the wet sound spectrum and the dry sound spectrum of the frame.

[0107] In one possible implementation of this application, for spectral information, the rapidly changing components of spectral amplitude and phase in the time-frequency domain are extracted, and an adaptive tracking algorithm (such as a Kalman filter-based spectrum tracking algorithm) is used to track the changes in spectral information in real time to ensure accurate capture of the spectral features of dry and wet audio. For spatial information, the focus is on the spatial relative relationships of signals in the time-frequency domain, such as the slowly changing features of phase difference and amplitude ratio between different channels. Methods such as moving average and low-pass filtering are used to process these features and slowly adjust the spatial relative relationships to avoid audio spatial positioning confusion caused by rapid changes in spatial information.

[0108] In one possible implementation, the step of decoupling and separating spectral and spatial information of the complex time spectrum of the frame using the wet and dry tones of the frame to obtain the wet and dry tones spectrum of the frame includes:

[0109] S3011: Obtain the system parameters corresponding to the dry and wet sound state of the frame;

[0110] Current mainstream audio separation technologies mostly use a single parameter or a uniform processing model to process all time segments, making it impossible to adjust parameter settings or switch processing strategies based on the current state of the signal (such as dry audio dominance, wet audio dominance, or a transitional state). This "one-size-fits-all" approach is prone to performance degradation under specific conditions, such as when signal characteristics are unstable during state transitions, and uniform parameters can lead to "fuzzy" estimations.

[0111] To address this issue, this application allows for the setting of different system parameters based on varying dry and wet sound conditions.

[0112] Different system parameters are defined for the dry tone stable state, dry tone transition state, wet tone stable state, and wet tone transition state. System parameters include, but are not limited to, filter coefficients (for spectral filtering), weighting factors (for fusion of dry and wet tone components), and spatial prediction parameters (for spatial prediction, etc.). For example, in the dry tone stable state, the filter coefficients are set to emphasize preserving the original spectral characteristics of the dry tone while suppressing the wet tone component; in the wet tone transition state, the weighting factor is dynamically adjusted, gradually increasing the weight of the wet tone component as the transition process progresses.

[0113] S3012: Based on the system parameters, the complex time spectrum of the frame is decoupled and separated into spectral and spatial information to obtain the wet spectrum and dry spectrum of the frame.

[0114] In one possible implementation, the step of decoupling and separating spectral and spatial information of the complex time spectrum of the frame according to the system parameters to obtain the wet and dry spectra of the frame includes: constructing a filter based on the system parameters; the filter includes a signal filter and a spatial filter; processing the spectral information of the complex time spectrum of the frame using the signal filter, and processing the spatial information of the complex time spectrum of the frame using the spatial filter to obtain the wet and dry spectra of the frame.

[0115] A common dry / wet audio separation scheme uses a filter to characterize the signal transmission path. For example, if the original signal X = dry audio signal D + wet audio signal A, we hope HX ≈ A; H can be understood as a filter. In this application, based on the idea of ​​decoupling spectral and spatial information, a spatial filter can be introduced, with H = Ha × Hs, where Ha is the signal filter and Hs is the spatial filter. This can reduce the impact on stable information.

[0116] In this application, in one possible implementation, the signal filter can be a filter using an adaptive tracking algorithm (such as Kalman filtering, Wiener filtering, etc.), and the spatial filter can be a filter using algorithms such as moving average or low-pass filtering. The specific settings can be based on the requirements of the scenario, and no limitations are imposed in this application.

[0117] For ease of understanding, assume that the original signal X = dry signal D + wet signal A, then the wet spectrum A = H×X, D = (1-H)×X, H=Ha×Hs, where Ha is the signal filter and Hs is the spatial filter.

[0118] S302: Perform time-domain transformation on the wet and dry spectra of the frame to obtain the dry and wet signals of the frame.

[0119] In this application, after separation, the dry and wet spectra are converted into time-domain signals by inverse short-time Fourier transform (ISTFT) to obtain dry and wet signals in stereo format.

[0120] In another embodiment of this application, the method further includes:

[0121] Detect whether the state of the frame has changed; when the state of the frame is changed, dynamically adjust the system parameters corresponding to the dry and wet sound states of the frame.

[0122] In this application, detecting whether the state of the frame has changed can be, for example, detecting whether the dry and wet sound state of the frame has changed from a dry sound stable state to a dry sound transition state, or from a wet sound stable state to a wet sound transition state, etc. This application does not impose any limitations on this.

[0123] When a state transition of the audio signal frame is detected (such as from a stable dry tone state to a transitional dry tone state), dynamic adjustment of system parameters is triggered.

[0124] In one possible implementation, the rules and rate of parameter adjustment between different states can be determined by constructing a state transition matrix or a state transition function. For example, the state transition function can calculate the step size of parameter adjustment based on the characteristic differences between the current state and the target state, so that the parameters transition smoothly and avoid problems such as distortion and noise in the audio signal caused by abrupt parameter changes, thus ensuring the continuity and stability of the dry and wet audio separation process.

[0125] In this application, the stability / transition state and its transition of the signal can be sensed in real time, and multi-frame system parameters can be adaptively adjusted, rather than using fixed system parameter settings. This allows the system to adapt to different effects types, parameter configurations and music styles, significantly improving the adaptability and overall robustness of the solution.

[0126] In summary, this application introduces dry and wet tone states, and by determining the signal state, the separation strategy is made more closely aligned with the dynamic changes of the signal. Based on this, when implementing dry and wet tone separation, a spatiotemporal information decoupling technique is employed to process and collaboratively utilize rapidly changing spectral information and relatively stable spatial information separately. This avoids interference from spatial aliasing on spectral separation and optimizes the separation results through spatial correlation verification, addressing the problem of insufficient utilization of stereo spatial information. This fully exploits the spatial information of stereo and improves the overall separation effect.

[0127] In addition, this application can also sense the stability / transition state and its transition of the signal in real time, and perform multi-frame system parameter adaptive adjustment instead of using fixed system parameter settings, so that it can adapt to different effects types, parameter configurations and music styles, significantly improving the adaptability and overall robustness of the solution.

[0128] The above describes a method for separating wet and dry audio signals according to embodiments of this application. The following describes the apparatus for performing the above-described method for separating wet and dry audio signals.

[0129] Please see Figure 2 , Figure 2 This is a schematic diagram of a dry and wet sound separation device provided in an embodiment of this application. Figure 2 As shown, the dry and wet sound separation device includes:

[0130] Acquisition unit 11 is used to acquire stereo audio signals;

[0131] Preprocessing unit 12 is used to preprocess the stereo audio signal to obtain the complex time spectrum of each frame in the stereo audio signal;

[0132] The dry / wet sound state determination unit 13 is used to determine the dry / wet sound state of each frame in the stereo audio signal based on the complex time spectrum of each frame in the stereo audio signal.

[0133] The separation processing unit 14 is used to decouple and separate the spectral information and spatial information of the complex time spectrum of each frame in the stereo audio signal by utilizing the dry and wet sound states of each frame in the stereo audio signal, so as to obtain the dry sound signal and the wet sound signal.

[0134] In one possible implementation, the dry / wet sound state determination unit includes:

[0135] The feature extraction unit is used to extract features from the complex time-frequency spectrum of each frame to obtain the multidimensional features of the frame.

[0136] The dry / wet sound state determination subunit is used to determine the dry / wet sound state of each frame in the stereo audio signal based on the multidimensional features of each frame.

[0137] In one possible implementation, the dry / wet sound state determining subunit is specifically used for:

[0138] For each frame, a reference frame is determined, which is a preset number of frames in the right field of view of the frame.

[0139] Based on the multidimensional features of the frame and the multidimensional features of the reference frame, the dry and wet sound state of the frame is detected to obtain the dry and wet sound state of the frame.

[0140] In one possible implementation, the separation processing unit includes:

[0141] The separation processing subunit is used to perform decoupling and separation processing of spectral information and spatial information of the complex time spectrum of each frame in the stereo audio signal by utilizing the dry and wet state of the frame, so as to obtain the wet spectrum and dry spectrum of the frame.

[0142] The time-domain conversion unit is used to perform time-domain conversion on the wet and dry spectra of the frame to obtain the dry and wet signals of the frame.

[0143] In one possible implementation, the separation processing subunit is specifically used for:

[0144] Obtain the system parameters corresponding to the dry and wet sound states of the frame;

[0145] Based on the system parameters, the complex time spectrum of the frame is decoupled and separated into spectral and spatial information to obtain the wet and dry spectra of the frame.

[0146] In one possible implementation, the separation processing subunit is specifically used for:

[0147] A filter is constructed based on the system parameters; the filter includes a signal filter and a spatial filter.

[0148] The complex time spectrum of the frame is processed using the signal filter to obtain spectral information, and the complex time spectrum of the frame is processed using the spatial filter to obtain spatial information, thereby obtaining the wet spectrum and dry spectrum of the frame.

[0149] In one possible implementation, the device further includes:

[0150] The system parameter dynamic adjustment unit is used to detect whether the state of the frame has changed; when the state of the frame is changed, the system parameters corresponding to the dry and wet sound states of the frame are dynamically adjusted.

[0151] Each unit in the aforementioned dry and wet sound separation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each unit.

[0152] This application also provides an electronic device in its embodiments. (See reference...) Figure 3 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 3The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0153] like Figure 3 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0154] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0155] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the dry and wet sound separation methods provided in this application.

[0156] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the dry and wet sound separation methods provided in this application.

[0157] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0159] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0160] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for separating wet and dry audio, characterized in that, include: Acquire stereo audio signals; The stereo audio signal is preprocessed to obtain the complex time spectrum of each frame in the stereo audio signal; Based on the complex time spectrum of each frame in the stereo audio signal, the dry and wet sound states of each frame in the stereo audio signal are determined. By utilizing the dry and wet tones of each frame in the stereo audio signal, the complex time-spectrum of each frame in the stereo audio signal is decoupled and separated from the spectral and spatial information to obtain the dry tone signal and the wet tone signal.

2. The method according to claim 1, characterized in that, Determining the dry / wet tone state of each frame in the stereo audio signal based on the complex time-spectrum of each frame includes: For each frame, feature extraction is performed on the complex time-spectrum of that frame to obtain the multidimensional features of that frame; Based on the multidimensional features of each frame, the dry and wet sound states of each frame in the stereo audio signal are determined.

3. The method according to claim 2, characterized in that, The determination of the dry / wet tone state of each frame in the stereo audio signal based on the multidimensional features of each frame includes: For each frame, a reference frame is determined, which is a preset number of frames in the right field of view of the frame. Based on the multidimensional features of the frame and the multidimensional features of the reference frame, the dry and wet sound state of the frame is detected to obtain the dry and wet sound state of the frame.

4. The method according to claim 1, characterized in that, The process of decoupling and separating spectral and spatial information of the complex time-spectrum of each frame in the stereo audio signal by utilizing the dry and wet tones of each frame to obtain dry and wet tones includes: For each frame in the stereo audio signal, the complex time spectrum of the frame is decoupled and separated from the spectral information and spatial information using the wet and dry states of the frame to obtain the wet spectrum and dry spectrum of the frame. The wet and dry spectra of the frame are transformed in the time domain to obtain the dry and wet signals of the frame.

5. The method according to claim 4, characterized in that, The process of decoupling and separating spectral and spatial information of the complex time spectrum of the frame using the wet and dry tones of the frame to obtain the wet and dry tones spectrum of the frame includes: Obtain the system parameters corresponding to the dry and wet sound states of the frame; Based on the system parameters, the complex time spectrum of the frame is decoupled and separated into spectral and spatial information to obtain the wet and dry spectra of the frame.

6. The method according to claim 5, characterized in that, The step of decoupling and separating the spectral and spatial information of the complex time spectrum of the frame according to the system parameters to obtain the wet spectrum and dry spectrum of the frame includes: A filter is constructed based on the system parameters; the filter includes a signal filter and a spatial filter. The complex time spectrum of the frame is processed using the signal filter to obtain spectral information, and the complex time spectrum of the frame is processed using the spatial filter to obtain spatial information, thereby obtaining the wet spectrum and dry spectrum of the frame.

7. The method according to claim 5, characterized in that, The method further includes: Detect whether the state of the frame has transitioned; When a state transition of the frame is detected, the system parameters corresponding to the dry and wet sound states of the frame are dynamically adjusted.

8. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the dry and wet sound separation method as described in any one of claims 1 to 7.

9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the dry and wet sound separation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the dry and wet sound separation method as described in any one of claims 1 to 7.