Audio processing method and device, electronic equipment and storage medium
By combining frequency domain feature separation with time domain, the artifact problem caused by the reliance of audio separation on phase information in existing technologies is solved, computational complexity is reduced, and the quality of audio separation and user experience are improved.
Patent Information
- Application Number
- CN202410316730.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, time-frequency masking technology requires a large number of parameters in audio separation and is highly dependent on phase information, which leads to artifacts in the separation results and affects the separation quality. In addition, traditional methods have high computational complexity in real-time applications and it is difficult to achieve an ideal user experience.
Frequency domain features are used to separate human voice audio signals and non-human voice audio signals, and combined with time domain signals. Through frequency domain feature extraction, mask network and deep latent diffusion model, the number of parameters and computational complexity are reduced, and the combination of frequency domain and time domain is used to reduce information loss and improve separation performance.
It effectively avoids the loss in the audio signal separation process, reduces computational complexity, improves user experience, and improves the quality and efficiency of audio separation.
Smart Images

Figure CN120708641A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an audio processing method, device, electronic device, and storage medium. Background Art
[0002] With the continuous development of computer technology, the way of audio processing is constantly improving.
[0003] In the related art, time-frequency (TF) masking technology is used to perform audio separation on the mixed audio data to be processed containing human voices. However, the time-frequency masking technology involved in the related art often requires a large number of parameters and is highly dependent on phase information, which leads to artifacts in the separated sound source and affects the quality of the separation result. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides an audio processing method, device, electronic device and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio processing method, comprising: obtaining an audio signal to be processed, wherein the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal; extracting frequency domain features of the audio signal to be processed, and based on the frequency domain features, performing separation of the frequency domain human voice audio signal and the frequency domain non-human voice audio signal to obtain a frequency domain human voice audio signal and a frequency domain non-human voice audio signal; obtaining a time domain human voice audio signal and a time domain non-human voice audio signal that have been separated; and obtaining a human voice audio signal and a non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal, and the time domain non-human voice audio signal.
[0006] In one embodiment, the separating of frequency domain human voice audio signals and frequency domain non-human voice audio signals based on the frequency domain features to obtain frequency domain human voice audio signals and frequency domain non-human voice audio signals includes: performing time-series segmentation on the frequency domain features to obtain multiple frequency domain features; for each frequency domain feature in the multiple frequency domain features, separating the frequency domain human voice audio signals and frequency domain non-human voice audio signals based on the frequency domain features to obtain multiple frequency domain human voice audio signals and multiple frequency domain non-human voice audio signals; the obtaining of human voice audio signals and non-human voice audio signals based on the frequency domain human voice audio signals, the frequency domain non-human voice audio signals, the time domain human voice audio signals, and the time domain non-human voice audio signals includes: synthesizing the multiple frequency domain human voice audio signals according to the time sequence of the time domain human voice audio signals to obtain a human voice audio signal, and synthesizing the multiple frequency domain non-human voice audio signals according to the time sequence of the time domain non-human voice audio signals to obtain a non-human voice audio signal.
[0007] In one embodiment, the separation of frequency domain human voice audio signals and frequency domain non-human voice audio signals based on frequency domain features includes: determining the signal hiding features of the audio signal in the frequency domain based on the frequency domain features; determining the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal based on the signal hiding features and the masking network; and obtaining the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal.
[0008] In one embodiment, the determining of the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal based on the signal hidden features and the mask network includes: determining the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal based on the signal hidden features and the mask network; using a jump connection method to determine the attention weight for distinguishing the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the attention mechanism; based on the attention weight, weighting the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal to obtain the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal.
[0009] In one embodiment, the determining of the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal based on the signal hidden features and the mask network includes: activating the signal hidden features based on the hyperbolic tangent activation function; separating the signal hidden features based on the normalized exponential function to obtain the separated frequency domain human voice audio signal hidden features and frequency domain non-human voice audio signal hidden features; determining the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal based on the frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features, and the mask network.
[0010] In one embodiment, determining the signal hiding features of the audio signal in the frequency domain based on the frequency domain features includes: inputting the frequency domain features into multiple gated recurrent unit (GRU) block processing layers, and using the multiple gated recurrent unit (GRU) block processing layers to determine the signal hiding features of the audio signal to be processed in the frequency domain using sequence information.
[0011] In one embodiment, the method further includes: converting the human voice audio signal into the frequency domain, and extracting a single-frame human voice audio signal from the human voice audio signal in the frequency domain; and repairing the human voice audio signal based on the single-frame human voice audio signal.
[0012] In one embodiment, repairing the human voice audio signal based on the single-frame human voice audio signal includes: compressing the single-frame human voice audio signal based on an autoencoder model composed of multiple perceptrons to obtain a single-frame spectrogram; repairing the human voice audio signal based on the single-frame spectrogram and a Unet diffusion model with a stable diffusion structure; wherein the Unet diffusion model with a stable diffusion structure is obtained by enhancing the channel dimension and the time domain dimension through a cross-attention mechanism.
[0013] In one embodiment, the repairing of the human voice audio signal based on the single-frame spectrogram and the Unet diffusion model of the Stable Diffusion structure includes: taking the human voice fundamental frequency as a feature condition, inputting it into the Unet diffusion model of the Stable Diffusion structure, guiding the denoising process of the single-frame spectrogram to obtain a denoised single-frame spectrogram; and obtaining a repaired human voice audio signal based on the denoised spectrogram and the network model composed of dilated convolution.
[0014] According to a second aspect of an embodiment of the present disclosure, an audio processing device is provided, including: an acquisition unit, configured to acquire an audio signal to be processed, wherein the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal; acquiring a time-domain human voice audio signal and a time-domain non-human voice audio signal that have been separated; a separation unit, configured to extract frequency-domain features of the audio signal to be processed, and to separate the frequency-domain human voice audio signal and the frequency-domain non-human voice audio signal based on the frequency-domain features to obtain a frequency-domain human voice audio signal and a frequency-domain non-human voice audio signal; and a processing unit, configured to obtain a human voice audio signal and a non-human voice audio signal based on the frequency-domain human voice audio signal, the frequency-domain non-human voice audio signal, the time-domain human voice audio signal, and the time-domain non-human voice audio signal.
[0015] In one embodiment, the separation unit separates the frequency domain human voice audio signal and the frequency domain non-human voice audio signal in the following manner to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal: performing time-series segmentation on the frequency domain features to obtain multiple frequency domain features; for each frequency domain feature in the multiple frequency domain features, performing frequency domain human voice audio signal and frequency domain non-human voice audio signal separation based on the frequency domain features to obtain multiple frequency domain human voice audio signals and multiple frequency domain non-human voice audio signals; obtaining the human voice audio signal and the non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal and the time domain non-human voice audio signal includes: synthesizing the multiple frequency domain human voice audio signals according to the time sequence of the time domain human voice audio signal to obtain the human voice audio signal, and synthesizing the multiple frequency domain non-human voice audio signals according to the time sequence of the time domain non-human voice audio signal to obtain the non-human voice audio signal.
[0016] In one embodiment, the separation unit separates the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain features in the following manner: based on the frequency domain features, determining the signal hidden features of the audio signal in the frequency domain; based on the signal hidden features and the mask network, determining the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal; based on the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal, obtaining the frequency domain human voice audio signal and the frequency domain non-human voice audio signal.
[0017] In one embodiment, the separation unit determines the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal based on the signal hidden features and the mask network in the following manner: based on the signal hidden features and the mask network, the initial frequency domain human voice audio signal significant features and the initial frequency domain non-human voice audio signal significant features are determined; using a jump connection method, the attention weight for distinguishing the frequency domain human voice audio signal and the frequency domain non-human voice audio signal is determined based on the attention mechanism; based on the attention weight, the initial frequency domain human voice audio signal significant features and the initial frequency domain non-human voice audio signal significant features are weighted to obtain the frequency domain human voice audio signal significant features and the frequency domain non-human voice audio signal significant features.
[0018] In one embodiment, the separation unit determines the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal based on the signal hidden features and the mask network in the following manner: based on the hyperbolic tangent activation function, the signal hidden features are activated; based on the normalized exponential function, the signal hidden features are separated to obtain the separated frequency domain human voice audio signal hidden features and frequency domain non-human voice audio signal hidden features; based on the frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features, and the mask network, the initial frequency domain human voice audio signal significant features and the initial frequency domain non-human voice audio signal significant features are determined.
[0019] In one embodiment, the separation unit determines the signal hiding features of the audio signal in the frequency domain based on the frequency domain features in the following manner: the frequency domain features are input into a plurality of gated recurrent unit (GRU) block processing layers, and the plurality of gated recurrent unit (GRU) block processing layers use sequence information to determine the signal hiding features of the audio signal to be processed in the frequency domain.
[0020] In one embodiment, the processing unit is further used to: convert the human voice audio signal into the frequency domain, and extract a single-frame human voice audio signal from the human voice audio signal in the frequency domain; and repair the human voice audio signal based on the single-frame human voice audio signal.
[0021] In one embodiment, the processing unit repairs the human voice audio signal based on the single-frame human voice audio signal in the following manner: compressing the single-frame human voice audio signal based on an autoencoder model composed of multiple perceptrons to obtain a single-frame spectrogram; repairing the human voice audio signal based on the single-frame spectrogram and a Unet diffusion model with a stable diffusion structure; wherein the Unet diffusion model with a stable diffusion structure is obtained by enhancing the channel dimension and the time domain dimension through a cross-attention mechanism.
[0022] In one embodiment, the processing unit repairs the human voice audio signal based on the single-frame spectrogram and the Unet diffusion model of the StableDiffusion structure in the following manner: the human voice fundamental frequency is used as a feature condition and input into the Unet diffusion model of the StableDiffusion structure to guide the denoising process of the single-frame spectrogram to obtain a denoised single-frame spectrogram; based on the denoised spectrogram and the network model composed of dilated convolution, a repaired human voice audio signal is obtained.
[0023] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to execute the audio processing method in the first aspect or any one of the implementations of the first aspect.
[0024] According to the fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by the processor of the terminal, the terminal device is able to execute the audio processing method in the first aspect or any one of the implementations of the first aspect.
[0025] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: separating human voice audio signals and non-human voice audio signals according to frequency domain characteristics, and combining the existing human voice audio signals and non-human voice audio signals in the time domain to reduce the processing of audio signals in the time domain, thereby reducing the parameters in the audio signal processing process, reducing the computational complexity, and combining the frequency domain human voice audio signals and frequency domain non-human voice audio signals separated in the frequency domain with the time domain human voice audio signals and time domain non-human voice audio signals separated in the time domain, effectively avoiding the loss of audio signals during the separation process and improving the user experience.
[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0028] Figure 1 This is a flowchart of an audio processing method according to an exemplary embodiment.
[0029] Figure 2 The figure is a flowchart of an audio segmentation method according to an exemplary embodiment.
[0030] Figure 3 The figure is a flowchart of an audio separation method according to an exemplary embodiment.
[0031] Figure 4 The present invention is a flowchart showing a method for determining significant features of an audio signal according to an exemplary embodiment.
[0032] Figure 5 The flowchart of a method for determining significant features of an initial audio signal is shown according to an exemplary embodiment.
[0033] Figure 6 The figure is a flow chart showing a method of obtaining hidden features according to an exemplary embodiment.
[0034] Figure 7 The figure is a schematic diagram of an audio signal separation network according to an exemplary embodiment.
[0035] Figure 8 The figure is a flowchart of a vocal restoration method according to an exemplary embodiment.
[0036] Figure 9 The figure is a flowchart of single-frame vocal restoration according to an exemplary embodiment.
[0037] Figure 10 The figure is a flowchart showing an application of a single-frame vocal restoration model according to an exemplary embodiment.
[0038] Figure 11 The figure is a schematic diagram of music source separation according to an exemplary embodiment.
[0039] Figure 12 The figure is a block diagram of an audio processing device according to an exemplary embodiment.
[0040] Figure 13 The present invention is a block diagram of a device for audio processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0041] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present disclosure.
[0042] The audio processing method provided in the embodiments of the present disclosure is applied to mixed audio processing scenarios. The audio processing method involved in the embodiments of the present disclosure is mainly applied to scenarios for separating music audio.
[0043] Targeting the real-time application needs of mobile devices, especially development on end-side devices such as terminals and tablets. The time-frequency domain methods in related technologies often rely on complex phase information when processing audio signals, which not only increases the difficulty of estimation, but may also introduce artifacts in the separated sources due to inaccurate information, thereby limiting the quality of audio processing. In addition, time-frequency masking technology often involves a large number of parameters, which not only makes model training difficult, but also leads to slow operation in real-time applications, making it difficult to achieve an ideal user experience. At the same time, in the construction of time series information, although the traditional bidirectional long short-term memory (BLSTM) network can effectively process time domain information, it requires a large number of model parameters and a long input sequence, which leads to delay problems when processing long sequences and occupies a large amount of memory. It does not meet the requirements of real-time applications for response speed and resource consumption, and has low computational complexity.
[0044] In the related art, the technology for sound source separation generally adopts the time-frequency masking method to separate the music source of the processed audio signal. However, since the time-frequency masking method adopted in the related art usually requires modeling the time domain and the two dimensions in the frequency domain separately, it requires a huge number of parameters and is highly dependent on phase information, resulting in artifacts in the separated sound source.
[0045] Furthermore, related technologies use a method of separating the audio signal to be processed in the time domain, such as a basic network based on signal estimation and a U-Net-style encoder / decoder network. However, for example, TasNet, as a basic signal estimation network, can effectively estimate the principal components of the input and masking values, but because it processes very short input blocks, it is difficult to fully utilize information of long time intervals. In addition, the method of separating the audio signal to be processed in the time domain requires extremely high accuracy of the separated audio data, which leads to data loss in the separated audio data.
[0046] In addition, for the human voice source separated from the audio signal to be processed, related technologies use computer algorithms and signal processing technologies to repair the separated damaged human voice, and apply it to speech enhancement, speech noise reduction, speech separation and speech reconstruction.
[0047] However, the current vocal restoration technology cannot achieve the expected restoration effect in complex speech scenes and environments. In addition, the current vocal restoration technology in the processed audio has excessive processing, resulting in a decrease in the quality of the speech signal.
[0048] In view of this, the present disclosure provides a method for audio processing, which separates the audio data of the audio signal to be processed in the frequency domain to obtain separated human voice data and non-human voice audio data, and repairs the human voice based on the separated human voice data based on the deep latent diffusion model, thereby reducing the number and complexity of model parameters of the deepest latent layer, and utilizing deep latent masking to significantly improve the source separation performance.
[0049] In addition, the audio features are processed in a combined time-frequency domain manner. In the time dimension, modeling is performed based on the continuity of time in the time domain, and in the energy (or amplitude) dimension, processing is performed based on the frequency domain. This ensures that modeling is performed based on the single dimension in the frequency domain and the validity and continuity of time in the time domain, thereby reducing the parameters of model training. In addition, the time domain characteristics of audio data are combined to reduce the information loss and delay of the processed audio data.
[0050] In an exemplary embodiment of the present disclosure, a method for separating an audio signal in the frequency domain and combining the separated frequency domain audio signal with an existing audio signal in the time domain to obtain a separated audio signal is adopted. Figure 1 The way shown, Figure 1 The flowchart of an audio processing method according to an exemplary embodiment is shown. The audio processing method includes the following steps.
[0051] In step S11, an audio signal to be processed is obtained.
[0052] In one implementation, the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal.
[0053] In step S12, the frequency domain features of the audio signal to be processed are extracted, and based on the frequency domain features, the frequency domain human voice audio signal and the frequency domain non-human voice audio signal are separated to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal.
[0054] In one embodiment, the frequency domain features of the audio signal to be processed are extracted based on the frequency domain features of the audio signal, and the obtained frequency domain human voice audio signal and frequency domain non-human voice audio signal are separated based on the frequency domain features of the audio signal to be processed using a frequency domain audio signal separation method.
[0055] In step S13, the separated time-domain human voice audio signal and the time-domain non-human voice audio signal are acquired.
[0056] In one implementation, according to a method for separating time-domain audio signals, a time-domain human voice audio signal and a time-domain non-human voice audio signal that have been separated in the time domain are obtained.
[0057] The present disclosure does not limit the method of separating the time-domain human voice audio signal and the time-domain non-human voice audio signal according to the time-domain audio signal separation method.
[0058] In step S14, a human voice audio signal and a non-human voice audio signal are obtained based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal and the time domain non-human voice audio signal.
[0059] In one embodiment, the frequency domain human voice audio signal and the frequency domain non-human voice audio signal are converted into time domain features, and the audio signal converted into time domain features is combined with the time domain human voice audio signal and the time domain non-human voice audio signal to obtain a human voice audio signal and a non-human voice audio signal.
[0060] In one embodiment, the characteristics of frequency domain audio processing are utilized to separate the audio data of the audio signal to be processed in the frequency domain according to a single dimension in the frequency domain, so as to obtain human voice data and non-human voice audio data separated in the frequency domain, thereby reducing the calculation parameters in the audio separation process, and combining the human voice data and non-human voice audio data separated in the frequency domain with the human voice data and non-human voice audio data separated in the time domain, thereby avoiding the loss of separated audio data caused by only using the time domain to separate the audio data or only using the frequency domain to separate the audio data, improving the performance of audio separation, reducing the calculation complexity, and improving the user experience.
[0061] In an exemplary embodiment of the present disclosure, when separating audio signals in the frequency domain, the audio signal to be processed is first segmented according to its frequency domain characteristics. The audio signal segmentation method is as follows: Figure 2 As shown, Figure 2 The flowchart of an audio segmentation method according to an exemplary embodiment includes the following steps.
[0062] In step S21, the frequency domain features are time-series segmented to obtain a plurality of frequency domain features.
[0063] In one implementation, time series segmentation is performed based on the frequency domain characteristics of the audio signal to be processed, thereby obtaining a plurality of frequency domain characteristics of the audio signal to be processed after time series segmentation.
[0064] Here, time segmentation can be understood as dividing the audio signal to be processed into multiple time segments based on the time sequence information of the audio signal. For each time segment, the frequency domain features are extracted to obtain the frequency domain features of multiple audio signals to be processed.
[0065] In step S22, for each of the multiple frequency domain features, the frequency domain human voice audio signal and the frequency domain non-human voice audio signal are separated based on the frequency domain features to obtain multiple frequency domain human voice audio signals and multiple frequency domain non-human voice audio signals.
[0066] In one implementation, frequency domain features of a plurality of audio signals to be processed are separated to obtain a plurality of separated frequency domain human voice audio signals and frequency domain non-human voice audio signals.
[0067] In one embodiment, multiple frequency domain human voice audio signals are synthesized according to the time sequence of the time domain human voice audio signals to obtain a human voice audio signal, and multiple frequency domain non-human voice audio signals are synthesized according to the time sequence of the time domain non-human voice audio signals to obtain a non-human voice audio signal.
[0068] In one implementation, audio signals for different time segments are processed independently using a time segmentation approach, thereby improving processing flexibility.
[0069] In an exemplary embodiment of the present disclosure, after obtaining a plurality of frequency domain audio signals to be processed, the plurality of frequency domain audio signals to be processed are separated and processed respectively, wherein the separation processing method is as follows: Figure 3 As shown, Figure 3 The flowchart of an audio separation method according to an exemplary embodiment includes the following steps.
[0070] In step S31, based on the frequency domain features, the signal hiding features of the audio signal in the frequency domain are determined.
[0071] In one implementation, the frequency domain features are input into an encoder, and the signal hidden features of the audio signal to be processed in the frequency domain are obtained according to the encoding process.
[0072] Among them, the hidden features of the signal can be understood as the result of the audio signal being processed by feature mapping.
[0073] In step S32, based on the signal hidden features and the mask network, the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal are determined.
[0074] In one embodiment, based on the obtained signal hidden features and the mask network, separation processing is performed to obtain the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal in a single audio.
[0075] The mask network can be understood as a network consisting of multiple gated recurrent unit (GRU) block processing layers and different activation functions, which is used to calculate the mask of the audio signal that is the same as the original audio signal.
[0076] In step S33, based on the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal, a frequency domain human voice audio signal and a frequency domain non-human voice audio signal are obtained.
[0077] In one embodiment, the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal in the obtained single audio are decoded and processed, and the significant features of the human voice audio signal and the significant features of the frequency domain non-human voice audio signal are separated in the decoder to obtain the separated frequency domain human voice audio signal and the frequency domain non-human voice audio signal.
[0078] In one embodiment, based on the frequency domain characteristics of the audio signal, in the frequency domain, the characteristics of a single dimension of the audio signal, such as the amplitude characteristics, are used to perform audio signal separation operations, reduce calculation parameters, reduce calculation complexity, and exist in the form of a single audio before separating the final frequency domain human voice audio signal and the frequency domain non-human voice audio signal, thereby reducing the channels for audio separation and improving the efficiency of audio separation.
[0079] In an exemplary embodiment of the present disclosure, in the separation operation of the audio signal, the method of determining the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal by using the signal hidden features and the mask network is as follows: Figure 4 As shown, Figure 4 The present invention is a flowchart showing a method for determining significant features of an audio signal according to an exemplary embodiment, which includes the following steps.
[0080] In step S41, based on the signal hidden features and the mask network, the significant features of the initial frequency-domain human voice audio signal and the significant features of the initial frequency-domain non-human voice audio signal are determined.
[0081] In one embodiment, signal hiding techniques are used to extract hidden features from audio signals. These hidden features can reflect the intrinsic properties of the audio signal. Then, a masking network is combined to preprocess the audio signal to remove redundant information and highlight the signal's primary features. Based on this, the salient features of the initial frequency-domain human voice audio signal and the salient features of the initial frequency-domain non-human voice audio signal are determined.
[0082] In step S42, a skip connection method is adopted to determine the attention weight for distinguishing the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the attention mechanism.
[0083] In one implementation, feature information at different levels is fused through skip connections, and then the attention weight is calculated based on the attention mechanism.
[0084] Among them, the attention weight can be understood as reflecting the importance of different features in distinguishing human voice audio signals from non-human voice audio signals.
[0085] In step S43, based on the attention weight, the initial frequency domain human voice audio signal salient features and the initial frequency domain non-human voice audio signal salient features are weighted to obtain the frequency domain human voice audio signal salient features and the frequency domain non-human voice audio signal salient features.
[0086] In one embodiment, the significant features of the initial frequency-domain human voice audio signal and the initial frequency-domain non-human voice audio signal are weighted according to the attention weights. This weighting can highlight important features that distinguish human voice audio signals from non-human voice audio signals and suppress secondary features, thereby obtaining more accurate significant features of the frequency-domain human voice audio signal and the frequency-domain non-human voice audio signal.
[0087] In one embodiment, the signal hidden features and mask network are used to accurately extract the intrinsic properties and main features of the audio signal, thereby improving the accuracy of feature extraction. By introducing the attention mechanism and jump connection, the importance of different features in distinguishing between human voice audio signals and non-human voice audio signals can be automatically learned, further improving the accuracy of feature distinction. The weighted processing can highlight important features and suppress secondary features, making the extracted significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal more representative.
[0088] In an exemplary embodiment of the present disclosure, in the frequency domain non-human voice audio signal obtained by separating the frequency domain human voice audio signal, the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal are first separated, wherein the method for determining the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal by using signal hidden features and a mask network is as follows: Figure 5 As shown, Figure 5 The flowchart of a method for determining significant features of an initial audio signal according to an exemplary embodiment includes the following steps.
[0089] In step S51, the signal hidden features are activated based on the hyperbolic tangent activation function.
[0090] In one embodiment, the nonlinear characteristics of the hyperbolic tangent activation function are used to activate the signal hidden features, ensuring that the activated signal hidden features can better reflect the essential properties and characteristics of the audio signal.
[0091] In step S52, based on the normalized exponential function, the signal hidden features are separated to obtain separated frequency domain human voice audio signal hidden features and frequency domain non-human voice audio signal hidden features.
[0092] In one embodiment, the activated signal hidden features are probabilistically processed through the normalized exponential Softmax function, and the features are mapped to the probability space, thereby achieving the separation of the signal hidden features, and obtaining the separated frequency domain human voice audio signal hidden features and frequency domain non-human voice audio signal hidden features.
[0093] In step S53, based on the frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features, and the mask network, the initial frequency domain human voice audio signal salient features and the initial frequency domain non-human voice audio signal salient features are determined.
[0094] In one embodiment, the powerful learning and representation capabilities of the mask network are utilized to further process the separated hidden features of the frequency domain human voice audio signal and the hidden features of the frequency domain non-human voice audio signal to extract characteristics that can more accurately describe the frequency domain human voice audio signal and the non-human voice audio signal.
[0095] In one implementation, the hyperbolic tangent activation function is used to activate the signal's hidden features, improving their expressiveness and richness. A normalized exponential Softmax function is then used to probabilistically process the activated signal's hidden features, enabling accurate separation of these features. A mask network is then used to further process the separated hidden features, extracting more representative and significant features, thereby improving the accuracy and efficiency of audio signal processing.
[0096] In the exemplary embodiment of the present disclosure, before performing the audio signal separation operation, the signal hidden features of the audio signal to be processed in the frequency domain are obtained according to the frequency domain features of the audio signal. The specific implementation method is as follows: Figure 6 As shown, Figure 6 A flowchart of obtaining hidden features according to an exemplary embodiment is shown, which includes the following steps.
[0097] In step S61, the frequency domain features are input to multiple GRU block processing layers.
[0098] In one embodiment, the audio signal to be processed is converted into frequency domain features and input as input data to multiple gated recurrent unit (GRU) block processing layers.
[0099] In step S62, based on multiple GRU block processing layers, the sequence information is used to obtain the signal hidden features of the audio signal to be processed in the frequency domain.
[0100] In one embodiment, the GRU can control the transmission and forgetting of information through an internal gating mechanism, thereby effectively utilizing the sequential information of the audio signal in the frequency domain. After processing by multiple GRU block processing layers, the hidden features of the audio signal to be processed in the frequency domain can be extracted.
[0101] In one embodiment, the characteristics of the gated recurrent unit (GRU) can be used to capture the temporal dependencies and hidden features of audio signals in the frequency domain, thereby improving the accuracy and effectiveness of audio signal feature extraction. In addition, by superimposing multiple GRU block processing layers, the ability to extract hidden features of audio signals in the frequency domain can be further enhanced, thereby improving the richness and representativeness of the features.
[0102] In one embodiment, the separation operation of a mixed audio signal as an audio signal to be processed is taken as an example to specifically describe the separation of the mixed audio signal to be processed using a frequency domain separation network. Figure 7 As shown, Figure 7 The figure is a schematic diagram of an audio signal separation network according to an exemplary embodiment.
[0103] In one embodiment, the mixed audio signal is input to the encoding layer (Encoder) of the separation network.
[0104] The mixed audio signal to be processed is pre-converted into a frequency domain audio signal, the converted mixed audio signal is time-series segmented to obtain a target source of a unit mixed audio signal, the target source of the unit mixed audio signal is input into a coding layer, and a masking vector after encoding the target source is obtained using the coding layer.
[0105] The masking vector can be understood as performing feature mapping processing on the frequency domain audio signal using the coding layer, and the frequency domain audio signal processed by the feature mapping is called the masking vector.
[0106] In one embodiment, the mask vector after target source encoding is input to the processing layer of the frequency domain separation network.
[0107] In the processing layer of the separation network, the normalization (i.e., BN&LN) method is used in this layer, and the normalized data is input into the processing layer based on multiple gated recurrent units (GRU) blocks, where the masking vector is processed using sequence information to obtain the masked features of the masked vector, and the obtained masked features are input into the MLP layer based on the multi-layer perceptron. Through the separation processing in the MLP layer, the target hidden features contained in the masked features and the mixed hidden features in the potential domain are obtained.
[0108] The target hidden features in a single audio data obtained after Softmax activation processing in the MLP layer and the mixed hidden features in the potential domain are input into the Attention module of the MLP layer.
[0109] Before inputting into this module, a mask network composed of a GRU block processing layer and an MLP layer is used to calculate the target hidden features of the obtained single audio data and the mixed hidden features in the latent domain. The target hidden features and the mixed hidden features in the latent domain are multiplied by the target source of the unit mixed audio signal respectively to obtain the masks of the target hidden features and the mixed hidden features in the latent domain.
[0110] In one embodiment, the target hidden features and the mixed hidden features in the latent domain are respectively multiplied with the target source of the unit mixed audio signal to be processed, and the masks of the obtained target hidden features and the mixed hidden features in the latent domain are input into the Attention module of the MLP layer, and masking processing is performed in the form of dot multiplication in the matrix to obtain the salient features of the target source of the unit mixed audio signal, and the salient features of the target source exist in the form of a single audio data.
[0111] Among them, the masking processing in the form of dot product in the matrix can be understood as multiplying the masking vector after the target source is encoded with the mask of the target hidden feature and the mask of the mixed hidden feature in the latent domain, thereby obtaining the salient features of the target source of the unit mixed audio signal.
[0112] Among them, the MLP layer consists of three fully connected FC layers and an Attention module, and uses hyperbolic tangent as the activation function of the first and second layers, and Softmax as the activation function of the third layer. The Attention module adopts the attention mechanism to obtain weights according to the jump connection. The target hidden feature can be the human voice feature in the mixed audio signal to be processed, and the mixed hidden feature in the potential domain can be the accompaniment feature in the mixed audio signal to be processed.
[0113] In one embodiment, the salient features processed by the processing layer are used to reconstruct the unit audio signal through a decoding layer (Decoder).
[0114] The salient features of the target source of the unit mixed audio signal (i.e., human voice features and non-human voice features) are converted into audio signals in the decoding layer, and the converted audio signals are combined with the time domain audio signal to obtain the reconstructed unit audio signal.
[0115] In the exemplary embodiment of the present disclosure, after the frequency domain separation network is used to perform separation operation on the audio signal to be processed, the separated human voice is restored according to the separation result. The specific process is as follows: Figure 8 As shown, Figure 8 FIG. 4 is a flowchart of a vocal restoration method according to an exemplary embodiment, and the specific steps are as follows.
[0116] In step S81 , the human voice audio signal is converted into a frequency domain, and a single-frame human voice audio signal is extracted from the human voice audio signal in the frequency domain.
[0117] In one embodiment, the acquired human voice audio data is subjected to Fourier transform to be converted into the frequency domain, and the audio data converted into the frequency domain is subjected to audio signal extraction to obtain a single-frame human voice audio signal.
[0118] In step S82, the human voice audio signal is repaired based on the single-frame human voice audio signal.
[0119] In one implementation, a single-frame signal is repaired based on a single-frame human voice audio signal, and the repaired single-frame signal is synthesized to obtain a repaired human voice.
[0120] In one implementation, a restoration operation of a single-frame human voice audio signal is utilized to improve processing flexibility.
[0121] In the exemplary embodiment of the present disclosure, when a single-frame human voice audio signal is used to repair the human voice, the following method is used: Figure 9 The way shown, Figure 9 FIG. 4 is a flowchart of single-frame vocal restoration according to an exemplary embodiment, and the specific steps are as follows.
[0122] In step S91, based on the autoencoder model composed of multiple perceptrons, a single-frame human voice audio signal is compressed to obtain a single-frame spectrogram.
[0123] In one embodiment, a single-frame human voice audio signal is compressed using an autoencoder model composed of multiple perceptrons to extract the main features of the audio signal and generate a corresponding single-frame spectrogram.
[0124] In step S92, the human voice audio signal is restored based on the single-frame spectrogram and the Unet diffusion model of the Stable Diffusion structure.
[0125] The Unet diffusion model with a Stable Diffusion structure is obtained by enhancing the channel and time domain dimensions through a cross-attention mechanism. In one embodiment, the extracted single-frame spectrogram is input into a Unet diffusion model based on the Stable Diffusion structure for repair processing. Furthermore, a cross-attention mechanism is introduced to enhance the Stable Diffusion Unet diffusion model.
[0126] In one embodiment, a single-frame human voice audio signal is compressed and processed according to an autoencoder model composed of multiple perceptrons to obtain a spectrogram. The spectrogram can reflect the frequency distribution and energy of the audio signal, providing an important basis for subsequent human voice audio restoration. In addition, by utilizing the characteristics of the Stable Diffusion structure, the distorted or damaged parts in the spectrogram are gradually restored through a diffusion process, thereby realizing the restoration of human voice audio. Through the cross-attention mechanism, the model can better capture the correlation and dependency between different frequency components in the spectrogram, thereby more accurately repairing the distorted or damaged parts in the human voice audio, improving the restoration quality, accuracy and efficiency.
[0127] In the exemplary embodiment of the present disclosure, the operation of audio restoration and vocal restoration is performed according to the Unet diffusion model of the Stable Diffusion structure, such as Figure 10 As shown, Figure 10 This is a flowchart showing an application of a single-frame vocal restoration model according to an exemplary embodiment, and the specific steps are as follows.
[0128] In step S101, the fundamental frequency of human voice is used as a feature condition and input into the Unet diffusion model of the Stable Diffusion structure to guide the denoising process of the single-frame spectrogram to obtain the denoised single-frame spectrogram.
[0129] In one embodiment, the fundamental frequency information of the human voice audio is extracted as a characteristic condition, and then the characteristic condition is input into the Unet diffusion model based on the Stable Diffusion structure to ensure that the model denoises the spectrum graph according to the guidance of the human voice fundamental frequency characteristic condition during the denoising generation process, thereby obtaining the denoised spectrum graph.
[0130] In step S102, a restored human voice audio signal is obtained based on the denoised spectrogram and the network model formed by dilated convolution.
[0131] In one embodiment, the denoised spectrogram is used as input, and a network composed of dilated convolutions is used to capture feature information of different scales in the spectrogram, thereby extracting single-frame spectrum dimensions. The obtained single-frame spectrum dimensions are converted back into time domain signals through inverse Fourier transform, thereby obtaining the restored human voice.
[0132] In one implementation, a Unet diffusion model with a Stable Diffusion structure is used for denoising, accurately removing noise components from the spectrogram and improving restoration quality. The fundamental frequency of the human voice is input into the model as a feature condition to guide the denoising process, making the restored voice closer to the original sound. A network composed of dilated convolutions is introduced to capture feature information at different scales in the spectrogram, further improving the accuracy and efficiency of restoration.
[0133] In one implementation, the operation of voice restoration is described by taking a human voice signal separated from a mixed audio signal to be processed as an example.
[0134] In one embodiment, the extracted vocal features are received and a single-frame FFT transformation is performed on the vocal features to obtain a spectrogram feature of the single-frame audio.
[0135] The separated human voice is subjected to a single-frame Fourier operation to convert it into a frequency domain audio signal, and the spectrogram features of the single-frame audio in the frequency domain audio signal are obtained, and the spectrogram features are compressed using an automatic encoding model.
[0136] In this embodiment, the automatic encoding model includes multiple perceptrons, and compresses the spectrum graph features through a linear layer to reduce the computational complexity.
[0137] In one embodiment, the compressed spectrogram features are fed into a Unet diffusion model based on a stable diffusion fusion structure for diffusion generation. During the diffusion generation process, the fundamental frequency features of the human voice are used as conditions to guide the Unet to perform denoising generation. The latent features output by the Unet are restored to the original single-frame spectrum dimensions through a network composed of dilated convolutions, and the restored spectrum dimensions are inverse Fourier transformed to obtain a repaired single-frame audio signal.
[0138] Among them, the Unet model is enhanced by a cross-attention mechanism to perform conditional generation more flexibly.
[0139] In one embodiment, the repaired vocal signal or the accompaniment signal is output, or the repaired vocal signal and the accompaniment signal are synthesized and then output.
[0140] In an exemplary embodiment of the present disclosure, the processing of music source in an audio signal is taken as an example to specifically describe the operations of audio separation and audio restoration involved in this application. Figure 11 As shown, Figure 11 The figure is a schematic diagram of music source separation according to an exemplary embodiment.
[0141] In one embodiment, a piece of original music sound is obtained and input into a separation network for separation processing. The operation of separation processing according to the separation network is described in detail. Figure 7 The music data after separation is separated into vocals and accompaniment / single-track instruments. The separated vocals are repaired using a diffusion repair network. The specific operation of repairing using the diffusion repair network is detailed in the above description of the vocal repair operation. The repaired vocals are obtained. According to the requirements of different solutions, the repaired vocals or accompaniment / single-track instruments are output as single-track audio, that is, a single-track audio stream is output. Alternatively, the repaired vocals and accompaniment / single-track instruments are re-synthesized to output the repaired music.
[0142] In the disclosed embodiment, the processed audio signal is post-processed according to different requirements. The post-processing includes directly outputting multi-track audio and outputting the repaired synthesized audio. If the output is used for post-processing of streaming media or other sound balance processing, there is no need to synthesize the multiple tracks, and the audio can be directly output to the audio stream of other applications. If the need is to simply repair the vocals of a musical piece, the generated multiple tracks or accompaniment are re-synthesized with the repaired vocals into a singing voice, and the result is output.
[0143] In one embodiment, the audio signal processing process mentioned above can be processed based on a model, wherein the training of the model is based on Figure 7 In the process of separating the music source in the manner shown, the following formula is mainly used:
[0144] According to the loss phenomenon in the time domain and frequency domain during the separation operation, the multi-domain loss (MDL) is used to calculate the loss in the time domain or frequency domain respectively, where the redundant loss is defined as follows:
[0145]
[0146] In addition, since there is a loss of vocals and accompaniment in the process of audio separation and vocal restoration, the combined loss is used for calculation, where the combined loss is defined as follows:
[0147]
[0148] Where N represents the number of sound sources in the target signal. The combination loss reflects all possible combination losses of the estimated signals and prevents the estimated target signal from leaking into other signals. Multi-domain loss is implemented using multi-resolution FFT loss as the frequency domain criterion and MSE loss as the time domain criterion. The FFT size of the multi-resolution STFT is set to (256, 512, 1024). The contrast loss L-con is also used as the latent embedding feature in the bottleneck module and is defined as follows:
[0149]
[0150] Where z represents the hidden feature. In addition, m∈M≡1,2,····,N×K, P(m) is the set of positive samples of the remaining blocks belonging to the same target, and N(i) is the set of negative samples. The positive and negative pairs are obtained as follows: the input audio signal W∈2×L is divided into K blocks of equal size, and each target source produces K blocks. For the purpose of contrastive learning, each target embedding can be regarded as having K-1 positive samples and N×(K-1) negative samples, where N is the number of target sources. A simple convolutional network is used to convert the latent embedding to the new domain, which makes it easier to apply contrastive loss. This loss function is only used during training. The total training loss is as follows:
[0151] Loss = α·L1 + (1-α)·Lcon
[0152] In the experiment, we set α = 0.95
[0153] In one embodiment, the operation of separating the sound source in the frequency domain ensures that the model used in the separation operation is smaller, reduces the computational complexity and reduces the model parameters of the model required for the separation operation, improves the separation performance, and improves the efficiency of audio processing for the real-time separation and repair operations of audio.
[0154] In the disclosed embodiment, the audio processing method is used to reduce the parameters required for audio processing, achieve real-time processing capabilities while maintaining a small model size, improve classification performance, and the separation results meet the requirements of subjective evaluation. At the same time, the model can run in real time, reducing computational complexity and the number of model parameters, achieving an optimal balance between music source separation performance and efficiency. As for the repair effect, a good repair effect can be achieved for 1-2 seconds of sound loss, meeting the needs of music repair. At the same time, the repaired vocals and separated multi-track audio can also meet the downstream task requirements of other applications, improving the user experience.
[0155] In the disclosed embodiments, audio processing can help remove noise, improve recording quality, and enhance the sound quality of songs, thereby enhancing the user's musical experience. Furthermore, when background music or background music is required for video production, the sound quality can be improved, thereby enhancing the video's interest.
[0156] In the above embodiments of the present disclosure, the data set involved may be MUSDB18. The MUSDB18 data set evaluates the proposed model of the present technology. The data set contains 100 training set tracks and 50 test set tracks, all sampled at a rate of 44.1kHz. For training, 75 tracks were randomly selected from the training set, and the remaining tracks were used for verification purposes. In order to evaluate the performance of the model, the signal-to-distortion ratio (SDR) calculated using the BSS evaluation metric was used. The above data is for illustrative purposes only and is not specifically limited in this disclosure.
[0157] Based on the same concept, an embodiment of the present disclosure also provides an audio processing device.
[0158] It is understandable that the audio processing device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of the various examples disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.
[0159] Figure 12 FIG. 1 is a block diagram of an audio processing device according to an exemplary embodiment. Figure 12 The device 100 includes an acquisition unit 101, a separation unit 102 and a processing unit 103.
[0160] The acquisition unit 101 is used to acquire an audio signal to be processed, where the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal; and acquire a time-domain human voice audio signal and a time-domain non-human voice audio signal that have been separated.
[0161] The separation unit 102 is used to extract the frequency domain features of the audio signal to be processed, and based on the frequency domain features, separate the frequency domain human voice audio signal and the frequency domain non-human voice audio signal to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal.
[0162] The processing unit 103 is configured to obtain a human voice audio signal and a non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal, and the time domain non-human voice audio signal.
[0163] In one embodiment, the separation unit 102 separates the frequency domain human voice audio signal and the frequency domain non-human voice audio signal in the following manner to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal: performing time-series segmentation on the frequency domain features to obtain multiple frequency domain features; for each frequency domain feature in the multiple frequency domain features, performing frequency domain human voice audio signal and frequency domain non-human voice audio signal separation based on the frequency domain features to obtain multiple frequency domain human voice audio signals and multiple frequency domain non-human voice audio signals; obtaining the human voice audio signal and the non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal and the time domain non-human voice audio signal, including: synthesizing the multiple frequency domain human voice audio signals according to the time sequence of the time domain human voice audio signal to obtain the human voice audio signal, and synthesizing the multiple frequency domain non-human voice audio signals according to the time sequence of the time domain non-human voice audio signal to obtain the non-human voice audio signal.
[0164] In one embodiment, the separation unit 102 separates the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain features in the following manner: based on the frequency domain features, determining the signal hidden features of the audio signal in the frequency domain; based on the signal hidden features and the masking network, determining the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal; based on the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal, obtaining the frequency domain human voice audio signal and the frequency domain non-human voice audio signal.
[0165] In one embodiment, the separation unit 102 determines the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal based on the signal hidden features and the mask network in the following manner: based on the signal hidden features and the mask network, the initial frequency domain human voice audio signal significant features and the initial frequency domain non-human voice audio signal significant features are determined; using a jump connection method, the attention weight for distinguishing the frequency domain human voice audio signal and the frequency domain non-human voice audio signal is determined based on the attention mechanism; based on the attention weight, the initial frequency domain human voice audio signal significant features and the initial frequency domain non-human voice audio signal significant features are weighted to obtain the frequency domain human voice audio signal significant features and the frequency domain non-human voice audio signal significant features.
[0166] In one embodiment, the separation unit determines the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal based on the signal hidden features and the mask network in the following manner: based on the hyperbolic tangent activation function, the signal hidden features are activated; based on the normalized exponential function, the signal hidden features are separated to obtain the separated frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features; based on the frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features, and the mask network, the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal are determined.
[0167] In one embodiment, the separation unit 102 determines the signal hidden features of the audio signal in the frequency domain based on the frequency domain features in the following manner: the frequency domain features are input into a plurality of gated recurrent unit (GRU) block processing layers, and the plurality of gated recurrent unit (GRU) block processing layers use sequence information to determine the signal hidden features of the audio signal to be processed in the frequency domain.
[0168] In one embodiment, the processing unit 103 is further configured to: convert the human voice audio signal into the frequency domain, extract a single-frame human voice audio signal from the human voice audio signal in the frequency domain; and repair the human voice audio signal based on the single-frame human voice audio signal.
[0169] In one embodiment, the processing unit 103 repairs the human voice audio signal based on a single-frame human voice audio signal in the following manner: compressing the single-frame human voice audio signal based on an autoencoder model composed of multiple perceptrons to obtain a single-frame spectrogram; repairing the human voice audio signal based on the single-frame spectrogram and a Unet diffusion model with a stable diffusion structure; wherein the Unet diffusion model with a stable diffusion structure is obtained by enhancing the channel dimension and the time domain dimension through a cross-attention mechanism.
[0170] In one embodiment, the processing unit 103 repairs the human voice audio signal based on a single-frame spectrogram and a Unet diffusion model with a Stable Diffusion structure in the following manner: the fundamental frequency of the human voice is used as a feature condition and input into the Unet diffusion model with a Stable Diffusion structure to guide the denoising process of the single-frame spectrogram to obtain a denoised single-frame spectrogram; based on the denoised spectrogram and a network model composed of dilated convolution, a repaired human voice audio signal is obtained.
[0171] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0172] Figure 13 FIG2 is a block diagram of an apparatus 200 for audio processing according to an exemplary embodiment. For example, the apparatus 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0173] Reference Figure 13 , apparatus 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .
[0174] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.
[0175] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0176] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 200.
[0177] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0178] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0179] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0180] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect changes in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0181] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0182] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0183] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, which can be executed by the processor 220 of the apparatus 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0184] It is understood that in this disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of related objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0185] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.
[0186] It can be further understood that the terms "center", "longitudinal", "lateral", "front", "back", "up", "down", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation.
[0187] It is further understood that, unless otherwise specified, “connection” includes a direct connection where there are no other components between the two elements, and also includes an indirect connection where there are other elements between the two elements.
[0188] It is further understood that although operations are described in a particular order in the drawings in the embodiments of the present disclosure, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.
[0189] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0190] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.
Claims
1. An audio processing method, characterized in that: The method comprises: Acquire an audio signal to be processed, wherein the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal; Extracting frequency domain features of the audio signal to be processed, and separating the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain features to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal; Obtaining the separated time-domain human voice audio signal and time-domain non-human voice audio signal; A human voice audio signal and a non-human voice audio signal are obtained based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal, and the time domain non-human voice audio signal.
2. The method according to claim 1, characterized in that The separating of the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain features to obtain the frequency domain human voice audio signal and the frequency domain non-human voice audio signal includes: Performing time series segmentation on the frequency domain features to obtain multiple frequency domain features; For each of the multiple frequency domain features, separate the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain feature to obtain multiple frequency domain human voice audio signals and multiple frequency domain non-human voice audio signals; The obtaining of the human voice audio signal and the non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal, and the time domain non-human voice audio signal includes: The multiple frequency domain human voice audio signals are synthesized according to the time sequence of the time domain human voice audio signal to obtain a human voice audio signal, and the multiple frequency domain non-human voice audio signals are synthesized according to the time sequence of the time domain non-human voice audio signal to obtain a non-human voice audio signal.
3. The method according to claim 1 or 2, characterized in that The separating of the frequency domain human voice audio signal and the frequency domain non-human voice audio signal based on the frequency domain features includes: Determining a hidden signal feature of the audio signal in the frequency domain based on the frequency domain feature; Determining significant features of a frequency-domain human voice audio signal and significant features of a frequency-domain non-human voice audio signal based on the signal hidden features and the masking network; Based on the significant features of the frequency domain human voice audio signal and the significant features of the frequency domain non-human voice audio signal, a frequency domain human voice audio signal and a frequency domain non-human voice audio signal are obtained.
4. The method according to claim 3, characterized in that The determining of significant features of a frequency domain human voice audio signal and significant features of a frequency domain non-human voice audio signal based on the signal hidden features and the mask network includes: Determining significant features of an initial frequency-domain human voice audio signal and significant features of an initial frequency-domain non-human voice audio signal based on the signal hidden features and the mask network; Using a skip connection method, the attention weights for distinguishing frequency-domain human voice audio signals from frequency-domain non-human voice audio signals are determined based on the attention mechanism. Based on the attention weight, the initial frequency-domain human voice audio signal salient features and the initial frequency-domain non-human voice audio signal salient features are weighted to obtain the frequency-domain human voice audio signal salient features and the frequency-domain non-human voice audio signal salient features.
5. The method according to claim 4, characterized in that The determining, based on the signal hidden features and the mask network, the significant features of the initial frequency domain human voice audio signal and the significant features of the initial frequency domain non-human voice audio signal, comprises: activating the signal hidden features based on a hyperbolic tangent activation function; Separating the signal hidden features based on a normalized exponential function to obtain separated frequency-domain human voice audio signal hidden features and frequency-domain non-human voice audio signal hidden features; Based on the frequency domain human voice audio signal hidden features and the frequency domain non-human voice audio signal hidden features, and a masking network, the initial frequency domain human voice audio signal salient features and the initial frequency domain non-human voice audio signal salient features are determined.
6. The method according to claim 3, characterized in that The determining, based on the frequency domain feature, a signal hiding feature of the audio signal in the frequency domain includes: The frequency domain features are input into a plurality of gated recurrent unit (GRU) block processing layers, and the plurality of gated recurrent unit (GRU) block processing layers use sequence information to determine the signal hidden features of the audio signal to be processed in the frequency domain.
7. The method according to claim 1, characterized in that The method further comprises: Converting the human voice audio signal into a frequency domain, and extracting a single-frame human voice audio signal from the human voice audio signal in the frequency domain; Based on the single-frame human voice audio signal, the human voice audio signal is repaired.
8. The method according to claim 7, characterized in that The repairing of the human voice audio signal based on the single-frame human voice audio signal includes: Based on an autoencoder model composed of multiple perceptrons, compress the single-frame human voice audio signal to obtain a single-frame spectrogram; Repairing the human voice audio signal based on the single-frame spectrogram and the Unet diffusion model of the Stable Diffusion structure; Among them, the Unet diffusion model of the Stable Diffusion structure is obtained by enhancing the channel dimension and time domain dimension through the cross attention mechanism.
9. The method according to claim 8, characterized in that The repairing of the human voice audio signal based on the single-frame spectrogram and the Unet diffusion model of the StableDiffusion structure includes: The fundamental frequency of the human voice is used as a feature condition and input into a Unet diffusion model with a stable diffusion structure to guide the denoising process of the single-frame spectrogram to obtain a denoised single-frame spectrogram; Based on the denoised spectrum and the network model composed of dilated convolution, a restored human voice audio signal is obtained.
10. An audio processing device, characterized in that: The device comprises: An acquiring unit is configured to acquire an audio signal to be processed, wherein the audio signal to be processed includes a human voice audio signal and a non-human voice audio signal; and acquire a time-domain human voice audio signal and a time-domain non-human voice audio signal that have been separated; a separation unit, configured to extract frequency domain features of the audio signal to be processed, and perform frequency domain human voice audio signal separation from frequency domain non-human voice audio signal based on the frequency domain features, thereby obtaining a frequency domain human voice audio signal and a frequency domain non-human voice audio signal; A processing unit is configured to obtain a human voice audio signal and a non-human voice audio signal based on the frequency domain human voice audio signal, the frequency domain non-human voice audio signal, the time domain human voice audio signal, and the time domain non-human voice audio signal.
11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the method according to any one of claims 1 to 9.
12. A storage medium, characterized in that: The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 9.