Audio processing method, audio processing device, electronic equipment and storage medium

By using dilated convolution, multi-scale fusion, and separation layers of multi-channel attention mechanisms in the audio processing method, the difficult problem of sound source separation in mixed audio signals is solved, and efficient and real-time audio separation effects are achieved on portable devices.

CN120656473APending Publication Date: 2025-09-16BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410302942.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, it is difficult to effectively separate the audio information of multiple sound sources contained in a mixed audio signal, resulting in the listener being unable to clearly hear the required audio information. In particular, the computational cost is high on portable mobile devices, making it difficult to achieve real-time stable operation.

Method used

An audio processing method consisting of multiple separation layers is adopted. The separation layers sequentially perform processing methods based on dilated convolution, multi-scale fusion, and multi-channel attention mechanisms. The audio features are separated and restored through the separator to obtain the audio features corresponding to each sound source.

Benefits of technology

It achieves efficient sound source separation on portable mobile devices, reduces the demand for device performance, and ensures good separation effect and real-time operation capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656473A_ABST
    Figure CN120656473A_ABST
Patent Text Reader

Abstract

The invention relates to an audio processing method, an audio processing device, electronic equipment and a storage medium. The audio processing method comprises the following steps: acquiring an audio to be processed, and acquiring audio features of the audio to be processed; the audio features are subjected to audio separation through a separator, target audio features are obtained, the separator comprises a plurality of separation layers, and each separation layer sequentially executes a processing mode based on expansion convolution, a processing mode based on multi-scale fusion and a processing mode based on a multi-channel attention mechanism; the target audio feature comprises an audio feature corresponding to each sound source in the plurality of sound sources; and performing audio reduction on the target audio feature to obtain a target audio, the target audio comprising an audio corresponding to each sound source in the plurality of sound sources. According to the sound source separation method and device, when sound source separation is executed, a good separation effect is guaranteed, meanwhile, the requirement for equipment performance is lowered, and a sound source separation algorithm can stably run in portable mobile equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure relates to the field of audio processing, and in particular to an audio processing method, an audio processing device, an electronic device, and a storage medium. Background Art

[0002] The audio information contained in the videos and audios recorded by users on a daily basis is generally mixed audio information. The mixed audio information may include sounds from multiple sound sources, as well as background sounds and noise, causing the listener to have difficulty in clearly hearing the required audio information.

[0003] Based on the problems existing in the above-mentioned mixed audio, it is necessary to adopt sound separation technology (sound source separation technology) to extract audio signals corresponding to multiple sound sources from the mixed audio signal, and perform further audio processing based on the extracted audio signals and output the final audio to ensure that the listener can hear the audio information to be obtained and improve the listening experience. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides an audio processing method, an audio processing device, an electronic device and a storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio processing method, comprising: obtaining audio to be processed and obtaining audio features of the audio to be processed, wherein the audio to be processed includes audio emitted by multiple sound sources; performing audio separation on the audio features through a separator to obtain target audio features, wherein the separator comprises a preset number of separation layers, wherein each of the preset number of separation layers sequentially performs a first process, a second process, and a third process, wherein the first process is a process based on dilated convolution, the second process is a process based on multi-scale fusion, and the third process is a process based on a multi-channel attention mechanism, wherein the target audio features include audio features corresponding to each of the multiple sound sources; performing audio restoration on the target audio features to obtain target audio, wherein the target audio includes audio corresponding to each of the multiple sound sources.

[0006] In one embodiment, obtaining the audio features of the audio to be processed includes: segmenting the time domain signal of the audio to be processed according to a preset length to obtain multiple audio segments with a length of the preset length, the preset length corresponding to a first number of audio sampling points, and there is an overlapping area between adjacent audio segments in the multiple audio segments, the overlapping area corresponds to a second number of audio sampling points, and the second number is less than the first number; encoding the multiple audio segments through a one-dimensional convolution operation and a first activation function to obtain a high-dimensional mixed feature corresponding to the audio to be processed, and determining the high-dimensional mixed feature as the audio feature.

[0007] In one embodiment, the preset number of separation layers are multiple separation layers executed in sequence; the audio separation of the audio features by the separator to obtain the target audio features includes: for the first separation layer in the preset number of separation layers, using the audio features as input features, and performing data processing on the input features to obtain output features; for other separation layers other than the first separation layer in the preset number of separation layers, using the output features of the previous separation layer as input features, and performing data processing on the input features to obtain output features; for the last separation layer in the preset number of separation layers, performing sound source separation processing on the output features of the last separation layer to obtain the target audio features.

[0008] In one embodiment, for each separation layer in the preset number of separation layers, data processing is performed on the input feature to obtain an output feature, including: performing a first processing on the input feature to obtain a first audio feature, the first audio feature is the input feature after the first processing, and the first processing is a processing method based on dilated convolution; performing a second processing on the first audio feature to obtain a second audio feature, the second audio feature is the first audio feature after the second processing, and the second processing is a processing method based on multi-scale fusion; performing a third processing on the second audio feature to obtain a third audio feature, the third audio feature is the second audio feature after the third processing, and the third processing is a processing method based on a multi-channel attention mechanism; and determining the sum of the third audio feature and the input feature as the output feature.

[0009] In one embodiment, the first processing of the input features to obtain the first audio features includes: performing batch normalization processing on the input features, and performing dilated convolution processing on the input features after batch normalization processing based on a preset dilation rate to obtain a first intermediate feature, wherein the preset dilation rate represents the interval between convolution kernels; performing one-dimensional convolution processing on the first intermediate features, and performing a second activation function on the first intermediate features after the one-dimensional convolution processing to obtain the first audio feature.

[0010] In one embodiment, the second processing of the first audio feature to obtain the second audio feature includes: determining the current time step and determining a third number based on the current time step, where the third number is the number of times the multi-scale fusion processing is performed; based on the third number, performing multi-scale fusion processing on the first audio feature one by one, and determining the first audio feature after the multi-scale fusion processing is performed one by one as a second intermediate feature; and obtaining the second audio feature based on the second intermediate feature and the first audio feature.

[0011] In one embodiment, the performing a second processing on the first audio feature to obtain the second audio feature includes: performing a second processing on the first audio feature after each time step; the obtaining the second audio feature based on the second intermediate feature and the first audio feature includes: in response to the current time step being the first time step, obtaining the second audio feature based on the second intermediate feature and the first audio feature through a first preset method; in response to the current time step not being the first time step, determining a feature set, and obtaining the second audio feature based on the second intermediate feature, the feature set and the first audio feature through a first preset method, the feature set including the second audio features outputted respectively at each time step before the current time step.

[0012] In one embodiment, the third processing of the second audio feature to obtain the third audio feature includes: performing maximum pooling on the second audio feature to obtain a maximum pooling layer, and performing average pooling on the second audio feature to obtain an average pooling layer; determining a third intermediate feature according to the first kernel size, the maximum pooling layer, the average pooling layer and the second audio feature through a second preset method; and determining the third audio feature according to the second kernel size and the third intermediate feature through a third preset method.

[0013] According to a second aspect of an embodiment of the present disclosure, an audio processing device is provided, comprising: an acquisition unit, configured to acquire audio to be processed and acquire audio features of the audio to be processed, wherein the audio to be processed includes audio emitted by multiple sound sources; a processing unit, configured to perform audio separation on the audio features through a separator to obtain target audio features, wherein the separator comprises a preset number of separation layers, wherein each of the preset number of separation layers sequentially performs a first process, a second process, and a third process, wherein the first process is a process based on dilated convolution, the second process is a process based on multi-scale fusion, and the third process is a process based on a multi-channel attention mechanism, wherein the target audio features include audio features corresponding to each of the multiple sound sources; and a restoration unit, configured to perform audio restoration on the target audio features to obtain target audio, wherein the target audio includes audio corresponding to each of the multiple sound sources.

[0014] In one embodiment, the acquisition unit acquires the audio features of the audio to be processed in the following manner: dividing the time domain signal of the audio to be processed according to a preset length to obtain multiple audio segments with a length of the preset length, the preset length corresponding to a first number of audio sampling points, and overlapping areas between adjacent audio segments in the multiple audio segments, the overlapping areas corresponding to a second number of audio sampling points, and the second number being smaller than the first number; encoding the multiple audio segments through a one-dimensional convolution operation and a first activation function to obtain high-dimensional mixed features corresponding to the audio to be processed, and determining the high-dimensional mixed features as the audio features.

[0015] In one embodiment, the preset number of separation layers are multiple separation layers executed in sequence; the processing unit performs audio separation on the audio features through a separator in the following manner to obtain target audio features: for the first separation layer in the preset number of separation layers, the audio features are used as input features, and data processing is performed on the input features to obtain output features; for other separation layers other than the first separation layer in the preset number of separation layers, the output features of the previous separation layer are used as input features, and data processing is performed on the input features to obtain output features; for the last separation layer in the preset number of separation layers, sound source separation processing is performed on the output features of the last separation layer to obtain target audio features.

[0016] In one embodiment, for each separation layer in the preset number of separation layers, the processing unit performs data processing on the input feature in the following manner to obtain an output feature: performing a first processing on the input feature to obtain a first audio feature, where the first audio feature is the input feature after the first processing, and the first processing is a processing method based on dilated convolution; performing a second processing on the first audio feature to obtain a second audio feature, where the second audio feature is the first audio feature after the second processing, and the second processing is a processing method based on multi-scale fusion; performing a third processing on the second audio feature to obtain a third audio feature, where the third audio feature is the second audio feature after the third processing, and the third processing is a processing method based on a multi-channel attention mechanism; and determining the sum of the third audio feature and the input feature as the output feature.

[0017] In one embodiment, the processing unit performs a first processing on the input features in the following manner to obtain a first audio feature: performing batch normalization processing on the input features, and performing dilated convolution processing on the input features after batch normalization processing based on a preset dilation rate to obtain a first intermediate feature, wherein the preset dilation rate represents the interval between convolution kernels; performing one-dimensional convolution processing on the first intermediate features, and performing a second activation function on the first intermediate features after the one-dimensional convolution processing to obtain the first audio feature.

[0018] In one embodiment, the processing unit performs a second processing on the first audio feature to obtain a second audio feature in the following manner: determining a current time step and determining a third number based on the current time step, where the third number is the number of times the multi-scale fusion processing is performed; based on the third number, performing multi-scale fusion processing on the first audio feature one by one, and determining the first audio feature after the multi-scale fusion processing is performed one by one as a second intermediate feature; and obtaining the second audio feature based on the second intermediate feature and the first audio feature.

[0019] In one embodiment, the processing unit performs a second processing on the first audio feature to obtain a second audio feature in the following manner: performing a second processing on the first audio feature after each time step; obtaining the second audio feature based on the second intermediate feature and the first audio feature includes: in response to the current time step being the first time step, obtaining the second audio feature based on the second intermediate feature and the first audio feature through a first preset method; in response to the current time step not being the first time step, determining a feature set, and obtaining the second audio feature based on the second intermediate feature, the feature set and the first audio feature through a first preset method, the feature set including the second audio features outputted respectively at each time step before the current time step.

[0020] In one embodiment, the processing unit performs a third processing on the second audio feature in the following manner to obtain a third audio feature: performing maximum pooling on the second audio feature to obtain a maximum pooling layer, and performing average pooling on the second audio feature to obtain an average pooling layer; determining a third intermediate feature according to the first kernel size, the maximum pooling layer, the average pooling layer and the second audio feature through a second preset method; determining the third audio feature according to the second kernel size and the third intermediate feature through a third preset method.

[0021] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute the audio processing method described in the first aspect or any one of the embodiments of the first aspect.

[0022] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor, the processor is enabled to execute the audio processing method described in the first aspect or any one of the embodiments of the first aspect.

[0023] The disclosure provided by the embodiments of the present disclosure may include the following beneficial effects: when performing sound source separation, obtaining the audio to be processed and its audio features, the audio to be processed includes audio emitted by multiple sound sources; performing audio separation on the audio features through a separator including multiple separation layers to obtain target audio features, the target audio features including the audio features corresponding to each of the multiple sound sources, wherein each of the multiple separation layers sequentially performs a first processing based on dilated convolution, a second processing based on multi-scale fusion, and a third processing based on a multi-channel attention mechanism; performing audio restoration on the target audio features to obtain target audio, the target audio including the audio corresponding to each of the multiple sound sources. Through the present disclosure, when performing sound source separation, a good balance between computational cost and required performance is achieved, ensuring a good separation effect while reducing the demand for device performance, so that the sound source separation algorithm can run stably in portable mobile devices.

[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0026] Figure 1 The figure is a flowchart of an audio processing method according to an exemplary embodiment.

[0027] Figure 2 The figure is a flowchart of a method for obtaining audio features of audio to be processed according to an exemplary embodiment.

[0028] Figure 3 The present invention is a flowchart of a method for performing audio separation on audio features to obtain target audio features according to an exemplary embodiment.

[0029] Figure 4The present invention is a flowchart of a method for performing data processing on input features to obtain output features according to an exemplary embodiment.

[0030] Figure 5 The present invention is a flowchart of a method for performing a first processing on an input feature to obtain a first audio feature according to an exemplary embodiment.

[0031] Figure 6 The present invention is a flowchart of a method for performing a second processing on a first audio feature to obtain a second audio feature according to an exemplary embodiment.

[0032] Figure 7 is a flowchart of a method for obtaining a second audio feature according to yet another exemplary embodiment.

[0033] Figure 8 The present invention is a flowchart of a method for performing a third processing on a second audio feature to obtain a third audio feature according to an exemplary embodiment.

[0034] Figure 9 The figure is a flowchart of an audio processing method according to an exemplary embodiment of the present disclosure.

[0035] Figure 10 The figure is a block diagram of an audio processing device according to an exemplary embodiment.

[0036] Figure 11 The present invention is a block diagram showing an apparatus for audio processing according to an exemplary embodiment. DETAILED DESCRIPTION

[0037] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present disclosure.

[0038] The audio processing method provided by the embodiments of the present disclosure is applied to a scenario where audio corresponding to a specific sound source is separated from mixed audio.

[0039] The audio information contained in the audio that users listen to through playback devices is generally mixed audio information. The mixed audio information may contain sounds from multiple sound sources, and may also contain background sounds and noise, causing the listener to not clearly hear the required audio information when listening to the audio. The audio that users listen to through playback devices includes but is not limited to the audio contained in the videos that users record daily, the audio recorded directly by the users, the audio collected during live broadcasts, the audio contained in downloaded videos, etc. Based on the above-mentioned problems with mixed audio, it is necessary to adopt sound separation technology (or sound source separation technology) to extract the audio signals corresponding to the multiple sound sources from the mixed audio signal, and perform further audio processing based on the extracted audio signals and output the final audio to ensure that the listener can clearly hear the audio information they want to obtain and enhance the listening experience.

[0040] Sound separation technology involves separating audio signals corresponding to different sound sources from mixed audio signals and is an important research area in the field of audio processing. In the field of speech signal processing, multi-sound separation technology has become an important foundation for applications such as speech recognition, speech enhancement, and speech synthesis. Multi-sound separation can improve audio quality and audience experience, while also enhancing teaching effectiveness and meeting efficiency. The purpose of sound separation is to separate the desired sound (such as human voices or musical instruments) from surrounding noise (such as environmental noise and other sounds). Depending on the number of audio channels or the number of microphones recording the audio, sound separation is generally divided into multi-channel sound separation and single-channel sound separation. Depending on whether a neural network model is used, sound separation can be divided into sound separation that relies on deep learning and sound separation that relies on deep learning.

[0041] Prior to the era of deep learning, traditional sound separation methods that didn't rely on deep learning included non-negative matrix factorization (NMF), computational auditory scene analysis (CASA), and support vector machines (SVM). These models often struggled to represent the nonlinear connections between sound data, severely limiting their practical applications and impacting sound separation effectiveness.

[0042] In related technologies, with the emergence of deep learning methods, large amounts of data have been used to build models that can separate the sound of the target sound source from mixed sounds. Deep learning-based sound separation methods are mainly divided into two categories:

[0043] (1) End-to-end time-domain sound separation method based on time-domain signals. Among the time-domain sound separation methods, there are multi-scale fusion methods based on convolution and methods based on transformer models. Both have achieved excellent performance in sound separation tasks by learning features at different time scales. However, convolution-based models often encounter the problem of limited receptive field, which causes them to fail to reach the theoretical upper limit of the model. Transformer-based models usually require huge computational costs and long training time, are not suitable for practical applications, and are difficult to run in mobile terminals. In addition, few methods in sound separation focus on channel attention, which may make it difficult to capture the channel correlation between input features, thereby limiting performance.

[0044] (2) Time domain-Frequency domain (TF) domain sound separation method that requires time-frequency domain conversion. In the time-frequency sound separation method, the input signal is first transformed into frequency domain features using short-time Fourier transform (STFT), and then the amplitude spectrum of the mixed signal is mathematically multiplied with the predicted masking value to obtain the estimated amplitude spectrum target signal. Finally, the enhanced amplitude spectrum and the original phase spectrum are separated by inverse Fourier transform (ISTFT). However, due to the difficulty of waveform phase reconstruction, the predicted source waveform is usually synthesized by mixing the initial phase, which reduces the performance upper limit of this method. In addition, the time domain sound separation method uses a convolutional neural network-based encoding and decoding framework to directly model the mixed waveform. The separation part consists of multiple convolution filters to achieve time domain sound separation.

[0045] In summary, traditional sound separation methods in related technologies that do not rely on deep learning have difficulty in representing the nonlinear connections between sound data, which severely limits their practical application and affects the sound separation effect. Although the end-to-end time-domain sound separation method and time-frequency domain sound separation method that rely on deep learning in related technologies can represent the nonlinear connections between sound data and have better sound separation effects than traditional methods, the sound separation method based on deep learning builds a sound separation model based on the Transformer model, which requires huge computing costs and long training time. When performing sound separation, the algorithm complexity is too high, and it consumes a lot of computing resources. It also places too high requirements on device performance and is not suitable for real-time and stable operation on mobile portable devices such as mobile phones.

[0046] In view of this, the present disclosure proposes an audio processing method, which, when performing sound source separation, obtains audio to be processed that includes audio emitted by multiple sound sources, and obtains audio features of the audio to be processed; performs audio separation on the audio features through a separator including multiple separation layers to obtain target audio features. The target audio features include audio features corresponding to each of the multiple sound sources, and each of the multiple separation layers sequentially performs a first processing based on dilated convolution, a second processing based on multi-scale fusion, and a third processing based on a multi-channel attention mechanism. After obtaining the target audio, the target audio features are audio restored to obtain the target audio, and the target audio includes audio corresponding to each of the multiple sound sources. Through the present disclosure, when performing sound source separation, a good balance between computational cost and required performance is achieved, ensuring a good separation effect while reducing the demand for device performance, so that the sound source separation algorithm can run stably in portable mobile devices.

[0047] Figure 1 FIG. 1 is a flow chart of an audio processing method according to an exemplary embodiment. Figure 1 As shown, the method includes steps S101 to S103.

[0048] In step S101, audio to be processed is obtained, and audio features of the audio to be processed are obtained. The audio to be processed includes audio emitted by multiple sound sources.

[0049] In step S102, the audio features are separated by a separator to obtain target audio features. The separator includes a preset number of separation layers. Each of the preset number of separation layers performs a first process, a second process, and a third process in sequence. The first process is a process based on dilated convolution, the second process is a process based on multi-scale fusion, and the third process is a process based on a multi-channel attention mechanism. The target audio features include audio features corresponding to each of multiple sound sources.

[0050] In step S103, audio restoration is performed on the target audio features to obtain target audio, where the target audio includes audio corresponding to each of the multiple sound sources.

[0051] In the embodiments of the present disclosure, the audio to be processed is mixed audio, that is, audio containing sounds emitted by multiple sound sources. The sources of the audio to be processed include, but are not limited to, audio collected in real time during user live broadcasts, audio contained in videos recorded daily by users, audio recorded directly by users, and audio contained in downloaded videos (such as movies and short videos). In one example, the audio processing method proposed in the present disclosure performs sound separation for the live broadcast scene on the current mobile device, and can realize the real-time separation function of the host's voice, sound effects, and music.

[0052] The audio processing method in the embodiment of the present disclosure is implemented based on an audio separation model. The basic network architecture of the audio separation model is based on a stable encoder-decoder form. A separator for separating audio features is provided between the encoder and the decoder. When performing sound separation, the audio features of the audio to be processed are obtained through the encoder. The audio features output by the encoder are further processed by the separator, and the processed audio features are separated and processed to obtain multiple audio features corresponding to different sound sources (i.e., target audio features). The audio features output by the separator are restored by the decoder to obtain multiple audios corresponding to different sound sources and output them to complete the sound separation. Among them, the separator includes multiple separation layers, and the audio features are further processed by the multiple separation layers. Each separation layer includes an expanded convolution block, a multi-scale fusion block, and a multi-channel attention mechanism block, which are used to sequentially perform a first processing based on expanded convolution, a second processing based on multi-scale fusion, and a third processing based on the multi-channel attention mechanism on the input features input to the separation layer.

[0053] In the embodiments of the present disclosure, a sound separation model for executing an audio processing method and achieving sound separation combines technical methods such as dilated convolution, multi-scale fusion, and a multi-channel attention mechanism. This overcomes the problems of limited receptive field and high computational cost in related deep learning-based sound separation models, and can achieve real-time separation on mobile devices. Through this disclosure, a good balance is achieved between computational cost and required performance when performing sound source separation, ensuring good separation results while reducing the demand for device performance, enabling the sound source separation algorithm to run stably on portable mobile devices.

[0054] In an exemplary embodiment of the present disclosure, the present disclosure focuses on real-time effects for mobile devices such as mobile phones and tablets, and the real-time effect meets the need for sound source separation in mobile phone live broadcasts. For example, when performing sound separation in a live broadcast scenario, the audio processing method in the present disclosure is used to obtain the sounds emitted by each sound source in real time (such as the host's voice, live broadcast sound effects, and live broadcast background music). Based on the multiple sounds separated, the user can freely choose to retain the host's voice (including the voices of multiple people), live broadcast sound effects, or live broadcast background music. In one example, the sound sources targeted by the audio processing method in the present disclosure are people, musical instruments, and sound effects, and the application scenario is a live broadcast scenario.

[0055] In the embodiment of the present disclosure, the audio to be processed is input into the encoder in the form of a one-dimensional time domain waveform. The encoder extracts the audio features of the audio to be processed and inputs the audio features into the separator. The following embodiment of the present disclosure describes the method for obtaining the audio features of the audio to be processed.

[0056] Figure 2 FIG. 1 is a flow chart showing a method for obtaining audio features of audio to be processed according to an exemplary embodiment. Figure 2As shown, the method includes steps S201 to S202.

[0057] In step S201 , the time domain signal of the audio to be processed is segmented according to a preset length to obtain a plurality of audio segments of the preset length.

[0058] The preset length corresponds to a first number of audio sampling points, and there is an overlapping area between adjacent audio segments in the multiple audio segments. The overlapping area corresponds to a second number of audio sampling points, and the second number is smaller than the first number.

[0059] In step S202, multiple audio clips are encoded using a one-dimensional convolution operation and a first activation function to obtain high-dimensional mixed features corresponding to the audio to be processed, and the high-dimensional mixed features are determined as audio features.

[0060] In an embodiment of the present disclosure, for the encoder, the mixed sound input to the encoder (that is, the audio to be processed in the form of a time domain waveform) is segmented according to a preset length, and the audio to be processed is segmented into multiple adjacent segments with overlapping relationships. It can be understood that the time domain waveform is composed of multiple sampling points, so the preset length is a specific number (first number) of sampling points, and the overlapping area between adjacent audio segments is also a specific number (second number) of sampling points. In one example, for the audio to be processed with a length of 50 sampling points (p0-p49), the preset length is set to 10 sampling points, and the length corresponding to the overlapping area is set to 5 sampling points. The multiple audio segments obtained after segmentation are (p0-p9, p5-p14, p10-p19..., p40-p49). The multiple audio segments obtained by segmentation in the present disclosure can be expressed as follows: X∈s×H. Wherein, s=1,..., N represents the segment index, N is the total number of input segments, and H represents a single audio segment with a preset length.

[0061] In the embodiment of the present disclosure, after completing the audio segmentation and obtaining multiple audio clips, the obtained audio clips are encoded into high-dimensional mixed features through a trainable one-dimensional convolution operation and an SMU activation function (i.e., the first activation function): that is, Y = SMU (Conv1D (X)). Among them, Y is the high-dimensional mixed feature obtained by encoding, Conv1D represents the one-dimensional convolution operation, SMU represents the activation operation of the activation function, and X is the audio clips obtained by multiple segmentations of the audio to be processed. The high-dimensional mixed feature can be Y∈C×L (C represents the number of channels of the encoder, and L is the signal length generated by the one-dimensional convolution). The high-dimensional mixed feature finally obtained is the audio feature of the audio to be processed.

[0062] In the embodiments of the present disclosure, the preset number of separation layers included in the separator are multiple separation layers executed sequentially. The audio features output by the encoder are processed and transmitted in the multiple separation layers, and are ultimately output by the last separation layer. The sound source separation is performed at the output of the last separation layer, ultimately obtaining the target audio features. The following embodiments of the present disclosure illustrate the method for obtaining the target audio features.

[0063] Figure 3 FIG. 1 is a flow chart showing a method for performing audio separation on audio features to obtain target audio features according to an exemplary embodiment. Figure 3 As shown, the method includes steps S301 to S303.

[0064] In step S301, for the first separation layer among the preset number of separation layers, the audio feature is used as the input feature, and data processing is performed on the input feature to obtain the output feature.

[0065] In step S302 , for the other separation layers except the first separation layer among the preset number of separation layers, the output features of the previous separation layer are used as input features, and data processing is performed on the input features to obtain output features.

[0066] In step S303, for the last separation layer among the preset number of separation layers, sound source separation processing is performed on the output features of the last separation layer to obtain target audio features.

[0067] In the disclosed embodiment, the multiple separation layers included in the separator of the audio separation model are multiple separation layers that perform feature processing tasks in sequence. For the first separation layer, the audio features output by the encoder are the input of the first separation layer. For the separation layers executed after the first separation layer, the input of the separation layer is the output of the previous separation layer. For the last separation layer, the output of the last separation layer is the final output of all separation layers. By performing sound source separation processing on the output features of the last separation layer, the target audio features can be obtained.

[0068] It is understood that the greater the number of separation layers, the higher the computing performance requirements. Therefore, the present disclosure can be optimized for different terminal devices, analyzing the terminal performance and setting the number of separation layers (i.e., the preset number) that matches its performance. This achieves a good balance between computing cost and required performance on the device, ensuring good separation results while reducing the demand for device performance, allowing the sound source separation algorithm to run stably on portable mobile devices.

[0069] In the disclosed embodiment, the output features of the last separation layer are subjected to sound source separation processing to obtain the target audio features, which are masks for different sound sources. The masks for different sound sources are decoded and restored to obtain audio corresponding to different sound sources.

[0070] In the disclosed embodiment, the decoder performs sound feature restoration for the output masks corresponding to different sound sources in the output features of the last separation layer. The decoder's one-dimensional transposed convolution process uses the same step size and convolution kernel as the encoder. The decoder input is obtained by element-by-element multiplication between the output of the encoder X and the mask Di of the i-th source. The decoder transformation can be expressed as follows:

[0071] Y i =ConvTrans(D i *X)

[0072] Where Yi∈1×T is the final waveform (time domain waveform of the target audio) for the i-th speaker of the electronic device. After the decoding operation, the real-time separation output of the audio for different sound sources can be obtained.

[0073] Through the present disclosure, multiple separation layers are used to further process the audio features extracted by the decoder, and the features finally output by the multiple separation layers are subjected to audio separation processing to obtain audio features corresponding to different sound sources (i.e., target audio features), thereby ensuring that the acquired audio features correspond to the corresponding sound sources and do not contain features of other sound sources, thereby ensuring the accuracy of sound source separation.

[0074] In an embodiment of the present disclosure, the separator in the sound separation model includes multiple separation layers executed sequentially. Each of the multiple separation layers includes a dilated convolution block, a multi-scale fusion block, and a multi-channel attention mechanism block, which sequentially performs a first processing based on dilated convolution, a second processing based on multi-scale fusion, and a third processing based on the multi-channel attention mechanism on the input features input to the separation layer. The following embodiment of the present disclosure describes a method for obtaining output features for each of the multiple separation layers.

[0075] Figure 4 FIG. 1 is a flow chart showing a method for processing input features to obtain output features according to an exemplary embodiment. Figure 4 As shown, the method includes steps S401 to S404.

[0076] In step S401, a first processing is performed on the input feature to obtain a first audio feature, where the first audio feature is the input feature after the first processing, and the first processing is a processing method based on dilated convolution.

[0077] In step S402, a second processing is performed on the first audio feature to obtain a second audio feature. The second audio feature is the first audio feature after the second processing. The second processing is a processing method based on multi-scale fusion.

[0078] In step S403, the second audio feature is subjected to a third processing to obtain a third audio feature. The third audio feature is the second audio feature after the third processing. The third processing is a processing method based on a multi-channel attention mechanism.

[0079] In step S404, the sum of the third audio feature and the input feature is determined as the output feature.

[0080] In the disclosed embodiment, for each of the multiple separation layers, after obtaining input features, the input features are processed using a dilated convolution method to obtain a first audio feature. The second audio features are processed using a multi-scale fusion method to obtain a second audio feature. The second audio features are processed using a multi-channel attention mechanism to obtain a third audio feature.

[0081] In this disclosed embodiment, a residual connection is introduced between the input of the separation layer (input features) and the output of the multi-channel attention mechanism block (third audio features). The sum of the third audio features and the input features is used as the output feature. This residual connection enhances the model's representational capabilities and alleviates the vanishing gradient problem.

[0082] In the embodiment of the present disclosure, local and global features are learned by using dilated convolutions with gradually increasing dilation values, and they are fused in adjacent stages, so that the sound separation model can learn rich feature content. At the same time, by adding a multi-channel attention module to the model, the model can extract channel weights, thereby allowing the network to focus on specific areas, thereby enhancing the network's feature representation and discrimination capabilities, and ultimately improving its expressiveness and robustness. The present disclosure makes the sound features derived in the sound feature separation stage achieved by the separation layer more comprehensive through the application of dilated convolution, multi-scale fusion and multi-channel attention, which can better utilize the information between contexts and ensure the processing effect of sound separation.

[0083] The following embodiments of the present disclosure illustrate a method for obtaining a first audio feature.

[0084] Figure 5 FIG. 1 is a flow chart showing a method for performing a first processing on an input feature to obtain a first audio feature according to an exemplary embodiment. Figure 5 As shown, the method includes steps S501 to S502.

[0085] In step S501, batch normalization processing is performed on the input features, and dilation convolution processing is performed on the input features after batch normalization processing based on a preset dilation rate to obtain a first intermediate feature, where the preset dilation rate represents the interval between convolution kernels.

[0086] In step S502, one-dimensional convolution processing is performed on the first intermediate feature, and the first intermediate feature after the one-dimensional convolution processing is processed by a second activation function to obtain a first audio feature.

[0087] In the disclosed embodiment, the high-dimensional mixed features obtained by the separator (i.e., the output of the encoder: the audio features of the audio to be processed) are fed into a preset number of separation layers to achieve audio feature separation for different sound sources. After the output of the encoder is passed to the multiple separation layers in the separator, for the first separation layer among the multiple separation layers, the audio features output by the encoder are the input features of the first separation layer. For the separation layers that are executed after the first separation layer, the input features of the separation layer are the output features of the previous separation layer. For each separation layer in the multiple separation layers, after obtaining the input features, batch normalization (BN) is performed on the input features. Dilated convolution processing is performed on the input features after batch normalization based on a preset dilation rate. One-dimensional convolution operation and global layer normalization (GLN) processing are performed on the input features after the dilated convolution processing (first intermediate features), and activation processing is performed using an activation function (PReLU, i.e., the second activation function) to obtain the first audio features.

[0088] In this disclosure, dilated convolutions at lower layers are used to extract local detail features, while dilated convolutions at higher layers are used to extract global semantic features. This dilated convolution process enables the dilated convolution layer to achieve a sufficiently large receptive field, fully extracting audio features. This disclosure also inserts GLN between each dilated convolution operation to ensure more stable output data.

[0089] In an exemplary embodiment of the present disclosure, in order to allow higher-level dilated convolution layers to obtain a sufficiently large receptive field, dilation values ​​are used that increase exponentially from bottom to top in powers of 2. GLN is inserted between each dilated convolution operation to ensure more stable output data.

[0090] The following embodiments of the present disclosure further illustrate the method for obtaining the second audio feature.

[0091] Figure 6 FIG. 1 is a flow chart showing a method for performing a second processing on a first audio feature to obtain a second audio feature according to an exemplary embodiment. Figure 6 As shown, the method includes steps S601 to S603.

[0092] In step S601 , the current time step is determined, and a third number is determined according to the current time step, where the third number is the number of times the multi-scale fusion process is performed.

[0093] In step S602, multi-scale fusion processing is performed on the first audio features one by one according to the third quantity, and the first audio features after the multi-scale fusion processing is performed one by one are determined as second intermediate features.

[0094] In step S603, a second audio feature is obtained according to the second intermediate feature and the first audio feature.

[0095] In the disclosed embodiment, after the dilated convolution block in the separation layer outputs the first audio feature, a second processing based on multiscale fusion is performed on the first audio feature by executing a multiscale fusion block (M-fusion, where M is the number of channels in the sound separation model) that follows the dilated convolution block. When performing the second processing, the number of times (a third number) to perform the multiscale fusion processing is first determined based on the time step of the current audio processing, and the multiscale fusion processing is performed on the first audio feature a corresponding number of times.

[0096] In each multi-scale fusion processing process in the embodiment of the present disclosure, the input (first audio feature) flows from bottom to top in the multi-scale fusion block, and different receptive field information is obtained at different stages, with a total of preset number (such as J) stages. And each convolution layer in the multi-scale fusion block is followed by a GLN function for global layer normalization and a PReLU function for activation processing, and global layer normalization processing and activation processing are performed, thereby improving the expressiveness and robustness of the sound separation model. The present disclosure obtains information of different scales through multi-scale fusion processing, and by fusing information of different scales, the sound separation model learns comprehensive features and improves its performance. Compared with the widely used dual-path network for learning local and global features, the sound separation model proposed in the present disclosure has a shorter calculation time of the multi-scale fusion block, is more suitable for practical applications, and is convenient for real-time and stable operation in portable mobile terminals.

[0097] It is understood that the present disclosure can be applied to real-time separation of acquired audio. The multiscale fusion block in the present disclosure performs a second multiscale fusion-based processing on the first audio feature at each time step. The following embodiments of the present disclosure further illustrate the method for obtaining the second audio feature.

[0098] Figure 7 FIG. 1 is a flow chart showing a method for obtaining a second audio feature according to another exemplary embodiment. Figure 7 As shown, the method includes step S701, step S702A and step S702B.

[0099] In step S701, a first processing is performed on the input feature to obtain a first audio feature.

[0100] In step S702A, in response to the current time step being the first time step, a second audio feature is obtained according to the second intermediate feature and the first audio feature in a first preset manner.

[0101] In step S702B, in response to the current time step not being the first time step, a feature set is determined, and a second audio feature is obtained based on the second intermediate feature, the feature set and the first audio feature through a first preset method. The feature set includes the second audio feature outputted respectively at each time step before the current time step.

[0102] In the embodiment of the present disclosure, in order to better fuse the information between the multi-scale fusion blocks and prevent the gradient from disappearing, the output of the current block is integrated with the output of all previous blocks and the input of the initial multi-scale fusion block before being passed to the subsequent block, and a one-dimensional convolution is added before each input of the multi-scale fusion block. Each convolution layer is followed by a GLN function for global layer normalization and a SMU activation function for activation. When the current time step is not the first time step, this method (i.e., the first preset method) is described as follows:

[0103] M(t+1)=Fun(conv(M(t)+M(t-1)+...+row))

[0104] When the current time step is the first time step, the method is described as follows:

[0105] M(2)=Fun(conv(M(1))+row)

[0106] Where t is the time step, M(t) represents the output of the multiscale fusion block at time step t [similarly for M(t-1), M(t+1), M(2), and M(1)], Fun represents the operation performed by the multiscale fusion block, conv represents the one-dimensional convolution operation combined with GLN and SMU activation, and row represents the input of the initial multiscale fusion block (i.e., the first audio feature). M(1) to M(t-1) correspond to the above feature set.

[0107] The following embodiments of the present disclosure illustrate a method for obtaining the third audio feature.

[0108] Figure 8 FIG. 1 is a flow chart showing a method for performing a third processing on a second audio feature to obtain a third audio feature according to an exemplary embodiment. Figure 8 As shown, the method includes steps S801 to S803.

[0109] In step S801, maximum pooling is performed on the second audio feature to obtain a maximum pooling layer, and average pooling is performed on the second audio feature to obtain an average pooling layer.

[0110] In step S802, a third intermediate feature is determined according to the first kernel size, the maximum pooling layer, the average pooling layer and the second audio feature in a second preset manner.

[0111] In step S803, a third audio feature is determined according to the second kernel size and the third intermediate feature in a third preset manner.

[0112] In the embodiment of the present disclosure, after the multi-scale fusion block completes the second processing based on multi-scale fusion and outputs the second audio feature, the multi-channel attention mechanism block performs a third processing based on the multi-channel attention mechanism on the second audio feature. The maximum pooling and average pooling in the second audio feature output by the multi-scale fusion block are applied to the input along the last dimension. A one-dimensional convolution with a kernel size of a preset value (a first kernel size, such as 5) is applied to generate a weight vector for each channel. The two weight vectors are added and multiplied with the input feature, each channel is weighted and a new feature map Fea1 (corresponding to the third intermediate feature) is generated. Further, the maximum pooling layer and the average pooling layer are applied to the new feature vector along their channel dimensions respectively. A one-dimensional convolution with a kernel size of another preset value (a second kernel size, such as 6) is applied to generate a weight vector for the spatial position, and the weight vector is multiplied with the input to generate the final feature Fea2 (corresponding to the third audio feature). In addition, by incorporating the residual connection into the channel attention module, the gradient disappearance is alleviated and the model stability is improved. The formula for the third processing performed based on the multi-channel attention mechanism is as follows:

[0113] Fea1=σ(H 5 (Ave(Y)))+H 5 (Map(Y))*σ(Y)

[0114] Fea2=σ(H 16 (Concat(Ave(Fea1),Map(Fea1))))*σ(Fea1)+Fea1

[0115] Here, Fea1 is the third intermediate feature, Fea2 is the final feature (i.e., the third audio feature), and Y is the input to the multi-scale fusion block (i.e., the second audio feature). Ave and Map represent average pooling and maximum pooling operations, respectively. Thus, Ave(Y) is the average pooling layer, and Map(Y) is the maximum pooling layer. σ represents the sigmoid activation operation. H5 represents a one-dimensional convolution operation with a kernel size of 5, 1 input channel, and 1 output channel. H16 represents a one-dimensional convolution operation with a kernel size of 16, 2 input channels, and 1 output channel. Concat represents concatenation along the channel dimension.

[0116] In the embodiment of the present disclosure, the third processing based on the multi-channel attention mechanism block is performed through the multi-channel attention mechanism block to further improve the generalization ability of the sound separation model and enable the sound separation model to focus on important features.

[0117] In an exemplary embodiment of the present disclosure, the audio processing method proposed in the present disclosure may be applicable to the following scenarios:

[0118] (1) Live music performance scenario: In a music performance, multiple instruments are played simultaneously. If the sounds of all the instruments are mixed together, the sound quality will be degraded, making it difficult to hear the sound of each instrument. The audio processing method proposed in this disclosure can separate the sound of each instrument, resulting in better sound quality and allowing the audience to better appreciate the music performance.

[0119] (2) Multi-person live broadcast scenario: In a multi-person live broadcast scenario, real-time speech recognition of the host or guest's speech is required. If everyone's voices are mixed together, the recognition accuracy will decrease. The audio processing method proposed in this disclosure can separate each person's voice, improving the speech recognition accuracy.

[0120] (3) Personal live broadcast scenario: In a personal live broadcast scenario, background music and the live broadcaster's voice need to be played simultaneously. If the live broadcaster's voice is mixed with the background music, the viewer will not be able to hear the human voice clearly. The audio processing method proposed in this disclosure can separate the background music and the live broadcaster's voice, allowing the viewer to hear the live broadcaster's voice more clearly and also better appreciate the background music.

[0121] (4) In the online education live broadcast scenario: When the teacher is giving a lecture while playing a teaching video, it is necessary to play the teaching video and the teacher's voice at the same time. If the teacher's voice is mixed with the sound of the teaching video, it will make it difficult for the viewer to understand the teaching content. Through sound separation technology, the sound of the teaching video and the teacher's voice can be separated, so that students can hear the teacher's voice more clearly and watch the teaching video better.

[0122] In an exemplary embodiment of the present disclosure, the following method is used to obtain training data and verification data for training a sound separation model: multiple publicly accessible audio datasets are used. 30 hours of training data and 10 hours of verification data are generated based on the audio dataset. In particular, when training a sound separation model for separating human voices, the training data and verification data are created by adding different environmental noise samples to the mixed sounds of different speakers in the dataset. In this example, when performing sound separation, it is necessary not only to separate the speaker signal but also to reduce the background noise. Therefore, random noise with an intensity within (-5db, 15db) (db [decibel]) is added to the audio dataset for interference to generate training data and verification data. The random noise samples in the present disclosure are added to the sound mixture and evenly distributed. The loudness of the noise samples is between (-38LUFS, -30LUFS) (LUFS, Loudness). All sound signals are resampled at a frequency of 16kHz. When training a sound separation model for separating music, the music in the music dataset is used as initialization data for training the separation of music sounds.

[0123] In an embodiment of the present disclosure, the parameters of the audio separation model used for sound separation can be customized based on requirements, thereby achieving a balance between power consumption and performance for terminals with different performance capabilities. The parameters of the audio separation model include the number of channels, kernel size, stride size, number of channels and separation layers in the encoder and decoder, number of stages in the multiscale fusion block, dilation value, kernel size, and stride size of dilated convolution, model training period, training audio segment length, and optimizer used. In an exemplary embodiment of the present disclosure, the sound separation model is configured as follows: the encoder and decoder are configured with 512 channels, a kernel size of 16, and a stride size of 4. The number of separation layers in the separator is set to 7, and the number of channels is set to 512. The number of stages (J value) in the multiscale fusion block is set to 5, the dilation value of the dilated convolution is set from low to high to {1, 2, 4, 8, 16}, the kernel size is 5, and the stride size is 2. The model is trained for 128 training cycles (epochs). The training sound segment length was set to 4 seconds, and AdamW was used as the optimizer. During training, the initial learning rate was set to 0.0005. If the validation accuracy did not improve after two consecutive training runs, the learning rate was reduced to 0.95. The scale-invariant source-distortion ratio (SI-SDR) was used as the training objective to maximize the model's signal fidelity.

[0124] In an exemplary embodiment of the present disclosure, Figure 9As shown in the flowchart of the method for audio processing in the audio processing, the sound separation is performed in the following manner: in response to the input audio to be processed (audio signal / third-party audio stream), the audio features of the audio to be processed are obtained through the encoder. That is, the input audio is divided into multiple (such as M) overlapping segments with a preset length, and the multiple overlapping segments are encoded through a one-dimensional convolution operation and an activation function, and the multiple overlapping segments are encoded into high-dimensional mixed features, wherein the high-dimensional mixed features are the audio features of the audio to be processed. The high-dimensional mixed features are separated by the separator to obtain the audio features corresponding to the multiple sound sources. As shown in FIG. Figure 9 As shown, the separator consists of eight sequentially executed separation layers (with the specific structure and simplified representation of each separation layer *7 given). Each of the multiple separation layers has, in order of execution, variable-stretch convolution blocks (BN & Conv1d * M & GLN), multi-scale fusion blocks (M-fusion), and channel attention blocks (M-channel). For the first separation layer, the audio features output by the encoder serve as the input to the first separation layer. For subsequent separation layers, the input to the previous separation layer serves as the output. For the last separation layer, the output of the last separation layer serves as the final output of all separation layers. Source separation is performed on the output features of the last separation layer to obtain the target audio features. For each of the multiple separation layers, the variable-stretch convolution blocks, multi-scale fusion blocks, and channel attention blocks sequentially perform processing based on variable-stretch convolution, multi-scale fusion, and multi-channel attention on the input features to the separation layer, generating intermediate features. The intermediate features are summed (SUM) with the input features (Conv1D) to obtain the output features and input them to the next layer. The output features of the last separation layer are processed for sound source separation. Specifically, different convolution processes (Convld*2 & SMU) and activation processes based on the activation function (pReLu) are performed on the output features of the last separation layer to obtain masks (Mask1, 2, 3) corresponding to different sound sources. The decoder performs sound feature restoration (ConyTrans1D) on the output masks corresponding to different sound sources. This results in audio (S1, S2, S3) corresponding to different sound sources.

[0125] Through the present disclosure, when performing sound source separation, the combination of a multi-channel attention module enables the model to pay more attention to important features, and its expressive power and generalization performance are relatively high, meeting the separation requirements of various sound sources (such as human voices, musical instruments, and sound effects). The sound separation effect of multiple speakers in complex acoustic environments such as reverberation conditions and noisy environments is excellent, meeting application requirements. It meets the live broadcast requirements of third-party application APPs and can accurately separate human voices and music. A good balance between computing cost and required performance is achieved, ensuring good separation effects while reducing the demand for device performance, so that the sound source separation algorithm can run stably in portable mobile devices.

[0126] Based on the same concept, the embodiment of the present disclosure further provides an audio processing device 100 .

[0127] It is understandable that the audio processing device 100 provided in the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of the various examples disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints disclosed. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the disclosure of the embodiment of the present disclosure.

[0128] Figure 10 FIG. 1 is a block diagram of an audio processing device 100 according to an exemplary embodiment. Figure 10 The device includes an acquisition unit 101, a processing unit 102 and a restoration unit 103.

[0129] The acquisition unit 101 is configured to acquire audio to be processed and acquire audio features of the audio to be processed. The audio to be processed includes audio emitted by multiple sound sources.

[0130] The processing unit 102 is used to perform audio separation on the audio features through a separator to obtain target audio features. The separator includes a preset number of separation layers. Each of the preset number of separation layers performs a first process, a second process, and a third process in sequence. The first process is a process based on dilated convolution, the second process is a process based on multi-scale fusion, and the third process is a process based on a multi-channel attention mechanism. The target audio features include audio features corresponding to each of multiple sound sources.

[0131] The restoration unit 103 is configured to perform audio restoration on the target audio feature to obtain target audio, where the target audio includes audio corresponding to each of the multiple sound sources.

[0132] In one embodiment, the acquisition unit 101 acquires audio features of the audio to be processed by segmenting the time domain signal of the audio to be processed according to a preset length to obtain multiple audio segments of a preset length, where the preset length corresponds to a first number of audio sampling points, and adjacent audio segments in the multiple audio segments have overlapping regions, where the overlapping regions correspond to a second number of audio sampling points, where the second number is less than the first number. The multiple audio segments are encoded using a one-dimensional convolution operation and a first activation function to obtain high-dimensional mixed features corresponding to the audio to be processed, and the high-dimensional mixed features are determined as audio features.

[0133] In one embodiment, the preset number of separation layers is a plurality of separation layers that are executed sequentially. The processing unit 102 performs audio separation on the audio features through a separator in the following manner to obtain target audio features: for the first separation layer in the preset number of separation layers, the audio features are used as input features, and data processing is performed on the input features to obtain output features. For other separation layers other than the first separation layer in the preset number of separation layers, the output features of the previous separation layer are used as input features, and data processing is performed on the input features to obtain output features. For the last separation layer in the preset number of separation layers, sound source separation processing is performed on the output features of the last separation layer to obtain target audio features.

[0134] In one embodiment, for each separation layer in a preset number of separation layers, the processing unit 102 performs data processing on the input features in the following manner to obtain output features: performing a first processing on the input features to obtain a first audio feature, where the first audio feature is the input feature after the first processing, and the first processing is a processing method based on dilated convolution. Performing a second processing on the first audio feature to obtain a second audio feature, where the second audio feature is the first audio feature after the second processing, and the second processing is a processing method based on multi-scale fusion. Performing a third processing on the second audio feature to obtain a third audio feature, where the third audio feature is the second audio feature after the third processing, and the third processing is a processing method based on a multi-channel attention mechanism. The sum of the third audio feature and the input feature is determined as the output feature.

[0135] In one embodiment, the processing unit 102 performs a first processing on the input features to obtain the first audio features by performing batch normalization on the input features and performing dilated convolution on the batch normalized input features based on a preset dilation rate to obtain first intermediate features, where the preset dilation rate represents the spacing between convolution kernels. One-dimensional convolution is performed on the first intermediate features, and the first intermediate features after the one-dimensional convolution are processed using a second activation function to obtain the first audio features.

[0136] In one embodiment, the processing unit 102 performs a second processing on the first audio feature to obtain a second audio feature in the following manner: determining a current time step and, based on the current time step, determining a third number, where the third number is the number of times the multi-scale fusion process is performed. Based on the third number, the multi-scale fusion process is performed on the first audio feature one after another, and the first audio feature after each multi-scale fusion process is determined as a second intermediate feature. The second audio feature is obtained based on the second intermediate feature and the first audio feature.

[0137] In one embodiment, the processing unit 102 performs a second processing on the first audio feature to obtain a second audio feature in the following manner: the first audio feature is subjected to a second processing for each time step. The second audio feature is obtained based on the second intermediate feature and the first audio feature, including: in response to the current time step being the first time step, the second audio feature is obtained based on the second intermediate feature and the first audio feature through a first preset method. In response to the current time step not being the first time step, a feature set is determined, and the second audio feature is obtained based on the second intermediate feature, the feature set, and the first audio feature through a first preset method, the feature set including the second audio features outputted respectively at each time step before the current time step.

[0138] In one embodiment, the processing unit 102 performs a third processing on the second audio feature to obtain a third audio feature by performing maximum pooling on the second audio feature to obtain a maximum pooling layer, and performing average pooling on the second audio feature to obtain an average pooling layer. A third intermediate feature is determined using a second preset method based on the first kernel size, the maximum pooling layer, the average pooling layer, and the second audio feature. A third audio feature is determined using a third preset method based on the second kernel size and the third intermediate feature.

[0139] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0140] Figure 11 FIG2 is a block diagram illustrating an apparatus 200 for audio processing according to an exemplary embodiment. The apparatus 200 may be provided as a terminal. For example, the apparatus 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0141] Reference Figure 11, apparatus 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0142] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0143] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0144] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 200.

[0145] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0146] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0147] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0148] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect changes in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0149] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0150] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0151] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, which can be executed by the processor 220 of the apparatus 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0152] It is understood that in this disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of related objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0153] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0154] It can be further understood that the terms "center", "longitudinal", "lateral", "front", "back", "up", "down", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation.

[0155] It is further understood that, unless otherwise specified, “connection” includes a direct connection where there are no other components between the two elements, and also includes an indirect connection where there are other elements between the two elements.

[0156] It is further understood that although operations are described in a particular order in the drawings in the embodiments of the present disclosure, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.

[0157] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0158] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.

Claims

1. An audio processing method, characterized in that: include: Acquire audio to be processed and obtain audio features of the audio to be processed, where the audio to be processed includes audio emitted by multiple sound sources; Performing audio separation on the audio features through a separator to obtain target audio features, the separator comprising a preset number of separation layers, each of the preset number of separation layers sequentially performing a first process, a second process, and a third process, the first process being based on dilated convolution, the second process being based on multi-scale fusion, and the third process being based on a multi-channel attention mechanism, the target audio features comprising audio features corresponding to each of the multiple sound sources; Audio restoration is performed on the target audio feature to obtain target audio, where the target audio includes audio corresponding to each of the multiple sound sources.

2. The method according to claim 1, characterized in that The obtaining of audio features of the audio to be processed includes: Segmenting the time domain signal of the audio to be processed according to a preset length to obtain a plurality of audio segments having the preset length, wherein the preset length corresponds to a first number of audio sampling points, and adjacent audio segments in the plurality of audio segments have overlapping regions, wherein the overlapping regions correspond to a second number of audio sampling points, and the second number is smaller than the first number; The multiple audio segments are encoded using a one-dimensional convolution operation and a first activation function to obtain high-dimensional mixed features corresponding to the audio to be processed, and the high-dimensional mixed features are determined as the audio features.

3. The method according to claim 1, characterized in that The preset number of separation layers is a plurality of separation layers executed sequentially; The step of performing audio separation on the audio feature by a separator to obtain a target audio feature includes: For a first separation layer among the preset number of separation layers, using the audio feature as an input feature, and performing data processing on the input feature to obtain an output feature; For each separation layer other than the first separation layer in the preset number of separation layers, the output feature of the previous separation layer is used as an input feature, and the input feature is processed to obtain an output feature; For the last separation layer of the preset number of separation layers, sound source separation processing is performed on the output features of the last separation layer to obtain target audio features.

4. The method according to claim 3, characterized in that For each separation layer in the preset number of separation layers, data processing is performed on the input features to obtain output features, including: Performing a first processing on the input feature to obtain a first audio feature, where the first audio feature is the input feature after the first processing, and the first processing is based on a dilated convolution processing method; Performing a second processing on the first audio feature to obtain a second audio feature, where the second audio feature is the first audio feature after the second processing, and the second processing is based on multi-scale fusion; Performing a third processing on the second audio feature to obtain a third audio feature, where the third audio feature is the second audio feature after the third processing, and the third processing is a processing method based on a multi-channel attention mechanism; The sum of the third audio feature and the input feature is determined as the output feature.

5. The method according to claim 4, characterized in that The performing a first processing on the input feature to obtain the first audio feature includes: Performing batch normalization on the input features, and performing dilated convolution on the batch normalized input features based on a preset dilation rate to obtain a first intermediate feature, wherein the preset dilation rate represents the spacing between convolution kernels; One-dimensional convolution processing is performed on the first intermediate feature, and the first intermediate feature after the one-dimensional convolution processing is subjected to a second activation function to obtain the first audio feature.

6. The method according to claim 4, characterized in that The performing a second processing on the first audio feature to obtain a second audio feature includes: Determine a current time step, and determine a third number based on the current time step, where the third number is the number of times the multi-scale fusion process is performed; performing multi-scale fusion processing on the first audio features one by one according to the third quantity, and determining the first audio features after the multi-scale fusion processing is performed one by one as second intermediate features; The second audio feature is obtained according to the second intermediate feature and the first audio feature.

7. The method according to claim 6, characterized in that The performing a second processing on the first audio feature to obtain a second audio feature includes: Performing a second processing on the first audio feature at each time step; The obtaining the second audio feature according to the second intermediate feature and the first audio feature includes: In response to the current time step being the first time step, obtaining the second audio feature according to the second intermediate feature and the first audio feature in a first preset manner; In response to the current time step not being the first time step, a feature set is determined, and the second audio feature is obtained according to the second intermediate feature, the feature set and the first audio feature through a first preset method, wherein the feature set includes the second audio feature outputted respectively at each time step before the current time step.

8. The method according to claim 4, characterized in that The performing a third processing on the second audio feature to obtain a third audio feature includes: Performing maximum pooling on the second audio feature to obtain a maximum pooling layer, and performing average pooling on the second audio feature to obtain an average pooling layer; Determining a third intermediate feature according to the first kernel size, the maximum pooling layer, the average pooling layer, and the second audio feature in a second preset manner; The third audio feature is determined according to the second kernel size and the third intermediate feature in a third preset manner.

9. An audio processing device, characterized in that: include: an acquisition unit, configured to acquire audio to be processed and acquire audio features of the audio to be processed, wherein the audio to be processed includes audio emitted by multiple sound sources; a processing unit, configured to perform audio separation on the audio features through a separator to obtain target audio features, wherein the separator includes a preset number of separation layers, and each of the preset number of separation layers sequentially performs a first process, a second process, and a third process, wherein the first process is based on dilated convolution, the second process is based on multi-scale fusion, and the third process is based on a multi-channel attention mechanism, and the target audio features include audio features corresponding to each of the multiple sound sources; The restoration unit is configured to perform audio restoration on the target audio feature to obtain target audio, where the target audio includes audio corresponding to each of the multiple sound sources.

10. The device according to claim 9, characterized in that The acquiring unit acquires the audio features of the audio to be processed in the following manner: Segmenting the time domain signal of the audio to be processed according to a preset length to obtain a plurality of audio segments having the preset length, wherein the preset length corresponds to a first number of audio sampling points, and adjacent audio segments in the plurality of audio segments have overlapping regions, wherein the overlapping regions correspond to a second number of audio sampling points, and the second number is smaller than the first number; The multiple audio segments are encoded using a one-dimensional convolution operation and a first activation function to obtain high-dimensional mixed features corresponding to the audio to be processed, and the high-dimensional mixed features are determined as the audio features.

11. The device according to claim 9, characterized in that The preset number of separation layers is a plurality of separation layers executed sequentially; The processing unit performs audio separation on the audio feature through a separator to obtain a target audio feature in the following manner: For a first separation layer among the preset number of separation layers, using the audio feature as an input feature, and performing data processing on the input feature to obtain an output feature; For each separation layer other than the first separation layer in the preset number of separation layers, the output feature of the previous separation layer is used as an input feature, and the input feature is processed to obtain an output feature; For the last separation layer of the preset number of separation layers, sound source separation processing is performed on the output features of the last separation layer to obtain target audio features.

12. The device according to claim 11, characterized in that For each separation layer in the preset number of separation layers, the processing unit processes the input feature in the following manner to obtain an output feature: Performing a first processing on the input feature to obtain a first audio feature, where the first audio feature is the input feature after the first processing, and the first processing is based on a dilated convolution processing method; Performing a second processing on the first audio feature to obtain a second audio feature, where the second audio feature is the first audio feature after the second processing, and the second processing is based on multi-scale fusion; Performing a third processing on the second audio feature to obtain a third audio feature, where the third audio feature is the second audio feature after the third processing, and the third processing is a processing method based on a multi-channel attention mechanism; The sum of the third audio feature and the input feature is determined as the output feature.

13. The device according to claim 12, characterized in that The processing unit performs a first processing on the input feature to obtain a first audio feature in the following manner: Performing batch normalization on the input features, and performing dilated convolution on the batch normalized input features based on a preset dilation rate to obtain a first intermediate feature, wherein the preset dilation rate represents the spacing between convolution kernels; One-dimensional convolution processing is performed on the first intermediate feature, and the first intermediate feature after the one-dimensional convolution processing is subjected to a second activation function to obtain the first audio feature.

14. The device according to claim 13, characterized in that The processing unit performs second processing on the first audio feature to obtain a second audio feature in the following manner: Determine a current time step, and determine a third number based on the current time step, where the third number is the number of times the multi-scale fusion process is performed; performing multi-scale fusion processing on the first audio features one by one according to the third quantity, and determining the first audio features after the multi-scale fusion processing is performed one by one as second intermediate features; The second audio feature is obtained according to the second intermediate feature and the first audio feature.

15. The device according to claim 14, characterized in that The processing unit performs second processing on the first audio feature to obtain a second audio feature in the following manner: Performing a second processing on the first audio feature at each time step; The obtaining the second audio feature according to the second intermediate feature and the first audio feature includes: In response to the current time step being the first time step, obtaining the second audio feature according to the second intermediate feature and the first audio feature in a first preset manner; In response to the current time step not being the first time step, a feature set is determined, and the second audio feature is obtained according to the second intermediate feature, the feature set and the first audio feature through a first preset method, wherein the feature set includes the second audio feature outputted respectively at each time step before the current time step.

16. The device according to claim 12, characterized in that The processing unit performs third processing on the second audio feature to obtain a third audio feature in the following manner: Performing maximum pooling on the second audio feature to obtain a maximum pooling layer, and performing average pooling on the second audio feature to obtain an average pooling layer; Determining a third intermediate feature according to the first kernel size, the maximum pooling layer, the average pooling layer, and the second audio feature in a second preset manner; The third audio feature is determined according to the second kernel size and the third intermediate feature in a third preset manner.

17. An electronic device, characterized in that: include: processor: a memory for storing processor-executable instructions; The processor is configured to execute the audio processing method according to any one of claims 1 to 8.

18. A storage medium, characterized in that The storage medium stores instructions. When the instructions in the storage medium are executed by a processor, the processor is enabled to execute the audio processing method according to any one of claims 1 to 8.