Audio separation method, training method, device, equipment, storage medium and product

CN117012223BActive Publication Date: 2026-09-04伟光有限公司(CN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210472271.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2026-09-04
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

[0004]然而,采用相关技术中方案进行音频分离时,通常需使用较长一段时间内音频数据对应信息进行分离,其分离延迟较高,无法实现实时音频分离

Benefits of technology

[0028] In this embodiment, when performing audio separation on the i-th audio spectrum of the i-th audio segment, audio separation is performed based on the features of the (i-1)-th audio segment during the audio separation process and the i-th audio spectrum, resulting in track masks for each audio track. This separates the audio spectrum into the corresponding track spectra for each track. In other words, by introducing the audio features obtained from the previous audio segment's separation process during each audio segment separation, the problem of insufficient information due to short-window input can be avoided, improving the accuracy of the separation network in separating audio segments with short-window input. Furthermore, using the method provided in this embodiment, when performing audio separation using the separation network, only audio segments input with a short time window can be separated each time, reducing separation latency and enabling real-time audio separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117012223B_ABST
    Figure CN117012223B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an audio separation method, a training method, an apparatus, a device, a storage medium and a product, relating to the field of artificial intelligence. The method comprises: obtaining an i-1th audio feature of an i-1th audio segment in an audio separation process; inputting the i-1th audio feature and an i th audio spectrum of an i th audio segment into a separation network to perform audio separation, to obtain a track mask of each track in the i th audio segment, the i-1th audio segment being a previous segment of the i th audio segment in a target audio; and performing spectrum extraction on the i th audio spectrum by using the track mask of each track, to obtain a track spectrum of each track. The method provided in the embodiments of the present application can reduce separation delay and improve the accuracy of the separation network in separating a short-window input audio segment, by introducing an audio feature obtained in the separation process of a previous segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to an audio separation method, training method, apparatus, device, storage medium, and product. Background Technology

[0002] Audio separation technology refers to the technique of extracting and separating the original audio tracks, such as human voices and instrument sounds, from audio.

[0003] In related technologies, audio separation algorithms based on artificial intelligence (AI) can be used for audio separation. In this process, a separation network is used to separate audio data over a period of time to obtain the audio spectrum corresponding to each audio track.

[0004] However, when using related technologies for audio separation, it is usually necessary to use audio data corresponding information over a relatively long period of time for separation, resulting in high separation latency and making real-time audio separation impossible. Summary of the Invention

[0005] This application provides an audio separation method, training method, apparatus, device, storage medium, and product. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide an audio separation method, the method comprising:

[0007] Obtain the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process;

[0008] The (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment are input into a separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio.

[0009] The audio spectrum of the i-th audio track is extracted using the track mask of each audio track to obtain the track spectrum of each audio track.

[0010] On the other hand, embodiments of this application provide a training method for a separation network, the method comprising:

[0011] Acquire sample track-by-track audio data, and perform mixing processing on the sample track-by-track audio data to obtain mixed audio data;

[0012] The mixed audio spectrum corresponding to the mixed audio data is input into the separation network for audio separation to obtain the predicted audio track mask for each audio track;

[0013] The spectrogram of the mixed audio is extracted using the predicted audio track masks to obtain the predicted audio track spectrogram of each track.

[0014] The separation network is updated and trained based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data.

[0015] On the other hand, embodiments of this application provide an audio separation device, the device comprising:

[0016] The acquisition module is used to acquire the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process;

[0017] An audio separation module is used to input the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into a separation network for audio separation, thereby obtaining the audio track mask of each audio track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio.

[0018] The spectrum extraction module is used to extract the spectrum of the i-th audio spectrum using the track mask of each audio track, so as to obtain the track spectrum of each audio track.

[0019] On the other hand, embodiments of this application provide a training apparatus for a separation network, the apparatus comprising:

[0020] The acquisition module is used to acquire sample track-by-track audio data and to perform mixing processing on the sample track-by-track audio data to obtain mixed audio data.

[0021] The audio separation module is used to input the mixed audio spectrum corresponding to the mixed audio data into the separation network for audio separation, and obtain the predicted audio track mask for each audio track;

[0022] The spectrum extraction module is used to extract the spectrum of the mixed audio using each of the predicted audio track masks to obtain the predicted audio track spectrum of each track.

[0023] The training module is used to update and train the separation network based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data.

[0024] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the audio separation method or the training method of the separation network as described above.

[0025] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the audio separation method or the training method of the separation network as described above.

[0026] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio separation method or the training method for the separation network provided in the various optional implementations of the above aspects.

[0027] The technical solution provided in this application can bring the following beneficial effects:

[0028] In this embodiment, when performing audio separation on the i-th audio spectrum of the i-th audio segment, audio separation is performed based on the features of the (i-1)-th audio segment during the audio separation process and the i-th audio spectrum, resulting in track masks for each audio track. This separates the audio spectrum into the corresponding track spectra for each track. In other words, by introducing the audio features obtained from the previous audio segment's separation process during each audio segment separation, the problem of insufficient information due to short-window input can be avoided, improving the accuracy of the separation network in separating audio segments with short-window input. Furthermore, using the method provided in this embodiment, when performing audio separation using the separation network, only audio segments input with a short time window can be separated each time, reducing separation latency and enabling real-time audio separation. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A flowchart of an audio separation method provided in an exemplary embodiment of this application is shown;

[0031] Figure 2 A schematic diagram illustrating two forms of discrete networks is shown in an exemplary embodiment of this application;

[0032] Figure 3 A schematic diagram of an audio separation process provided in an exemplary embodiment of this application is shown;

[0033] Figure 4 A flowchart of an audio separation method provided by another exemplary embodiment of this application is shown;

[0034] Figure 5 A schematic diagram of a time-domain feature filling process provided in an exemplary embodiment of this application is shown;

[0035] Figure 6 A schematic diagram of a buffer provided in an exemplary embodiment of this application is shown;

[0036] Figure 7 A flowchart of an audio separation method provided by another exemplary embodiment of this application is shown;

[0037] Figure 8 A schematic diagram of the network structure of a discrete network provided in an exemplary embodiment of this application is shown;

[0038] Figure 9 A schematic diagram of a causal convolution process provided in an exemplary embodiment of this application is shown;

[0039] Figure 10 A flowchart illustrating a training method for a separate network provided in another exemplary embodiment of this application is shown;

[0040] Figure 11 A flowchart illustrating a training method for a separate network provided in another exemplary embodiment of this application is shown;

[0041] Figure 12 A schematic diagram of a separate network training process provided in an exemplary embodiment of this application is shown;

[0042] Figure 13 This paper shows a structural block diagram of an audio separation device provided in one embodiment of the present application;

[0043] Figure 14 A structural block diagram of an audio separation device provided in another embodiment of this application is shown;

[0044] Figure 15 A structural block diagram of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0046] In related technologies, to ensure accuracy during audio separation, audio segments are typically input into the separation network over a relatively long time window. For example, a 12-second audio data segment might be used as the input window. However, this approach introduces a significant separation delay due to the longer input time window. For instance, with a 12-second input window, 12 seconds of audio data must be received, making real-time audio separation impossible. Conversely, inputting audio segments into the separation network with a shorter time window results in reduced accuracy due to the smaller data volume.

[0047] In this embodiment of the application, in order to achieve real-time audio separation and ensure the accuracy of separation during the real-time audio separation process, the separation features of the previous audio segment are introduced during the audio separation process of each audio segment to obtain more information for audio separation. This can improve the accuracy of the separation network when performing audio separation on input audio in a short time window, thereby achieving real-time audio separation.

[0048] The method provided in this application can be applied to any scenario that requires audio separation. For example, during music playback, the audio separation method can be used to separate the music into audio tracks corresponding to different audio tracks, thereby improving the stereo effect of music playback.

[0049] The method provided in this application can be applied to computer devices, which are electronic devices that provide audio separation functionality. These electronic devices can be mobile terminals such as smartphones, tablets, and laptops, or terminals such as desktop computers and projector computers; this application does not limit the specific type. When audio separation is required, the target audio can be input to the computer device, which then uses a separation network to perform audio separation on the target audio, obtaining the audio data corresponding to each audio track in the target audio.

[0050] Please refer to Figure 1 This document illustrates a flowchart of an audio separation method provided in an exemplary embodiment of this application. The embodiment uses the application of this method to a computer device as an example for illustration. The method includes:

[0051] Step 101: Obtain the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process.

[0052] In one possible implementation, during the audio separation process using a separation network, the target audio is divided into multiple audio segments, and each audio segment is separated sequentially. That is, after separating the (i-1)th audio segment, the ith audio segment is then separated.

[0053] To achieve real-time audio separation, the audio segments input to the separation network each time need to be short to avoid significant latency; that is, input should be done within a short time window. However, if only short audio segments are input to the separation network each time, the amount of data contained in the short segments will be small, affecting the separation accuracy. Therefore, in one possible implementation, during the audio separation of the i-th audio segment, the computer device acquires the (i-1)-th audio features of the (i-1)-th audio segment during the audio separation process, thereby fusing the features extracted in the previous audio segment separation process to improve the audio separation accuracy.

[0054] Step 102: Input the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into the separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio.

[0055] When performing audio separation on the i-th audio segment, a time-frequency transformation must first be performed on the i-th audio segment. This time-frequency transformation uses a Short-Time Fourier Transform (STFT) to transform the audio data to the frequency domain, obtaining the complex spectrum of the i-th audio segment, i.e., the i-th audio spectrum. The time-frequency transformation method is as follows:

[0056] X = STFT(x)

[0057] Where x is the audio data corresponding to the i-th audio segment, and X is the audio spectrum of the i-th audio segment.

[0058] Optionally, after obtaining the i-th audio spectrum, the computer device inputs the i-th audio spectrum into the separation network, and simultaneously inputs the (i-1)-th audio features into the separation network. The separation network is then used to perform audio separation on the i-th audio spectrum to obtain the audio track masks for each audio track.

[0059] Simultaneously, the separation network will output the i-th audio feature of the i-th audio spectrum during the separation process, which will be used for the audio separation process of the (i+1)-th audio segment.

[0060] The processing procedure for the separation network is shown in the following equation:

[0061] Net(X, H) i-1 ) = ([m0, m1, m2, ..., m N-1 ], H i )

[0062] Among them, H i-1 This represents the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process, H. iLet m represent the i-th audio feature of the i-th audio segment during the audio separation process. N-1 This represents the track mask corresponding to the (N-1)th audio track.

[0063] In one possible implementation, the separation network can be a single separation network, such as... Figure 2 As shown, the audio spectrum can be input into the separation network 201 for audio separation, and the separation network can output audio track masks for each audio track. Alternatively, the separation network can also be a network containing multiple separation subnetworks, such as... Figure 2 As shown, the audio spectrum is input into the separation network 202, and audio separation is performed using each separation sub-network in the separation network 202. Each separation sub-network can output the audio track mask corresponding to each audio track.

[0064] Furthermore, during the audio separation process of the target audio, the entire audio data corresponding to the target audio can be subjected to time-frequency transformation to obtain the complex spectrum of the target audio, which is then input into the separation network. The separation network decomposes the complex spectrum of the target audio into complex spectra corresponding to multiple audio segments, thereby performing audio separation on the complex spectrum of each audio segment separately. Alternatively, in another possible implementation, time-frequency transformation can be performed on each audio segment separately to obtain the complex spectrum of each audio segment, which is then input into the separation network for audio separation. This application embodiment does not limit this approach.

[0065] Step 103: Use the track mask of each audio track to extract the spectrum of the i-th audio spectrum to obtain the track spectrum of each audio track.

[0066] After obtaining the track masks for each audio track, the computer device can use the track masks to extract the spectrum of the i-th audio frequency. The track mask is also in complex form. Specifically, when extracting the spectrum using the track mask, multiplying the track mask by the spectrum of the i-th audio frequency yields the spectrum corresponding to each track. The spectrum extraction process is as follows:

[0067]

[0068]

[0069] in, This represents the audio spectrum corresponding to the i-th audio track.

[0070] Optionally, after obtaining the audio spectrum of each track, an inverse time-frequency transform is performed on the spectrum of each track. This inverse time-frequency transform can be achieved using the Inverse Short-Time Fourier Transform (ISTFT) to transform the audio spectrum into the time domain, thus obtaining the audio data corresponding to each track. That is:

[0071]

[0072] in, This refers to the audio data corresponding to the i-th audio track.

[0073] In one possible implementation, the audio separation process for the i-th audio segment is as follows: Figure 3 As shown, the computer device first performs a time-frequency transformation 301 on the i-th audio segment to obtain the i-th audio spectrum. Then, the i-th audio spectrum and the (i-1)-th audio feature are input into the separation network 302 to obtain the audio track masks corresponding to tracks 0 to N-1. Simultaneously, the separation network 302 outputs the i-th audio feature, which is used for the audio separation process of the (i+1)-th audio segment. Afterward, the computer device multiplies the i-th audio spectrum with the audio track masks of each track to obtain the audio track spectra corresponding to tracks 0 to N-1. Then, it performs an inverse time-frequency transformation 303 on each audio track spectrum to obtain the audio data corresponding to tracks 0 to N-1.

[0074] In summary, in this embodiment, when performing audio separation on the i-th audio spectrum of the i-th audio segment, audio separation is performed based on the features of the (i-1)-th audio segment during the audio separation process and the i-th audio spectrum, resulting in track masks for each audio track. This separates the audio spectrum into the corresponding track spectra for each track. That is, by introducing the audio features obtained from the previous audio segment's separation process during each audio segment separation, the problem of insufficient information due to short-window input can be avoided, improving the accuracy of the separation network in separating audio segments with short-window input. Furthermore, using the method provided in this embodiment, when performing audio separation using the separation network, only audio segments input with a short time window can be separated each time, reducing separation latency and enabling real-time audio separation.

[0075] In one possible implementation, the separation network includes multiple convolutional layers, wherein a computer device utilizes these layers to perform convolutional processing to obtain audio features corresponding to the audio spectrum. In this embodiment, during the convolutional processing in the convolutional layers, audio features from the previous segment in the separation process are introduced for causal convolution processing, thereby fusing audio features from historical separation processes to achieve audio separation. The causal convolution process will be described exemplarily below.

[0076] Please refer to Figure 4 This document illustrates a flowchart of an audio separation method provided in another exemplary embodiment of this application. The embodiments of this application use an application of this method to a computer device as an example. The method includes:

[0077] Step 401: Obtain the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process.

[0078] The implementation method for this step can refer to step 101 above, and will not be repeated in this embodiment.

[0079] Step 402: Determine each (i-1)th layer audio feature in the (i-1)th audio feature, wherein different (i-1)th layer audio features correspond to different convolutional layers in the separation network.

[0080] In one possible implementation, the separation network includes multiple convolutional layers to perform convolutional processing on the audio features. In this embodiment, to enable the separation network to accept input within a short time window (i.e., each audio segment separation is relatively short), the computer device transfers information during the separation of adjacent audio segments. This information transfer process is forward-propagating, meaning that information from the previous audio segment's separation process is passed to the current audio segment's separation process, allowing the current audio segment's separation process to perform audio separation based on more information.

[0081] The information transmission process is essentially a transmission of convolutional states, where each convolutional state represents the (i-1)th audio feature during the audio segment separation process. In one possible implementation, during information transmission, the convolutional state of the convolutional layer for the (i-1)th audio segment during audio segment separation is transmitted to the corresponding convolutional layer during the audio segment separation process for the ith audio segment. That is, the (i-1)th audio feature contains the convolutional states corresponding to different convolutional layers during audio segment separation, and the convolutional states corresponding to different convolutional layers constitute different (i-1)th layer audio features. After obtaining the (i-1)th audio feature, it is necessary to determine the (i-1)th layer audio features corresponding to each convolutional layer, and then input them into the convolutional layer corresponding to the separation network.

[0082] Step 403: Input the audio features of each (i-1)th layer into the audio features of the i-th audio spectrum in each convolutional layer and perform causal convolution processing to obtain the audio track mask of each audio track.

[0083] To ensure the continuity of information transmission, this embodiment employs causal convolution to transfer the (i-1)th audio feature to the audio segment separation process of the ith audio segment. Optionally, this step may include the following:

[0084] Step 403a: Based on the (i-1)th layer audio features corresponding to the kth convolutional layer, feature filling is performed on the kth audio input features of the kth convolutional layer to obtain the kth audio filling features.

[0085] During convolution processing, to ensure that the input and output feature maps of the convolutional layer have the same size, the input features are typically padded to ensure the output size matches the input size. Taking a 3×3 convolutional kernel as an example, a column needs to be padded to the left and right sides of the feature map, and a row needs to be padded to the top and bottom to ensure the output size. In one possible implementation, the computer device uses the (i-1)th layer's audio features to padded the audio input features of the convolutional layer, obtaining padded audio features, and then performs convolution processing on these padded features in the convolutional layer. The process of padded the k-th audio input features of the k-th convolutional layer may include the following steps:

[0086] Step 1: Use the (i-1)th layer audio features corresponding to the kth convolutional layer to perform temporal feature filling on the kth audio input features.

[0087] In audio separation, two-dimensional convolutional layers are typically used for feature extraction in both the temporal and frequency domains. Therefore, feature padding must be performed in both the temporal and frequency domains to ensure the output size. For example, if each input channel has a size of N*T (where N represents N frequency points and T represents T time frames) and the convolutional kernel is 3×3, the final input size of the convolutional layer after feature padding is (N+2)*(T+2).

[0088] In one possible implementation, when performing causal convolution processing, the temporal padding method will be changed, and the computer device will use the (i-1)th layer audio features corresponding to the kth convolutional layer to perform temporal feature padding on the kth audio input features.

[0089] In related technologies, temporal padding is typically performed within historical and future time frames, with padding values ​​of 0. For example, with a 3×3 convolutional kernel, one column is filled on the left and right sides of the feature map, representing information from one historical and one future time frame. However, this method only ensures the output size of the convolutional layer, but the padding information is meaningless. This embodiment employs causal convolution, where the padding value is no longer 0, but rather the convolution state from the forward propagation is used. Furthermore, padding is performed within historical time frames, thus decomposing the convolution operation on a longer input feature map into several convolution operations on shorter feature maps. This allows the separation network to use a short time window as input for audio separation.

[0090] Optionally, within a historical time frame, the (i-1)th layer audio features corresponding to the kth convolutional layer are used to perform temporal feature filling on the kth audio input features.

[0091] In this implementation, the feature size of the (i-1)th layer audio feature is the same as the difference in temporal feature size between the input and output features of the convolutional layer. In one possible implementation, when storing the (i-1)th layer audio features of the (i-1)th audio segment during the audio separation process, the computer device determines the features to be stored based on the difference in temporal feature size between the input and output features of the convolutional layer in the separation network. Taking a 3×3 convolutional kernel as an example, when performing temporal feature padding, padding needs to be performed within the last two historical time frames. Therefore, the (i-1)th layer audio feature will include the input audio features of the last two time frames of the convolutional layer. Figure 5 As shown, the k-th audio input feature contains features from frames 0 to T-1 in the time domain. The (i-1)-th layered audio features will be used to fill the time domain features within the two historical time frames 501, causing the feature size to change to T+2 in the time domain. Furthermore, for the current k-th audio input feature, the i-th layered audio features from the last two time frames 502 will be stored and passed to the audio separation process for the (i+1)-th audio segment.

[0092] Step 2: Perform frequency domain feature filling on the k-th input audio feature after time domain feature filling to obtain the k-th audio filled feature.

[0093] In one possible implementation, the frequency domain feature filling method remains unchanged, that is, frequency domain feature filling is performed on the k-th input audio feature after time domain feature filling at high frequency points and low frequency points respectively.

[0094] Optionally, a padding value of 0 can be used to supplement information for higher and lower frequency points. For example... Figure 5 As shown, when the convolution kernel is 3×3, zeros are filled in the high-frequency point 503 and the low-frequency point 504, so that the feature size in the frequency domain changes to N+2.

[0095] The above method is illustrated by taking the time-domain feature filling followed by the spectral feature filling as an example. However, in the actual separation process, the frequency-domain feature filling can also be performed first, followed by the time-domain feature filling. This embodiment only describes the filling method and does not limit the filling timing.

[0096] Step 403b: Perform causal convolution processing on the k-th audio fill feature to obtain the k-th audio output feature.

[0097] After the k-th audio input feature is filled and the k-th audio filled feature is obtained, the computer device performs two-dimensional causal convolution processing on the k-th audio filled feature in the k-th convolutional layer to obtain the k-th audio output feature of the k-th convolutional layer.

[0098] Step 403c: Based on the nth audio output features of the nth convolutional layer, determine the audio track mask for each audio track, where n≥k.

[0099] After performing causal convolution processing using multiple convolutional layers, the audio track mask for each track can be determined based on the nth audio output feature of the nth convolutional layer. The nth convolutional layer can be the last convolutional layer contained in the final output layer of the separation network.

[0100] Step 404: Store the i-th audio feature of the i-th audio segment in the buffer during the audio separation process.

[0101] After performing audio separation on the i-th audio segment, the computer device can store the i-th audio features from the separation process. These i-th audio features also include the i-th layer audio features corresponding to each convolutional layer. (Illustrative example follows.) Figure 5 As shown, the audio features within the last two time frames are the audio features of the i-th layer corresponding to the k-th convolutional layer.

[0102] In one possible implementation, the audio features of each i-th layer can be merged and stored. During the next audio separation process, the i-th audio features are then segmented again to obtain the individual i-th layer audio features, which are then input into the corresponding convolutional layers.

[0103] In another possible implementation, the computer device incorporates a neural network accelerator to accelerate the audio separation process of the separation network. To further accelerate the audio separation process, a buffer can be set up within the neural network accelerator to store audio features that need to be passed to the next step during audio separation. Each convolutional layer has a corresponding buffer partition, and audio features corresponding to different convolutional layers can be stored in different buffer partitions. The separation network can read the convolutional state (i.e., audio features) generated in the previous audio separation process from the buffer partitions corresponding to each convolutional layer, thereby performing temporal feature filling on the input feature map of the current convolutional layer. In this approach, there is no need to merge and segment the stored audio features, reducing additional overhead and power consumption, and further accelerating the audio separation efficiency of the separation network.

[0104] Indicative, such as Figure 6 As shown, a buffer 602 is provided inside the neural network accelerator 601, which contains buffer partitions corresponding to convolutional layers 1 to M.

[0105] Optionally, during the storage of the i-th audio feature, the i-th layer audio features corresponding to each convolutional layer in the i-th audio feature are stored in each buffer partition, with different buffer partitions corresponding to different convolutional layers.

[0106] Furthermore, when storing each i-th layer audio feature in the buffer partition, since the buffer partition originally stored the (i-1)-th layer audio features, the (i-1)-th layer audio features need to be cleared, and the i-th layer audio features are stored in the cleared buffer partition. That is, the i-th layer audio features replace the (i-1)-th layer audio features in the buffer partition.

[0107] In this embodiment, a causal convolution processing method is adopted to forward the information of the audio separation process. That is, the features of the previous audio separation process are used to fill the features of the current audio separation process. This can split the input features of the long window into the input features of the short window. Therefore, when using the separation network to separate audio segments with short time windows, the separation accuracy caused by insufficient information is avoided. This helps to improve the accuracy of audio separation, reduce separation latency, and achieve real-time audio separation.

[0108] Furthermore, in this embodiment, a dedicated buffer is set up for the audio features that need to be passed to the next audio separation process in each audio separation process. This avoids merging and segmenting the hierarchical audio features corresponding to all convolutional layers, which can improve the separation efficiency of the separation network for audio separation and help reduce power consumption.

[0109] In one possible implementation, the separation network is a U-Net network, which includes an input layer, an encoding block, a decoding block, and an output layer, wherein each layer contains different convolutional layers. In this embodiment, the input of the separation network includes audio features from the previous audio separation process, which are respectively input into the convolutional layers corresponding to the input layer, encoding block, decoding block, and output layer. Exemplary embodiments will be described below.

[0110] Please refer to Figure 7 This document illustrates a flowchart of an audio separation method provided in another exemplary embodiment of this application. The embodiments of this application use an application of this method to a computer device as an example. The method includes:

[0111] Step 701: Obtain the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process.

[0112] The implementation method for this step can refer to step 101 above, and will not be repeated in this embodiment.

[0113] Step 702: Determine each (i-1)th layer audio feature in the (i-1)th audio feature, wherein different (i-1)th layer audio features correspond to different convolutional layers in the separation network.

[0114] The implementation method for this step can refer to step 402 above, and will not be repeated in this embodiment.

[0115] Step 703: Extract the audio features of the (i-1)th layer corresponding to the input layer and the audio spectrum of the i-th layer into the input layer to obtain the initial audio features.

[0116] In one possible implementation, the input layer of the separation network consists of one or more two-dimensional convolutional layers. The input layer first extracts features from the i-th audio spectrum of the input, obtaining initial audio features, i.e., the initial input feature map. Since the input layer contains convolutional layers, the input of the input layer includes the (i-1)-th layer audio features corresponding to the (i-1)-th audio segment during the audio separation process. Furthermore, when the input layer contains multiple convolutional layers, the (i-1)-th layer audio features corresponding to the input layer include the (i-1)-th layer audio features from multiple convolutional layers. During the input process, the computer device inputs each (i-1)-th layer audio feature into the convolutional layer corresponding to the input layer.

[0117] Step 704: Input the (i-1)th layer audio features and the initial audio features corresponding to the coding block into the coding block for feature encoding to obtain the audio coding features.

[0118] After obtaining the initial audio features, the computer device inputs these features into N coding blocks for encoding. Each coding block consists of an encoder and a downsampling unit. The encoder consists of one or more two-dimensional convolutional layers. Each convolutional layer can be followed by a Rectified Linear Unit (ReLU) activation layer, i.e.:

[0119] y = max(0, x)

[0120] Downsampling, on the other hand, downsamples audio features in both the temporal and frequency domains. For example, when downsampling in the temporal domain, features from two time frames can be merged into features from a single time frame, thus coarsening the coding scale. Because the coding scale increases from fine to coarse, different numbers of channels can be set in the convolutional layers of different coding blocks. When the coding scale is coarser, more channels can be used to learn more feature information, reducing the probability of inaccurate separation caused by information loss during downsampling.

[0121] The encoding process for each coded block is shown in the following formula:

[0122] encoding i =Encoder(x i )

[0123] x i+1 =Downsample(encoding) i )

[0124] Among them, encoding iIt is the encoded information output by the i-th coded block, while x i+1 It is the input provided to the next encoded block.

[0125] After encoding using each coding block, the output features are input into the bottleneck layer. The bottleneck layer is an encoder with the coarsest coding scale and the smallest data volume. After passing through the bottleneck layer, the encoding is completed, and the audio coded features are obtained.

[0126] Since each coding block contains different convolutional layers, and the bottleneck layer also contains convolutional layers, the inputs of each coding block and the bottleneck layer contain the corresponding (i-1)th layer of audio features. (Illustrative example follows.) Figure 8 As shown, the i-th audio spectrum is input into input layer 801, and simultaneously, the (i-1)-th layer audio features corresponding to the input layer are input into the input layer to obtain initial audio features, which are then input into coding block 802. Each coding block 802 simultaneously inputs the corresponding (i-1)-th layer audio features. After encoding using N coding blocks, the output features are input into bottleneck layer 803 for encoding, while simultaneously inputting the corresponding (i-1)-th layer audio features. Bottleneck layer 803 outputs the final audio encoded features. In one possible implementation, the (i-1)-th layer audio features corresponding to each convolutional layer are obtained by segmenting the (i-1)-th audio features.

[0127] Step 705: Input the (i-1)th layer audio features and audio encoding features corresponding to the decoding block into the decoding block for feature decoding to obtain the audio decoding features.

[0128] After obtaining the audio encoded features, the computer device continues to input these features into N decoding blocks for feature decoding. Each decoding block contains an upsampler and a decoder. During decoding, because the upsampler performs upsampling in both the time and frequency domains, the decoding scale gradually recovers from coarse to fine. Correspondingly, the decoder also consists of one or more two-dimensional convolutional layers. Each decoding block upsamples the input features before decoding and inputting them into the next decoding block. Furthermore, since the encoding process includes downsampling, which results in the loss of some detail information, to recover this lost detail, the decoder's input features are concatenated with the encoder's output features of the corresponding size during decoding. This concatenation of the encoder's output features with the decoder's input features is then input into the decoder to recover the information lost due to downsampling, thereby improving the accuracy of audio separation. The decoding operation of the decoding block can be represented by the following formula:

[0129] y i+1 =Decoder(Concatenate(Upsample(y i ), encoding i)

[0130] Among them, y i This represents the input of the i-th decoded block, in relation to y. i After upsampling, the encoding information of the encoder in the i-th coding block is compared with the encoding information of the encoder. i Perform the splicing. i+1 This is the output of the i-th decoding block, which is the input of the (i+1)-th decoding block.

[0131] Similarly, the input of each decoding block contains the corresponding (i-1)th layer audio features.

[0132] like Figure 8 As shown, after the bottleneck layer 803 outputs the audio coding features, it is input into the decoding block 804, and the input of each decoding block 804 contains the audio features of the (i-1)th layer.

[0133] Step 706: Input the (i-1)th layer audio features and audio decoding features corresponding to the output layer into the output layer for feature separation to obtain the audio track mask for each audio track.

[0134] Finally, the audio decoding features, after being decoded by N decoding blocks, are input to the output layer for feature separation, yielding audio track masks for each track. The output layer also consists of one or more two-dimensional convolutional layers. Each convolutional layer in the output layer is followed by an activation layer, performing a tanh activation operation to control the output range between (-1, 1). The tanh activation operation is as follows:

[0135]

[0136] Here, x represents the output feature of the convolutional layer.

[0137] Correspondingly, the input to the output layer contains the corresponding (i-1)th layer's audio features. For example... Figure 8 As shown, the audio features of the (i-1)th layer corresponding to the output layer 805 and the audio decoding features after being decoded by N decoding blocks 804 are input into the output layer 805 to obtain the audio track mask of each track.

[0138] Furthermore, in the process of audio separation using the input layer, encoding block, bottleneck layer, decoding block, and output layer, each convolutional layer outputs the hierarchical audio features corresponding to the input, which are then used in the next audio separation process.

[0139] The feature size of the output layered audio features can be determined based on the size of the convolutional kernel of the convolutional layer. After output, the layered audio features can be merged and stored, or directly stored in the corresponding buffer partition.

[0140] like Figure 8As shown, the audio features of the i-th layer input from each convolutional layer in the input layer 801, encoding block 802, bottleneck layer 803, decoding block 804, and output layer 805 are merged to obtain the i-th audio feature, which is used for the audio separation process of the (i+1)-th audio segment.

[0141] As mentioned above, the convolution processes of the input layer, encoding block, decoding block, and each convolutional layer contained in the output layer of the U-Net network are all causal convolutions. Figure 9 As shown, the causal convolution process will be illustrated using three encoding blocks and three decoding blocks. Each convolutional layer has a 3×3 kernel, and since the frequency domain processing in causal convolution is the same as in ordinary convolution, Figure 9 The explanation focuses solely on the temporal processing. Each dot represents a time frame, and the lines connecting these dots represent the dependencies in the convolutional operations. During the audio separation of the i-th audio segment, the (i-1)-th audio feature 901 is input into the input of each convolutional layer, and temporal feature padding is performed on the input features of each convolutional layer. Furthermore, the input features of the last two frames of each convolutional layer are output to obtain the i-th audio feature 902, which is used for the audio separation of the (i+1)-th audio segment.

[0142] It should be noted that the diagram is only used to illustrate the presence of convolutional layers in the encoder and decoder. In the actual separation process, the downsampling and upsampling may also contain convolutional layers to perform causal convolution operations.

[0143] For an N-layer U-Net network, after N downsampling iterations, the bottleneck layer needs to retain at least one complete time frame. Therefore, the shortest input window length is 2. N Each time frame.

[0144] The duration of the i-th audio segment is determined by the number of layers in the U-Net network. Optionally, the duration of the i-th audio segment must be greater than or equal to 2. N The duration of each time frame, N is the number of network layers.

[0145] Indicative, such as Figure 9 As shown, when the U-Net network has 3 layers, the shortest input window is 8 time frames. If the sampling rate is 48kHz and each time frame has 1024 sampling points, the corresponding signal duration is 1024*8 / 48000 = 0.171s. That is, the method provided in this application embodiment can input audio data in 0.171s time windows each time, significantly reducing separation latency and making it suitable for real-time audio separation.

[0146] The above embodiments illustrate the process of audio separation using a separation network. The separation network is a pre-trained network, and the training process of the separation network will be illustrated below.

[0147] Please refer to Figure 10 It illustrates a flowchart of a training method for a split network provided in another exemplary embodiment of this application, which is applied to... Figure 1 Taking the computer device 110 shown as an example, the method includes:

[0148] Step 1001: Obtain sample track-by-track audio data and perform mixing processing on the sample track-by-track audio data to obtain mixed audio data.

[0149] In one possible implementation, the network parameters of the separation network are updated using sampled track-by-track audio data. First, a computer device performs a mixing process on the sampled track-by-track audio data, where the mixing is performed according to certain rules to obtain mixed audio data. The mixing process is shown in the following equation:

[0150]

[0151] Among them, s i Let α represent the sample track-by-track audio data, where i = 0, 1, 2...N-1. i This represents the mixing gain corresponding to the i-th audio track, which can be preset or randomly generated.

[0152] Step 1002: Input the mixed audio spectrum corresponding to the mixed audio data into the separation network for audio separation to obtain the predicted audio track mask for each audio track.

[0153] After obtaining the mixed audio data, the computer device performs time-frequency transformation on the mixed audio data, thereby transforming it to the frequency domain to obtain the corresponding mixed audio spectrum. Optionally, the STFT method is used to transform the mixed audio data to the frequency domain.

[0154] Next, a separation network is used to separate the mixed audio spectrum, obtaining the predicted audio track mask for each track. That is:

[0155] [m0, m1, m2, ..., m N-1 ]=Net(X)

[0156] Where X represents the mixed audio spectrum, and m N-1 This represents the predicted track mask for the (N-1)th track.

[0157] In one possible implementation, when using a separation network to perform audio separation on a mixed audio spectrum, the mixed audio spectrum is also split into multiple audio spectrum segments, which are then input into the separation network for audio separation in sequence.

[0158] Step 1003: Use the masks of each predicted audio track to extract the spectrum of the mixed audio spectrum and obtain the predicted audio track spectrum of each track.

[0159] Once the predicted audio track masks for each audio track are obtained, the spectrograms of the mixed audio can be extracted using these masks to obtain the predicted audio track spectrograms for each track. The spectrogram extraction process involves multiplying the predicted audio track mask by the mixed audio spectrum to obtain the predicted audio track spectrograms for each track.

[0160] Step 1004: Update and train the separation network based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data.

[0161] After obtaining the predicted audio spectrum for each audio track, the computer can use the predicted audio spectrum and the sample audio spectrum to update and train the separation network. The sample audio spectrum is a complex spectrum obtained by performing a time-frequency transformation on the audio data of each sample track.

[0162] In this embodiment, a large amount of sample track-by-track audio data is used to train the separation network, thereby improving the accuracy of audio separation. When using the separation network for short-window audio separation, separation accuracy can be ensured, thus achieving real-time audio separation.

[0163] Please refer to Figure 11 This document illustrates a flowchart of a training method for a separation network provided in another exemplary embodiment of this application. The embodiments of this application use the application of this method to a computer device as an example. The method includes:

[0164] Step 1101: Obtain sample track-by-track audio data and perform mixing processing on the sample track-by-track audio data to obtain mixed audio data.

[0165] The implementation method of this step can refer to step 1001 above, and will not be repeated here.

[0166] Step 1102: In each convolutional layer of the separation network, causal convolution processing is performed on the audio features of the mixed audio spectrum to obtain the audio track mask for each audio track.

[0167] In one possible implementation, the separation network is a U-Net network. During the audio separation process of the mixed audio spectrum using the separation network, the computer device utilizes the convolutional layers in the separation network to perform causal convolution processing on the corresponding audio features. Finally, after outputting through the output layer, the audio track masks for each track are obtained. The causal convolution processing may include the following steps:

[0168] Step 1102a: Perform feature padding on the k-th audio input feature of the k-th convolutional layer.

[0169] Optionally, the temporal feature padding method will change during causal convolution processing. The specific temporal feature padding method differs between the training and usage phases of the separation network. Since there is no latency constraint during training and real-time separation is not required, the input window during training can be a long window. When the input window is a long window, the impact of the padding value within a time frame on the separation result is negligible, and zero can be used directly for padding, meaning there is no need to use the audio features from the previous audio segment in the audio separation process. However, the convolution method remains causal convolution. Feature padding of the k-th audio input feature of the k-th convolutional layer may include the following steps:

[0170] Step 1: Within the historical time frame, perform temporal feature filling on the k-th audio input feature.

[0171] When padding in the temporal dimension, padding will be performed within historical time frames. Taking a 3×3 convolutional kernel as an example, padding will be performed in the left two columns of the feature map, that is, filling in information from two historical time frames. Since the audio features of a longer window have already been input, there is no need to use the audio features from the previous audio separation process; optionally, the padding value can be 0.

[0172] Step 2: At both high-frequency and low-frequency points, perform frequency domain feature filling on the k-th audio input feature.

[0173] Furthermore, frequency domain feature padding is performed on the k-th audio input feature at both high and low frequency points. Taking a 3×3 convolution kernel as an example, padding is performed on the upper and lower rows of the feature map, respectively filling in information at higher and lower frequency points.

[0174] Step 1102b: Perform causal convolution processing on the k-th audio input feature after feature filling.

[0175] After feature padding, in the convolutional layer, the computer device can perform causal convolution processing on the padded k-th audio input feature to obtain the audio output feature of the convolutional layer.

[0176] Finally, the output layer of the separation network will output the predicted audio track mask for each audio track.

[0177] Step 1103: Use the masks of each predicted audio track to extract the spectrum of the mixed audio spectrum and obtain the predicted audio track spectrum of each track.

[0178] The implementation method for this step can refer to step 1003 above, and will not be repeated in this embodiment.

[0179] Step 1104: Determine the contrast loss based on the predicted audio track spectrum and the sample audio track spectrum.

[0180] Once the predicted audio track spectrum is obtained, a loss function can be calculated between it and the sample audio track spectra of each track to obtain the contrast loss between the two. The contrast loss is determined as follows:

[0181]

[0182] in, S represents the predicted audio track spectrum. i This represents the spectrum of the sample audio track.

[0183] In the above method, the mean-square error (MSE) is used as the loss function to determine the contrast loss. In other possible methods, loss functions such as L1 and L2 can also be used to determine the contrast loss; this embodiment does not limit the loss function.

[0184] Step 1105: Based on the contrastive loss, perform reverse update training on the separation network.

[0185] Once the contrast loss for each audio track is determined, the computer can use this loss to backpropagate and update the separation network. Optionally, the network parameters can be updated using a gradient backpropagation algorithm until the loss function converges. The trained separation network can then be used for real-time audio separation.

[0186] In one possible implementation, the training process of the separation network is as follows: Figure 12 As shown, the sample track-segmented audio data is mixed 1201 to obtain mixed audio data. Then, the mixed audio data is subjected to time-frequency transformation 1202 to obtain the mixed audio spectrum. After that, the computer device inputs the mixed audio spectrum into the separation network 1203 for audio separation to obtain the predicted audio track mask for each track. Then, the mixed audio spectrum is multiplied by each predicted audio track mask to obtain the predicted audio track spectrum corresponding to each track. Then, the loss function is calculated 1204 using the sample audio track spectrum corresponding to the sample track-segmented audio data and the predicted audio track spectrum to obtain the contrast loss. The contrast loss is then used to update the separation network 1203 in reverse, realizing the update and training of the separation network.

[0187] In this embodiment, during the audio separation process of the mixed audio spectrum using the separation network, each convolutional layer adopts a causal convolution method, thereby making the separation network suitable for separating audio segments with short window input, reducing separation latency, and achieving real-time separation.

[0188] Please refer to Figure 13 This illustrates a structural block diagram of an audio separation device provided in one embodiment of this application. Figure 13 As shown, the device may include:

[0189] The acquisition module 1301 is used to acquire the (i-1)th audio feature of the (i-1)th audio segment during the audio separation process;

[0190] The audio separation module 1302 is used to input the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into the separation network for audio separation, and obtain the audio track mask of each audio track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio.

[0191] The spectrum extraction module 1303 is used to extract the spectrum of the i-th audio spectrum using the track mask of each audio track to obtain the track spectrum of each audio track.

[0192] Optionally, the audio separation module 1302 is further configured to:

[0193] Determine each (i-1)th layer audio feature in the (i-1)th audio feature, wherein different (i-1)th layer audio features correspond to different convolutional layers in the separation network;

[0194] Each of the (i-1)th layer audio features is input into each of the convolutional layers and causally convolved with the audio features of the i-th audio spectrum to obtain the audio track mask for each audio track.

[0195] Optionally, the audio separation module 1302 is further configured to:

[0196] Based on the (i-1)th layer audio features corresponding to the kth convolutional layer, feature filling is performed on the kth audio input features of the kth convolutional layer to obtain the kth audio filling features;

[0197] The k-th audio fill feature is subjected to causal convolution to obtain the k-th audio output feature;

[0198] Based on the nth audio output features of the nth convolutional layer, the audio track mask for each audio track is determined, where n ≥ k.

[0199] Optionally, the audio separation module 1302 is further configured to:

[0200] Temporal feature filling is performed on the k-th audio input feature using the (i-1)th layer audio features corresponding to the k-th convolutional layer;

[0201] The k-th input audio feature, after being filled with time-domain features, is filled with frequency-domain features to obtain the k-th audio filled feature.

[0202] Optionally, the audio separation module 1302 is further configured to:

[0203] Within a historical time frame, the (i-1)th layer audio features corresponding to the kth convolutional layer are used to fill the kth audio input features with temporal features;

[0204] The step of performing frequency domain feature filling on the k-th input audio feature after time-domain feature filling includes:

[0205] At high-frequency and low-frequency points, frequency domain feature filling is performed on the k-th input audio feature after time-domain feature filling.

[0206] Optionally, the separation network is a U-Net network, and each of the convolutional layers includes the input layer, encoding block, decoding block, and convolutional layer in the output layer of the separation network;

[0207] The audio separation module 1302 is also used for:

[0208] The audio features of the (i-1)th layer corresponding to the input layer and the i-th audio spectrum are input into the input layer for feature extraction to obtain the initial audio features;

[0209] The (i-1)th layer audio feature corresponding to the coding block and the initial audio feature are input into the coding block for feature encoding to obtain the audio coding feature;

[0210] The (i-1)th layer audio feature and the audio encoding feature corresponding to the decoding block are input into the decoding block for feature decoding to obtain the audio decoding feature;

[0211] The audio features of the (i-1)th layer corresponding to the output layer and the audio decoding features are input into the output layer for feature separation to obtain the audio track mask of each audio track.

[0212] Optionally, the duration of the i-th audio segment is determined based on the number of network layers in the U-Net network.

[0213] Optionally, the device further includes:

[0214] A storage module is used to store the i-th audio feature of the i-th audio segment in a buffer during the audio separation process.

[0215] Optionally, the storage module is further configured to:

[0216] The audio features of the i-th layer corresponding to each convolutional layer in the i-th audio feature are stored in each buffer partition, and different buffer partitions correspond to different convolutional layers.

[0217] In summary, in this embodiment, when performing audio separation on the i-th audio spectrum of the i-th audio segment, audio separation is performed based on the features of the (i-1)-th audio segment during the audio separation process and the i-th audio spectrum, resulting in track masks for each audio track. This separates the audio spectrum into the corresponding track spectra for each track. That is, by introducing the audio features obtained from the previous audio segment's separation process during each audio segment separation, the problem of insufficient information due to short-window input can be avoided, improving the accuracy of the separation network in separating audio segments with short-window input. Furthermore, using the method provided in this embodiment, when performing audio separation using the separation network, only audio segments input with a short time window can be separated each time, reducing separation latency and enabling real-time audio separation.

[0218] Please refer to Figure 14 This illustrates a structural block diagram of a training apparatus for a discrete network provided in another embodiment of this application. Figure 14 As shown, the device may include:

[0219] The acquisition module 1401 is used to acquire sample track-by-track audio data and to perform mixing processing on the sample track-by-track audio data to obtain mixed audio data.

[0220] The audio separation module 1402 is used to input the mixed audio spectrum corresponding to the mixed audio data into the separation network for audio separation to obtain the predicted audio track mask for each audio track;

[0221] The spectrum extraction module 1403 is used to extract the spectrum of the mixed audio using each of the predicted audio track masks to obtain the predicted audio track spectrum of each track.

[0222] The training module 1404 is used to update and train the separation network based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data.

[0223] Optionally, the audio separation module 1402 is further configured to:

[0224] In each convolutional layer of the separation network, causal convolution processing is performed on the audio features of the mixed audio spectrum to obtain the audio track mask for each audio track.

[0225] Optionally, the audio separation module 1402 is further configured to:

[0226] Feature imputation is performed on the k-th audio input feature of the k-th convolutional layer;

[0227] The k-th audio input feature after feature filling is subjected to causal convolution processing.

[0228] Optionally, the audio separation module 1402 is further configured to:

[0229] Within the historical time frame, the k-th audio input feature is filled with temporal features;

[0230] Frequency domain feature filling is performed on the k-th audio input feature at both high and low frequency points.

[0231] Optionally, the training module 1404 is further configured to:

[0232] Based on the predicted audio track spectrum and the sample audio track spectrum, determine the contrast loss;

[0233] Based on the contrast loss, the separation network is updated and trained in reverse.

[0234] In this embodiment, a large amount of sample track-by-track audio data is used to train the separation network, thereby improving the accuracy of audio separation. When using the separation network for short-window audio separation, separation accuracy can be ensured, thus achieving real-time audio separation.

[0235] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0236] Please refer to Figure 15 This diagram illustrates a structural block diagram of a computer device 1500 provided in an exemplary embodiment of this application. The computer device 1500 in this application may include one or more components such as a memory 1520 and a processor 1510.

[0237] Processor 1510 may include one or more processing cores. Processor 1510 connects to various parts within the computer device 1500 using various interfaces and lines, and performs various functions and processes data of the computer device 1500 by running or executing instructions, programs, code sets, or instruction sets stored in memory 1520, and by calling data stored in memory 1520. Optionally, processor 1510 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 1510 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on display screen 1530; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 1510 and may be implemented separately through a communication chip.

[0238] The memory 1520 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1520 may include a non-transitory computer-readable storage medium. The memory 1520 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1520 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described above, etc. The operating system may be an Android system (including systems deeply developed based on the Android system), an iOS system developed by Apple Inc. (including systems deeply developed based on the iOS system), or other systems. The data storage area may also store data created by the computer device 1500 during use (such as phone books, audio and video data, chat log data, etc.).

[0239] In addition, those skilled in the art will understand that the structure of the computer device 1500 shown in the above figures does not constitute a limitation on the computer device 1500. The computer device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the computer device 1500 also includes radio frequency circuits, imaging components, sensors, audio circuits, Wireless Fidelity (WiFi) components, power supplies, Bluetooth components, etc., which will not be described in detail here.

[0240] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the audio separation method or the training method for the separation network provided in any of the above exemplary embodiments.

[0241] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a printing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio separation method or the training method for the separation network provided in the optional implementation described above.

[0242] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0243] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio separation method, characterized in that, The method includes: Obtain the (i-1)th audio feature of the (i-1)th audio segment in the audio separation process. The (i-1)th audio feature includes the (i-1)th hierarchical audio features corresponding to different convolutional layers in the separation network. The (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment are input into the separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio, and the audio track mask is in complex form. The audio spectrum of the i-th audio track is extracted using the track mask of each audio track to obtain the track spectrum of each audio track; The step of inputting the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into the separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment includes: Based on the (i-1)th layer audio features corresponding to the kth convolutional layer, the i-th audio spectrum is filled with the k-th audio input features of the k-th convolutional layer to obtain the k-th audio filled features. The k-th audio fill feature is subjected to causal convolution to obtain the k-th audio output feature; Based on the nth audio output features of the nth convolutional layer, the audio track mask for each audio track is determined, where n ≥ k.

2. The method according to claim 1, characterized in that, The method of using the (i-1)th layer audio features corresponding to the kth convolutional layer to fill the kth audio input features of the i-th audio spectrum in the kth convolutional layer to obtain the kth audio filled features includes: Temporal feature filling is performed on the k-th audio input feature using the (i-1)th layer audio features corresponding to the k-th convolutional layer; The k-th input audio feature, after being filled with time-domain features, is filled with frequency-domain features to obtain the k-th audio filled feature.

3. The method according to claim 2, characterized in that, The step of using the (i-1)th layer audio features corresponding to the kth convolutional layer to perform temporal feature filling on the kth audio input features includes: Within a historical time frame, the (i-1)th layer audio features corresponding to the kth convolutional layer are used to fill the kth audio input features with temporal features; The step of performing frequency domain feature filling on the k-th input audio feature after time-domain feature filling includes: At high-frequency and low-frequency points, frequency domain feature filling is performed on the k-th input audio feature after time-domain feature filling.

4. The method according to any one of claims 1 to 3, characterized in that, The separation network is a U-Net network, and each of the convolutional layers includes the input layer, encoding block, decoding block, and convolutional layer in the output layer of the separation network. The step of inputting the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into the separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment includes: The audio features of the (i-1)th layer corresponding to the input layer and the i-th audio spectrum are input into the input layer for feature extraction to obtain the initial audio features; The (i-1)th layer audio feature corresponding to the coding block and the initial audio feature are input into the coding block for feature encoding to obtain the audio coding feature; The (i-1)th layer audio feature and the audio encoding feature corresponding to the decoding block are input into the decoding block for feature decoding to obtain the audio decoding feature; The audio features of the (i-1)th layer corresponding to the output layer and the audio decoding features are input into the output layer for feature separation to obtain the audio track mask of each audio track.

5. The method according to claim 4, characterized in that, The duration of the i-th audio segment is determined based on the number of network layers in the U-Net network.

6. The method according to any one of claims 1 to 3, characterized in that, After inputting the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into a separation network for audio separation to obtain the track masks of each audio track in the i-th audio segment, the method further includes: The i-th audio feature of the i-th audio segment during the audio separation process is stored in a buffer.

7. The method according to claim 6, characterized in that, The step of storing the i-th audio feature of the i-th audio segment in the buffer during the audio separation process includes: The audio features of the i-th layer corresponding to each convolutional layer in the i-th audio feature are stored in each buffer partition, and different buffer partitions correspond to different convolutional layers.

8. A training method for a separation network, characterized in that, The method includes: Acquire sample track-by-track audio data, and perform mixing processing on the sample track-by-track audio data to obtain mixed audio data; The mixed audio spectrum corresponding to the mixed audio data is input into the separation network, and causal convolution processing is performed on the audio features of the mixed audio spectrum in each convolutional layer of the separation network to obtain the predicted audio track mask for each audio track. The predicted audio track mask is in complex form. The spectrogram of the mixed audio is extracted using the predicted audio track masks to obtain the predicted audio track spectrogram of each track. Based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data, the separation network is updated and trained. Specifically, in each convolutional layer within the separation network, causal convolution processing is performed on the audio features of the mixed audio spectrum, including: Feature imputation is performed on the k-th audio input feature of the k-th convolutional layer; The k-th audio input feature after feature filling is subjected to causal convolution processing.

9. The method according to claim 8, characterized in that, The feature padding of the k-th audio input features of the k-th convolutional layer includes: Within the historical time frame, the k-th audio input feature is filled with temporal features; Frequency domain feature filling is performed on the k-th audio input feature at both high and low frequency points.

10. The method according to claim 8 or 9, characterized in that, The step of updating and training the separation network based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track-separated audio data includes: Based on the predicted audio track spectrum and the sample audio track spectrum, determine the contrast loss; Based on the contrast loss, the separation network is updated and trained in reverse.

11. An audio separation device, characterized in that, The device includes: The acquisition module is used to acquire the i-1th audio feature of the i-1th audio segment in the audio separation process. The i-1th audio feature includes the i-1th layered audio features corresponding to different convolutional layers in the separation network. An audio separation module is used to input the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into a separation network for audio separation, thereby obtaining the audio track mask of each track in the i-th audio segment. The (i-1)th audio segment is the previous segment of the i-th audio segment in the target audio, and the audio track mask is in complex form. The spectrum extraction module is used to extract the spectrum of the i-th audio spectrum using the track mask of each audio track to obtain the track spectrum of each audio track. The step of inputting the (i-1)th audio feature and the i-th audio spectrum of the i-th audio segment into the separation network for audio separation to obtain the audio track mask of each track in the i-th audio segment includes: Based on the (i-1)th layer audio features corresponding to the kth convolutional layer, the i-th audio spectrum is filled with the k-th audio input features of the k-th convolutional layer to obtain the k-th audio filled features. The k-th audio fill feature is subjected to causal convolution to obtain the k-th audio output feature; Based on the nth audio output features of the nth convolutional layer, the audio track mask for each audio track is determined, where n ≥ k.

12. A training device for a discrete network, characterized in that, The device includes: The acquisition module is used to acquire sample track-by-track audio data and to perform mixing processing on the sample track-by-track audio data to obtain mixed audio data. An audio separation module is used to input the mixed audio spectrum corresponding to the mixed audio data into a separation network, so that causal convolution processing is performed on the audio features of the mixed audio spectrum in each convolutional layer of the separation network to obtain the predicted audio track mask of each audio track, wherein the predicted audio track mask is in complex form; The spectrum extraction module is used to extract the spectrum of the mixed audio using each of the predicted audio track masks to obtain the predicted audio track spectrum of each track. The training module is used to update and train the separation network based on the predicted audio track spectrum and the sample audio track spectrum corresponding to the sample track audio data. Specifically, in each convolutional layer within the separation network, causal convolution processing is performed on the audio features of the mixed audio spectrum, including: Feature imputation is performed on the k-th audio input feature of the k-th convolutional layer; The k-th audio input feature after feature filling is subjected to causal convolution processing.

13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the audio separation method as described in any one of claims 1 to 7, or to implement the training method for the separation network as described in any one of claims 8 to 10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the audio separation method as described in any one of claims 1 to 7, or to implement the training method of the separation network as described in any one of claims 8 to 10.

15. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, which a processor reads from and executes to implement the audio separation method as described in any one of claims 1 to 7, or the training method for the separation network as described in any one of claims 8 to 10.

Citation Information

Patent Citations

  • Reduced latency streaming dynamic noise suppression using convolutional neural networks

    US20220084535A1