Audio processing method and device, electronic equipment and computer readable medium
By extracting the complex spectrum characteristics of the audio in the audio processing system and using the tracked music extraction model, the problems of computing power and delay in the audio source separation task of the existing audio processing system are solved, and real-time application scenarios suitable for holographic audio/spatial audio are realized.
Patent Information
- Application Number
- CN202311599955.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
The existing audio processing system has a large model in the audio source separation task, high computing power requirements, and obvious delays, which cannot meet the real-time application scenarios of holographic audio/spatial audio.
An audio processing method is proposed. By obtaining the target data segment of the audio to be processed, the complex spectrum features are input using the sub-track music extraction model to obtain the complex spectrum features of the target audio track, and the sub-audio bands of the target audio track are extracted based on this. This method simplifies the model architecture and avoids unnecessary delays and computing power consumption.
It realizes simplified model architecture in the audio processing system, reduces delay and computing power consumption, and is suitable for real-time application scenarios of holographic audio/spatial audio, and improves the sense of participation and playability of user-defined sound effects.
Smart Images

Figure CN120050589A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and more specifically, to an audio processing method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] With the continuous development of mobile terminal audio technology, spatial audio has broken through the two-channel limitation of stereo audio, opening up a new dimension of music and greatly enhancing the user's listening experience. In the music scenario, sound separation technology can adjust the sound field orientation and volume, or extract specific sound objects; provide users with a custom function that has the ability to independently create sound effect styles, improve user participation and playability, and expand the experience boundary of audio. Summary of the Invention
[0003] This application proposes an audio processing method, apparatus, electronic device, and computer-readable medium to improve the above-mentioned defects.
[0004] In a first aspect, this application provides an audio processing method, including: obtaining a target data segment of the audio to be processed, where the target data segment includes audio data of a continuous first number of frames; inputting the complex spectral features corresponding to the target data segment into a multi-track music extraction model to obtain the complex spectral features of the target track as the first spectral features; and obtaining a corresponding sub-audio segment of the target track based on the first spectral features.
[0005] In a second aspect, this application further provides an audio processing apparatus, including: an obtaining unit, an extracting unit, a determining unit, a looping unit, and a synthesizing unit. The obtaining unit is configured to obtain a target data segment of the audio to be processed, where the target data segment includes audio data of a continuous first number of frames. The extracting unit is configured to input the complex spectral features corresponding to the target data segment into a multi-track music extraction model to obtain the complex spectral features of the target track as the first spectral features. The determining unit is configured to obtain a corresponding sub-audio segment of the target track based on the first spectral features.
[0006] In a third aspect, this application further provides an electronic device, including: one or more processors; a memory; and one or more applications, where the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the above method.
[0007] In a fourth aspect, this application further provides a computer-readable medium, where the readable storage medium stores program code executable by a processor, and when the program code is executed by the processor, the processor is caused to execute the above method.
[0008] The audio processing method, apparatus, electronic device, and computer-readable medium provided by this application obtain a target data segment of the audio to be processed, where the target data segment includes audio data of a continuous first number of frames; input the complex spectral features corresponding to the target data segment into a multi-track music extraction model to obtain the complex spectral features of the target track as the first spectral features; and obtain the corresponding sub-audio segment of the target track based on the first spectral features. Therefore, through the multi-track music extraction model, the sub-music frequency band corresponding to the target track of the target data segment of the audio to be processed can be extracted, and by using the complex spectral features of the audio as the input, both the amplitude and phase are predicted simultaneously, avoiding unnecessary time delay and computing power consumption.
[0009] Other features and advantages of this application will be described in the subsequent description, and part of them will become obvious from the description or be understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings. Brief Description of the Drawings
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0011] Figure 1 Shows a schematic diagram of the audio processing system provided by the embodiments of this application;
[0012] Figure 2 Shows a flowchart of the audio processing method provided by an embodiment of this application;
[0013] Figure 3 Shows a flowchart of the audio processing method provided by another embodiment of this application;
[0014] Figure 4 Shows a schematic diagram of the sound source separation system with a streaming architecture provided by the embodiments of this application;
[0015] Figure 5 Shows an architecture diagram of the multi-track music extraction model provided by the embodiments of this application;
[0016] Figure 6 Shows an architecture diagram of TFC-TDF provided by the embodiments of this application;
[0017] Figure 7 Shows a state change diagram of the multi-track music extraction model provided by the embodiments of this application;
[0018] Figure 8 It shows the latency analysis diagram provided by the embodiments of the present application;
[0019] Figure 9 It shows the module block diagram of the audio processing device provided by the embodiments of the present application;
[0020] Figure 10 It shows the structural block diagram of the electronic device provided by the embodiments of the present application;
[0021] Figure 11 It shows the storage unit for storing or carrying the program code for implementing the audio processing method according to the embodiments of the present application. Detailed implementation manners
[0022] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0023] It should be noted that: similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0024] With the continuous development of mobile terminal audio technology, spatial audio has broken through the two-channel limitation of stereo audio, opened up a new dimension of music, and greatly improved the user's listening experience. In the music scene, sound separation technology can adjust the sound field orientation and volume, or extract specific sound objects; provide users with a custom function with the ability to independently create sound effect styles, improve the user's sense of participation and playability, and expand the experience boundary of audio. In the long run, the separation technology will support a wider range of sound object separation, and its importance is analogous to image segmentation in the field of images, and will provide a precondition for subsequent immersive spatial rendering and various sound effects of sound object post-processing.
[0025] Music Source Separation (MSS) is the task of decomposing music into its constituent parts. For example, separating music into tracks such as vocals, bass, drums, and other. Currently, mainstream MSS tasks use AI technology to extract different tracks from regular stereo music. The separated tracks can be processed individually using track enhancement or spatial rendering algorithms and then remixed again to build various applications. In the recreated music, the extracted tracks can be automatically placed at different positions in a virtual space to provide an immersive listening experience.
[0026] However, the inventors found in their research that although the current mainstream MSS systems can achieve good results, the models are relatively large, have high requirements for computing power, and the latency of the algorithms is also obvious. Whether from the perspective of space occupancy, computing power consumption, or algorithm latency, they cannot be applied to real-time application scenarios of holographic audio / spatial audio.
[0027] Therefore, to overcome the above defects, the embodiments of the present application provide an audio processing method that can simplify the model architecture and avoid unnecessary latency and computing power consumption.
[0028] It should be noted that this audio processing method is applied to an audio processing system. Before introducing this audio processing method below, the audio processing system will be introduced first, as Figure 1 shown, the sound source separation system 100 includes a preprocessing model 101, a multi-track music extraction model 102, and a post-processing model 103. The input of the sound source separation system 100 is mixed audio with 2 channels, and the output includes independent track audio with 8 channels. For example, if the mixed audio is a stereo song, then the audio corresponding to vocals, drums, bass, and other tracks is output. Exemplarily, the preprocessing model 101 is used to convert the time-domain signal of the input stereo music into a frequency-domain signal, and this feature is then sent to the multi-track music extraction model 102 to generate complex spectral features of 4 independent tracks. Then, the post-processing model 103 reconstructs the time-domain signal of each track.
[0029] Please refer to Figure 2 , Figure 2 shows an audio processing method provided by the embodiments of the present application. This method is applied to the above-mentioned sound source separation system 100. Specifically, this method includes: S201 to S205.
[0030] S201: Obtain a target data segment of the audio to be processed, where the target data segment includes audio data of a continuous first number of frames.
[0031] As an implementation, the starting data segment of the audio to be processed can be the audio data of the first quantity of consecutive frames including the first audio frame of the audio to be processed. For example, if the first quantity is 8, the starting data segment is the first 8-frame audio data of the audio to be processed. It can be understood that in the audio processing method provided by the embodiments of the present application, every 8-frame audio data is processed once.
[0032] S202: Input the complex spectral features corresponding to the target data segment into the multi-track music extraction model to obtain the complex spectral features of the target track, which are used as the first spectral features.
[0033] As an implementation, the target data segment is processed by the aforementioned preprocessing model to obtain the complex spectral features of the target data segment. Exemplarily, the preprocessing model converts the time-domain signal of the target data segment into a frequency-domain signal through short-time Fourier transform (STFT) to obtain the complex spectral features of the target data segment. The complex spectral features corresponding to the target data segment are input into the multi-track music extraction model, and the multi-track music extraction model can extract the complex spectral features of the target track from the complex spectral features corresponding to the target data segment. The complex spectral features of the target track are used as the first spectral features. Among them, the multi-track music extraction model can be a deep learning model, a non-negative matrix factorization model, etc. For example, a track separation model based on deep learning uses structures such as convolutional neural network (CNN), recurrent neural network (RNN), or transformer, and learns the representations of different tracks in the audio signal through training to achieve track separation; the non-negative matrix factorization technology can be used not only for sound source separation but also for track separation. By performing NMF decomposition on the audio signal, matrices of different tracks can be obtained, thereby achieving track separation.
[0034] It should be noted that the number of the target tracks is not limited in the embodiments of the present application. It can be one or multiple. If the target tracks are multiple, each track included in the target tracks is named a specified track, and the obtained first spectral features refer to the complex spectral features corresponding to each specified track. That is to say, after the complex spectral features corresponding to the target data segment are input into the multi-track music extraction model, the multi-track music extraction model outputs the complex spectral features corresponding to each specified track.
[0035] S203: Based on the first spectral features, obtain the corresponding sub-audio segment of the target track.
[0036] It can be understood that the sub-audio segment corresponding to the target track is a time-domain signal. That is to say, the implementation of obtaining the corresponding sub-audio segment of the target track based on the first spectral features can be to convert the first spectral features into a time-domain signal, thereby obtaining the corresponding sub-audio segment of the target track.
[0037] As described above, there may be multiple target audio tracks, that is, the target audio tracks may include multiple specified audio tracks. After obtaining the plural spectral features corresponding to each specified audio track in S202, inverse short-time Fourier transform (ISTFT) is performed on the plural spectral features corresponding to each specified audio track to reconstruct the time-domain signal of each audio track.
[0038] It should be noted that when performing S202 and S203, the relevant data for processing the previous target data segment may still be used, which will be specifically introduced in subsequent embodiments.
[0039] It can be understood that through the above processing of the target data segment, the sub-audio segment corresponding to the target audio track can be separated from the target data segment. Therefore, the above steps can be repeatedly executed to process each data segment in the audio to be processed, so as to complete the extraction of the audio data of the target audio track of the entire audio to be processed.
[0040] As an implementation manner, after S203, the method further includes: taking the next segment of the target data segment as a new target data segment, and returning to execute the step of inputting the plural spectral features corresponding to the target data segment into the multi-track music extraction model and subsequent steps until the sub-audio segments corresponding to the target audio tracks of each data segment of the audio to be processed are obtained. Based on the sub-audio segments corresponding to the target audio tracks of each data segment, the audio of the target audio track after separation of the audio to be processed is obtained.
[0041] After completing the operation of separating the audio tracks of the starting data segment, that is, obtaining the sub-audio segment corresponding to the target audio track in the starting data segment, the next consecutive first number of frames of audio data is obtained as a new target data segment, and the foregoing operation is returned to be executed until the sub-audio segments corresponding to the target audio tracks of each data segment are obtained.
[0042] It should be noted that the next segment of the target data segment refers to the next consecutive first number of frames of audio data of the target data segment. Here, the next refers to the next data segment of the target data segment based on the playback timestamp of the audio to be processed. Taking the first number as 8 as an example, based on the playback timestamp of the audio to be processed, the next consecutive 8 frames of audio data of the target data segment are used as a new target time period. Therefore, it can be seen that in the process of determining the sub-audio segments corresponding to the target audio tracks of each data segment, the sub-audio segments corresponding to the target audio tracks in multiple data segments (that is, consecutive first number of frames of audio data) are sequentially and continuously obtained according to the playback timestamp of the audio to be processed.
[0043] It can be understood that when performing audio track separation on the audio data of each consecutive first number of frames of the audio to be processed, each data segment processed is operated in sequence according to the playback timestamp of the audio to be processed. Therefore, each obtained sub-audio segment also corresponds to the playback timestamp of the audio to be processed. Thus, based on the timestamp range corresponding to each sub-audio segment, the sub-audio segments are integrated into a complete audio, thereby separating the audio corresponding to the target track from the audio to be processed.
[0044] As an implementation manner, there are multiple target tracks. For example, the target tracks include vocals, drums, bass, and guitar. Through the above S201 to S204, multiple sub-audio segments corresponding to vocals, drums, bass, and guitar respectively are obtained. That is to say, for each data segment of the audio to be processed, sub-audio segments corresponding to vocals, drums, bass, and guitar respectively are extracted. If there is no certain track in a data segment, the signal amplitude of that track in the sub-audio segment is zero, or a stable small value. That is to say, sub-audio segments of each target track can be obtained in each data segment, and then the audio of each target track is integrated. The audio of each target track has the same duration as the audio to be processed.
[0045] Therefore, in the embodiment of the present application, for the audio to be processed, an audio track separation operation is performed every other consecutive first number of frames of audio data, reducing the complexity of processing data each time and reducing the latency. Moreover, taking the complex spectral features of the audio as the input and predicting both the amplitude and the phase simultaneously avoids unnecessary latency and computing power consumption.
[0046] Please refer to Figure 3 , Figure 3 FIG. shows an audio processing method provided by an embodiment of the present application. This method is applied to the above-mentioned sound source separation system 100. It should be noted that since the sound source separation system is a streaming architecture, part of the model state cached in the previous inference also needs to be used as the input, and part of the model state of each inference is cached. Specifically, this method includes: S301 to S309.
[0047] S301: Obtain a target data segment of the audio to be processed, where the target data segment includes consecutive first number of frames of audio data.
[0048] S302: Determine whether the target data segment is a starting data segment.
[0049] As described above, the sound source separation system in the embodiments of the present application is a streaming architecture. Therefore, when processing the previous data segment, the network state of the multi-track music extraction model will continue to the processing of this data segment, that is, the model state of the multi-track music extraction model cached in the previous inference is also used as the input of the multi-track music extraction model in this processing. Therefore, it is necessary to determine whether the target data segment being processed is the starting data segment, that is, whether there are other data segments processed before the target data segment being processed currently, that is, whether there is a model state cached in the previous inference.
[0050] It should be noted that each frame of audio data of the audio to be processed corresponds to a timestamp. Therefore, the continuous first number of frames of audio data corresponding to the first timestamp is used as the starting data segment, and when processing this audio to be processed, it will start from this starting data segment.
[0051] If the target data segment is the starting data segment, then execute S305; if the target data segment is not the starting data segment, then execute S303.
[0052] S303: Obtain the model state of the multi-track music extraction model when processing the previous data segment of the target data segment as the historical model state.
[0053] As an implementation manner, the model state of the multi-track music extraction model can also be understood as the convolutional state (conv state) of this multi-track music extraction model. Exemplarily, this multi-track music extraction model is a deep neural network. Generally, in a neural network, the convolutional state usually refers to the internal state of the convolutional layer (Convolutional Layer). The convolutional layer is a commonly used neural network layer for processing input data with local correlation, such as images or audio.
[0054] The internal state of the convolutional layer can include convolutional kernel weights, feature maps, and bias terms. Among them, the convolutional kernel weights are a set of learnable convolutional kernels (also called filters) used by the convolutional layer, and each convolutional kernel has a set of weights. These weights are used to perform local convolutional operations on the input data and generate output feature maps. The feature maps can be the output of the convolutional layer, and each feature map represents the result of the convolutional operation of the corresponding convolutional kernel on the input. These feature maps can capture different features of the input data. The role of the bias term is to add a constant offset after the convolutional operation.
[0055] As an implementation, when the multi-track music extraction model processes the previous data segment, the model state of the multi-track music extraction model can be saved, for example, as a model state output buffer (Conv State OutputBuffer). Then, when the next data segment is processed, the model state is input into the model as a model state input buffer (ConvState Input Buffer).
[0056] S304: Input the historical model state and the complex spectral features corresponding to the target data segment into the multi-track music extraction model to obtain the complex spectral features of the target audio track as the first spectral features.
[0057] It should be noted that the input method of inputting the model state and the complex spectral features into the multi-track music extraction model can be that the complex spectral features are a feature matrix. By expanding in a certain row or column direction of the feature matrix, the model state and the complex spectral features are fused to obtain a new input feature matrix, and the new input feature matrix is input into the multi-track music extraction model.
[0058] S305: Obtain the initial model state.
[0059] As an implementation, for the starting data segment, since there is no corresponding historical model state, that is, there is no data segment processed last time before the starting data segment, an initial model state can be set for use when processing the starting data segment. For example, the initial model state can be an initial value, and the initial value can be 0.
[0060] S306: Input the initial model state and the complex spectral features corresponding to the target data segment into the multi-track music extraction model to obtain the complex spectral features of the target audio track as the first spectral features.
[0061] That is to say, when processing the starting data segment, the historical model state used is the initial model state.
[0062] Exemplarily, please refer to Figure 4 , the sound source separation system with a streaming architecture in the embodiments of the present application is as Figure 4 shown.
[0063] As an implementation manner, before executing S306 and S304, the complex spectral features corresponding to the target data segment need to be obtained. As described above, the preprocessing model 101 converts the time-domain signal of the input stereo music into a frequency-domain signal through preprocessing operations. Exemplarily, the target data segment can be converted into complex spectral features. Exemplarily, the preprocessing operations may include Framing, Analysis Window, and Fast Fourier Transform (FFT). Specifically, the Framing operation refers to framing the target data segment based on a window length of 2048 points and a hop length of 512 points to obtain audio frames. In the Analysis Window operation, each frame of signal is windowed with a 2048-point Hamming window. Then, in the Fast Fourier Transform, a 2048-point FFT complex spectrum is applied to the windowed signal to obtain the complex spectral features corresponding to the target data segment.
[0064] Correspondingly, the postprocessing model 103 reconstructs the time-domain signal of each track, and its processing operations correspond to the preprocessing operations of the preprocessing model 101. Specifically, the postprocessing model 103 reconstructs the complex spectral features of each target track output by the track separation music extraction model 102 through postprocessing operations to obtain the time-domain signal of each track. The postprocessing operations include Inverse Fast Fourier Transform (IFFT), Synthesis Window, and Overlap-Add. Specifically, the Inverse Fast Fourier Transform performs a 2048-point IFFT operation on the complex spectrum signal of each track. In the Analysis Window operation, the signal after the Inverse Fast Fourier Transform is windowed with a 2048-point Hamming window. Then, the signal is subjected to Overlap-Add, and finally the track signal is output.
[0065] As an implementation manner, the track separation music extraction model is a neural network model based on the U-NET structure. The track separation music extraction model includes an encoder, an intermediate layer, and a decoder. It can be understood that the U-Net architecture consists of an encoder layer, an intermediate layer, and a decoder layer. Among them, the encoder consists of multiple convolutional layers and pooling layers. It is responsible for gradually reducing the spatial dimension of the input data and extracting high-level feature representations. Each convolutional layer usually includes convolutional operations, activation functions (such as ReLU), and batch normalization, etc. The intermediate layer is usually a convolutional layer with a large receptive field, which converts the input data into a higher-level feature representation. The intermediate layer plays a role in information transmission in the entire network. The decoder is opposite to the encoder. It gradually restores the spatial dimension of the input data through upsampling and convolutional operations and generates the final output. In addition, through the connection between the encoder and the decoder in the U-Net architecture, the network can capture features at different levels and perform gradual restoration and reconstruction in the decoding stage. This structure makes U-Net perform well in many segmentation tasks, especially suitable for fields such as medical image segmentation.
[0066] It can be seen that since the split-track music extraction model includes an encoder, an intermediate layer, and a decoder, its model state also needs to correspond to the model states of the encoder, the intermediate layer, and the decoder. Therefore, the implementation of S304 may include: obtaining a first model state, a second model state, and a third model state based on the historical model state; obtaining a first eigenvalue through a first operation of the encoder on the target data segment based on the first model state and the complex spectral features corresponding to the target data segment; obtaining a second eigenvalue through a second operation of the intermediate layer on the first eigenvalue based on the second model state and the first eigenvalue; and obtaining a first spectral feature through a third operation of the decoder based on the third model state and the second eigenvalue.
[0067] That is to say, the historical model state may include a first model state, a second model state, and a third model state. Among them, the first model state corresponds to the encoder, and the encoder also corresponds to a convolutional layer. Therefore, the first model state may be the convolutional state of the encoder. Similarly, both the intermediate layer and the decoder correspond to convolutional layers, so they also correspond to convolutional states. Therefore, the first model state is the model state when the encoder processes the previous data segment of the target data segment, the second model state is the model state when the intermediate layer processes the previous data segment of the target data segment, and the third model state is the model state when the decoder processes the previous data segment of the target data segment.
[0068] It should be noted that taking the first model state and the complex spectral features corresponding to the target data segment as the first input feature and inputting them into the encoder, the first operation of the encoder on the target data segment may include an extraction operation of useful features, that is, an extraction operation of target track features, and a dimensionality reduction operation. The second operation of the intermediate layer may include an information fusion and transfer operation. Among them, the information fusion operation may include capturing features in different frequency ranges, time domains, and frequency domains of the mixed audio signal, such as information on the fundamental frequency, energy, harmonics, etc. of the target track, and then transferring them to the decoder. The third operation of the decoder includes gradually restoring the spatial dimension of the input data through upsampling and convolutional operations and generating the final output.
[0069] S307: Obtain the corresponding sub-audio segment of the target track based on the first spectral feature.
[0070] It should be noted that in the streaming architecture, the prediction output of the previous inference (frequency spectrum input buffer) and the prediction output of this inference (frequency spectrum output buffer) need to be concatenated to meet the overlapping phase calculation in ISTFT. Specifically, the implementation manner of obtaining the corresponding sub-audio segment of the target audio track based on the first spectral feature may include: concatenating the first spectral feature and the specified spectral feature to obtain the concatenated spectral feature, performing a reconstruction operation on the concatenated spectral feature to obtain the corresponding sub-audio segment of the target audio track, and updating the specified spectral feature to the first spectral feature. Among them, when the target data segment is the starting data segment, the specified spectral feature is the initial feature, and when the target data segment is not the starting data segment, the specified spectral feature is the spectral feature output by the multi-track music extraction model for processing the previous data segment of the target data segment.
[0071] As mentioned above, the overlap-and-add operation is included in the post-processing operation. In ISTFT, this overlap-and-add operation is also called the overlapping phase calculation, which refers to processing the phase information of the spectrum between overlapping window functions to achieve signal synthesis. Since STFT usually uses overlapping windows to improve the time-frequency resolution, when performing ISTFT, it is necessary to consider how to reasonably process the phase information of the overlapping part to avoid introducing discontinuous or non-smooth jumps during signal synthesis. The main goal of the overlapping phase calculation is to ensure that the impact of the phase jumps caused by the overlapping windows on signal recovery is minimized during ISTFT. Generally speaking, common overlapping phase calculation methods include phase averaging, weighted averaging, and phase interpolation, etc. These overlapping phase calculation methods are all aimed at dealing with the impact of the overlapping windows on the spectrum phase to ensure that the time-domain characteristics of the original signal can be restored as accurately as possible during the inverse STFT operation.
[0072] Therefore, in the overlapping phase calculation, concatenating the predicted output of the previous inference and the predicted output of the current inference is to correctly handle the phase information brought by the overlapping window. When performing the inverse STFT operation, an overlapping window function is usually used to synthesize the spectrum to obtain the reconstruction of the time-domain signal. Due to the overlap of the windows, some data in the spectrum output buffer will overlap with some data in the spectrum input buffer. In this case, correctly handling the phase information in the overlapping region is very important for the accurate recovery of the signal. Concatenating the predicted output of the previous inference and the predicted output of the current inference can ensure that the phase information in the overlapping region is correctly processed. By concatenating these two buffers, the spectrum information of adjacent frames can be merged in the overlapping region, and the spectrum phase can be correctly processed by combining the method of overlapping phase calculation (such as phase averaging, weighted averaging or phase interpolation), thus avoiding introducing discontinuous or non-smooth phase jumps in the overlapping part.
[0073] Therefore, concatenating the predicted outputs of the previous and current inferences is to utilize the spectrum information of the previous and current inferences during the overlapping phase calculation to ensure that the time-domain characteristics of the original signal can be restored as accurately as possible when performing the inverse STFT operation.
[0074] It should be noted that for the first frame of input, that is, when the current target data segment being processed is the starting data segment, there are no specified spectrum characteristics before this starting data segment, that is, there are no spectrum characteristics output from the previous data segment processed by the multi-track music extraction model for the target data segment, that is, there is no previous output to concatenate. Therefore, some initialization data, such as a vector of all zeros, needs to be stored in the buffer. At the same time, for the last frame of output, due to the possible existence of an incomplete overlapping region, special processing is required, such as using bilinear interpolation or other techniques to repair the boundary effect.
[0075] In addition, the way of concatenating the first spectrum feature and the specified spectrum feature can be to obtain the average value, weighted average or perform interpolation processing on the spectrum data in the overlapping region, which is not limited here.
[0076] S308: Use the next segment of the target data segment as the new target data segment, and return to execute the step of inputting the complex spectrum feature corresponding to the target data segment into the multi-track music extraction model and subsequent steps until the sub-audio segments corresponding to the target tracks of each data segment of the audio to be processed are obtained.
[0077] S309: Based on the sub-audio segments corresponding to the target tracks of each data segment, obtain the audio of the target tracks after separating the audio to be processed.
[0078] The architecture of the multi-track music extraction model is described in detail below. As Figure 5As shown in the figure, in the embodiment of the present application, the multi-track music extraction model is designed based on the U-NET structure, which mainly consists of an encoder, an intermediate layer, and a decoder. There are 3 TFC-TDFs and 3 downsampling layers (DownSample) at the encoder end, and 3 TFC-TDFs and 3 upsampling layers (UpSample) at the decoder end, and there is also an intermediate layer between them. Among them, the encoder of U-NET has 3 times of downsampling to obtain high-level semantic information, and the decoder naturally corresponds to 3 times of upsampling to restore the resolution. In order to reduce the loss of spatial information caused by the downsampling process, skip connections are introduced, so that the feature maps restored by upsampling contain more low-level semantic information, making the result finer. After the streaming transformation, the model is input with the complex spectrum features of the music signal and the previous conv state, that is, the historical model state, and outputs the complex spectrum features of the multi-track music signal and the conv state for the next inference.
[0079] It can be understood that in the semantic segmentation model of the U-Net architecture, skip connections, also known as jump connections, refer to the mechanism of connecting the feature maps of corresponding levels between the encoder (downsampling path) and the decoder (upsampling path). The U-Net model usually consists of two parts: an encoder and a decoder. The encoder is used to gradually extract the high-level semantic features of the input image, while the decoder is used to map these features back to the size of the original input image and generate pixel-level semantic segmentation results. The introduction of skip connections enables U-Net to better capture features of different scales, thereby improving the accuracy of semantic segmentation.
[0080] Specifically, skip connections connect the feature map of a certain layer in the encoder with the feature map of the corresponding level in the decoder, so that the decoder can directly access the earlier feature representations in the encoder. This can retain detailed information. Since there is a direct connection between the encoder and the decoder, the decoder can obtain the feature representations in the encoder that contain more detailed information, which helps to improve the detail accuracy of the segmentation result. In addition, it can also solve the problem of information loss. Usually in the encoding-decoding structure, as the size of the feature map shrinks, some information is likely to be lost during the encoding process. Skip connections can solve this problem to a certain extent because the decoder can obtain larger-sized feature maps through skip connections, thereby retaining more information.
[0081] Exemplarily, the downsampling is implemented using Conv2d with a kernel size of (2, 2) and a stride of (2, 2), which ensures that the downsampling will reduce both the time domain and frequency domain dimensions of the input features to half of the original, and no padding processing is required.
[0082] That is to say, when downsampling is performed using a 2D convolution (Conv2d) with a kernel size of (2, 2) and a stride of (2, 2), both the time domain and frequency domain dimensions of the input features are reduced by half, and no padding process is required. This is because a kernel size of (2, 2) means the size of the convolution kernel is 2x2, and a stride of (2, 2) means the convolution operation moves 2 pixels each time for calculation. Since the stride is (2, 2), the convolution operation skips 2 pixels each time for calculation, which reduces the size of the output feature map by half. In the time domain, after the input features pass through the convolution operation, one is taken every two time steps, so the time domain dimension is reduced by half. In the frequency domain, the convolution operation calculates adjacent pixels in space, which is equivalent to downsampling the frequency domain and also reduces the frequency domain dimension by half. Therefore, the downsampling operation implemented using a 2D convolution with a kernel size of (2, 2) and a stride of (2, 2) can reduce both the time domain and frequency domain dimensions of the input features to half of the original, and no padding process is required.
[0083] It can be seen that the encoder includes N feature extraction network layers and N downsampling layers. The feature extraction network layer is the aforementioned TFC-TDF. TFC-TDF is a network model composed of time-frequency convolution and time-domain fully-connected network. Among them, time-frequency convolution (Time-Frequency Convolutions, TFC) considers both the time domain and frequency domain dimensions, and the time-domain fully-connected network (Time-Distributed Fully-connected network, TDF) is a time-domain distributed fully-connected network, and the goal is to extract useful features to make the audio separation effect of the target audio track better.
[0084] In the embodiment of the present application, for the encoder, in the way of sequentially connecting one feature extraction network layer and one downsampling layer, the N feature extraction network layers and N downsampling layers are cascaded into a first network string, and the first model state includes the model state corresponding to each feature extraction network layer. As Figure 5It can be seen that each TFC-TDF receives a model state. In the embodiments of the present application, the model states input by each TFC-TDF can be different. Combining the first model state for the encoder mentioned above, the first model state includes the model states corresponding to each feature extraction network layer. The model state corresponding to each feature extraction network layer can be understood as follows: when the multi-track music extraction model processes the previous data segment, each feature extraction network layer of the encoder will perform corresponding data processing operations. Correspondingly, convolution states will be generated. When the multi-track music extraction model processes the next data segment, the previous convolution states of each feature extraction network layer are applied to the next processing of that feature extraction network layer.
[0085] As Figure 5 shown, in the first network string of the encoder, along the input direction of the target data, a TFC-TDF is connected to a downsampling in sequence, then the downsampling is connected to the next TFC-TDF, and the next TFC-TDF is connected to the next downsampling, and so on until the last downsampling is connected.
[0086] Therefore, in the first network string, the model state corresponding to each feature extraction network layer and the audio data features input corresponding to that feature extraction network layer are combined into the input feature value corresponding to that feature extraction network layer. Among them, the audio data features input corresponding to the first feature extraction network layer in the first network string are the complex spectrum features corresponding to the target data segment, and the audio data features input corresponding to other feature extraction network layers are the feature value sequences output by the previous downsampling layer; the input feature values of each feature extraction network layer are input into that feature extraction network layer to obtain a first intermediate value and are input into the downsampling layer connected to that feature extraction network layer; the feature value sequence output by the last downsampling layer of the first network string is obtained as the first feature value. Among them, the last specified number of feature values of the feature value sequence output by each downsampling layer are used as the model state corresponding to the previous feature extraction network layer connected to that downsampling layer and are saved.
[0087] Similarly, the decoder includes N feature extraction network layers and N upsampling layers. By connecting a feature extraction network layer and an upsampling layer in sequence, the N feature extraction network layers and the N upsampling layers are concatenated into a second network string. The third model state includes the model state corresponding to each feature extraction network layer. Based on the third model state and the second eigenvalue, through the third operation of the decoder, a first spectral feature is obtained, including: in the second network string, the model state corresponding to each feature extraction network layer and the audio data feature corresponding to the input of this feature extraction network layer are merged into the input eigenvalue corresponding to this feature extraction network layer. Among them, the audio data feature corresponding to the input of the first feature extraction network layer in the second network string is the second eigenvalue, and the audio data features corresponding to the inputs of other feature extraction network layers are the eigenvalue sequences output by the previous upsampling layer; the input eigenvalue of each feature extraction network layer is input into this feature extraction network layer to obtain a second intermediate value, and the second intermediate value is input into the upsampling layer connected to this feature extraction network layer; the eigenvalue sequence output by the last upsampling layer of the second network string is obtained as the second eigenvalue. In addition, the last specified number of eigenvalues of the eigenvalue sequence output by each upsampling layer are used as the model state corresponding to the previous feature extraction network layer connected to this upsampling layer and saved. Exemplarily, N is 3.
[0088] The structure of this TFC-TDF can be referred to Figure 6 , such as Figure 6 shown. It can be seen that the TFC-TDF mainly includes several time-frequency convolutions (TFC) and time-distributed fully-connected networks (TDF). Among them, TFC is mainly composed of batch normalization (BatchNorm), Gaussian error linear unit (GeLU), and causal 2D convolution (Casual 2d Conv), and TDF is mainly composed of batch normalization, Gaussian error linear unit, and linear neural networks (LNN). The role of batch normalization is to normalize the input data, making the mean of each feature close to 0 and the variance close to 1, thereby accelerating the training process of the network and improving the robustness and generalization ability of the model.
[0089] To achieve the effect of streaming inference and reduce the algorithm latency, each TFC block uses causal convolution (Casual 2dconv). Specifically, for Casual 2d conv, in the frequency domain dimension, one dimension is padded at both the top and the bottom; in the time domain dimension, to ensure that future information is not used, two dimensions are padded on the left side and no padding is performed on the right side. This ensures that the network structure of this solution only uses past information and does not use future information, thus achieving the effect of streaming inference. Since three downsampling layers (Downsample) are used in the embodiments of this application, and each downsampling layer reduces the time domain dimension to half of the original, 8 frames of music signals are required for each inference.
[0090] It can be understood that in the frequency domain dimension, one dimension is padded at the top and one dimension is padded at the bottom respectively. This is done to keep the size of the feature map unchanged in the frequency domain so that the input and output sizes of each TFC block are consistent. In the time domain dimension, to ensure that the model does not use future information, two dimensions are padded on the left side and no padding is performed on the right side. This is the key point of causal convolution, which ensures that the output of the current time step only depends on the input of the current time step and the previous ones, and is not affected by future information.
[0091] For example, assume the input feature is
[0092] Then, after padding the dimensions in the frequency domain and time domain, the input becomes
[0093] where s is the aforementioned historical model state.
[0094] As Figure 7 shown, it is the state change diagram between the encoder and the decoder. This state change diagram can characterize the changes in the input and output between the encoder and the decoder. It can be seen that each TFC-TDF corresponds to a historical model state. Figure 7 Shows the relationship between time ranges when using causal 2d conv, taking 3x3 conv as an example. The conv state containing the last 2 frames of the previous inference is split to fit each layer and padded to the left side of the current feature of each layer. After the inference is completed, the last two frames of the output feature of each layer are merged together as the input of the model state for the next inference. As Figure 7 shown, in the features corresponding to the output of each TFC-TDF, the last two frames enclosed by the dashed box are used as the model state corresponding to this TFC-TDF.
[0095] Exemplarily, assume that the music to be separated has a sampling rate of 44100HZ. The multi-track music extraction model accepts the complex spectrum of a 2-channel 8-frame signal as input (the frame shift is 512 / 44100 = 11.6ms, and the signal length is 8 * 11.6 = 92.8ms), and outputs the complex spectrum of an 8-channel 8-frame signal (each audio track is 2 channels and 8 frames). As Figure 8 shown, after inference, the ISTFT reconstruction has an additional 3-frame delay (3 * 11.6 = 34.8ms). Assume that the model inference time is 10 milliseconds, and the expected delay between input and output is 137.6 milliseconds (8 * 11.6 + 3 * 11.6 + 10 = 137.6ms).
[0096] Therefore, the streaming sound source separation system (MSS) proposed in this application can expand the multi-track mixing framework of holographic audio, and output high-quality audio multi-track data of different instruments to the downstream spatial rendering module in real time; enable holographic audio / spatial audio to support richer and more flexible audio spatialization effects, and enhance the user's sense of participation and playability in customizing sound effects.
[0097] Please refer to Figure 9 , which shows a structural block diagram of an audio processing device 900 provided by an embodiment of this application. The device may include: an acquisition unit 901, an extraction unit 902, and a determination unit 903.
[0098] The acquisition unit 901 is configured to acquire a target data segment of the audio to be processed, where the target data segment includes audio data of a continuous first number of frames.
[0099] The extraction unit 902 is configured to input the complex spectral features corresponding to the target data segment into a multi-track music extraction model to obtain the complex spectral features of the target audio track as the first spectral features.
[0100] Further, the extraction unit 902 is further configured to, if the target data segment is not the starting data segment, acquire the model state of the multi-track music extraction model when processing the previous data segment of the target data segment as the historical model state; input the historical model state and the complex spectral features corresponding to the target data segment into the multi-track music extraction model to obtain the complex spectral features of the target audio track as the first spectral features.
[0101] Further, the multi-track music extraction model includes an encoder, an intermediate layer, and a decoder. The extraction unit 902 is further configured to obtain a first model state, a second model state, and a third model state based on the historical model state, where the first model state is the model state when the encoder processes the previous data segment of the target data segment, the second model state is the model state when the intermediate layer processes the previous data segment of the target data segment, and the third model state is the model state when the decoder processes the previous data segment of the target data segment; based on the first model state and the complex spectrum features corresponding to the target data segment, through a first operation of the encoder on the target data segment, a first eigenvalue is obtained; based on the second model state and the first eigenvalue, through a second operation of the intermediate layer on the first eigenvalue, a second eigenvalue is obtained; based on the third model state and the second eigenvalue, through a third operation of the decoder, a first spectrum feature is obtained.
[0102] Further, the encoder includes N feature extraction network layers and N downsampling layers. By connecting one feature extraction network layer and one downsampling layer in sequence, the N feature extraction network layers and the N downsampling layers are cascaded into a first network cascade. The first model state includes the model states corresponding to each feature extraction network layer. The extraction unit 902 is further configured to, in the first network cascade, combine the model state corresponding to each feature extraction network layer and the audio data features corresponding to the input of this feature extraction network layer into the input eigenvalue corresponding to this feature extraction network layer, where the audio data features corresponding to the input of the first feature extraction network layer in the first network cascade are the complex spectrum features corresponding to the target data segment, and the audio data features corresponding to the input of other feature extraction network layers are the eigenvalue sequences output by the previous downsampling layer; input the input eigenvalue of each feature extraction network layer into this feature extraction network layer to obtain a first intermediate value, and input it into the downsampling layer connected to this feature extraction network layer; obtain the eigenvalue sequence output by the last downsampling layer of the first network cascade as the first eigenvalue.
[0103] Further, the extraction unit 902 is further configured to use the last specified number of eigenvalues of the eigenvalue sequence output by each downsampling layer as the model state corresponding to the previous feature extraction network layer connected to this downsampling layer and save it.
[0104] Further, the decoder includes N feature extraction network layers and N upsampling layers. By connecting a feature extraction network layer and an upsampling layer in sequence, the N feature extraction network layers and the N upsampling layers are concatenated into a second network string. The third model state includes the model state corresponding to each feature extraction network layer. The extraction unit 902 is further configured to, in the second network string, merge the model state corresponding to each feature extraction network layer and the audio data feature corresponding to the input of this feature extraction network layer into the input feature value corresponding to this feature extraction network layer. Among them, the audio data feature corresponding to the input of the first feature extraction network layer in the second network string is the second feature value, and the audio data features corresponding to the inputs of other feature extraction network layers are the feature value sequences output by the previous upsampling layer; input the input feature value of each feature extraction network layer into this feature extraction network layer to obtain a second intermediate value, and input the second intermediate value into the upsampling layer connected to this feature extraction network layer; obtain the feature value sequence output by the last upsampling layer of the second network string as the second feature value.
[0105] Further, the extraction unit 902 is further configured to use the feature values of the last specified number of digits of the feature value sequence output by each upsampling layer as the model state corresponding to the previous feature extraction network layer connected to this upsampling layer and save it.
[0106] Further, both the encoder and the decoder include feature extraction network layers, and the feature extraction network layer is a TFC-TDF model.
[0107] The determination unit 903 is configured to obtain the corresponding sub-audio segment of the target audio track based on the first spectral feature.
[0108] Further, the determination unit 903 is further configured to splice the first spectral feature and the specified spectral feature to obtain the spliced spectral feature. Among them, when the target data segment is the starting data segment, the specified spectral feature is the initial feature; when the target data segment is not the starting data segment, the specified spectral feature is the spectral feature output by the previous data segment processed by the multi-track music extraction model for the target data segment; perform a reconstruction operation on the spliced spectral feature to obtain the corresponding sub-audio segment of the target audio track, and update the specified spectral feature to the first spectral feature.
[0109] In addition, the audio processing device further includes a loop unit and a synthesis unit.
[0110] The loop unit is configured to use the next segment of the target data segment as the new target data segment, and return to execute the step of inputting the complex spectral feature corresponding to the target data segment into the multi-track music extraction model and subsequent steps until the corresponding sub-audio segments of the target audio tracks of each data segment of the to-be-processed audio are obtained.
[0111] The synthesizing unit is configured to obtain the audio of the target track after separating the to-be-processed audio based on the sub-audio segments corresponding to the target track of each data segment.
[0112] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0113] In several embodiments provided in the present application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0114] In addition, in each embodiment of the present application, the various functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0115] Please refer to Figure 10 , which shows a structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 can be an electronic device such as a smart phone, a tablet computer, an e-book, etc. that can run application programs. The electronic device 100 in the present application can include one or more of the following components: a processor 110, a memory 120, and one or more application programs, where one or more application programs can be stored in the memory 120 and configured to be executed by one or more processors 110, and one or more programs are configured to execute the methods described in the foregoing method embodiments.
[0116] The processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire electronic device 100 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by invoking the data stored in the memory 120, it performs various functions of the electronic device 100 and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.
[0117] The memory 120 may include random access memory (RAM) and may also include read-only memory. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device 100 (such as phone book, audio-video data, chat record data, etc.).
[0118] Please refer to Figure 11 , which shows a structural block diagram of a computer-readable medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 1100, and the program code can be called by the processor to execute the methods described in the above method embodiments.
[0119] The computer-readable medium 1100 can be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable medium 1100 includes a non-transitory computer-readable storage medium. The computer-readable medium 1100 has a storage space for program code 1110 that executes any of the method steps in the above-described methods. These program codes can be read from or written to one or more computer program products. The program code 1110 can be compressed in an appropriate form, for example.
[0120] In summary, the embodiments of the present application have the following effects:
[0121] a. Taking the plural spectral features of music as input and predicting both the amplitude and phase simultaneously simplifies the model architecture and avoids unnecessary time delay and computing power consumption.
[0122] b. Adopting the TFC network structure and performing convolutions in both the time dimension and the frequency domain to extract effective high-dimensional features
[0123] c. All convolutional networks in the TFC network adopt Casual 2d Conv to ensure that future time information is not used, realizing true streaming computing
[0124] d. By reducing the number of DownSample layers in the encoder, while reducing the time delay, the number of model parameters is also reduced, thus ensuring that the model can run on the edge side
[0125] e. In the data preprocessing stage, pitch shift and time stretch are used for data augmentation to improve the robustness of the model.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio processing method, It is characterized in that include: Acquire a target data segment of audio to be processed, wherein the target data segment includes a first number of consecutive frames of audio data; Inputting the complex spectral features corresponding to the target data segment into the track-by-track music extraction model to obtain the complex spectral features of the target track as the first spectral features; Based on the first frequency spectrum feature, a corresponding sub-audio segment of the target audio track is obtained.
2. The method according to claim 1, It is characterized in that The step of inputting the complex spectral features corresponding to the target data segment into the track-by-track music extraction model to obtain the complex spectral features of the target track as the first spectral features includes: If the target data segment is not the starting data segment, obtaining the model state of the sub-track music extraction model when the sub-track music extraction model processes the previous data segment of the target data segment as the historical model state; The complex spectral features corresponding to the historical model state and the target data segment are input into the track-by-track music extraction model to obtain the complex spectral features of the target track as the first spectral features.
3. The method according to claim 2, It is characterized in that The track-by-track music extraction model includes an encoder, an intermediate layer, and a decoder. The complex spectral features corresponding to the historical model state and the target data segment are input into the track-by-track music extraction model to obtain the complex spectral features of the target track as the first spectral features, including: Based on the historical model state, a first model state, a second model state and a third model state are obtained, wherein the first model state is a model state when the encoder processes a data segment before the target data segment, the second model state is a model state when the intermediate layer processes a data segment before the target data segment, and the third model state is a model state when the decoder processes a data segment before the target data segment; Based on the first model state and the complex spectral features corresponding to the target data segment, obtaining a first feature value by performing a first operation on the target data segment by the encoder; Based on the second model state and the first eigenvalue, obtaining a second eigenvalue by performing a second operation on the first eigenvalue by the intermediate layer; Based on the third model state and the second feature value, a first spectrum feature is obtained through a third operation of the decoder.
4. The method according to claim 3, It is characterized in that The encoder includes N feature extraction network layers and N downsampling layers, and the N feature extraction network layers and the N downsampling layers are serially connected into a first network string by sequentially connecting one feature extraction network layer and one downsampling layer, and the first model state includes a model state corresponding to each feature extraction network layer; the first feature value is obtained by performing a first operation on the target data segment by the encoder based on the first model state and the complex spectrum feature corresponding to the target data segment, including: In the first network string, the model state corresponding to each feature extraction network layer and the audio data features corresponding to the input of the feature extraction network layer are merged into the input feature value corresponding to the feature extraction network layer, wherein the audio data features corresponding to the input of the first feature extraction network layer in the first network string are the complex spectrum features corresponding to the target data segment, and the audio data features corresponding to the input of other feature extraction network layers are the feature value sequences output by the previous downsampling layer; Inputting the input feature value of each feature extraction network layer into the feature extraction network layer to obtain a first intermediate value, and inputting the first intermediate value into a downsampling layer connected to the feature extraction network layer; Obtain a sequence of eigenvalues output by the last downsampling layer of the first network string as the first eigenvalue.
5. The method according to claim 4, It is characterized in that Also includes: The eigenvalue of the last specified number of bits of the eigenvalue sequence output by each downsampling layer is used as the model state corresponding to the previous feature extraction network layer connected to the downsampling layer, and is saved.
6. The method according to claim 3, It is characterized in that The decoder includes N feature extraction network layers and N upsampling layers, and the N feature extraction network layers and the N upsampling layers are connected in series to form a second network string by connecting one feature extraction network layer and one upsampling layer in sequence, and the third model state includes a model state corresponding to each feature extraction network layer; Based on the third model state and the second characteristic value, obtaining a first spectrum characteristic through a third operation of the decoder includes: In the second network string, the model state corresponding to each feature extraction network layer and the audio data features corresponding to the input of the feature extraction network layer are merged into the input feature value corresponding to the feature extraction network layer, wherein the audio data features corresponding to the input of the first feature extraction network layer in the second network string are the second feature values, and the audio data features corresponding to the input of other feature extraction network layers are the feature value sequences output by the previous upsampling layer; Inputting the input feature value of each feature extraction network layer into the feature extraction network layer to obtain a second intermediate value, and inputting the second intermediate value into an upsampling layer connected to the feature extraction network layer; Obtain a sequence of eigenvalues output by the last upsampling layer of the second network string as the second eigenvalue.
7. The method according to claim 6, It is characterized in that Also includes: The eigenvalue of the last specified number of bits of the eigenvalue sequence output by each upsampling layer is used as the model state corresponding to the previous feature extraction network layer connected to the upsampling layer, and is saved.
8. The method according to claim 3, It is characterized in that The encoder and the decoder both include a feature extraction network layer, and the feature extraction network layer is a TFC-TDF model.
9. The method according to claim 4 or 6, It is characterized in that The N is 3.
10. The method according to claim 1, It is characterized in that The obtaining, based on the first spectrum feature, a corresponding sub-audio segment of the target audio track includes: The first spectrum feature and the designated spectrum feature are concatenated to obtain a concatenated spectrum feature, wherein, when the target data segment is a starting data segment, the designated spectrum feature is an initial feature, and when the target data segment is not a starting data segment, the designated spectrum feature is a spectrum feature output by the track-by-track music extraction model after processing a data segment before the target data segment; A reconstruction operation is performed on the concatenated frequency spectrum feature to obtain a corresponding sub-audio segment of the target audio track, and the designated frequency spectrum feature is updated to the first frequency spectrum feature.
11. The method according to claim 1, It is characterized in that After obtaining the corresponding sub-audio segment of the target audio track based on the first spectrum feature, the method further includes: The next segment of the target data segment is used as a new target data segment, and the step of inputting the complex spectral features corresponding to the target data segment into the track-by-track music extraction model and subsequent steps are returned until the sub-audio segment corresponding to the target audio track of each data segment of the audio to be processed is obtained; Based on the sub-audio segments corresponding to the target audio track of each data segment, the audio of the target audio track after the audio to be processed is separated is obtained.
12. An audio processing device, It is characterized in that include: An acquisition unit, configured to acquire a target data segment of the audio to be processed, wherein the target data segment includes a first number of consecutive frames of audio data; An extraction unit, configured to input the complex spectral features corresponding to the target data segment into a track-by-track music extraction model to obtain the complex spectral features of the target track as the first spectral features; The determining unit is configured to obtain a corresponding sub-audio segment of the target audio track based on the first frequency spectrum feature.
13. An electronic device, It is characterized in that include: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1-11.
14. A computer readable medium, It is characterized in that The computer-readable medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the method according to any one of claims 1 to 11.
Citation Information
Cited By
Music source separation method and wearable device
CN120913585A
Music source separation method and wearable device
CN120913585B