Audio processing method, training method of sound source separation model and electronic equipment
By adding the frequency band interaction layer and the frequency point interaction layer in the two-dimensional convolution feature extraction network, the problem of large amount of calculation and time delay of the sound source separation method is solved, and high-precision sound source separation is achieved, reducing spectrum holes and improving tone quality.
Patent Information
- Application Number
- CN202411240438.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The existing sound source separation method has too much calculation and long delay, which cannot meet the needs of online real-time processing, and the separation accuracy is insufficient, resulting in tone distortion.
On the basis of the two-dimensional convolution feature extraction network, the frequency band interaction layer and the frequency point interaction layer are added, and the frequency band interaction characteristics between multiple subbands and the global frequency point interaction characteristics within each subband are extracted respectively to make up for the insufficient global interactive feature extraction capability of the two-dimensional convolution network.
With the advantage of retaining the small amount of calculation of the two-dimensional convolution feature extraction network, the separation accuracy is improved, the processing delay is shortened, the spectrum hole is reduced, and the tone quality of sound source separation is ensured.
Smart Images

Figure CN120472922A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio processing method, a training method for a sound source separation model, and an electronic device. Background Art
[0002] Sound source separation involves isolating the audio signal of a specific source from a mixed audio stream. For example, it can be used to separate (or extract) the audio signals of one or more instruments from a multi-instrument ensemble, or to isolate the human voice from a song with accompaniment. Sound source separation is often used in online processing scenarios, which places high demands on latency. However, traditional sound source separation methods, due to their high computational complexity and long latency, cannot meet these requirements. Summary of the Invention
[0003] The present application provides an audio processing method, a training method for a sound source separation model, and an electronic device, which can shorten processing delay while ensuring separation accuracy.
[0004] In a first aspect, an audio processing method is provided, which includes: obtaining audio data to be processed, wherein the audio data to be processed includes audio signals of multiple sound sources; using a sound source separation model to process the audio data to be processed to obtain an audio signal of at least one target sound source; the at least one target sound source is at least one sound source among the multiple sound sources; the sound source separation model includes a feature extraction network, a frequency band interaction layer and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network, and is used to extract features of the audio data to be processed to obtain a first audio feature set; the frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be processed from the first audio feature set, thereby obtaining a second audio feature set; the frequency point interaction layer is used to extract global frequency point interaction features within each sub-band in the second audio feature set, thereby obtaining a third audio feature set; the sound source separation model is also used to separate the audio signal of at least one target sound source from the third audio feature set.
[0005] In the technical solution of the present application, a frequency band interaction layer and a frequency point interaction layer are added on the basis of the two-dimensional convolution feature extraction network, and the frequency band interaction features between multiple sub-bands and the global frequency point interaction features within each sub-band are extracted respectively. In this way, while retaining the advantage of the two-dimensional convolution feature extraction network in terms of small computational complexity, the defect of the two-dimensional convolution feature extraction network in terms of insufficient extraction ability of global interaction features is compensated, so that the entire sound source separation process can achieve relatively high separation accuracy and relatively short processing delay.
[0006] In conjunction with the first aspect, in certain implementations of the first aspect, the multiple sub-bands are obtained by dividing the audio data to be processed along the frequency dimension, or the multiple sub-bands are obtained by dividing the first extraction result output by the first network layer along the frequency dimension, where the first network layer is any layer of the feature extraction network. In this implementation, the multiple sub-bands corresponding to the audio data to be processed can be divided before entering the feature extraction network, during feature extraction by the feature extraction network, or before inputting the extracted features into the frequency band interaction layer after the feature extraction network completes feature extraction, thereby enabling flexible adjustment as needed based on actual application scenarios.
[0007] In combination with the first aspect, in certain implementations of the first aspect, when the feature extraction network has the ability to extract frequency band interaction features between multiple sub-bands so that the first audio feature set includes frequency band interaction features between multiple sub-bands, the sound source separation model does not include a frequency band interaction layer; and / or, when the feature extraction network has the ability to extract global frequency point interaction features for each frequency point corresponding to the audio data to be processed so that the first audio feature set includes global frequency point interaction features for each frequency point, the sound source separation model does not include a frequency point interaction layer.
[0008] For some special feature extraction networks, combined with segmentation operations (operations that divide the frequency bands into multiple sub-bands), some traditional feature extraction networks can also have the ability to extract global interaction features. Therefore, for such feature extraction networks, the network structure can be appropriately streamlined, thereby further reducing the amount of calculation and shortening the delay without affecting the feature extraction effect.
[0009] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: performing a reshaping operation to match the target dimension with the active dimension of the underlying network used to construct the frequency band interaction layer; and / or performing a reshaping operation to match the target dimension of the frequency point interaction layer with the active dimension of the underlying network used to construct the frequency point interaction layer. In this implementation, the reshaping operation enables different types of underlying networks to act on corresponding target dimensions.
[0010] In combination with the first aspect, in certain implementations of the first aspect, when using a sound source separation model to process the audio data to be processed to obtain an audio signal of at least one target sound source, it can include: when the first target sound source belongs to a preset sound source category, using the frequency band interaction layer to extract frequency band interaction features between multiple sub-bands from the first audio feature set, thereby obtaining a second audio feature set, using the frequency point interaction layer to extract frequency point interaction features within each sub-band in the second audio feature set, thereby obtaining a third audio feature set, and using the sound source separation model to separate the audio signal of the first target sound source from the third audio feature set; the first target sound source is any one of the at least one target sound source; or, when the first target sound source does not belong to the preset sound source category, using the sound source separation model to separate the audio signal of the first target sound source from the first audio feature set.
[0011] In this implementation method, only the target sound sources belonging to the preset sound source category will have frequency band interaction features and frequency point interaction features extracted, while the target sound sources that are not in the preset sound source category will not have frequency band interaction features and frequency point interaction features extracted. The extraction of frequency band interaction features and frequency point interaction features for target sound sources that are not in the preset sound source category can be adaptively omitted, thereby further shortening the processing delay.
[0012] In one example, the preset sound source category is used to represent the full-band sound source category; or, the output value of the discriminant network corresponding to the sound source in the preset sound source category is less than or equal to a preset threshold. That is to say, the possible sound source can be confirmed as a full-band sound source category or a non-full-band sound source category through prior knowledge, thereby determining the preset sound source category in an enumerated manner; the feature vector of the sound source can also be input into the discriminant network so that the discriminant network outputs the corresponding output value, and a preset threshold is set. Only the sound source category with an output value less than or equal to the preset threshold is the preset sound source category. This example gives an example of how to set the preset sound source category. The sound source categories can be enumerated or determined by discriminant scoring.
[0013] In combination with the first aspect, in some implementations of the first aspect, the frequency band interaction layer and / or the frequency point interaction layer are constructed using a two-dimensional convolutional basic network, an attention mechanism (attention) network or a fully connected (FC) layer.
[0014] It should be understood that although the attention network in this implementation is more computationally intensive than a two-dimensional convolutional network, since it is only a network layer added to the feature extraction network, the overall computational complexity of the entire sound source separation model is still far less than the feature extraction network constructed entirely from a one-dimensional convolutional network in traditional solutions. This implementation provides examples of frequency band interaction layers and frequency point interaction layers. By adding a frequency band interaction layer, frequency band interaction features between multiple sub-bands can be extracted, while by adding a frequency point interaction layer, frequency point interaction features within each sub-band can be extracted. It should also be understood that since feature extraction is first performed through the frequency band interaction layer and then through the frequency point interaction layer, although the frequency point interaction layer extracts frequency point interaction features within each sub-band, these frequency point features have been influenced by the previous feature extraction network and frequency point interaction layer, generating new global features. Therefore, the frequency point interaction features extracted here are also global features.
[0015] In one example, the frequency band interaction layer is constructed using a 1x1 convolutional layer; and the frequency point interaction layer is constructed using an FC layer.
[0016] In this example, a 1x1 convolution is used to construct a frequency band interaction layer for channel fusion. By changing the proportion of each sub-band in the channel, the channel features are fused, and the nonlinear characteristics between the sub-bands are learned through the activation function. In addition, the computational complexity of 1x1 convolution is very small. The FC layer is used to construct a frequency point interaction layer. With the help of the fully connected characteristics of the FC layer, the interaction information between all frequency points in each sub-band can be learned in each time step. With the help of the FC layer, each frequency point refers to the information of other frequency points in the current time step, that is, the global frequency point interaction information (features) is learned.
[0017] In a second aspect, a method for training a sound source separation model is provided, the training method comprising: obtaining training data, the training data comprising audio data to be trained and an audio signal label of at least one known sound source corresponding to the audio data to be trained, the audio data to be trained being a mixed audio of multiple sound sources synthesized by the audio signal of at least one known sound source and other audio signals; inputting the audio data to be trained into an initial sound source separation model, and updating the weight parameters of the initial sound source separation model according to the difference between the predicted audio signal of at least one known sound source output by the initial sound source separation model and the audio signal label of at least one known sound source, thereby obtaining a trained sound source separation model ; The initial sound source separation model includes a feature extraction network, a frequency band interaction layer and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network, and is used to extract features of the audio data to be trained to obtain a first audio feature set; the frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be trained from the first audio feature set, thereby obtaining a second audio feature set; the frequency point interaction layer is used to extract frequency point interaction features within each sub-band in the second audio feature set, thereby obtaining a third audio feature set; the sound source separation model is also used to separate a predicted audio signal of at least one known sound source from the third audio feature set.
[0018] The trained sound source separation model obtained by the training method of the second aspect can be applied to the audio processing method of the first aspect. It should also be understood that the description of the sound source separation model in the first aspect can be referenced in the second aspect. The difference between the two is that the first aspect uses the trained model for inference, while the second aspect is the model training phase. For the sake of brevity, this will not be repeated here.
[0019] In combination with the second aspect, in some implementations of the second aspect, the frequency band interaction layer and / or the frequency point interaction layer are constructed using a two-dimensional convolutional basic network, an attention mechanism network or a fully connected layer.
[0020] In combination with the second aspect, in some implementations of the second aspect, the multiple sub-bands are obtained by dividing the training audio data along the frequency dimension, or the multiple sub-bands are obtained by dividing the first extraction result output by the first network layer along the frequency dimension, and the first network layer is any layer of the feature extraction network.
[0021] In a third aspect, a sound source separation device is provided, which includes a unit composed of software and / or hardware for executing any one of the methods in the first aspect.
[0022] In a fourth aspect, a training device is provided, which includes a unit composed of software and / or hardware for executing any one of the methods of the second aspect.
[0023] In a fifth aspect, an electronic device is provided, comprising: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, the one or more processors calling the computer instructions to enable the electronic device to implement any one of the methods of the first and second aspects.
[0024] In a sixth aspect, a chip system is provided, which is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device can implement any one of the methods of the first and second aspects.
[0025] Optionally, the chip system also includes a memory, which is electrically connected to the processor.
[0026] Optionally, the chip system may further include a communication interface.
[0027] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes instructions, and when the instructions are executed on an electronic device, the electronic device can implement any one of the methods of the first aspect and the second aspect.
[0028] In an eighth aspect, a computer program product is provided, which includes a computer program, and when the computer program is executed by an electronic device, it can implement any one of the methods of the first aspect and the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of an applicable scenario of an embodiment of the present application.
[0030] Figure 2 It is a schematic diagram of the execution process of an audio processing method according to an embodiment of the present application.
[0031] Figure 3 It is a structural diagram of a feature extraction network in an embodiment of the present application.
[0032] Figure 4 This is a schematic flowchart of an audio processing method according to an embodiment of the present application.
[0033] Figure 5 It is a schematic diagram of the execution process of feature extraction using a sound source separation model in an embodiment of the present application.
[0034] Figure 6 This is a schematic diagram of a frequency band interaction layer processing process in an embodiment of the present application.
[0035] Figure 7This is a schematic diagram of another frequency band interaction layer processing process of an embodiment of the present application.
[0036] Figure 8 This is a schematic diagram of another frequency band interaction layer processing process of an embodiment of the present application.
[0037] Figure 9 This is a schematic diagram of a frequency point interaction layer processing process in an embodiment of the present application.
[0038] Figure 10 This is a schematic diagram of another frequency interaction layer processing process of an embodiment of the present application.
[0039] Figure 11 This is a schematic diagram of another frequency interaction layer processing process of an embodiment of the present application.
[0040] Figure 12 It is a schematic diagram of the execution process of an audio processing method according to an embodiment of the present application.
[0041] Figure 13 This is a comparison chart of the results of separating drum sounds from the same mixed audio using different schemes.
[0042] Figure 14 Schematic diagram of a non-full-band sound source according to an embodiment of the present application.
[0043] Figure 15 This is a schematic flowchart of another audio processing method according to an embodiment of the present application.
[0044] Figure 16 It is a schematic flowchart of a method for training a sound source separation model in an embodiment of the present application.
[0045] Figure 17 This is a schematic diagram of the software architecture of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0046] The following describes the solutions of the embodiments of the present application with reference to the accompanying drawings.
[0047] Figure 1 This is a schematic diagram of an applicable scenario of the embodiment of the present application. Figure 1 As shown in (a), suppose that a piece of original music is a mixture audio, which includes vocals, drums, bass, piano, guitar and other musical instruments. Any musical instrument is a sound source, and vocals are also a sound source. That is, the mixed audio of the original music contains audio signals from multiple sound sources. The so-called sound source separation is to separate the audio signals of one or several sound sources in the mixed audio, for example Figure 1 In (a), we take the example of separating human voice and drum sound. Human voice has harmonic characteristics and is a non-full-band sound source, while drum sound is a full-band sound source. However, it should be understood that there is no limitation on the number of separated sound source categories and the types of sound sources. For example, the audio signal to be separated can also be the audio signal of a guitar, which has both full-band characteristics and harmonic characteristics. The above-mentioned sound source types can all use the solution of this application to separate their audio signals from the mixed audio. Other situations will not be listed one by one.
[0048] Assume that a piece of original music is used as a movie interlude, such as Figure 1 In the example of sound source separation shown in (b), the original movie audio is a mixed audio, which includes human voice, background music, sound effects, ambient sound, noise and other sounds. The background music includes audio signals of drums, bass, piano, guitar and other instruments. Figure 1 In (b), the example of separating the vocals from the original movie audio and the drum sounds from the background music in the original movie audio is used. However, it should be understood that there is no limitation on the number of separated sound source categories, the types of sound sources, or whether to separate (extract) from the movie audio or further separate from the background music.
[0049] That is to say, sound source separation can be to separate the audio signal of a specified sound source from a mixed audio composed of audio signals of multiple sound sources, or to separate the audio signal of a specified sound source from a mixed audio composed of multiple sound sources and further mixed audio composed of other sound sources.
[0050] The separated audio signal can be further applied to other audio processing. Figure 1 In (c), the example of using the audio signal separated from the sound source for spatial audio is used. However, it should be understood that the separated audio signal can also be played directly or used for other purposes without limitation. Figure 1 As shown in (c), after the vocals and drums are amplified separately, the vocals are further separated into the center vocals and background vocals (such as the lead singer and backing vocals in a song), and the drums are combined with the time domain data of the mixed audio to calculate the time domain data of the other sound source audio, that is, the part of the time domain data of the mixed audio excluding the vocals and drums is calculated.
[0051] The time domain data of other audio is combined with the background vocals and drums to expand the sound field, and then combined with the center vocals to obtain the processed audio. This processing process needs to be performed online while the song is playing, that is, it is necessary to continuously separate the sound sources in real time, extract the audio signal of the specified sound source, and then use Figure 1 The execution process shown in (c) generates spatial audio (processed audio).
[0052] Therefore, if the sound source separation is not fast enough, it will cause a long delay and the spatial audio cannot be generated in time, affecting the user's auditory experience. In addition, if the separation accuracy is too low, the timbre of the separated audio signal will be distorted, affecting the user's auditory experience.
[0053] To address this problem, the present application proposes a new audio processing method that can achieve relatively high separation accuracy while relatively short processing delay, which is described in detail below in conjunction with the accompanying drawings.
[0054] Figure 2 FIG is a schematic diagram of the execution process of an audio processing method according to an embodiment of the present application. Figure 2 As shown, after the audio data to be processed is input into the sound source separation model, it can be processed by the sound source separation model to output the processed audio data. The audio data to be processed is a mixed audio containing audio signals from multiple sound sources, and the processed audio data is the audio signals of one or more sound sources separated from the input mixed audio.
[0055] The sound source separation model consists of a feature extraction network, a frequency band interaction layer, and a frequency point interaction layer. The feature extraction network extracts high-order spectral features of the mixed audio, the frequency band interaction layer extracts global features between different frequency bands, and the frequency point interaction layer extracts global features between different frequency points within the same frequency band.
[0056] It should be noted that, in the embodiment of the present application, global mainly refers to global in the frequency dimension, that is, full coverage in the frequency dimension, or can be understood as full frequency band.
[0057] In the traditional solution, the sound source separation model only includes a feature extraction network. When the feature extraction network is constructed entirely with a one-dimensional convolutional network as the basic network, since the one-dimensional convolutional network can cover the entire data range of the data to be processed, it has the ability to extract local features and global features. However, it is precisely because of this full coverage that the one-dimensional convolutional network has a large amount of computation each time, and the amount of computation in the entire processing process is too large, which brings about the problem of long delay. When the feature extraction network is constructed entirely with a two-dimensional convolutional network as the basic network, since each operation of the two-dimensional convolutional network is based on the size of the convolution kernel, the amount of computation in the entire processing process is greatly reduced. However, this also brings new problems, because the two-dimensional convolutional network is limited by the size of the convolution kernel, resulting in each operation being limited to the data range that the convolution kernel size can cover. Therefore, only local features can be extracted, but the ability to extract global interaction features (mainly full-band features in the frequency dimension in this application) is very poor, resulting in insufficient extracted global interaction features. Based on such insufficiently extracted features for separation, the separated sound source signal is prone to spectral holes in the entire frequency band, causing the timbre of the separated audio signal to change, affecting the user's auditory experience. For example, for percussion music such as the above-mentioned drum sound, the percussion features in the low-frequency band are relatively more obvious. If only local features are used for separation, it is easy to mistake the audio signals of other sound sources in the high-frequency band for the audio signals of the drum sound, or to mistake the audio signals belonging to the drum sound in the high-frequency band for the audio signals of other sound sources. If local features and global features are combined, the features of the audio signals in the high-frequency band can be combined with the characteristic excitation of the percussion music of the low-frequency drum sound to separate the features of the audio signal, thereby improving the separation accuracy. It is based on this that the sound source separation model of the present application is designed as a two-dimensional convolutional network + one-dimensional convolutional network structure, so as to fully extract local features and global features. In short, since the two-dimensional convolutional network is limited by the size of the convolution kernel, it is easier to extract local features, but it is difficult to extract global features, resulting in too low separation accuracy, spectral holes, and changes in the timbre of the sound source. In addition, since the human ear is more sensitive to sounds between 2 kilohertz (kHz) and 5 kHz, the spectrum holes in this range are more easily perceived by users, causing users to feel that the timbre of the audio signal is distorted (for example, it may be a drum sound, but it sounds like the sound of other instruments).
[0058] In addition, for a sound source with harmonic characteristics like human voice, the global features are not obvious, so only using a two-dimensional convolutional network to extract local features can also achieve relatively high accuracy. When the present application solution is adopted, the feature extraction network is constructed by a two-dimensional convolutional network, which can fully extract local features, that is, fully extract the features of non-full-band sound sources with harmonic characteristics such as human voice, and the subsequent frequency band interaction layer and frequency point interaction layer are used to supplement the extraction of global features after the local features have been fully extracted. Since the global features are not obvious, the global features extracted by the frequency band interaction layer and the frequency point interaction layer may be relatively few. Based on the feature vectors extracted by such a feature extraction process, after subsequent separation, a high-precision audio signal can still be separated. In other words, for non-full-band sound sources, the present application solution does not affect the accuracy of local feature extraction because the extraction of global features is a supplementary extraction after the local features are fully extracted, and only the number of global features extracted is relatively small because the global features are not obvious. Therefore, the separation accuracy can still be ensured, and the accuracy will not be reduced due to the extraction of global features.
[0059] Based on the above analysis, this application adds a frequency band interaction layer and a frequency point interaction layer on the basis of constructing a feature extraction network based on a two-dimensional convolutional network as the basic network to make up for the deficiency of the two-dimensional convolutional network in extracting global interaction features insufficiently. And because it is only an additional network layer, rather than running through the entire feature extraction network, the computational complexity of the sound source separation model is still much smaller than the computational complexity of the feature extraction network constructed entirely with a one-dimensional convolutional network as the basic network in the traditional solution. In addition, in order to be able to extract frequency band interaction features, it is also necessary to segment the audio data to be processed, and divide the audio data or the extraction results generated during the feature extraction network extraction into multiple sub-bands in the frequency dimension, and each sub-band is a sub-band of the full time dimension. For example, assuming that the duration corresponding to the audio data to be processed is 10 seconds, then the duration corresponding to all sub-bands is 10 seconds. It should be understood that the above numerical values are for understanding the solution and there is no limitation. Combined with Figure 13 The spectrum of the standard drum sound shown in (a) is divided into multiple segments along the vertical axis, and the length of each segment along the horizontal axis is the same. The vertical axis is the frequency dimension, and the horizontal axis is the time dimension. Figure 13 As shown in (a), an example of dividing three sub-bands along the frequency dimension is also given. It can be seen that the time dimensions of sub-bands 1 to 3 are completely consistent, with the same start and end times, while the frequency dimensions cover different frequency bands. It should be understood that this is only for the purpose of Figure 13 The sub-bands are described. Figure 13 The rest of the content will be expanded below and will not be repeated here.
[0060] The two-dimensional convolutional feature extraction network has a good ability to extract local features, so it can fully extract the harmonic features of different scales of the audio data to be processed.
[0061] It should be understood that the dimensions of the sub-bands in the frequency dimension can be the same or different, that is, when divided along the frequency dimension, they can be divided equally or unequally. Equal division can make it more convenient to integrate the processing data of all sub-bands in the future, and unequal division is conducive to flexible adjustment according to different situations of different audio data to be processed. For example, relatively important frequency bands can be divided more finely, and relatively unimportant frequency bands can be divided more roughly. For example, assuming that the characteristics of a certain sound source are mainly concentrated in the high frequency band, the high frequency band can be divided into N1 sub-bands along the frequency dimension, and the low frequency band can be divided into N2 sub-bands along the frequency dimension. The dimension of the frequency dimension of each sub-band in the N1 sub-bands is greater than the dimension of the frequency dimension of each sub-band in the N2 sub-bands.
[0062] In some implementations, the sound source separation model may also only include a feature extraction network and a frequency point interaction layer. The reason is that the present application solution will segment the audio data to be processed, so when the feature extraction network itself has a network layer that can extract the interaction features between frequency bands, the interaction features between frequency bands can be extracted under the premise of segmentation, so that the frequency band interaction layer can be removed. However, it should be understood that in this case, not canceling the frequency band interaction layer will only make the feature extraction more sufficient and will not bring about degradation. It should also be understood that in traditional solutions, even if the feature extraction network used itself has a network layer that can extract the interaction features between frequency bands, because the audio data to be processed is not segmented, it is still impossible to extract the interaction features between frequency bands. In addition, the present application solution addresses the problem of poor extraction capability of the traditional feature extraction network based on two-dimensional convolution, by adding a frequency band interaction layer and a frequency point interaction layer to supplement the extraction of features that the traditional feature extraction network based on two-dimensional convolution cannot extract. On this basis, during the experiment, it was found that after removing the frequency band interaction layer, the separation effect of some special feature extraction networks was still good. Then, this phenomenon was analyzed and it was found that these special feature extraction networks have a common feature that they all have their own network layer that can extract the interaction features between frequency bands. Based on the principle analysis, since the audio data to be processed will be segmented in the early stage of this application, such a special feature extraction network can play the ability to extract features between frequency bands. Therefore, after removing the frequency band interaction layer, the feature extraction effect is still relatively good. Therefore, this feature extraction network that extracts both audio features and frequency band interaction features is still different from the traditional solution where only the feature extraction network extracts audio features.
[0063] According to similar logic as above, in other implementations, when the feature extraction network is capable of extracting global frequency interaction features, the sound source separation model may also only include the feature extraction network and the frequency band interaction layer.
[0064] The above method of omitting some network layers in the sound source separation model based on the performance analysis of the feature extraction network can simplify the model without affecting the adequacy of the extracted features, further reduce the amount of calculation and shorten the delay.
[0065] The audio data to be processed may be frequency domain data obtained by performing Fourier transform on time domain audio data.
[0066] The convolution kernel size of a one-dimensional convolution is proportional to the number of features. In the sound source separation task, the number of features of a one-dimensional convolution is usually 1024 resolution, where the resolution refers to the number of fast Fourier transform (FFT) points. The convolution kernel size of a two-dimensional convolution is independent of the number of features. Taking a 3x3 convolution kernel as an example, when the number of features of the two-dimensional convolution is equal to that of the one-dimensional convolution, the computational complexity of the two-dimensional convolution is approximately 3 / F times that of the one-dimensional convolution. It can be seen that the computational complexity of the two-dimensional convolution is much smaller than that of the one-dimensional convolution. For this reason, although the present application solution adds a frequency band interaction layer and a frequency point interaction layer on the basis of the two-dimensional convolutional feature extraction network, the computational complexity of the sound source separation model is still much smaller than that of the one-dimensional convolutional feature extraction network. It is a solution that effectively improves the extraction capability of the two-dimensional convolutional feature extraction network without increasing the computational complexity. It can also be understood as a solution that greatly reduces the computational complexity while ensuring feature extraction capabilities equivalent to those of the one-dimensional convolutional feature extraction network.
[0067] Figure 3 It is a structural diagram of a feature extraction network in an embodiment of the present application. Figure 3 This is an example of the structure of the feature extraction network of this application. The example structure is called U-Net structure. The U-Net structure includes a contraction path and an expansion path. The contraction path includes feature extraction and down sampling (down sample), which is equivalent to distilling and refining the features. The expansion path includes feature map up sampling (up sample) and recovery. Here, the U-Net structure includes three down sampling and three up sampling as an example. Figure 3As shown in the figure, the U-Net structure downsamples the input data three times before entering the center, and then outputs the processed data after upsampling three times. Each downsampling or upsampling is preceded by a base network processing. Taking the two-dimensional convolutional base network as an example, in the U-Net structure, it includes, in order: the two-dimensional convolutional base network, downsampling layer, the two-dimensional convolutional base network, downsampling layer, the two-dimensional convolutional base network, downsampling layer, the two-dimensional convolutional base network, center layer, the two-dimensional convolutional base network, upsampling layer, the two-dimensional convolutional base network, upsampling layer, the two-dimensional convolutional base network, upsampling layer, the two-dimensional convolutional base network, upsampling layer, and the two-dimensional convolutional base network.
[0068] In traditional solutions, a U-Net structured feature extraction network may be constructed based on a one-dimensional convolutional base network. It can be seen that the one-dimensional convolutional base network needs to be reused multiple times for calculations, which leads to a further dramatic increase in the amount of computation, resulting in a greater overall computational load. Therefore, the present application constructs a U-Net structured feature extraction network based on a two-dimensional convolutional base network, making the computational load far less than that of the U-Net structured feature extraction network constructed based on the one-dimensional convolutional base network. The resulting weakening of the ability to extract global interactive features is overcome by adding frequency band interaction layers and frequency point interaction layers, thereby achieving the effect of reducing the computational load and shortening the processing delay of sound source separation while ensuring separation accuracy.
[0069] Figure 4 This is a schematic flow chart of an audio processing method according to an embodiment of the present application. Figure 4 Each step is described below.
[0070] S401: Acquire audio data to be processed.
[0071] The audio data to be processed includes audio signals from various sound sources, such as a musical instrument or human voice. A sound source may also be referred to as a sound source, a sound origin, a sound type, or a sound category.
[0072] The audio data to be processed may be frequency domain data obtained by Fourier transforming a time domain audio signal.
[0073] The method of obtaining the audio data to be processed can be an online acquisition method, for example, after collecting the audio signal in real time through a sound receiving sensor such as a microphone, performing Fourier transform to obtain the audio data to be processed in the frequency domain; it can also be an offline acquisition method, for example, the audio data to be processed can be read from a storage module, or downloaded from the network through a communication interface, etc. There is no limitation.
[0074] S402: Process the audio data to be processed using a sound source separation model to obtain an audio signal of at least one target sound source.
[0075] The at least one target sound source is at least one sound source among multiple sound sources. The target sound source can be understood as the sound source that is to be separated, for example Figure 1 Taking the example of separating vocals and drum sounds, the vocals and drum sounds can be examples of at least one target sound source. This example can also be an example where the at least one target sound source includes multiple target sound sources, and these multiple target sound sources are vocals and drum sounds. However, it should be understood that there is no limitation on the specific type or sources of the at least one target sound source, nor on the number of sound sources included. It should also be understood that the number of the at least one target sound source needs to be less than the number of multiple sound sources of the audio data to be processed. If the two numbers are equal, it is equivalent to not performing sound source separation.
[0076] In one implementation, the sound source separation model includes a feature extraction network, a frequency band interaction layer and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network, and is used to extract features of the audio data to be processed to obtain a first audio feature set; the frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be processed from the first audio feature set, thereby obtaining a second audio feature set; the frequency point interaction layer is used to extract global frequency point interaction features within each sub-band in the second audio feature set, thereby obtaining a third audio feature set; the sound source separation model is also used to separate the audio signal of at least one target sound source from the third audio feature set. Figure 2 and Figure 5 The sound source separation model in can be an example of the sound source separation model here.
[0077] The multiple sub-bands are all-time sub-bands, that is, the time dimensions of the multiple sub-bands are the same. For the content of division along the frequency dimension, please refer to the above description, which will not be repeated for the sake of brevity.
[0078] In the solution of the present application, by dividing the frequency sub-bands, the frequency band interaction layer in the sound source separation model can extract the frequency band interaction features between multiple frequency sub-bands.
[0079] In one implementation, the multiple sub-bands are obtained by partitioning the processed audio data along the frequency dimension, or by partitioning the first extraction result output by the first network layer along the frequency dimension, where the first network layer is any layer of the feature extraction network. In this implementation, the multiple sub-bands corresponding to the processed audio data can be partitioned before entering the feature extraction network, during feature extraction by the feature extraction network, or before inputting the extracted features into the frequency band interaction layer after the feature extraction network completes feature extraction, thereby enabling flexible adjustment based on actual application scenarios.
[0080] As mentioned above, for some special feature extraction networks, combined with segmentation operations (operations that divide the frequency bands into multiple sub-bands), some traditional feature extraction networks can also have the ability to extract global interaction features. Therefore, for such feature extraction networks, the network structure can be appropriately streamlined, thereby further reducing the amount of calculation and shortening the delay without affecting the feature extraction effect.
[0081] In one implementation, when the feature extraction network has the ability to extract frequency band interaction features between multiple sub-bands so that the first audio feature set includes frequency band interaction features between multiple sub-bands, the sound source separation model does not include a frequency band interaction layer; and / or, when the feature extraction network has the ability to extract global frequency point interaction features for each frequency point corresponding to the audio data to be processed so that the first audio feature set includes global frequency point interaction features for each frequency point, the sound source separation model does not include a frequency point interaction layer.
[0082] In one example, a sound source separation model includes a feature extraction network and a frequency interaction layer. The feature extraction network is constructed using a two-dimensional convolutional base network, and the feature extraction network includes a 1x1 convolutional layer for extracting features from multiple sub-bands input to the feature extraction network to obtain a first audio feature set. The frequency interaction layer is used to extract global frequency interaction features within each sub-band in the first audio feature set, thereby obtaining a third audio feature set. The sound source separation model is further used to separate the audio signal of at least one target sound source from the third audio feature set. In this implementation, since the feature extraction network itself includes a 1x1 two-dimensional convolutional layer, and the input to the feature extraction network is a segmented sub-band, the 1x1 two-dimensional convolutional layer in the feature extraction network can already function as a frequency band interaction layer, capable of extracting frequency band interaction features between multiple sub-bands. Therefore, in this case, the frequency band interaction layer in the sound source separation model can be omitted, and this implementation can include only the feature extraction network and the frequency band interaction layer. This implementation method can further simplify the model structure when the feature extraction network contains a 1x1 convolution layer, further reduce the amount of calculation, and thus shorten the processing delay of sound source separation. It should also be understood that if a feature extraction network has multiple 1x1 convolution layers, and the sub-bands are divided based on the output results of the network layer at the front of the feature extraction network, then these multiple sub-bands can still be used under the action of the subsequent 1x1 convolution layer to extract the frequency band interaction features. Figure 12 , Figure 12 The sound source separation model in includes a feature extraction network ( Figure 12 Take U-Net as an example), and the sub-bands are divided before the feature extraction network U-Net. If U-Net itself contains 1x1 convolution, then during the feature extraction of U-Net, both local features and frequency band interaction features can be extracted. Therefore, under this premise, the frequency band interaction layer can be removed, which means that Figure 12 The sound source separation model in only includes U-Net and frequency interaction layer. Figure 12 Explain under what circumstances the frequency band interaction layer can be simplified. Figure 12 The rest of the content will be explained in detail below and will not be repeated here.
[0083] In one implementation, the frequency band interaction layer and / or frequency point interaction layer are constructed using a two-dimensional convolutional base network, an attention network, or a fully connected (FC) layer. It should be understood that while the attention network in this implementation is computationally more complex than a two-dimensional convolutional network, since it is merely a layer added to the feature extraction network, the overall computational complexity of the sound source separation model is still significantly less than that of a feature extraction network constructed entirely from a one-dimensional convolutional network in traditional solutions. This implementation method gives examples of frequency band interaction layers and frequency point interaction layers. By adding a frequency band interaction layer, the frequency band interaction features between multiple sub-bands can be supplemented and extracted. By adding a frequency point interaction layer, the frequency point interaction features in each sub-band can be supplemented and extracted. It should also be understood that since the feature extraction is first performed through the frequency band interaction layer and then the frequency point interaction layer is used for feature extraction, although the frequency point interaction layer extracts the frequency point interaction features in each sub-band, these frequency point features have been affected by the previous feature extraction network and the frequency point interaction layer to generate new global features. Therefore, the frequency point interaction features extracted here are also global frequency point interaction features.
[0084] Figure 6-Figure 8 is an example of a band interaction layer, Figures 9-11 This is an example of a frequency interaction layer. For details, see below. For the sake of brevity, we will not elaborate on this here. The frequency interaction features extracted by the frequency interaction layer can be understood as follows: the calculation process of the frequency interaction layer ensures that the calculation of each sub-band is combined with the data of other sub-bands. The features obtained by considering other sub-bands are called frequency interaction features, and the frequency interaction layer is used to extract such features. The frequency interaction layer extracts global frequency interaction features. It can be understood as follows: the calculation process of the frequency interaction layer ensures that the calculation of each frequency point in each sub-band is combined with the data of other frequency points. Since the data of each frequency point in this sub-band has been combined with the data of other sub-bands after the action of the previous frequency interaction layer, the data of each frequency point in this sub-band is combined with the data of other sub-bands, so the frequency interaction layer extracts global frequency interaction features. The frequency interaction features obtained by considering all frequencies are called global frequency interaction features.
[0085] In one example, the frequency band interaction layer is constructed using a 1x1 convolution layer, and the frequency point interaction layer is constructed using an FC layer. In this example, the frequency band interaction layer is constructed using a 1x1 convolution layer to perform channel fusion. By changing the proportion of each sub-band in the channel, channel feature fusion is performed, thereby learning the nonlinear characteristics between sub-bands through the activation function. In addition, the computational complexity of the 1x1 convolution is very small. The frequency point interaction layer is constructed using an FC layer. With the help of the fully connected nature of the FC layer, the interaction information between all frequency points in each sub-band can be learned at each time step. With the help of the FC layer, each frequency point refers to the information of other frequency points in the current time step, that is, the global frequency point interaction information is learned.
[0086] It should be understood that a time step refers to a time unit of spectral data of audio.
[0087] It should be noted that when using different types of basic networks to construct the frequency band interaction layer and / or frequency point interaction layer, since the order of dimensions acted on by different types of basic networks is different, and the dimensions that need to be acted on are also different, it may be necessary to perform a reshape operation so that different types of basic networks can act on the target dimension. In the present application scheme, the target dimension of the frequency band interaction layer is the channel dimension, and the target dimension of the frequency point interaction layer is the frequency dimension.
[0088] In one implementation, the above method also includes: when the target dimension of the frequency band interaction layer and / or the frequency point interaction layer does not match the action dimension of the basic network for constructing the frequency band interaction layer and / or the frequency point interaction layer, a reshape operation is performed to make the target dimension match the action dimension.
[0089] For example, the attention network works on the last dimension, and the frequency band interaction layer needs to work on the channel dimension (an example of the target dimension). However, the channel dimension in the initial input data of the frequency band interaction layer is not in the last dimension. Therefore, the initial dimension needs to be reshaped to adjust the channel dimension to the last dimension, so that the attention network can work on the channel dimension. For other cases, please refer to Figures 6-11 The relevant introduction will not be repeated here.
[0090] It should also be understood that the effective dimension of a 1x1 convolutional layer is the second dimension, while the effective dimension of an attention network and a FC layer is the last dimension. Therefore, in a frequency-interaction layer, when a 1x1 convolutional layer is used, the effective dimension of the 1x1 convolutional layer is the second dimension, which is exactly the target dimension (channel dimension) of the frequency-interaction layer. Therefore, no reshape operation is required. In a frequency-interaction layer, when an attention network or FC layer is used, the effective dimension is the last dimension, not the target dimension (channel dimension) of the frequency-interaction layer, but the frequency dimension. Therefore, a reshape operation is required to adjust the channel dimension to the last dimension before processing with the frequency-interaction layer. After processing, a reshape operation is performed to restore the dimension. In a frequency-point interaction layer, when a 1x1 convolutional layer is used, the effective dimension of the 1x1 convolutional layer is the second dimension, not the target dimension (frequency dimension) of the frequency-point interaction layer. Therefore, a reshape operation is required to adjust the frequency dimension to the second dimension before processing with the frequency-point interaction layer. After processing, a reshape operation is performed to restore the dimension. In the frequency interaction layer, when using an attention network or FC layer, the last dimension is the target dimension of the frequency interaction layer (the frequency dimension). Since the two dimensions are consistent, no reshape operation is required. Furthermore, because the 1x1 convolution operates on the second dimension, which matches the dimension of the feature extraction network, the frequency interaction layer can be omitted when the feature extraction network includes a 1x1 convolution.
[0091] In actual applications, some sound sources have hollow characteristics, that is, they have harmonic characteristics, which makes them non-full-band sound sources. For example, human voice is a sound source with harmonic characteristics. Figure 14 As can be seen from the schematic diagram of the non-full-band sound source, this type of non-full-band sound source has holes in its own spectrum. For example, human voice is a non-full-band sound source with harmonic characteristics. Therefore, it is not necessary to further extract the frequency band interaction features and frequency point interaction features in the audio data to be processed for this type of sound source. Because the global features of this type of sound source are not obvious, combined with the spectrum, it is like Figure 14 The paragraphs shown are incoherent and almost unrelated. Figure 13 The vertical line runs from beginning to end along the vertical axis (frequency dimension), just like the spectrum of the drum sound (an example of a full-band sound source) in [1]. Therefore, there is no need to analyze the continuity of such line segments to make them more coherent. In other words, there is no need to extract frequency band interaction features and global frequency point interaction features to make the global frequency dimension more complete. This is reflected in the spectrum, that is, there is no need to make the frequency dimension lines more coherent and complete, because non-full-band sound sources refer to sound sources that do not occupy the entire dimensional range in the frequency dimension.
[0092] In one implementation, step S402 may include: when the first target sound source belongs to a preset sound source category, extracting frequency band interaction features between multiple sub-bands from the first audio feature set using a frequency band interaction layer to obtain a second audio feature set, extracting global frequency point interaction features within each sub-band in the second audio feature set using a frequency point interaction layer to obtain a third audio feature set, and separating the audio signal of the first target sound source from the third audio feature set using a sound source separation model; the first target sound source is any one of the at least one target sound source; or, when the first target sound source does not belong to the preset sound source category, separating the audio signal of the first target sound source from the first audio feature set using the sound source separation model. In this implementation, frequency band interaction features and global frequency point interaction features are extracted only for target sound sources that belong to the preset sound source category, while frequency band interaction features and global frequency point interaction features are not extracted for target sound sources that do not belong to the preset sound source category. This can adaptively omit the extraction of frequency band interaction features and frequency point interaction features for target sound sources that do not belong to the preset sound source category, further shortening processing latency. Figure 12 This is an example of this implementation method, which will not be described here for the sake of brevity.
[0093] In one example, the preset sound source category is used to represent the full-band sound source category; or, the output value of the discriminant network corresponding to the sound source in the preset sound source category is less than or equal to a preset threshold. That is to say, the possible sound source can be confirmed as a full-band sound source category or a non-full-band sound source category through prior knowledge, thereby determining the preset sound source category in an enumerated manner; the feature vector of the sound source can also be input into the discriminant network so that the discriminant network outputs the corresponding output value, and a preset threshold is set. Only the sound source category with an output value less than or equal to the preset threshold is the preset sound source category. This example gives an example of how to set the preset sound source category. The sound source categories can be enumerated or determined by discriminant scoring. Figure 15 This is an example of this, which will not be repeated here for the sake of brevity.
[0094] Figure 4 The method shown adds a frequency band interaction layer and a frequency point interaction layer on the basis of the two-dimensional convolutional feature extraction network, and extracts the frequency band interaction features between multiple sub-bands and the frequency point interaction features within each sub-band respectively. Thus, while retaining the advantage of the two-dimensional convolutional feature extraction network in terms of small computational complexity, it also makes up for the defect of the two-dimensional convolutional feature extraction network in its insufficient ability to extract global interaction features, so that the entire sound source separation process can achieve relatively high separation accuracy and relatively short processing delay.
[0095] Figure 5 It is a schematic diagram of the execution process of feature extraction using a sound source separation model in an embodiment of the present application. Figure 5 This is an example of the execution process of obtaining the third audio feature set in step S402. The sound source separation model includes a two-dimensional convolutional feature extraction network, a frequency band interaction layer and a frequency point interaction layer. Figure 5 As shown, a two-dimensional convolutional feature extraction network can be used to extract the spectral features of the audio data to be processed, which can be high-order spectral features. It should be understood that the two-dimensional convolutional feature extraction network can fully extract the harmonic features of the audio data to be processed at different scales, that is, extract local features.
[0096] After that, the high-order spectral features or spectrum can be input into the frequency band interaction layer. However, before inputting into the frequency band interaction layer, the spectrum or high-order spectral features of the audio data to be processed need to be divided into multiple sub-bands along the frequency dimension. Here, the division into sub-bands (band) 1-sub-band N is taken as an example, where N is an integer greater than 1. It can be seen that Figure 5 As shown in the figure, we can first use the feature extraction network to perform preliminary two-dimensional feature extraction and then divide the frequency bands, or we can divide the original spectrum (that is, the spectrum of the audio data to be processed) into sub-bands, or we can divide the intermediate results (feature vectors) output by any layer during the feature extraction process of the feature extraction network into sub-bands, as long as the sub-band division is completed before inputting into the frequency band interaction layer.
[0097] After entering the frequency band interaction layer, the frequency band interaction layer will extract the frequency band interaction features between multiple sub-bands, thereby obtaining the frequency band interaction results output by N channels (i.e., channel 1-channel N). Extracting the global interaction features between different sub-bands can avoid the information fragmentation between each sub-band. Figure 13 The three sub-bands shown in (a) are equivalent to preventing the isolation of the feature extraction of the three sub-bands, resulting in relatively sufficient extraction for a certain sub-band, but the data between different sub-bands cannot be combined, especially the obvious separation near the boundaries of different sub-bands. It should be understood that the feature extraction of the three sub-bands is isolated from each other. Figure 13 This is to illustrate that information fragmentation will cause obvious boundaries between sub-bands. Other contents will be explained below and will not be repeated here.
[0098] After entering the frequency interaction layer, for any input of channel i, the global frequency interaction features between the frequencies within the sub-band corresponding to the channel i can be extracted. The dimension of the frequency interaction result is the same as that of a single sub-band, thus ensuring that the details of the global features are not lost.
[0099] Figure 5This is an example of obtaining the third audio feature set in step S402. First, the local features are extracted using the feature extraction network (two-dimensional convolution). Then, the local feature set (first feature set) is divided into N sub-bands along the frequency dimension, and the N sub-bands are input into the band interaction layer. Each sub-band in the band interaction layer corresponds to a channel, thereby extracting the interaction features between different sub-bands. Then, these band interaction features ( Figure 5 The frequency band interaction results in the frequency band interaction layer are combined with the previously obtained local feature set to form a new feature set (the first feature set above), and then input into the frequency point interaction layer in the same sub-band division method. The frequency point interaction layer performs global frequency point interaction feature ( Figure 5 The frequency interaction results in the [frequency interaction results] are extracted and combined with the obtained frequency interaction features from the second feature set to obtain the aforementioned third feature set. The sound source separation model can then perform sound source separation based on the third feature set through channel selection to obtain the audio signal of the target sound source.
[0100] Figure 5 The frequency band interaction layer in Figure 6-Figure 8 Any one of them, or other appropriate network layer construction. Figure 5 The frequency interaction layer in Figures 9-11 Any one of the above, or other suitable network layer construction. No more listing.
[0101] Figure 6 This is a schematic diagram of a frequency band interaction layer processing process in an embodiment of the present application. Figure 6 Take the construction of a frequency band interaction layer based on 1x1 two-dimensional convolution as an example. With the help of 1x1 convolution, channel fusion characteristics can be achieved, and the proportion of each sub-band in the output channel can be adaptively selected, thereby learning the interaction characteristics between sub-bands. Figure 6 As shown in the figure, the sub-band features (sub-band 1 - sub-band N) are subjected to a 1x1 convolution, and the activation function relu(x) = max(0, x) is used to output the band interaction results. The output band interaction results can satisfy the formula below.
[0102] Where n∈[1,N] represents the channel number, and N represents the total number of channels. The frequency band interaction layer contains N convolution kernels, each of size 1x1 and dimension Nx1x1, with a total of N output channels. relu(x) is the activation function, and w and b represent the weight vector and bias vector of the convolution kernel, respectively. represents the i-th row and j-th column element of the k-th subband of the input, Represents the element in row i and column j of the nth channel of the band interaction layer output. It should be understood that i and j are random values for the row and column, and the range of values depends on the size of each subband. Furthermore, the process of training a model refers to updating the model parameters, which in this case refers to updating the parameter values in the weight and bias vectors.
[0103] Figure 7 This is a schematic diagram of another frequency band interaction layer processing process of an embodiment of the present application. Figure 7 Take the example of building a frequency band interaction layer based on the attention network. The characteristics of the attention network make it easier to fully learn the interaction features between sub-bands. Although its computational complexity is relatively large, since it is only an additional network layer added outside the feature extraction network, the increase in the computational complexity of the entire sound source separation model is not obvious. The overall computational complexity of the transliteration sound source separation model is still much smaller than the traditional solution of completely using the attention network to build a feature extraction network. In addition, since the last dimension of the input sub-band dimension is the frequency dimension, and the frequency band interaction features need to be extracted on the channel dimension, it is also necessary to deform (reshape) the sub-band dimension before extracting the frequency band interaction features, and after obtaining the frequency band interaction features, reshape again to restore the sub-band dimension. Figure 7 As shown in the figure, before the subband is input to the frequency band interaction layer, the initial dimensions of the subband are (b, c*N, t, f), where b represents the number of samples input to the frequency band interaction layer, also known as the batch size (batch_size), t represents the number of time steps, c represents the number of channels before the subband is divided, and N represents the number of subbands after the subband is divided. Q, K, and V are the calculation parameters of the attention mechanism, representing the query, key, and value, respectively. The corresponding weight dimensions are all (c*N, c*N). Therefore, the attention network needs to be applied to the last dimension, that is, the channel dimension. Therefore, the initial dimensions (b, c*N, t, f) need to be reshaped to (b, t, f, c*N). However, it should be understood that the reshaping here only needs to ensure that the last dimension is the channel dimension. There is no limitation on how the previous dimensions are set. For example, it can also be reshaped to (b, f, t, c*N).
[0104] like Figure 7 As shown, Q = input * W Q +b Q , K=input*W K +b K , V=input*W V +b VInput represents the data input to the attention layer. The output data can satisfy the following formula.
[0105] Among them, output represents the output data, T is the transposition operator, and softmax is the regression function.
[0106] Figure 7 The self-attention mechanism is used in the attention mechanism.
[0107] like Figure 7 As shown in the figure, after obtaining the output data of the attention layer, it is reshaped again to restore it to the same dimension as the initial one.
[0108] It should be noted that in the embodiment of the present application, the first reshape is to make the dimension of the selected basic network act on the target dimension required by the network layer, and the second reshape is to restore the dimension to better connect with the subsequent network structure. Figure 7 The basic network shown in the figure is reshaped so that the network can act on the channel dimension, but in order to better connect with the subsequent network, it is reshaped again after this layer to restore the dimension.
[0109] Figure 8 This is a schematic diagram of another frequency band interaction layer processing process of an embodiment of the present application. Figure 8 Taking the construction of the frequency band interaction layer based on FC as an example, the construction based on the FC layer is similar to the construction based on 1x1 convolution. However, since the last dimension of the input sub-band dimension is the frequency dimension, and the frequency band interaction features need to be extracted on the channel dimension, this method also requires reshaping the sub-band dimension before extracting the frequency band interaction features, and after obtaining the frequency band interaction features, reshape again to restore the sub-band dimension. Figure 8 As shown, before the sub-band is input to the frequency band interaction layer, the initial dimension of the sub-band is (b, c*N, t, f). For related explanations, refer to Figure 7 Therefore, we need to reshape the initial dimensions (b, c*N, t, f) to (b, t, f, c*N). However, it should be understood that the reshape only needs to ensure that the last dimension is the channel dimension. There is no restriction on how to set the first few dimensions. For example, we can also reshape to (b, f, t, c*N).
[0110] like Figure 8As shown in the figure, the reshaped sub-band features after processing in the channel dimension by the FC layer can satisfy the formula output = relu (input * W FC +b FC ), where output represents the output data, input represents the data input to the FC layer, and W FC is a two-dimensional matrix with dimensions (c*N, c*N), b FC The dimension is c*N.
[0111] like Figure 8 As shown in the figure, after obtaining the data output by the FC layer, it is reshaped again to restore it to the same dimension as the initial one.
[0112] Figures 6 to 8 Examples of three types of frequency band interaction layers are given respectively. Constructing the frequency band interaction layer based on 1x1 convolution or FC layer has the advantages of low computational complexity and fast operation. Constructing the frequency band interaction layer based on the attention mechanism requires relatively large computational complexity, but the learned frequency band interaction features will also be more complete. Although the computational complexity is relatively large, since only the frequency band interaction layer has a large computational complexity, for the entire sound source separation model, the overall computational complexity is still much smaller than the traditional solution where all the computations are based on one-dimensional convolutional networks.
[0113] Figure 9 This is a schematic diagram of the processing process of a frequency interaction layer in an embodiment of the present application. The frequency interaction layer needs to learn the global interaction information of the frequency points, that is, the global dependency information, so it can be constructed using the FC layer. Figure 9 This is an example of using the FC layer to construct a frequency point interaction layer. The FC layer is applied to each time step of the sub-band to calculate the output result. In the output result, each frequency point refers to the information of other frequency points in the current time step through the FC layer, thereby learning the global interaction information. Figure 9 Any sub-band (sub-band n) shown in the figure passes through the FC layer and uses the activation function relu(x) = max(0,x) to output the frequency point interaction result within the sub-band. The output frequency point interaction result can satisfy the formula below.
[0114] Where n∈[1,N] is used to represent the subband number, N is the total number of subbands, and F is the feature dimension, that is, the total dimension of the frequency feature. relu(x) is the activation function, w and b are the weight vector and bias vector of the FC layer respectively. ik Represents the element in the i-th row and j-th column of the weight matrix. The number of rows of w is only related to the input F, so w corresponds to i. represents the kth row and jth column element of the nth subband of the input, Represents the i-th row and j-th column element of the n-th channel output by the frequency band interaction layer.
[0115] Figure 10 This is a schematic diagram of another frequency interaction layer processing process of an embodiment of the present application. The frequency interaction layer can also be constructed using an attention network. Figure 10 This is an example of using the attention network to build a frequency interaction layer. Figure 10 For an introduction to the relevant parameters, please refer to Figure 7 Since the frequency interaction layer extracts features in the frequency dimension, Figure 10 The frequency interaction layer does not require reshape operation during processing. Figure 10 The weight dimensions of Q, K, and V in are (f, f), so attention will be performed on the last dimension of the input data (b, c*N, t, f), which is the frequency, so no reshape is required. Figure 10 As shown, the output data can satisfy the formula The relevant parameters are introduced above and will not be repeated here.
[0116] Figure 11 This is a schematic diagram of another frequency interaction layer processing process of an embodiment of the present application. The frequency interaction layer can also be constructed using a 1x1 convolutional layer. Figure 11 This is an example of using a 1x1 convolutional layer to construct a frequency interaction layer. This implementation is similar to the construction based on the FC layer. However, since the last dimension of the input sub-band dimension is the frequency dimension, and the 1x1 convolution will act on the second dimension when extracting the frequency interaction features, and since the second dimension of the initial dimension of the sub-band (b, c*N, t, f) is the channel dimension, not the frequency dimension, this method also requires reshaping the sub-band dimension before extracting the frequency interaction features, and after obtaining the frequency interaction features, reshape again to restore the sub-band dimension. Figure 11 As shown in the figure, before the subband is input to the frequency band interaction layer, the subband's initial dimensions (b, c*N, t, f) are reshaped to (b, f, t, c*N), and then a 1x1 convolution is performed. It should be understood that the reshaping here only needs to ensure that the second dimension is the frequency dimension f.
[0117] like Figure 11 As shown in the figure, the reshaped sub-band features after the 1x1 convolution layer processing output data can satisfy the formula output = relu (input * W conv +b conv ), where output represents the output data, input represents the data input to the 1x1 convolutional layer, and W convand b conv are the weight vector and bias vector respectively.
[0118] like Figure 11 As shown in the figure, after obtaining the data output by the 1x1 convolutional layer, it is reshaped again to restore it to the same dimension as the initial one.
[0119] Figures 9 to 11 Examples of three frequency interaction layers are given respectively. Constructing the frequency interaction layer based on 1x1 convolution or FC layer has the advantages of low computational complexity and fast operation. Constructing the frequency interaction layer based on the attention mechanism requires relatively large computational complexity, but the global frequency interaction features that can be learned will also be more complete. Although the computational complexity is relatively large, since only the frequency interaction layer has a large computational complexity, for the entire sound source separation model, the overall computational complexity is still much smaller than the traditional solution where all are constructed based on one-dimensional convolutional networks.
[0120] Figure 12 It is a schematic diagram of the execution process of an audio processing method according to an embodiment of the present application. Figure 12 yes Figure 4 An example of the execution process of the method shown. Figure 12 As shown, the input audio data is a mixture of audio signals from multiple sources. After Fourier transform, a 2048-point FFT spectrum data is obtained (an example of audio data to be processed). The 2048-point FFT spectrum data is divided into four subbands along the frequency dimension (an example of multiple subbands), subband 1 through subband 4, each with a 256-point FFT. The four subbands are then input into a U-Net structure (an example of a feature extraction network) constructed based on a two-dimensional convolutional base network. The output data (an example of a first audio feature set) is used to determine whether the target sound source is of a preset category (an example of a preset sound source category, such as a full-band sound source). If the target sound source is of the preset category, the U-Net is further processed to determine whether it contains a 1x1 convolutional layer. Otherwise, the output data is directly input into a channel selection module (for example, a 3x3 convolutional layer for two channels). When judging whether the U-Net contains a 1x1 convolution layer, if the judgment result is "yes", it means that the U-Net has the ability to extract the characteristics of the frequency band interaction layer, so the frequency band interaction layer is skipped ( Figure 12 In the example above, the frequency band interaction layer is constructed using a 1x1 convolutional layer, and the data is further input into the frequency point interaction layer ( Figure 12In the example above, the frequency interaction layer is constructed using the FC layer. Conversely, if the judgment result is "no", it means that U-Net does not have the ability to extract the frequency band interaction layer features. Therefore, the data is input into the frequency band interaction layer to extract the frequency band interaction features, and the obtained data (an example of the second audio feature set) is further input into the frequency interaction layer. After the frequency interaction layer extracts the global features of the frequency interaction, the obtained data (an example of the third audio feature set) is input into the channel selection module.
[0121] The channel selection module separates the audio signal corresponding to the target sound source from the third audio feature set, that is, outputs ratio data of the audio signal corresponding to the target sound source in the mixed audio. Figure 12 For example, if we want to separate the audio signals of two target sound sources, vocals and drums, the channel selection module corresponds to two channels. It should be understood that this means selecting two channels from N channels, each of which outputs one audio signal, while these two channels each output an audio signal of the target sound source. Therefore, by adjusting the proportion of each channel through the channel selection module, the audio signal separation can be achieved. For example, the proportion of the channel of the non-target sound source can be set to 0.
[0122] It should be understood that Figure 12 This is just a specific example of the audio processing method provided in this application. There is no limitation on the relevant data. For example, it can be 1024-point FFT instead of 2048-point FFT; for example, the number of divided sub-bands can be 3, 5 or other numbers instead of 4; for example, the number of channels selected can be 1 or 3 or more; for example, the convolution kernel size of the channel selection module can be other sizes such as 2x2 or 4x4 instead of 3x3. Other cases are not listed one by one.
[0123] The determination of whether the target sound source belongs to a preset category can be eliminated, requiring all audio signals from all sound sources to pass through both the frequency band interaction layer and the frequency point interaction layer. The determination of whether the feature extraction network contains 1x1 convolutions can even be eliminated. Since the feature extraction network can directly determine whether it contains a 1x1 convolution layer after it is determined, the decision to include a frequency band interaction layer can be made when constructing the initial sound source separation model. The feature extraction network can also utilize neural network structures other than the U-Net structure that can be used to extract audio features.
[0124] use Figure 12 The audio signals of human voice and drum sound extracted by the method shown can be played directly or processed later, for example, Figure 1 Spatial audio processing shown in (c).
[0125] Figure 13The following is a comparison chart of the results of separating drum sounds from the same mixed audio using different schemes. Figure 13 As shown in (a), the spectrum of the standard drum sound is a full-band audio with clear lines. The so-called full-band is reflected in the spectrum here. Figure 13 The lines along the vertical axis (frequency dimension) shown in (a) are continuous.
[0126] When the traditional solution of using a U-Net structure based on a two-dimensional convolutional network to extract audio features from the mixed audio and then separate the drum audio signal, since the two-dimensional convolution can only extract local features, the spectrum may only extract small line segments in the frequency dimension (that is, the vertical dimension). This will inevitably result in these small line segments not being able to be spliced into a coherent and complete line segment, resulting in many holes in the spectrum of the separated drum audio signal, such as Figure 13 The vertical lines shown in (b) are jumbled and discontinuous, e.g. Figure 13 In the S1 and S2 regions shown in (b), there are discontinuous vertical lines, indicating holes. These holes can cause the separated drum sound to be distorted (no longer sounding like a drum), making it possible for the human ear to perceive these holes as a distorted drum sound.
[0127] When the present invention is used, the feature extraction network constructed by the two-dimensional convolutional network is first used to extract local features, that is, to extract small line segments on the frequency dimension, and then the frequency band interaction layer and the frequency point interaction layer are used to extract global features, so that these small line segments can be accurately spliced into coherent and complete line segments, which can effectively reduce the spectrum holes and make the spectrum appear Figure 13 The effect of clear and continuous vertical lines shown in (c) is as follows, Figure 13 In the S1 and S2 regions shown in (c), the longitudinal lines are continuous and there are no holes, which is consistent with the Figure 13 The same area shown in (b) is in sharp contrast. Therefore, the drum sound is separated without distortion. Figure 13 It can also be seen from (c) that the frequency band interaction features and frequency point interaction features of this application are both intended to supplement the correlation between the extracted local features, and can be understood as the connection relationship between the above-mentioned small line segments.
[0128] It should be noted that the above-mentioned hypothetical two-dimensional convolutional network extracts a series of small vertical line segments, and the frequency band interaction layer and the frequency point interaction layer are used to extract the correlation between these small line segments in order to more vividly illustrate the effect of the present application scheme. However, it does not mean that the two-dimensional convolutional network extracts necessarily small vertical line segments. This depends on the spectral characteristics of the sound source itself and there is no limitation. It is only necessary to understand that the present application scheme first uses the two-dimensional convolutional network to extract local features, and then uses the frequency band interaction layer and the frequency point interaction layer to supplement the extraction of global features, so that the features of the entire frequency band are fully extracted.
[0129] It should also be understood that if the feature extraction network based on the one-dimensional convolutional network in the traditional solution is used, the spectrum effect similar to that of the present application solution can be achieved, that is, the separated drum audio signal can also show the same Figure 13 The effect is similar to that shown in (c). The reason is that one-dimensional convolution is a full-band operation, that is, each operation will cover the full frequency dimension. For ease of understanding, you can imagine that each operation is to extract a series of Figure 13 Multiple operations are repeated to extract the full-length (full-band) vertical line segments, which inevitably requires a very large amount of calculation. However, the present application solution retains the advantage of low computational complexity of two-dimensional convolution while making up for its shortcomings, thus achieving high separation accuracy with a small amount of computation.
[0130] Figure 14 Schematic diagram of a non-full-band sound source according to an embodiment of the present application. Figure 14 Mainly take human voice as an example, such as Figure 14 The spectrum of the human voice shown in the figure shows that the so-called harmonic characteristics are reflected in the spectrum as horizontal waves. The vertical axis direction (frequency dimension) does not have a continuous vertical line, and the horizontal axis direction (time dimension or data dimension) depends on whether the human voice has interruptions. Figure 14 In the S3 area, it can be seen that there is a vertical void. Figure 13 and Figure 14 As can be seen, a full-band sound source can be understood, to some extent, as a sound source with coherent lines in the frequency dimension, while a non-full-band sound source can be understood, to some extent, as a sound source without coherent lines in the frequency dimension. Therefore, for the latter, even if global features are not extracted, the coherence of the frequency-dimensional segments will not be affected, because non-full-band sound sources themselves have no coherent lines in the frequency dimension.
[0131] When the present invention is used to separate human voices, the feature extraction network is sufficient to extract Figure 14The local features in the frequency dimension shown, coupled with its non-full-band characteristics, do not require the use of the frequency band interaction layer and the frequency point interaction layer to further extract global features. Even if the solution of this application is adopted and the frequency band interaction features and the global frequency point interaction features are extracted after the feature extraction network, the features extracted by the frequency band interaction layer and the frequency point interaction layer will be very few, so its harmonic features will still be retained and will not result in the consequence of filling its own holes.
[0132] Figure 15 This is a schematic flowchart of another audio processing method according to an embodiment of the present application. Figure 15 yes Figure 4 An example of the method shown is to first divide the frequency bands and then use the feature extraction network to perform feature extraction, and add the judgment of whether to use the frequency band interaction layer and the frequency point interaction layer for feature extraction, so that for non-full-band sound sources, there is no need to go through the processing of the above two interaction layers, further reducing the amount of calculation and shortening the processing delay.
[0133] S1501: Obtain audio data to be processed.
[0134] This step is an example of step S401.
[0135] S1502: Extract audio features using a feature extraction network.
[0136] This step is an example of extracting features using the feature extraction network in the sound source separation model in step S403 to obtain the first audio feature set.
[0137] S1503. Obtain and determine whether the category of the target sound source belongs to the preset sound source category. When the determination result is "yes", execute step S1504; when the determination result is "no", execute step S1505.
[0138] In one implementation, a preset sound source category set may be created, which lists all-band sound source categories. Only target sound sources belonging to the preset sound source category set will be input into the frequency band interaction layer and the frequency point interaction layer.
[0139] In another implementation, a target sound source is created as a category feature vector, which is then fed into a discriminant network. The discriminant network's output is then evaluated. If the output is less than or equal to a preset threshold, the target sound source is considered to belong to the preset category. Otherwise, the target sound source does not belong to the preset category. The preset threshold can be, for example, 0.5, but it should be understood that other values, such as 0.6, may also be used. This is not a limitation.
[0140] S1504 : Extract frequency band interaction features and frequency point interaction features using the frequency band interaction layer and the frequency point interaction layer respectively.
[0141] S1505: Separate the audio signal of the target sound source through channel selection.
[0142] To facilitate understanding, the following description is given with reference to specific examples.
[0143] Assuming the target sound source is a human voice, which is harmonically non-full-band audio, then when executing step S1503, if the preset sound source category is a set of full-band sound source categories, the human voice is determined not to belong to the preset sound source category, and the process proceeds to step S1505 instead of executing step S1504. In this case, if the determination of whether the sound source belongs to the preset category is based on comparing the discriminant network with a preset threshold, then when executing step S1503, a feature vector for the sound source category of human voice is constructed, input into the discriminant network, and an output value is obtained. This output value is then compared with the preset threshold. In this case, since the corresponding output value of the human voice is greater than the preset threshold, it is determined not to belong to the preset sound source category. Based on this determination result, the process proceeds to step S1505 instead of executing step S1504.
[0144] Assume the target sound source is a drum sound, which is full-band audio. If, during step S1503, the preset sound source category is a set of full-band sound source categories, the drum sound is determined to belong to the preset sound source category, and the process proceeds to steps S1504 and S1505. In this case, if the determination of belonging to the preset sound source category is based on comparing the discriminant network with a preset threshold, then during step S1503, a feature vector for the drum sound category is constructed. This constructed feature vector is input into the discriminant network to obtain an output value, which is then compared with the preset threshold. Since the output value corresponding to the drum sound is less than the preset threshold, the drum sound is determined to belong to the preset sound source category. Based on this determination, the process proceeds to steps S1504 and S1505.
[0145] Figure 16 This is a schematic flow chart of a method for training a sound source separation model according to an embodiment of the present application. Figure 16 The steps shown are introduced.
[0146] S1601. Obtain training data.
[0147] The training data includes audio data to be trained and an audio signal label of at least one known sound source corresponding to the audio data to be trained. The audio data to be trained is a mixed audio of multiple sound sources synthesized by synthesizing the audio signal of at least one known sound source and other audio signals.
[0148] It should be understood that the training phase is supervised learning, which requires knowing the real data so that the predicted data output by the model can be compared with the real data, and the model weight parameters can be adjusted in reverse based on this, so that the predicted data output by the model becomes closer and closer to the real data.
[0149] Each audio data to be trained and the corresponding audio signal label of at least one known sound source can constitute a set of training data. In the present application scheme, each training is to compare the audio signal output by the model after processing the audio data to be trained (i.e., predicted data) with the audio signal label of at least one known sound source (real data), and inversely update the model weight parameters, so that the audio signal output by the model is closer and closer to the audio signal label of at least one known sound source.
[0150] S1602. Input the audio data to be trained into the initial sound source separation model, and update the weight parameters of the initial sound source separation model based on the difference between the predicted audio signal of at least one known sound source output by the initial sound source separation model and the audio signal label of at least one known sound source, thereby obtaining a trained sound source separation model.
[0151] In one implementation, the initial sound source separation model includes a feature extraction network, a frequency band interaction layer and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network, and is used to extract features of the audio data to be trained to obtain a first audio feature set; the frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be trained from the first audio feature set, thereby obtaining a second audio feature set; the frequency point interaction layer is used to extract frequency point interaction features within each sub-band in the second audio feature set, thereby obtaining a third audio feature set; the sound source separation model is also used to separate a predicted audio signal of at least one known sound source from the third audio feature set.
[0152] The relevant contents of the multiple sub-frequency bands can be referred to above and will not be elaborated on again.
[0153] It should be understood that the description of the sound source separation model structure in the inference stage above can be referred to the training stage, and for the sake of brevity, it will not be repeated.
[0154] In one implementation, the multiple sub-bands are obtained by dividing the training audio data along the frequency dimension, or the multiple sub-bands are obtained by dividing the first extraction result output by the first network layer along the frequency dimension, where the first network layer is any layer of the feature extraction network.
[0155] In one example, the initial sound source separation model includes a feature extraction network and a frequency interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network, and the feature extraction network includes a 1x1 convolution layer for extracting features from multiple sub-bands to obtain a first audio feature set; the frequency interaction layer is used to extract frequency interaction features within each sub-band in the first audio feature set, thereby obtaining a third audio feature set; the initial sound source separation model is also used to separate a predicted audio signal of at least one known sound source from the third audio feature set. From this example, it can be seen that for a feature extraction network with a 1x1 convolution layer, it can be decided to remove the frequency band interaction layer during the training stage. Therefore, when such a sound source separation model is applied to the audio processing method of this application (that is, the inference stage) after training, there is no need to judge whether it needs to be processed through the frequency band interaction layer.
[0156] In one implementation, the frequency band interaction layer and / or the frequency point interaction layer are constructed using a two-dimensional convolutional basic network, an attention mechanism network, or a fully connected layer.
[0157] In one example, the frequency band interaction layer is constructed using a 1x1 convolutional layer; and the frequency point interaction layer is constructed using an FC layer. Figure 12 An example of this is given.
[0158] The initial sound source separation model can be a completely untrained model, or it can be a model that has been trained (that is, a feature extraction network that already has the ability to extract local features in audio data) with frequency band interaction layers and frequency point interaction layers added to it. The latter can reduce the number of training rounds to a certain extent.
[0159] Figure 16 The method shown in the figure provides a new sound source separation model, which not only uses the feature extraction network with relatively small computational complexity to extract audio features, but also uses the frequency band interaction layer and the frequency point interaction layer to supplement the extraction of frequency band interaction features and frequency point interaction features respectively. By training the new sound source separation model, it has good sound source separation capabilities. The trained sound source separation model can be applied to Figure 4 The audio processing method shown.
[0160] The above mainly introduces the method of the embodiment of the present application in conjunction with the accompanying drawings. It should be understood that although the various steps in the flowcharts involved in the various embodiments described above are shown in sequence, these steps are not necessarily performed in sequence in the order shown in the figures. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the steps or stages in other steps. The device of the embodiment of the present application is introduced below in conjunction with the accompanying drawings.
[0161] The present application also provides a training device, which includes an acquisition unit and a training unit. The device can be integrated into an electronic device such as a server or a cloud server that can train a neural network model, that is, it can be integrated into a training device. The device can be used to execute any of the above-mentioned training methods for the sound source separation model. For example, the acquisition unit can be used to execute steps S1601-S1602, and the training unit can be used to execute step S1603. In one implementation, the device may further include a storage unit for storing relevant data. The storage unit may be integrated into any of the above-mentioned units, or it may be a unit independent of all the above-mentioned units.
[0162] The present application also provides a sound source separation device, which includes an acquisition unit and a processing unit. The device can be integrated into an electronic device such as a mobile phone, a laptop computer, a tablet computer, a car computer, a smart wearable device, etc. that can use a neural network model to process audio signals, that is, it can be integrated into an inference device. The device can be used to execute any of the above audio processing methods. For example, the acquisition unit can be used to execute step S401, and the processing unit can be used to execute step S402. The device can also be used to execute Figure 5-Figure 12 and Figure 15 In one implementation, the device may further include a storage unit for storing relevant data. The storage unit may be integrated into any one of the above units or may be a unit independent of all the above units.
[0163] Figure 17 This is a schematic diagram of the software architecture of an electronic device according to an embodiment of the present application. Figure 17 The execution process of the present application solution can be immediately introduced from the perspective of the internal execution process of the electronic device. Figure 17As shown, the application layer includes various applications, and here takes applications such as camera, video, browser and music as an example, but it should be understood that in practice only some of these applications may be included or other applications may also be included, and there is no limitation.
[0164] As long as it is an application that can play audio data, the present application solution can be used. For example, it can be used to listen to music online in a browser, or to create songs using a music application, etc., and these will not be listed one by one.
[0165] The framework layer includes the audio stream service module and the audio stream management module. The audio stream service module is mainly used to store and forward audio streams and realize cross-application transmission, while the audio stream management module is mainly used to manage related drivers such as speakers, set some management parameters, or make related settings for process scheduling.
[0166] It should be understood that Figure 17 The software architecture mainly shows the application layer and the framework layer, but in practice, the electronic device may also include other software layers, such as a native layer, a hardware abstraction layer (HAL layer), etc., without limitation.
[0167] The hardware layer includes a processor, a sound source separation model deployed on the hardware layer, and speakers.
[0168] Suppose the user selects the spatial audio effect in the music app and starts playing a song, then the music app will convert the audio stream of the song ( Figure 17 The audio stream #1 in the framework layer is sent to the audio stream service module of the framework layer, and is sent to the processor of the hardware layer via the audio stream service module. After receiving the audio stream, the processor executes the relevant steps of the present application scheme, including converting the time domain audio stream #1 into frequency domain data. Assuming that the obtained data is data #2 (an example of audio data to be processed), it is input into the sound source separation model, and is processed by the sound source separation model to obtain data #3, which is the separated audio signal. After obtaining data #3, the processor processes the spatial audio to obtain spatial audio data (data #4), and then sends data #4 to the audio stream service module, which is transmitted to the music application via the audio stream service module. The music application plays the song online under spatial audio effects. The music application can also call the speakers of the hardware layer through the audio stream management module and the audio stream service module to complete the playback.
[0169] It should be understood that the above is only an example of a processing flow for the purpose of understanding the solution. In actual applications, other scenarios may exist. For example, the processor may first perform frequency segmentation before inputting the frequency domain data of multiple sub-bands (that is, data #2 is the segmented data) into the sound source separation model. Other scenarios are not listed one by one.
[0170] The processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network process unit (NPU), other general-purpose processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. Different processing units may be independent devices or integrated into one or more processors.
[0171] The sound source separation model can be stored in memory, Figure 17 Not shown, the memory can be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device; or it can be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Optionally, the memory can include both an internal storage unit of the electronic device and an external storage device. The memory can be used to store an operating system, application programs, boot loaders, data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0172] The processor may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor 1 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.
[0173] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0175] An embodiment of the present application also provides an electronic device, which includes: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions so that the electronic device can execute the steps in any of the above methods.
[0176] An embodiment of the present application further provides a chip system, which is applied to an electronic device and includes one or more processors configured to invoke computer instructions to enable the electronic device to perform the steps of any of the above methods. Optionally, the chip system also includes a memory electrically connected to the processor. Optionally, the chip system may also include a communication interface.
[0177] The embodiment of the present application also provides a computer-readable storage medium, which stores instructions, and when the instructions are executed by an electronic device, any of the above methods can be implemented. The computer-readable medium may include at least: any entity or device capable of carrying computer program code (instructions) to a photographing device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard drive, a magnetic disk, or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0178] The present application also provides a computer program product, which includes a computer program that, when executed by an electronic device, can implement any of the above methods. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form.
[0179] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0180] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0181] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0182] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0183] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0184] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0185] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0186] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0187] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. An audio processing method, applied to electronic equipment, characterized in that: include: Acquiring audio data to be processed, wherein the audio data to be processed includes audio signals from multiple sound sources; Processing the audio data to be processed using a sound source separation model to obtain an audio signal of at least one target sound source; the at least one target sound source is at least one of the multiple sound sources; The sound source separation model includes a feature extraction network, a frequency band interaction layer, and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network and is used to extract features from the audio data to be processed to obtain a first audio feature set; The frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be processed from the first audio feature set, thereby obtaining a second audio feature set; The frequency interaction layer is used to extract global frequency interaction features in each sub-band in the second audio feature set, thereby obtaining a third audio feature set; The sound source separation model is further used to separate the audio signal of the at least one target sound source from the third audio feature set.
2. The method according to claim 1, characterized in that The multiple sub-bands are obtained by dividing the audio data to be processed along the frequency dimension, or the multiple sub-bands are obtained by dividing the first extraction result output by the first network layer along the frequency dimension, and the first network layer is any layer of the feature extraction network.
3. The method according to claim 1, characterized in that When the feature extraction network has the ability to extract frequency band interaction features between the multiple sub-bands so that the first audio feature set includes the frequency band interaction features between the multiple sub-bands, the sound source separation model does not include the frequency band interaction layer; and / or when the feature extraction network has the ability to extract global frequency point interaction features for each frequency point corresponding to the audio data to be processed so that the first audio feature set includes the global frequency point interaction features of each frequency point, the sound source separation model does not include the frequency point interaction layer.
4. The method according to claim 1, wherein The method further comprises: The target dimension of the frequency band interaction layer is matched with the effective dimension of the basic network for constructing the frequency band interaction layer by a reshape operation; and / or the target dimension of the frequency point interaction layer is matched with the effective dimension of the basic network for constructing the frequency point interaction layer by a reshape operation.
5. The method according to any one of claims 1 to 4, characterized in that The step of processing the audio data to be processed by using a sound source separation model to obtain an audio signal of at least one target sound source includes: In a case where the first target sound source belongs to a preset sound source category, the frequency band interaction layer is used to extract frequency band interaction features between the multiple sub-bands from the first audio feature set to obtain the second audio feature set, the frequency point interaction layer is used to extract global frequency point interaction features within each sub-band in the second audio feature set to obtain the third audio feature set, and the sound source separation model is used to separate the audio signal of the first target sound source from the third audio feature set; the first target sound source is any one of the at least one target sound source; or When the first target sound source does not belong to the preset sound source category, the sound source separation model is used to separate the audio signal of the first target sound source from the first audio feature set.
6. The method according to claim 5, characterized in that The preset sound source category is used to represent a full-band sound source category; or, an output value of a discriminant network corresponding to a sound source in the preset sound source category is less than or equal to a preset threshold.
7. The method according to any one of claims 1 to 4, characterized in that The frequency band interaction layer and / or the frequency point interaction layer are constructed using a two-dimensional convolutional basic network, an attention mechanism attention network or a fully connected FC layer.
8. A method for training a sound source separation model, characterized in that: include: Acquiring training data, the training data including audio data to be trained and an audio signal label of at least one known sound source corresponding to the audio data to be trained, wherein the audio data to be trained is a mixed audio of multiple sound sources obtained by synthesizing the audio signal of the at least one known sound source and other audio signals; Inputting the audio data to be trained into an initial sound source separation model, and updating the weight parameters of the initial sound source separation model based on the difference between the predicted audio signal of the at least one known sound source output by the initial sound source separation model and the audio signal label of the at least one known sound source, thereby obtaining a trained sound source separation model; the initial sound source separation model includes a feature extraction network, a frequency band interaction layer, and a frequency point interaction layer; the feature extraction network is constructed using a two-dimensional convolutional basic network and is used to extract features from the audio data to be trained to obtain a first audio feature set; The frequency band interaction layer is used to extract frequency band interaction features between multiple sub-bands corresponding to the audio data to be trained from the first audio feature set, thereby obtaining a second audio feature set; The frequency interaction layer is used to extract frequency interaction features in each sub-band in the second audio feature set, thereby obtaining a third audio feature set; The sound source separation model is further used to separate the predicted audio signal of the at least one known sound source from the third audio feature set.
9. The method according to claim 8, characterized in that The frequency band interaction layer and / or the frequency point interaction layer are constructed using a two-dimensional convolutional basic network, an attention mechanism attention network or a fully connected FC layer.
10. The method according to claim 8 or 9, characterized in that The multiple sub-bands are obtained by dividing the audio data to be trained along the frequency dimension, or the multiple sub-bands are obtained by dividing the first extraction result output by the first network layer along the frequency dimension, and the first network layer is any layer of the feature extraction network.
11. An electronic device, characterized in that: The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, wherein the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method as described in any one of claims 1 to 7, or 8 to 10.
12. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 7, or 8 to 10.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes instructions, and when the instructions are executed on an electronic device, the electronic device performs the method according to any one of claims 1 to 7 or 8 to 10.
Citation Information
Patent Citations
Audio sound source separation method based on improved self-attention mechanism and band-cross characteristics
CN111261186A
Single-channel voice separation method based on deep learning
CN111292762A
Electronic music classification method and system based on multi-sound-source separation
CN111488486A
Sound source separation method based on double attention mechanism and multi-stage hybrid convolutional network
CN113241092A
False voice detection method and device, electronic equipment and storage medium
CN114596879A