Method and system for instrument separation and reproduction for a mixed audio source
Patent Information
- Application Number
- DE602022017205
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-06
- Filing Date
- 2022-07-14
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2042-07-14
AI Technical Summary
Existing audio systems fail to enhance the sense of depth in sound field reproduction when playing music with multiple instruments, as they primarily support stereo or mono signal transmission, lacking multi-channel and low-latency audio capabilities.
A method and system using an instrument separation model to process a mixture audio source into separate instrument spectrograms, which are then modulated into multi-channel broadcast signals and reproduced by multiple speakers for synchronized playback.
Restores the original sound field effect by separating and reproducing instrument audio sources through multiple channels, enhancing sound quality, bandwidth efficiency, and data throughput with low latency.
Description
Technical Field
[0001] The present disclosure generally relates to audio source separation and playing. More particularly, the present disclosure relates to a method and a system for instrument separating and transmission for a mixture music audio source as well as reproducing same separately on multiple speakers.Background Art
[0002] In scenarios where better audio effects are required, multi-speaker playing can usually be used to enhance the live listening experience. Many speakers now support audio broadcasting. For example, several of JBL's portable speakers have an audio broadcasting function called Connect+, which can also be referred to as a 'Party Boost' function. Wireless connection to hundreds of Connect+-enabled speakers allows the multiple speakers to play the same signal synchronously, which may magnify the users' listening experience to an epic level and perfectly achieve stunning party effects.
[0003] However, existing speakers can only support stereo signal transmission at most during broadcasting, or even master devices can only broadcast mono signals to other slave devices, which helps to significantly increase the sound pressure level, but makes no contribution to the enhancement of the sense of depth of the sound field. For example, when music played by multiple instruments is played through speakers, the melody part is mainly reproduced, so the users' listening experience is more focused on the horizontal flow of the music, and it is difficult to identify the timbre between different instruments. On the other hand, based on the audio transmission characteristics of the existing speakers, the audio codec and single-channel transmission mechanisms thereof cannot meet the multi-channel and low-latency audio transmission requirements.
[0004] EP 3 608 903 A1 relates to a U-net architecture for neural network to perform instrument separation from an audio recording. The separation algorithm is trained to identify masks of instruments which may then be isolated and rendered by a plurality of speakers.
[0005] US 2015 / 063574 A1 relates to a multi-channel audio transformer for separating a plurality of sound source objects from an input multichannel mix.
[0006] WO 2016 / 140847 A1 defines a system for individually transmitting a plurality of audio stems to different reproduction devices and synchronously playing these stems.
[0007] Therefore, there is currently a need for a practical method to reproduce the timbre of different channels of an audio source by means of multiple speakers with better sound quality, higher bandwidth efficiency, and higher data throughput.Summary of the Invention
[0008] The present disclosure provides a method according to appended claim 1, for instrument separating and reproducing for a mixture audio source, comprising: obtaining a mixture audio source spectrogram based on the mixture audio source, wherein the mixture audio source comprises sound of a plurality of instruments;using an instrument separation model to sequentially obtain an instrument feature mask of each of the plurality of instruments from the mixture audio source; obtaining an instrument spectrogram of each of the plurality of instruments based on the respective instrument feature mask of each of the plurality of instruments; determining an instrument audio source of each of the plurality of instruments based on the respective instrument spectrogram; modulating the respective instrument audio sources of the plurality of instruments into a corresponding multi-channel broadcast audio signal; broadcasting the multi-channel broadcast audio signal to a plurality of speakers;demodulating the corresponding instrument audio sources of the plurality of instruments by the plurality of speakers; and synchronously reproducing each respective instrument audio source of the plurality of instruments accordingly by a respective one of the plurality of speakers.
[0009] The present disclosure also provides a non-transitory computer-readable medium according to appended claim 9, including instructions that, when executed by a processor, implement the method for instrument separating and reproducing for a mixture audio source.
[0010] The present disclosure also provides a system according to appended claim 11, for instrument separating and reproducing for a mixture audio source, comprising: a spectrogram conversion module configured to obtain a mixture audio source spectrogram based on the mixture audio source, wherein the mixture audio source comprises sound of a plurality of instruments; an instrument separation module comprising an instrument separation model, wherein the instrument separation model is configured to sequentially obtain an instrument feature mask of each of the plurality of instruments from the mixture audio source; an instrument extraction module configured to obtain an instrument spectrogram of each of the plurality of instruments based on the respective instrument feature mask of each of the plurality of instruments; and an instrument audio source rebuilding module configured to determine an instrument audio source of each of the plurality of instruments based on the respective instrument spectrogram, wherein the respective instrument audio sources of the plurality of instruments are modulated into a corresponding multi-channel broadcast audio signal, the multi-channel broadcast audio signal is broadcast to a plurality of speakers, the corresponding instrument audio sources of the plurality of instruments are demodulated by the plurality of speakers; and each respective instrument audio source of the plurality of instruments is synchronously reproduced accordingly by a respective one of the plurality of speakers. Brief Description of the Drawings
[0011] These and / or other features, aspects and advantages of the present invention will be better understood after reading the following detailed description with reference to the accompanying drawings, throughout which the same characters represent the same members, where: FIG. 1 shows an exemplary flow chart of a method for separating instruments from a mixture music audio source and reproducing same separately on multiple speakers according to one or more embodiments of the present disclosure; FIG. 2 shows a schematic diagram of a structure of an instrument separation model according to one or more embodiments of the present disclosure; FIG. 3 shows a schematic diagram of a structure of an upgraded instrument separation model according to one or more embodiments of the present disclosure; FIG. 4 shows a block diagram of a system for instrument separating and reproducing for a mixture audio source according to one or more embodiments of the present disclosure; and FIG. 5 shows a schematic diagram of disposing multiple speakers at designated positions according to one or more embodiments of the present disclosure. Detailed Description
[0012] The detailed description of the embodiment of the invention is as follows. However, it should be understood that the disclosed embodiments are merely exemplary, and may be embodied in various alternative forms. The drawings are not necessarily depicted on scale; and some features may be expanded or minimized to show details of specific components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as restrictive, but only as a representative basis for teaching those skilled in the art to variously employ the present disclosure.
[0013] Wireless connection allows multiple speakers to be connected to each other. For example, music audio streams can be played simultaneously through these speakers to obtain a stereo effect. However, the mechanism of playing mixture music audio streams simultaneously through the multiple speakers may not meet the multi-channel and low-latency audio transmission requirements; and it only increases the sound pressure level, but makes no contribution to the enhancement of the sense of depth of the sound field.
[0014] With the increasing demand for listening to music played via multiple instruments, users may wish to achieve better sound quality, higher bandwidth efficiency, and higher data throughput, as achieved by, for example, multi-channel sound systems, even with portable devices, while adopting a low-latency and reliable synchronous connection of multiple speakers to restore the original sound field effect during music recording, which can be achieved by, for example, treating the multiple speakers as a multi-channel system accordingly, and then reproducing the audio sources of various instruments restored in different channels by means of the different speakers.
[0015] Therefore, the present disclosure provides the method to reproduce the original sound field effect during music recording by first processing selected music through the instrument separation model to obtain the separate audio source of each instrument after separation, and then feeding the broadcast audio through multiple channels to different speakers for playing.
[0016] FIG. 1 shows an exemplary flow chart 100 of a method for separating instruments and reproducing music on multiple speakers in accordance with the present disclosure. Due to the different characteristics of the vibration of different objects, the basic three elements of sound (i.e., tone, volume and timbre) are related to the frequency, amplitude, and spectral structure of sound waves, respectively. A piece of music can express the magnitude of amplitude at a certain frequency at a certain point in time by means of a music audio spectrogram, and waveform data of sound propagating in a medium is represented by a two-dimensional image, which is a spectrogram. Differences in the distribution of energy between different instruments can be reflected in the radiating capacity of the sound produced by that instrument at different frequencies. The spectrogram is a two-dimensional graph represented by the time dimension and the frequency dimension, and the spectrogram can be divided into multiple pixels by, for example, taking the time unit as the abscissa and the frequency unit as the ordinate; and the different shades of colors of all the pixels can reflect the different amplitudes at corresponding time-frequencies. For example, bright colors denote higher amplitudes, and dark colors denote lower amplitudes.
[0017] Therefore, referring to the flow chart of the method for separating and reproducing the instruments shown in FIG. 1, firstly, in S102, a selected mixture music audio source is converted into a mixture music spectrogram. A mixture spectrogram image of a selected piece of music is formed by using the following method: x t = overlap input , 50 % x n t = windowing x t X n f = FFT X n t X nb f = X 1 f , X 2 f , … , X n f including: x (t): inputting a time domain of a mixture audio signal of the selected music; X (f): performing fast Fourier transform to achieve frequency domain representation of the mixture audio signal; X n (f): inputting a spectrogram of the signal from a time frame n; overlap(*) and windowing(*) are overlapping and windowing processing respectively, where an overlap coefficient is based on an experimental value, for example, adopting 50% of the experimental value; FFT means the fast Fourier transform; and | * | is an absolute value operator, which is equivalent to taking an amplitude value of sound waves. Therefore, the buffer X nb (f) of X n (f) represents a spectrogram of the mixture audio of the music x (t) to be input into an instrument separation model.
[0018] Next, in S104, an amplitude image of the spectrogram of the mixture audio is input into the instrument separation model to extract audio features of all the instruments separately.
[0019] The present disclosure provides the instrument separation model that enables the separation of different musical elements from selected original mixture music audio by machine learning. For example, spectrogram amplitude feature masks of different instrument audios are separated out from a mixture music audio by machine learning combined with instrument identification and masking. Although the present disclosure refers to the separation of music played by multiple instruments, it does not preclude the inclusion of the vocal portion of the mixture audio as equivalent to one instrument.
[0020] The instrument separation model provided by the present disclosure for separating instruments from a music audio source is shown in FIG. 2. The instrument separation model can be used for, for example, building an instrument sound source separation model generated based on a convolutional neural network. There are various network models of the convolutional neural network. In processing of images, the convolutional neural network can extract better features in the images due to its special organizational structure. Therefore, by processing the music audio spectrogram based on the instrument sound source separation model of the convolutional neural network provided by the present disclosure, the features of all kinds of instruments can be extracted, so that one and multiple instruments are separated out from the music audio played by mixed instruments, and subsequent separate reproduction is further facilitated.
[0021] The instrument sound source separation model of the present disclosure shown in FIG. 2 is divided into two parts, namely, a convolutional layer part and a deconvolutional layer part, where the convolutional layer part includes at least one two-dimensional (2D) convolutional layer, and the deconvolutional layer part includes at least one two-dimensional (2D) deconvolutional layer. The convolutional layers and the deconvolutional layers are used to extract features of images, and pooling layers (not shown) can also be disposed among the convolutional layers for sampling the features so as to reduce training parameters, and can reduce the overfitting degree of the network model at the same time. In the exemplary embodiment of the instrument sound source separation model of the present disclosure, there are six 2D convolutional layers (denoted as convolutional layer_0 to convolutional layer_5) available at the convolutional layer part, and there are correspondingly six 2D convolutional transposed layers (denoted as convolutional transposed layer_0 to convolutional transposed layer_5) available at the deconvolutional layer part. The first 2D convolutional transposed layer at the deconvolutional layer part is cascaded behind the last 2D convolutional layer at the convolutional layer part.
[0022] At the deconvolutional layer part, the result of each 2D convolutional transposition is further processed by a concatenate function and stitched with the feature result extracted from the corresponding previous 2D convolution at the convolutional layer part before entering the next 2D convolutional transposition. As shown, the result of the first 2D convolutional transposition_0 at the deconvolutional layer part is stitched with the result of the fifth 2D convolution_4 at the convolutional layer part, the result of the second 2D convolutional transposition_1 at the deconvolutional layer part is stitched with the result of the fourth 2D convolution_3 at the convolutional layer part, the result of the third 2D convolutional transposition_2 is stitched with the result of the third 2D convolution_2, the result of the fourth 2D convolutional transposition_3 is stitched with the result of the second 2D convolution_1, and the result of the fifth 2D convolutional transposition_4 is stitched with the result of the first 2D convolution_0.
[0023] Batch normalization layers are added between every two adjacent 2D convolutional layers at the convolutional layer part and every two adjacent 2D convolutional transposed layers at the deconvolutional layer part to renormalize the result of each layer, so as to provide good data for passing the next layer of neural network. In addition, a leaky rectified linear unit (Leaky_Relu) is further added between every two adjacent 2D convolutional layers, including Leaky_Relu function processing, and the function is expressed as f (x) = max (kx, 0). A rectified linear unit of Relu function processing is further added between every two adjacent 2D convolutional transposed layers, and the function is expressed as f(x) = max (0, x). Both of the two rectified linear units act to prevent gradient disappearance in the instrument separation model. In the exemplary embodiment of FIG. 2, three discard layers are also added for Dropout function processing, thus preventing overfitting of the instrument separation model. Then, after the last 2D convolutional transposition_5, the 1-2 layers are fully-connected layers, the fully-connected layers are responsible for connecting the extracted audio features and thus enabling same to be output from an output layer at the end of the model. In the exemplary embodiment of instrument separation model constructed in FIG. 1, the mixture music audio spectrogram amplitude graph is input into an input layer, and the spectrogram graph features of all instruments are extracted by the processing of the deep convolutional neural network in the model; and a softmax function classifier can be disposed at the output end as the output layer, and its function is to normalize the real number output into multiple types of probabilities, so that the audio spectrogram masks of the instruments can be extracted from the output layer of the instrument separation model.
[0024] For a newly established machine learning model, it is first necessary to use some databases as training data sets to train the model so as to adjust the parameters in the model. After the instrument separation model shown in FIG. 2 is built, an audio played by multiple instruments and having already contained respective sound track records of all the instruments can be selected, for example, from a database as the training data set to train the instrument separation model. In this case, some training data can be found from publicly available public music databases, such as the publicly available music database 'Musdb18' which contains more than 150 full-length pieces of music in different genres (lasting for about 10 hours), the separately recorded vocals, pianos, drums, bass, and the like that are corresponding to these pieces of music, as well as the audio sources of other sounds contained in the music. In addition, music such as vocals, pianos, and guitars with multi-sound track separately recorded in some other specialized databases can also be used as the training data sets.
[0025] When training the model, a set of training data sets are selected and sent to the neural network, and the model parameters are adjusted according to the difference between an actual output of the network and an expected output. That is to say, in this exemplary embodiment, music can be selected from a known music database, the mixture audio of this music can be converted into a mixture audio spectrogram image and then put into the input, all instrument audios of the music are respectively converted into characteristic spectrogram images of the instruments, and the obtained images are placed in the output of the instrument separation model as the expected output. By adopting the machine learning to try and try again, the instrument separation model can be trained, and the model features can be modified. For the instrument separation model based on a 2D convolutional neural network, the model features of the machine learning during the model training process can mainly include the weight and bias of a convolution kernel, the parameters of a batch normalization matrix, etc.
[0026] The training time of the model is usually based on offline processing, so it can be aimed at the model that provides the best performance regardless of computational resources. All the instruments included in the selected music in the training data set can be trained one by one to obtain the feature of each of the instruments, or the expected output of the multiple instruments can be placed in the output of the model to obtain the respective features thereof at the same time, so the trained instrument separation model has fixed model features and parameters. For example, the spectrogram of a mixture music audio of music selected from the music database 'Musdb 18' can be input into the input layer of the instrument separation model, and the spectrograms of vocal tracks, piano tracks, drum tracks and bass tracks of the music included in the database can be placed in the output layer of the instrument separation model, so that the vocal feature model parameters, piano feature model parameters, drum feature model parameters and bass feature model parameters of the model can be trained at the same time.
[0027] By using the trained instrument separation model to process a new music audio spectrogram amplitude input, an instrument feature mask of each of all the instruments can be obtained accordingly, that is, the probability that the spectrogram thereof accounts for the amplitude of the original mixture music audio spectrogram. The trained model should be expected to achieve more real-time processing capacity and better performance.
[0028] After being trained, the instrument separation model established in FIG. 2 can be loaded into a smart device (such as a smartphone, or other mobile devices, and audio play equipment) of a user to achieve the separation of music sources.
[0029] Returning to the flow chart shown in FIG. 1, in S104, the feature mask of a certain instrument can be extracted by inputting the mixture audio spectrogram of the selected music into the instrument separation model; and the feature mask of the certain instrument can mark the probability thereof in all pixels of the spectrogram, which is equivalent to a ratio of the amplitude of the certain instrument's voice to that of the original mixture music, so the feature mask of the certain instrument can be a real number ranging from 0 to 1, and the audio of the certain instrument can be distinguished from the mixture audio source accordingly. Then, in S106, the feature mask of the certain instrument is reapplied to the spectrogram of the original mixture music audio, so as to obtain the pixels thereof that are more prominent than the others and further stitch same into a feature spectrogram of the certain instrument; and the spectrogram of the certain instrument is subjected to inverse fast Fourier transform (iFFT), so that an individual sound signal of the certain instrument can be separated out, and an individual audio source thereof is thus obtained.
[0030] The above process can be described as: inputting an amplitude image X nb (f) of the mixture audio spectrogram of the selected piece of music x (t) into the instrument separation model for processing to obtain the feature masks X nbp (f) of the instruments, the type of instruments depending on instrument feature model parameters currently set in the instrument separation model of this input. For example, if trained piano feature model parameters are currently set in the instrument separation model, the output obtained by processing the input mixture audio spectrogram is a piano feature mask; and then, the piano feature model parameters are replaced with, for example, bass feature model parameters, and the mixture audio spectrogram is input again, so that the obtained output is a bass feature mask. Thus, different instrument feature masks can be replaced in turn; and each time the mixture audio spectrogram of the music is input, the respective feature masks of all the instruments can be obtained successively. The sounds in the music audio that cannot be separated out by the instrument separation model can be included in an extra sound feature output channel.
[0031] In addition, the original mixture audio source processed with the instrument separation model can be a mono audio source, a dual-channel audio source, or even a multi-channel stereo mixture audio source. In the exemplary embodiment shown in FIG. 2, the two spectrograms input into the input layer of the instrument separation model respectively represent spectrogram images of the left channel audio and right channel audio of a dual-channel mixture music stereo audio. For the processing of the instrument separation model, on the one hand, the audios of left and right channels can be processed separately, so that an instrument feature mask of the left channel and an instrument feature mask of the right channel are obtained respectively. On the other hand, alternatively, the instrument feature masks can be extracted after the audios of the left and right channels are mixed together.
[0032] Next, referring to the flow chart in FIG. 1, in S106, the obtained instrument feature mask X nbp (f) is reapplied to the mixture audio spectrogram of the music of the original input model, for example, firstly, smoothing is carried out to prevent distortion, the instrument feature masks predicted by the instrument separation model are multiplied with the mixture audio spectrogram of the original input music, and the spectrogram of the sound of the each of the instruments is then obtained by outputting. The smoothing can be expressed as: Y nb = X nb f * 1 − a f + X nbp f * a f where smoothing coefficient a (f) = sigmoid (instrument feature mask) * (perceptual frequency weighting).
[0033] The sigmoid function is defined as S x = 1 1 + e − x , where one of the parameters, the instrument feature mask, is the output of the instrument separation model, and the other parameter, the perceptual frequency weighting, is determined based on experimental values. Finally, the spectrograms of the instruments are transformed back to the time domain by using the iFFT and an overlap-add method, so that the reconstructed audio sources of the instrument sounds are obtained, as shown below: Y nbc f = Y nb f * e i * phase X nb f y b t = iFFT Y nbc f y n t = windowing y b t y t = overlap _ add y n t , 50 % where iFFT represents an inverse fast Fourier transform, and overlap_add(*) represents an overlap-add function.
[0034] Alternatively, the extraction of the spectrogram images from mixture music time domain signals x(t), and the reapplication of the instrument feature masks which are processed and output by the instrument separation model to the original input mixture music spectrogram for obtaining the spectrogram of the individual sound of the each instrument, the implementation of reconstruction for obtaining the audio sources y(t) of the sounds of the instruments, and the like, that are involved in the above instrument separation process, can also be regarded as newly added neural network layers in addition to the instrument separation model, so that the instrument separation model provided above can be upgraded. The upgraded instrument separation model can be described as including a 2D convolutional neural network-based instrument separation model and the above-mentioned newly added layers, as shown in FIG. 3. Therefore, the music signal processing features included in this upgraded instrument separation model, such as window shapes, frequency resolutions, time buffering and overlap percentages, can be modified by machine learning. After the upgraded instrument separation model is transformed into a real-time executable model, as long as the selected music is directly input into the upgraded instrument separation model, multiple maximized separate instrument audio sources, which are separately reconstituted from the mixture music audio source, of all the instruments can be output.
[0035] After being obtained, the multiple separate instrument audio sources are respectively fed to multiple speakers by means of signals through different channels, each channel including the sound of a type of instrument, and then all the instrument audio sources are played synchronously, which can reproduce or recreate an immersive sound field listening experience for users.
[0036] For example, after a piece of music to be played on a smart device of a user is input into the instrument separation model, and the separate audio sources of all the instruments are reconstructed, multiple speakers can be connected to the smart device of the user by a wireless technology, and the audio sources of all the instruments are played at the same time through different channels, so that the user who plays the music with the multiple speakers at the same time may get a listening experience with a better depth effect.
[0037] In an exemplary embodiment, for a portable Bluetooth speaker that is often used in conjunction with a smart device of a user, it is different from a mono stereo audio stream transmission mode of connecting a master speaker to the smart device of the user by means of, for example, classical Bluetooth, and then broadcasting to multiple other slave speakers by using the master speaker in a way of mono signals, the present disclosure adopts, for example, a Bluetooth low energy (BLE) audio technology, which enables multiple speakers (groups) to be regarded as a multi-channel system, so that the smart device of the user can be connected to the multiple speakers synchronously with low latency and reliable synchronization; and after being separated, the sounds of all instruments are transmitted to the speaker group that enables a broadcast audio function by means of multiple channel signals, then the different speakers receive the broadcast audio signals broadcasted by the smart device through multiple channels, audio sources of the different channels are modulated and demodulated, and all the instruments are synchronously reproduced, so that the sound field with an immersive listening effect is reproduced or restored.
[0038] FIG. 4 shows a block diagram of a system 400 for instrument separating and reproducing for a mixture audio source according to one or more embodiments of the present disclosure. In an exemplary embodiment of the present disclosure, the system for instrument separating and reproducing for a mixture audio source is positioned on a smart device of a user, and includes a mixture source conversion module 402, an instrument separation module 404, an instrument extraction module 406 and an instrument source rebuild module 408. When the system 400 is in use, firstly, a mixture music audio source is obtained from, for example, a memory (not shown) of the smart device, and is then converted into a mixture audio source spectrogram after being subjected to overlapping and windowing, fast Fourier transform, etc. in the mixture source conversion module 402. The mixture audio source spectrogram is then sent to the instrument separation module 404 including an instrument separation model, and the instrument feature masks of all instruments in the mixture audio source are sequentially obtained after feature extraction is performed on the mixture audio source spectrogram by means of the instrument separation model, and the feature masks of all the instruments are output into the instrument extraction module 406. The instrument feature masks are reapplied to the mixture audio source spectrogram in the instrument extraction module 406, which may include, for example, smoothing and then multiplying the instrument feature masks with the original mixture audio source spectrogram, so that the respective spectrograms of all the instrument sources are obtained. Finally, in the instrument source rebuild module 408, the respective spectrograms of all the instruments are processed by, for example, iFFT, overlapping, windowing, and the like so as to be converted into audio sources thereof, respectively. In the exemplary embodiment shown in FIG. 4, the instrument audio sources of all the instruments determined by the instrument source rebuild module 408 on the smart device may support the modulation of multiple audio streams corresponding to the multiple instruments onto multiple channels by a BLE connection, and are broadcast to multiple speakers (groups) by using a broadcast audio function in a form of multi-channel signals. It is understandable that, instrument sources or sounds that cannot be separated by the instrument separation module can also be modulated to one or more channels and sent to the corresponding speakers (groups) for playing. As shown in FIG. 4, the multiple speakers (such as the speaker 1, the speaker 2, the speaker 3, the speaker 4, ...... and the speaker N) that enable the broadcast audio function respectively receive broadcast audio signals (the signal X 1 , the signal X 2 , the signal X 3 , the signal X 4 , ......, and the signal X N ), and audio streams of the all the instruments are demodulated accordingly.
[0039] Due to the low power consumption and large transmission frequency of the BLE technology, the BLE technology can support wider bandwidth transmission to achieve faster synchronization; and a digital modulation technology or direct sequence spread spectrum is adopted, so that multi-channel audio broadcasting can be realized. In addition, the BLE technology can support transmission distances greater than 100 meters, so that the speakers can receive and synchronously reproduce audio sources within a larger range around the smart device of the user. Referring to S108 in the flow chart shown in FIG. 1 of the method, as the exemplary embodiment of the present disclosure, hundreds of speakers can be connected to the smart device of the user by BLE wireless connection, and the smart device broadcasts the respective reconstructed audio sources of all the instruments through multiple channels to all the speakers having the broadcast audio function. For example, separate audio sources of all instruments for playing mixed recorded symphony music can be separated out therefrom, and a sufficient number of speakers are used to reproduce the received and demodulated audio sources of all the instruments, which may amplify the user's listening experience to an epic level and further cause the user to achieve a perfect sound field shock effect.
[0040] In some cases, as shown in step S110 of FIG. 1, in order to reproduce or reconstruct the live performance of a band or achieve a magnificent sound field effect, the speakers playing different instrument audio sources may be placed at designated positions relative to listeners. Fig. 5 shows an exemplary embodiment of arranging speakers at the positions according to, for example, a layout required by a symphony orchestra for reproducing a symphony. The exemplary embodiment shows the reproduction of the different instruments for playing the symphonic work and even different parts thereof by using the multiple speakers, where the different instruments and all the parts of the reproduced music have first been separated out on the smart device of the user by means of an instrument separation model and modulated into multi-channel sound signals, and are then transmitted to the multiple speakers (groups) by audio broadcasting; and each or each group of speakers receive the audio broadcasting signals and demodulate same to obtain the audio source signals of all the instruments, thus being capable of respectively reproducing all the instruments and parts. For example, with a fixed separation order of all instruments known in the instrument separation model, a separate audio sources of each instrument can be transmitted correspondingly to the speaker at the designated position.
[0041] In this case, as mentioned previously in the present disclosure, when the original mixture music is divided into, for example, left channel audio sources and right channel audio sources and then input to the instrument separation model, the audio sources, which are reconstructed after the separation of the instrument separation model, of all the instruments are respectively modulated to different channels of the broadcast audio signals, each channel at this point may being, for example, but not limited to mono or binaural. The speakers receive the signals and demodulate same to obtain the audio source signals of the instruments. For example, the left channel audio sources and the right channel audio sources may be distinguished in the same speaker, or for example, the audio sources from a plurality of channels of the same instrument may be assigned to a plurality of speakers for playing.
[0042] In addition, as shown in FIG. 5, in one case, a first violin and a second violin are included in, for example, the symphony orchestra, they may be separated out as the same type of instruments from the mixture music audio source input into the instrument separation model, but audio sources of the same type of instruments can be broadcast, for example, with two or more speakers. Alternatively, in the other case of sounds played by string parts such as a viola and a cello, as well as chords played by the same type of instruments or different parts played by a plurality of the same type of instruments, these instruments or parts can also be assigned to multiple speakers, because the instrument separation model can distinguish different frequency components; although the separation of sounds made by the same type of instruments may not be as effective as that of sounds made by completely different types of instruments, but still does not affect the performance of the feeding to the one or more speakers for playing.
[0043] In accordance with the above description, those skilled in the art can understand that the above embodiments can be implemented in a way of being applied to a hardware platform by means of software. Accordingly, any combination of one or more computer-readable media can be used to perform the method provided by the present disclosure. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or equipment, or any suitable combination of the foregoing. More specific exemplary embodiments (non-exhaustive list) of the computer-readable storage media would, for example, include: electrical connections with one or more wires, portable computer floppy disks, hard disks, random access memory (RAM), read-read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In accordance with the context of the present disclosure, the computer-readable storage medium may be any tangible medium that may include or store programs used by or in combination with an instruction execution system, apparatus, or equipment.
[0044] The elements or steps referenced in a singular form and modified with the word 'a / an' or 'one' as used in the present disclosure shall be understood not to exclude being plural, unless such an exception is specifically stated. Further, the reference to the 'embodiments' or 'exemplary embodiments' of the present disclosure is not intended to be construed as exclusive, but also includes the existence of other embodiments of the enumerated features. The terms 'first', 'second', 'third', and the like are used only as identification, and are not intended to emphasize the number requirement or positioning order for their objects.
Claims
1. A method for instrument separating and reproducing for a mixture audio source, comprising: obtaining a mixture audio source spectrogram based on the mixture audio source, wherein the mixture audio source comprises sound of a plurality of instruments; using an instrument separation model to sequentially obtain an instrument feature mask of each of the plurality of instruments from the mixture audio source; obtaining an instrument spectrogram of each of the plurality of instruments based on the respective instrument feature mask of each of the plurality of instruments; determining an instrument audio source of each of the plurality of instruments based on the respective instrument spectrogram; modulating the respective instrument audio sources of the plurality of instruments into a corresponding multi-channel broadcast audio signal; broadcasting the multi-channel broadcast audio signal to a plurality of speakers; demodulating the corresponding instrument audio sources of the plurality of instruments by the plurality of speakers; and characterised by synchronously reproducing each respective instrument audio source of the plurality of instruments accordingly by a respective one of the plurality of speakers.
2. The method of claim 1, wherein the instrument separation model is based on a 2D convolutional neural network comprising multiple 2D convolutional layers and multiple 2D convolutional transposed layers for extracting the instrument feature masks of the plurality of instruments.
3. The method of claim 1 or 2, wherein the instrument separation model is pre-trained with a known training data set comprising mixture audios and their corresponding instrument separation audios of at least one instrument included.
4. The method of any preceding claim, wherein the mixture audio source is a stereo audio source comprising at least one channel, and the instrument separation model processes each of the at least one channel of the stereo audio source, separately.
5. The method of any preceding claim, wherein obtaining the instrument spectrogram of each of the plurality of instruments comprises multiplying the obtained instrument feature masks of the plurality of instruments with the mixture audio source spectrogram, separately.
6. The method of any preceding claim, wherein the multi-channel broadcast audio signal comprises the instrument audio source of a corresponding one of the plurality of instruments and / or the multi-channel broadcast audio signal is a stereo audio signal.
7. The method of any preceding claim, further comprising respectively disposing the plurality of speakers to designated positions, and reproducing the instrument audio sources, demodulated by each of the plurality of speakers, of the corresponding ones of the plurality of instruments, respectively.
8. The method of claim 7, wherein respectively disposing the plurality of speakers to designated positions comprises arranging the positions of each of the plurality of speakers according to a symphony orchestra layout.
9. A non-transitory computer-readable medium including instructions that, when executed by a processor, perform the following steps including: obtaining a mixture audio source spectrogram based on a mixture audio source, wherein the mixture audio source comprises sound of a plurality of instruments; using an instrument separation model to sequentially obtain an instrument feature mask of each of the plurality of instruments from the mixture audio source; obtaining an instrument spectrogram of each of the plurality of instruments based on the respective instrument feature mask of each of the plurality of instruments; determining an instrument audio source of each of the plurality of instruments based on the respective instrument spectrogram; modulating the respective instrument audio sources of the plurality of instruments into a corresponding multi-channel broadcast audio signal; broadcasting the multi-channel broadcast audio signal to a plurality of speakers; demodulating the corresponding instrument audio sources of the plurality of instruments by the plurality of speakers; and characterised by synchronously reproducing each respective instrument audio source of the plurality of instruments accordingly by a respective one of the plurality of speakers.
10. The non-transitory computer-readable medium of claim 9, wherein the instructions when executed by the processor perform the steps of a method as mentioned in any of claims 2 to 8.
11. A system for instrument separating and reproducing for a mixture audio source, comprising: a spectrogram conversion module configured to obtain a mixture audio source spectrogram based on the mixture audio source, wherein the mixture audio source comprises sound of a plurality of instruments; an instrument separation module comprising an instrument separation model, wherein the instrument separation model is configured to sequentially obtain an instrument feature mask of each of the plurality of instruments from the mixture audio source; an instrument extraction module configured to obtain an instrument spectrogram of each of the plurality of instruments based on the respective instrument feature mask of each of the plurality of instruments; and an instrument audio source rebuilding module configured to determine an instrument audio source of each of the plurality of instruments based on the respective instrument spectrogram, wherein the respective instrument audio sources of the plurality of instruments are modulated into a corresponding multi-channel broadcast audio signal, the multi-channel broadcast audio signal is broadcast to a plurality of speakers, the corresponding instrument audio sources of the plurality of instruments are demodulated by the plurality of speakers; and characterised in that each respective instrument audio source of the plurality of instruments is synchronously reproduced accordingly by a respective one of the plurality of speakers.
12. The system of claim 11, wherein the instrument separation model is based on a 2D convolutional neural network comprising multiple 2D convolutional layers and multiple 2D convolutional transposed layers for extracting the instrument feature masks of the plurality of instruments.
13. The system of claim 11 or 12, wherein the instrument separation model is pre-trained with a known training data set comprising mixture audios and their corresponding instrument separation audios of at least one instrument included.
14. The system of any of claims 11 to 13, wherein the mixture audio source is a stereo audio source comprising at least one channel, and the instrument separation model processes each of the at least one channel of the stereo audio source, separately.
15. The system of any of claims 11 to 14, wherein obtaining the instrument spectrogram of each of the plurality of instruments comprises multiplying the obtained instrument feature masks of the at least one instrument with the mixture audio source spectrogram, separately.
16. The system of claim 11, wherein the multi-channel broadcast audio signal comprises the instrument audio source of a corresponding one of the plurality of instruments and / or the multi-channel broadcast audio signal is a stereo audio signal.
17. The system of one of claims 11 to 16, further configured for disposing the plurality of speakers to designated positions, and reproducing the instrument audio sources, demodulated by each of the plurality of speakers, of the corresponding ones of the plurality of instruments, respectively.
18. The system of claim 17, wherein respectively disposing the plurality of speakers to designated positions comprises arranging the positions of the plurality of speakers according to a symphony orchestra layout.