Audio source separation and audio mixing processing
The method and system enhance audio quality in user-generated content by extracting and processing audio mix attributes to separate and process audio signals without prior knowledge, addressing the limitations of existing tools and improving audio quality for both professionals and non-professionals.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2024-04-26
- Publication Date
- 2026-05-21
AI Technical Summary
Existing audio processing tools require prior knowledge about the sound sources in an audio mix to improve quality, limiting their applicability to user-generated content of inferior quality.
A method and system for audio processing that extracts audio mix information, determines processing parameters based on semantic and signal attributes, and applies these parameters to separate and enhance audio signals without prior knowledge, using a combination of analysis and editing modules.
Enables automated and efficient enhancement of audio quality in user-generated content by separating and processing audio signals based on mix attributes, suitable for both professional and non-professional users.
Smart Images

Figure 2026516324000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of priority from Spanish Patent Application No. P202330336, filed on April 28, 2023, US Provisional Patent Application No. 63 / 512,218, filed on July 6, 2023, and European Patent Application No. 23183758.4, filed on July 6, 2023, each of which is hereby incorporated by reference in its entirety.
[0002] Technical field to which the invention pertains The present invention relates to separating an audio mix into audio signals representing separate audio sources and audio processing based on such source separation. The source separation may be general - purpose, i.e., it may not require prior knowledge about the sound sources in the audio mix.
Background Art
[0003] The Internet and social networks have made the sharing and consumption of user - generated media content (audio and video) a very popular form of entertainment and education. As a result, user - generated audio content is abundant and widespread. However, such user - generated audio content is often of inferior quality compared to professional commercial content, and tools that can improve audio quality are desirable.
[0004] The separation of different audio sources in an audio mix is recognized as an important element in audio quality improvement processes. In the past, source separation has been applied using previously acquired knowledge about the mix as a way to reproduce the intended audio mix. Such an approach is disclosed in Patent Document 1.
Patent Document Ⅰ
[0005] However, there is a need for more general audio processing tools that can improve the quality of any audio without requiring any prior knowledge. [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] The objective of the present invention is to provide efficient audio processing that can improve the perceived audio quality of an audio mix. [Means for solving the problem]
[0007] According to a first aspect of the present invention, this and other objectives are achieved by a method for processing an input audio mix, the method comprising: extracting at least two source audio signals from an input audio mix, each audio source signal representing a separate audio source; extracting audio mix information from the input audio mix, the audio mix information including at least one of audio mix semantic attributes and audio mix signal attributes; determining audio processing parameters based on the audio mix information; and processing the source audio signals based on the audio processing parameters to produce a processed audio mix.
[0008] This method can be embodied as a computer program executed by a computer processor.
[0009] According to a second aspect of the present invention, this and other objectives are achieved by a system for processing an input audio mix, the system comprising: a source separation module configured to extract at least two source audio signals from an input audio mix, each audio source signal representing a separate audio source; an analysis module configured to receive an input audio mix, extract audio mix information including at least one of audio mix semantic attributes and audio mix signal attributes, and determine audio processing parameters based on the audio mix information; and an editing module configured to receive the extracted audio signals and audio processing parameters, and process the source audio signals based on the audio processing parameters to produce a processed audio mix.
[0010] According to these aspects, the input audio mix is analyzed and audio mix information is extracted. This information is used to generate processing parameters that guide the processing in the editing module. Therefore, the processing of the sound source can be based on the attributes of the input audio mix before sound source separation, thereby enabling more automated sound source processing.
[0011] The semantic attributes of an audio mix may include at least one of the following: music genre, recording type, production style, identified sound sources in the mix, and characteristics of the identified sound sources. For example, the recording type or production style may influence automatically set processing parameters (e.g., target gain).
[0012] Audio mix signal attributes can include at least one of the following: loudness, dynamic range, average spectral power, and spectral power distribution. Such audio mix signal attributes may also affect automatically set processing parameters (e.g., target gain).
[0013] The analysis module may be further configured to receive a source audio signal, in which case the determined audio processing parameters may also be based on the source signal characteristics of the source audio signal. The source signal characteristics may include at least one of loudness, dynamic range, average spectral power, and spectral power distribution.
[0014] The combination of pre-separation audio mix information and post-separation signal attributes allows for highly generalized and automated determination of processing parameters. For example, the relative loudness of the sound source may be determined based on the audio mix information and compared to the target gain. The processing parameters can then indicate a set of gains to be applied to the source audio signal to obtain the desired mix.
[0015] In some implementations, the analysis module is further configured to determine source separation parameters based on audio mix information and / or source signal characteristics, and the source separation module is configured to receive the source separation parameters and extract the source audio signal based on the source separation parameters.
[0016] Information about the input audio mix, i.e., information before signal separation, can provide insights into appropriate methods for source separation. For example, audio mix information may indicate the presence of a specific type of sound source, such as speech, which can indicate the use of a speech separator in a source separation module.
[0017] The audio processing parameter sent to the editing module may represent a linear gain, and the editing module may be configured to mix the extracted audio signals by applying the linear gain to the source audio signal. This is a simple and straightforward type of editing that can benefit from the implementation of the present invention. The gain may vary over time.
[0018] In some implementations, the editing module is further configured to process each extracted audio signal individually before mixing by applying at least one of the following: dynamic range compression, equalization, dynamic equalization, or creative effects (e.g., panning).
[0019] Other relevant audio processing includes audio editing, volume rebalancing in mixes, audio zoom (emphasizing a sound source in a specific direction), or muting selected sounds, and these operations can be performed on the device or in the cloud. Such audio editing tools, when automated, can make audio editing easier for non-professional clients, yet they are powerful in the sense that they perform operations not available in typical professional workflows, and therefore they can also be used by professionals in less automated ways.
[0020] In some implementations, the editing module includes a user interface configured to receive user input for controlling audio processing. This allows for user interaction with the editing process ranging from fully manual control to system-assisted manual interaction. For example, the user interface may be configured to offer the user alternatives to a proposed processing and to receive user input related to said alternatives. Specifically, a selected sound source may be identified, and a slider may allow the user to increase or decrease (or completely mute) the loudness of this sound source.
[0021] A further aspect of the present invention relates to a method for general-purpose sound source separation (sometimes called general-purpose sound separation). This aspect may be advantageously combined with the method of the first aspect, or implemented in the sound source separation module of the second aspect. However, this further aspect is also a distinct inventive concept that offers distinct technical advantages independently of the first and second aspects.
[0022] The method according to this further aspect includes receiving an input audio mix, converting the input audio mix into a frequency domain audio mix spectrogram, binarizing the spectrogram by comparing each tile with a predetermined threshold to form a selection mask, applying the selection mask to the audio mix spectrogram to form a spectrogram selection, and inverse-transforming the spectrogram selection into the time domain to provide an output audio signal associated with one or more sound sources in the input audio mix. It is a method for general-purpose sound source separation.
[0023] Since this method does not rely on a deep learning model and only applies threshold analysis of the spectrogram, it is not computationally demanding and can be executed by a relatively lightweight device such as a smartphone.
[0024] When such a method is applied to sound source separation of the first or second aspect, the output audio signal may be used as the first sound source, and the remaining audio mix (residual audio) may be used as the second sound source.
Brief Description of the Drawings
[0025] The present invention will be described in more detail with reference to the accompanying drawings showing the current preferred embodiments of the present invention.
[0026] [Figure 1] It is a block diagram of a system according to an embodiment of the present invention.
[0027] [Figure 2] It shows remixing the separated sources based on a set of gains.
[0028] [Figure 3] It shows the principle of general-purpose sound source separation (sound separation).
[0029] [Figure 4]This is a flowchart of a process for general-purpose sound source separation (sound separation) according to one embodiment of the present invention.
[0030] [Figure 5A] This displays a spectrogram representing the input audio mix.
[0031] [Figure 5B] This shows a spectrogram representing the same audio mix after pre-emphasis filtering.
[0032] [Figure 5C] Figure 5B shows a spectrogram representing the selection of the spectrogram.
[0033] [Figure 5D] The spectrogram showing the residuals of the spectrogram in Figure 5C is shown.
[0034] [Figure 6A] Figure 5A shows the frequency spectrum representing the time slice of the spectrogram. [Figure 6B] Figure 5B shows the frequency spectrum representing the time slice of the spectrogram. [Modes for carrying out the invention]
[0035] The systems and methods disclosed herein may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, task division does not necessarily correspond to division into physical units; conversely, a single physical component may have multiple functions, and a single task may be performed collaboratively by several physical components.
[0036] Computer hardware may include, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify actions to be performed by such computer hardware. Furthermore, this disclosure relates to any set of computer hardware that execute instructions individually or collectively in order to perform any one or more of the concepts described herein.
[0037] Some or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions. When these instructions are executed by one or more of the processors, they perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) specifying an action to be taken is included. Thus, an example is a typical processing system (i.e., computer hardware) containing one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem, including a hard drive, an SSD, RAM, and / or ROM. A bus subsystem may be included for communication between components. Software may reside in the memory subsystem and / or within the processors while it is being executed by the computer system.
[0038] The one or more processors may operate as standalone devices or may be connected to other processors, for example, to a network. Such a network may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0039] Software may be distributed on computer-readable media, which may include computer storage media (or non-temporary media) and communication media (or temporary media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, various forms of physical (non-temporary) storage media, such as EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other media that can be used to store desired information and that a computer can access. Furthermore, as is well known to those skilled in the art, communication media (temporary) typically include any information delivery medium that embodies computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transport mechanisms.
[0040] The system 10 in Figure 1 consists of three parts: a sound source separation module 11, an editing module 12, and an audio analysis module 13.
[0041] The sound source separation module 11 is configured to reconstruct any mixture of different source audio signals ("sound sources") 15 present in the input audio mix 14. For artificially created mixes, the sound source separation module functions to estimate the original audio stream used to create the mix (e.g., a multi-track recording), or at least the most important one (e.g., vocals, drums, dominant source, etc.). For direct recordings of mixed sound sources (e.g., recordings of real events with multiple sound sources), the sound source separation module functions to estimate the sound contribution from each distinct sound source (e.g., voice, street noise, wind, bird song, etc.). The sound sources obtained by the sound source separation module can be single (individual) sound sources (e.g., speech, dog, car, piano), or compound (group) sound sources (e.g., guitar, children playing), or "stems" (e.g., a backing track or background contextual noise).
[0042] The sound source separation module 11 may employ various sound source separation models 16 depending on the situation. A first type of model may be configured to efficiently remove crosstalk (i.e., "contamination") from other sound sources, possibly at the cost of some signal degradation. A second type of model may be configured to preserve superior signal quality, possibly at the cost of some crosstalk. Both types may be applied in parallel. For example, a first type of model can be used as a front-end for loudness analysis of a separate sound source (a process in which the absence of residual components is more important than signal quality), while a second type of model can be used to obtain a sound source to which further processing and remixing are applied (a task in which perceptual quality is of paramount importance).
[0043] In some implementations, the sound source separation module 11 involves a deep learning model pre-trained to extract specific sound sources from an audio mix. Deep learning models can be highly efficient for separating specific sound sources, such as speech. However, deep learning models are also useful for "universal" sound source separation (sound separation), i.e., separating unknown sound sources given any audio mix. Such models are sometimes called "source-agnostic."
[0044] In an alternative implementation, the sound source separation module 11 engages in analytical signal processing rather than deep learning. Such methods may not have very high computational demands and can therefore be run on portable processing devices such as smartphones.
[0045] The editing module 12 is configured to process the extracted (separated) sound sources 15 and recombine them into the processed mix 17. For example, in a live recording of a concert, the vocals may sound distant (almost inaudible). By separating the vocal sound source in the sound source separation model 11, the editing module 12 can emphasize the vocal sound source to make it louder and more audible.
[0046] One operation that can be performed on the sound source 15 is a linear gain G corresponding to rebalancing the mix. This is shown in Figure 2. The gain G may be time-variable to ensure a consistent mix if the balance between the sound sources changes in an undesirable way across the audio mix (for example, between songs). More advanced processing such as dynamic range compression, equalization, dynamic equalization, or creative effects (e.g., reverb, delay, chorus, flanger, filter, etc.) can be applied. Once the estimated sound sources have been processed, they are remixed to provide the processed audio 17. In some implementations, the final enhancement process may be applied to the processed mix 17, for example, by using (automatic) audio / music mastering.
[0047] The audio analysis module 13 is configured to receive an input audio mix 14 and extract audio mix information that includes at least one of the following: audio mix semantic attributes and audio mix signal attributes. Audio mix semantic attributes (properties) may include music genre, recording type, production style, identified sound sources in the mix, and characteristics of the identified sound sources. Audio mix signal attributes may include at least one of the following: loudness, dynamic range, average spectral power, and spectral power distribution.
[0048] The audio analysis module 13 is further configured to determine audio processing parameters 18 based on audio mix information. The audio processing parameters are provided to the editing module 12 to guide the audio processing. The processing parameters 18 may be determined by a combination of audio analysis, predefined rules, artificial intelligence, and user-selected rules.
[0049] In the illustrated implementation, the analysis module 13 also receives the separated source audio signal 15 and extracts its source signal characteristics, thereby basing the audio processing parameters 18 on these signal characteristics as well.
[0050] The audio processing parameters 18 control the editing module 13. For example, referring to Figure 2, the mixed gains (G1, G2, ... GN) can be automatically generated based on the processing parameters 18 from the analysis module 13.
[0051] As an example, the gain G of a specific sound source 15 can be expressed by the following function. • The type of sound source (for example, speech). • A specified (relative) target level (for example, 20 dB higher than the background). • The type of recording identified by the analysis module (e.g., music, speech, or ambient sound).
[0052] In one implementation, the analysis module 13 analyzes the loudness L of each sound source 15. i Calculate the target loudness L (in dB) for each sound source. target,i A gain G is defined for each sound source 15. i =L target,i -L i This is calculated.
[0053] The pre-set target gain levels may be defined based on the scenario. For example, in speech recording, the background may be set at least 9 dB quieter than the speech to ensure that the speech is audible. The pre-set target gain may be based on a specific instrument, either in relation to the mix or to other instruments. For example, to emphasize the drums, the target level for the drums might be +3 dB.
[0054] As described above, the processing parameters 18 may be time-varying. For example, the time-varying processing parameters 18 may define gain and processing that are applied only to the time fragment of the isolated sound source 15, for example, only when a particular sound source 15 is actually present (for example, the vocal track is amplified only when vocal activity is detected).
[0055] In some implementations, the editing module may also apply dynamic range control, equalization, or dynamic equalization to further enhance each or some of the extracted sound sources individually before remixing them. One approach involves a “dynamic EQ” as disclosed in Patent Document 2 for each or some of the sound sources 15, where a specific target profile is selected according to the type of sound source (e.g., bass, speech) and the type of recording (e.g., rock music, podcast, environment). [Patent Document 2] U.S. Patent No. 11,430,463
[0056] Profiles for individual isolated sound sources 15 (e.g., bass, speech) and their relative target levels can be obtained by analyzing the same type of professional reference content using the same analysis and sound source separation models used in our audio editing pipeline. These target values can be obtained from specific audio (suitable for making input content sound "like" the reference audio) or from a statistical analysis of a collection of audio (suitable for making input content sound "correct" according to general professional standards).
[0057] The processing parameter 18 can also control the editing module 14 to process the individual source signals 15 before remixing them. For example, if the audio analysis module 13 identifies one sound source as a dialogue, the audio processing parameter can cause the editing module 13 to apply speech-specific signal processing to this sound source.
[0058] For some applications requiring greater artistic freedom in manipulating the sound source 15, more complex audio effects (such as reverb and delay) can be applied. Other creative effects may include spatial manipulation of content, such as: - By upmixing all or some of the sound sources, which potentially have different widths, you can, for example, ensure that the vocals remain positioned in a narrow, forward soundstage while other instruments are rendered in a wider soundstage. • Panning of sound sources, for example, Dolby Atmos® panning. Here, each sound source may be positioned at a different location, rendered with a different size, and / or moved along a predefined path.
[0059] Panning may be performed to correct audio image formation problems, for example, by re-panning the vocals to the center while maintaining the original positions of other instruments. Such re-panning of a sound source can be done in various ways. This involves maintaining the extracted audio source in its original multi-channel format and changing the balance between channels. For example, vocals can be extracted in stereo from a stereo song and then re-panned by changing the relative gain between the left and right channels. This involves downmixing the extracted audio sources to mono or stereo, feeding them into a channel-based panner along with specific position parameters, and rendering them to the target multi-channel format. For example, a mono recording of a car idling can be extracted and panned to the rear Ls and Rs channels of a 5.1 surround sound system. This involves downmixing the extracted audio to mono or stereo, adding positional metadata, and transforming them into object-based content that should be authored as layout-independent content. For example, a stereo piano from an original stereo track can be downmixed to mono and placed in the center of the ceiling with a specified size. It is then rendered in that position by each playback device, to the best of its capabilities.
[0060] In some implementations, the analysis module is configured to further determine source separation parameters 19 based on extracted audio mix information and / or source signal attributes. The source separation attributes 19 are provided to the source separation module 11 to guide the source separation process.
[0061] The sound source separation parameter 19 can influence the selection of the sound source separation model used in the sound source separation module 11. For example, if the audio analysis module 13 determines that the input audio mix 14 contains substantially two sound sources (e.g., vocals and piano), the sound source separation module 11 can use a model specifically adapted to separate these two sound sources. Similarly, if the audio analysis module 13 finds that a speech sound source is present, the sound source separation module can use a speech-specific model.
[0062] In other words, the audio analysis module 13 may be configured to guide both the separation module 11 and the editing module 12 based on characteristics extracted from the input audio mix 14 and / or the separated sound source 15. The sound source separation process may be sequential and iterative, and the signal attributes of the separated sound source 15 may function to further adapt and improve the sound source separation model being used.
[0063] In some implementations, the editing module 12 includes a user interface 21 configured to expose some or all of the editing controls to the user. Full manual control may be suitable for experts, while automated or semi-automatic workflows may be more suitable for non-experts. In a semi-automatic workflow, the audio analysis module 13 can perform analysis, make several editing decisions, and communicate these decisions to the editing module 12 as processing parameters 18. The editing module 12 then presents suggestions to the user via the user interface 21, allowing for user fine-tuning. Below are some examples of automatically generated suggestions and optional fine-tuning. • If the analysis module 13 finds a dominant sound source, it suggests editing to make the dominant sound source louder and exposes sliders for the user to fine-tune the volume of each sound source. Dominant sound sources could be, for example, speech, noisy cars, dog barking, guitar, etc. • If analysis module 13 finds a dominant sound source, it suggests an edit that mutes the dominant sound source (and exposes sliders for the user to fine-tune the volume of each sound source). • If analysis module 13 finds that there is background music, it will suggest editing to reduce or remove the music in order to avoid copyright infringement. • Propose editing that modifies the spatial characteristics of the audio (for example, by expanding the width). • If analysis module 13 finds that there is speech, it suggests editing to transform the speech (for example, transforming the voice in an interesting way and remixing it with the original background). • Suggest an edit that completely changes the background sound. For example, change the background sound of a captured cafeteria to a more relaxing beach background sound, or add background music to your recording.
[0064] In a specific implementation, the analysis module 13 receives the input audio mix 14 and identifies 1) the dominant sound source, 2) the spectral energy, and 3) the dynamic range of the audio mix 14. The information on the dominant sound source is used to select which sound sources from the sound source separation module 11 are suitable for individual processing (for example, if speech is not identified, the system will not attempt to rebalance the speech). The information on the dominant sound source, spectral profile, and loudness is used to set global processing parameters (for example, to determine the desired dynamic range and spectral profile of the processed audio 17).
[0065] In some implementations, a first optimization, such as dynamic range compression and equalization, is performed on the input mix 14 to obtain a signal that is more closely suited to professional content and to correct any obvious defects that may impair subsequent sound source separation modules. Such optimization may be performed on the input audio mix 14 before it is provided to the sound source separation module 11 and the speech analysis module 13. Alternatively, it may be an integrated part of these modules, in which case slightly different optimizations may be applied to each module 11, 13.
[0066] As described above, the sound source separation module 11 can include various sound source separation models 16 that are appropriate for different situations.
[0067] In most Western-style popular music, several sound sources (e.g., vocals, bass, and drums) appear consistently throughout the song. Consequently, most music sound source separation models are sound source-specific, separating vocals, bass, drums, and "other sound sources." In speech sound source separation, the different speakers in the mix are typically unknown beforehand. Therefore, most speech sound source separation models are speaker-independent. Such models, on the other hand, are specifically configured to separate utterances.
[0068] For some applications, a "general-purpose" sound source separation (sound separation) model is desirable, that is, a sound source separation model that is not source-specific but can separate any sound source given any audio mix. Such a general-purpose sound separation model can separate mixes such as user-generated telephone recordings that include animal and traffic noise. This is shown in Figure 3, where an audio mix 31 including traffic noise, wind noise, dog and bird noise is separated into four distinct source audio signals 33a-d by a general-purpose sound source separation module 32.
[0069] Furthermore, general-purpose sound source separation (or sound separation) is similar to speech sound source separation, because neither requires prior knowledge of the sound sources within the mix. However, speech separation models are constrained to a specific domain (speech), while general-purpose sound separation models are truly sound source-independent, and therefore can separate any sound source given any audio mix.
[0070] In some implementations, system 10 in Figure 1 benefits from such a general-purpose sound source isolation module. In particular, it may be desirable to provide such a general-purpose sound source isolation module that is not computationally intensive and can be run without significant processing power.
[0071] The following describes a general-purpose sound source separation model that can be run on lightweight processing devices such as smartphones. This general-purpose sound source separation model relies on an energy threshold-based approach rather than a deep learning model.
[0072] Now, referring to Figure 4, we will describe an implementation of a general-purpose sound separation method.
[0073] First, in step S1, a pre-emphasis filter is applied to the audio mix 14 in the waveform domain to emphasize higher frequencies. The pre-emphasis filter is a time-domain FIR filter P(z) = 1 - C·z -1It can be implemented as follows, where C is an arbitrary design parameter that can be set to 1.
[0074] After the pre-emphasis filter is applied, the signal is transformed into the frequency domain in step S2, in this case using the Short-Time Fourier Transform (STFT). The STFT transform decomposes the waveform signal into a set of complex sinusoidal basis vectors, and the output is a complex-valued spectrogram of the audio mix.
[0075] In step S3, an absolute spectrogram is obtained, representing the magnitude (or absolute value) of the complex values of the audio mix spectrogram. In other words, the phase information of the audio mix spectrogram is discarded. Figures 5a and 5b show examples of spectrograms 51 and 52 as time-frequency plots, with grayscale used to show the amplitude of each time-frequency tile. Spectrogram 51 in Figure 5a represents the input audio mix, and spectrogram 52 in Figure 5b represents the same audio mix after pre-emphasis filtering.
[0076] Next, in step S4, a binary selection mask is obtained by binarizing the absolute value spectrogram 52 according to a given threshold. That is, each value is set to 1 or 0 depending on whether its value is above or below the threshold T. As an example, the threshold T can be set to -45 dB. That is, values that are 45 dB or more below the maximum value are set to zero.
[0077] The thresholding process is shown in Figures 6a and 6b, which show the time slices 61 and 62 of the spectrograms 51 and 52, respectively. As shown in Figures 6a and 6b, pre-emphasis filtering works to include higher frequencies in the selection.
[0078] The selection mask is optionally smoothed along the time, frequency, or both dimensions in step S5. A preferred technique for smoothing the mask is by applying a low-pass filter. The resulting selection mask ranges from 0 to 1 and is the same size as the audio mix spectrogram.
[0079] The selection masks obtained in steps S3 to S5 indicate which time-frequency tiles in the spectrogram correspond to loud sound sources (having mask values close to 1) and which do not correspond to loud sound sources (having mask values close to 0). In step S6, the masks are multiplied element by element with the audio mix spectrogram to provide spectrogram selection. The spectrogram selection 53 of the filtered audio mix spectrogram 52 is shown in Figure 5c. Figure 5d shows the residual spectrogram 54, i.e., the audio mix spectrogram 52 excluding the spectrogram selection 53.
[0080] Finally, in step S7, the spectrogram selection 53 is converted back to the original time domain, here by applying the inverse STFT, to form the output signal 55 corresponding to the dominant sound in the input audio mix. In most situations, this dominant sound is associated with one or more audio sources. Similarly, the residual spectrogram 54 is converted back to the residual audio signal. These two signals, namely the output signal and the residual signal, can function as the two source audio signals 15 described above with respect to the system in Figure 1.
[0081] Unless otherwise specified, as is evident from the following description, any use of terms such as “process,” “calculate,” “calculate,” “determine,” and “analyze” throughout this disclosure is understood to refer to the actions and / or processes of computer hardware or computing systems, or similar electronic computing devices, that manipulate and / or transform data, which is expressed as an electronic or other physical quantity, into other data, which is similarly expressed as a physical quantity.
[0082] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description thereof for the purpose of improving the flow of disclosure and aiding in the understanding of one or more of the various aspects of the invention. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are explicitly described in each claim. Rather, as reflected in the following claims, the aspects of the invention are fewer than all the features of a single, aforementioned disclosed embodiment. Thus, the claims following the detailed description are explicitly incorporated into this detailed description, and each claim stands alone as a distinct embodiment of the invention. Furthermore, some embodiments described herein include some features included in other embodiments, but not others, and combinations of features of different embodiments are intended to be within the scope of the invention and to form different embodiments. This will be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0083] Furthermore, some embodiments are described herein as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing the function. Thus, a processor having instructions for performing such methods or elements of methods forms means for performing the methods or elements of methods. Note that if a method includes several elements, for example, several steps, the ordering of such elements is not depicted unless specifically stated. Furthermore, the elements of apparatus embodiments described herein are examples of means for performing the function performed by those elements for the purpose of performing embodiments of the present invention. Numerous specific details are described in the description provided herein. However, it is understood that embodiments of the present invention can be carried out without these specific details. On the other hand, well-known methods, structures, and arts are not shown in detail so as not to obscure the understanding of this description.
[0084] Those skilled in the art will understand that the present invention is by no means limited to the preferred embodiments described above. Rather, many modifications and variations are possible within the scope of the appended claims. For example, other types of audio sources other than those described above may be included in the audio mix. Also, many other types of audio processing may be applied to isolated audio sources.
[0085] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEEs). [EEE1] A method for processing an input audio mix (14), the method being: A step of extracting at least two source audio signals (15) from the input audio mix, wherein each audio source signal represents a separate audio source; A step of extracting audio mix information from the input audio mix (14), wherein the audio mix information includes at least one of the audio mix semantic attributes and the audio mix signal attributes; A step of determining audio processing parameters (18) based on the aforementioned audio mix information; A step of processing the source audio signal (15) based on the audio processing parameters (18) to generate a processed audio mix. Methods that include... [EEE2] The method according to EEE1, wherein the audio mix semantic attributes include at least one of music genre, recording type, production style, identified sound sources in the mix, and characteristics of the identified sound sources. [EEE3] The method according to EEE1 or 2, wherein the audio mix signal attributes include at least one of loudness, dynamic range, average spectral power, and spectral power distribution. [EEE4] The method according to any one of EEE1 to 3, wherein the audio processing parameter (18) is determined based on the source signal attributes of the source audio signal (15). [EEE5] The method according to EEE4, wherein the source signal attribute includes at least one of loudness, dynamic range, average spectral power, and spectral power distribution. [EEE6] Based on the aforementioned audio mix information, the sound source separation parameter (19) is determined; The further includes extracting the at least two source audio signals based on the sound source separation parameters, The method described in any one of EEE1 to EEE5. [EEE7] The method according to EEE6, wherein the sound source separation parameter is determined based on the source signal attributes of the source audio signal (15). [EEE8] The method according to any one of EEE1 to 7, wherein the audio processing parameter (18) represents a linear gain, and the method further comprises mixing the extracted audio signals by applying the linear gain to the source audio signal. [EEE9] The linear gain described above varies with time, as described in EEE8. [EEE10] The method according to any one of EEE1 to 9, further comprising the step of processing each source audio signal (15) individually before mixing by applying at least one of dynamic range compression, equalization, dynamic equalization, or creative effects. [EEE11] The stage of providing the user with proposed audio processing options; The step of receiving user input related to the aforementioned selections via the user interface. The method described in any one of EEE1 to 10, further including the method described in any one of EEE1 to 10. [EEE12] The method according to any one of EEE1 to 11, further comprising the step of applying dynamic range compression and / or equalization to the input audio mix (14) before extracting the source audio signal (15). [EEE13] A system for processing an input audio mix (14), wherein the system is: A sound source separation module (11) configured to extract at least two source audio signals (15) from the input audio mix, wherein each audio source signal represents a separate audio source; An analysis module (12) is configured to receive the input audio mix, extract audio mix information including at least one of the audio mix semantic attributes and audio mix signal attributes, and determine audio processing parameters (18) based on the audio mix information; An editing module (13) is configured to receive the extracted audio signal and the audio processing parameters (18), process the source audio signal (15) based on the audio processing parameters, and generate a processed audio mix (17). A system equipped with these features. [EEE14] The analysis module (13) is further configured to receive the source audio signal, and the determined audio processing parameters (18) are also based on the source signal attributes of the source audio signal, as described in EEE13. [EEE15] The analysis module is further configured to determine a sound source separation parameter (19) based on the audio mix information; The sound source separation module is configured to receive the sound source separation parameters and to extract the at least two source audio signals based on the sound source separation parameters. Systems as described in EEE13 or 14. [EEE16] The system according to EEE15, wherein the analysis module (13) is configured to determine the sound source separation parameter (19) based on the source signal attributes of the source audio signal (15). [EEE17] The system according to any one of EEE13 to 16, wherein the audio processing parameter (18) represents a linear gain, and the editing module is configured to mix the extracted audio signals by applying the linear gain to the source audio signal. [EEE18] The system according to any one of EEE13 to 17, wherein the editing module (12) is further configured to process each source audio signal (15) individually before mixing by applying at least one of dynamic range compression, equalization, dynamic equalization, or creative effects. [EEE19] The system according to any one of EEE13 to 18, wherein the editing module (12) includes a user interface (21) configured to receive user input for the audio processing. [EEE20] The system according to EEE19, wherein the user interface (21) is configured to provide the user with a selection of proposed processing options and to receive user input related to the selection. [EEE21] A computer program product comprising a portion of program code configured to perform the method described in any one of the EEE1 to 12 when executed on a computer processor. [EEE22] The system according to any one of EEE13 to 20, wherein the sound source separation module (11) is configured to apply dynamic range compression and / or equalization to the audio mix before extracting the source audio signal. [EEE23] A method for separating sound sources: The stage of receiving the input audio mix (14); The steps include converting the input audio mix into a frequency domain audio mix spectrogram (51); The steps include: (52) binarizing the spectrogram by comparing each tile with a predetermined threshold to form a selection mask (53); The steps include: applying the selection mask to the audio mix spectrogram (51) to form a spectrogram selection (54); The steps include: inversely transforming the spectrogram selection (54) into the time domain to provide an output audio signal (55) associated with one or more sound sources in the input audio mix; Methods that include... [EEE24] The method according to EEE23, further comprising the step of applying a pre-emphasis filter to the input audio mix (14) before the conversion step, thereby amplifying higher frequencies relative to lower frequencies. [EEE25] The method according to EEE23 or 24, further comprising the step of smoothing the selection mask along the time dimension and / or frequency dimension before applying the selection mask to the audio mix spectrum (51).
Claims
1. A method for processing an input audio mix (14), the method being: A step of extracting at least two source audio signals (15) from the input audio mix, wherein each audio source signal represents a separate audio source; A step of extracting audio mix information from the input audio mix (14), wherein the audio mix information includes at least one of the audio mix semantic attributes and the audio mix signal attributes; The steps include determining audio processing parameters (18) based on the aforementioned audio mix information; A step of processing the source audio signal (15) based on the audio processing parameters (18) to generate a processed audio mix. Methods that include...
2. The method according to claim 1, wherein the audio mix semantic attributes include at least one of music genre, recording type, production style, identified sound sources in the mix, and characteristics of the identified sound sources.
3. The method according to claim 1 or 2, wherein the audio mix signal attributes include at least one of loudness, dynamic range, average spectral power, and spectral power distribution.
4. The method according to any one of claims 1 to 3, wherein the audio processing parameter (18) is determined based on the source signal attributes of the source audio signal (15).
5. The method according to claim 4, wherein the source signal attribute includes at least one of loudness, dynamic range, average spectral power, and spectral power distribution.
6. Based on the aforementioned audio mix information, the sound source separation parameter (19) is determined; The further includes extracting the at least two source audio signals based on the sound source separation parameters, The method according to any one of claims 1 to 5.
7. The method according to claim 6, wherein the sound source separation parameter is determined based on the source signal attributes of the source audio signal (15).
8. The method according to any one of claims 1 to 7, wherein the audio processing parameter (18) represents a linear gain, and the method further comprises mixing the extracted audio signals by applying the linear gain to the source audio signal.
9. The method according to claim 8, wherein the linear gain changes over time.
10. The method according to any one of claims 1 to 9, further comprising the step of processing each source audio signal (15) individually before mixing by applying at least one of dynamic range compression, equalization, dynamic equalization, or creative effects.
11. The stage of providing the user with proposed audio processing options; The step of receiving user input related to the aforementioned options via the user interface. The method according to any one of claims 1 to 10, further comprising:
12. The method according to any one of claims 1 to 11, further comprising the step of applying dynamic range compression and / or equalization to the input audio mix (14) before extracting the source audio signal (15).
13. A system for processing an input audio mix (14), the system being: A sound source separation module (11) configured to extract at least two source audio signals (15) from the input audio mix, wherein each audio source signal represents a separate audio source; An analysis module (12) configured to receive the input audio mix, extract audio mix information including at least one of the audio mix semantic attributes and audio mix signal attributes, and determine audio processing parameters (18) based on the audio mix information; An editing module (13) is configured to receive the extracted audio signal and the audio processing parameters (18), process the source audio signal (15) based on the audio processing parameters, and generate a processed audio mix (17). A system that includes these features.
14. The system according to claim 13, wherein the analysis module (13) is further configured to receive the source audio signal, and the determined audio processing parameters (18) are also based on the source signal attributes of the source audio signal.
15. The analysis module is further configured to determine a sound source separation parameter (19) based on the audio mix information; The sound source separation module is configured to receive the sound source separation parameters and to extract the at least two source audio signals based on the sound source separation parameters. The system according to claim 13 or 14.
16. The system according to claim 15, wherein the analysis module (13) is configured to determine the sound source separation parameter (19) based on the source signal attributes of the source audio signal (15).
17. The system according to any one of claims 13 to 16, wherein the audio processing parameter (18) represents a linear gain, and the editing module is configured to mix the extracted audio signals by applying the linear gain to the source audio signal.
18. The system according to any one of claims 13 to 17, wherein the editing module (12) is further configured to process each source audio signal (15) individually before mixing by applying at least one of dynamic range compression, equalization, dynamic equalization, or creative effects.
19. The system according to any one of claims 13 to 18, wherein the editing module (12) includes a user interface (21) configured to receive user input for the audio processing.
20. The system according to claim 19, wherein the user interface (21) is configured to provide the user with a selection of proposed processes and to receive user input related to the selection.
21. A computer program product comprising a portion of program code configured to perform the method described in any one of claims 1 to 12 when executed on a computer processor.
22. The system according to any one of claims 13 to 20, wherein the sound source separation module (11) is configured to apply dynamic range compression and / or equalization to the audio mix before extracting the source audio signal.
23. A method for separating sound sources: The stage of receiving the input audio mix (14); The steps include: converting the input audio mix into a frequency-domain audio mix spectrogram (51); The steps include: (52) binarizing the spectrogram by comparing each tile with a predetermined threshold to form a selection mask (53); The steps include: applying the selection mask to the audio mix spectrogram (51) to form a spectrogram selection (54); The steps include: inversely transforming the spectrogram selection (54) into the time domain to provide output audio signals (55) associated with one or more sound sources in the input audio mix; Methods that include...
24. The method according to claim 23, further comprising the step of applying a pre-emphasis filter to the input audio mix (14) before the conversion step, thereby amplifying higher frequencies with respect to lower frequencies.
25. The method according to claim 23 or 24, further comprising the step of smoothing the selection mask along the time dimension and / or frequency dimension before applying the selection mask to the audio mix spectrum (51).