Techniques for assisting audio mixing of audio tracks
The electronic device and AI system assist in audio mixing by identifying and removing perceptually redundant tracks using perceptual models, simplifying project management and reducing hardware requirements.
Patent Information
- Application Number
- PCT/EP2025/057180
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-25
AI Technical Summary
Existing audio mixing techniques lack efficiency in identifying and removing perceptually redundant audio tracks, leading to increased complexity and difficulty in managing large music projects.
An electronic device and AI system that determine perceptual information loss by simulating human perception to identify and mute redundant audio tracks, using models like MP3 encoders and PEMO-Q to encode only audible sound components, and provide user interfaces for track removal suggestions.
Reduces project complexity by allowing users to remove perceptually redundant tracks without affecting the final mix quality, potentially reducing the need for expensive audio reproduction systems.
Smart Images

Figure EP2025057180_25092025_PF_FP_ABST
Abstract
Description
[0001] TECHNIQUES FOR ASSISTING AUDIO MIXING OF AUDIO TRACKS
[0002] TECHNICAL FIELD
[0003] The present disclosure generally pertains to the field of sound production, in particular to devices and methods for sound design and sound mixing for music.
[0004] TECHNICAL BACKGROUND
[0005] In music production, a “track” refers to a discrete audio recording or MIDI data performance that is organized in a linear fashion. For example, in a Digital Audio Workstation (DAW), audio tracks are represented as horizontal lanes in the timeline, with the recorded audio displayed as waveforms.
[0006] Audio tracks allow producers to layer different sounds (such as vocals, instruments, or sound effects) to create a complete musical composition. In sound design and sound mixing for music, movies, and games, it is customary to overlap multiple sound clips that are arranged in different tracks. A sound clip may for example be a sound recording or a sound created by a synthesizer. Sound clips that are arranged in tracks may be processed with audio effect units.
[0007] By overlapping multiple tracks, the complexity of the final production is increased, and the result of the session or project is made more interesting for the listener. At the same time, multiple tracks enable the sound designer, mixer or producer to more easily craft the target sound. This happens because a single tool (be it a synthesizer or a recording with one audio effect chain) likely cannot deliver directly the desired result. The procedure of crafting the desired sound, be it a short sound effect or a final mix of a song or movie, is an inherently iterative process, where the final result is obtained by superimposing and modifying multiple tracks until their mix (i.e., their simultaneous playback) is satisfactory. This means that mixing music is a progressive activity, starting with one track and gradually "building" a more complex sound by adding more tracks.
[0008] Although there exist techniques for assisting in audio mixing, there is a need for improving these existing techniques.
[0009] SUMMARY
[0010] According to a first aspect the disclosure provides an electronic device comprising circuity configured to determine, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks. According to a further aspect the disclosure provides a method comprising determining, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
[0011] According to a further aspect the disclosure provides an Al system that is configured to infer whether or not to mute a selected track of a predefined set of tracks, given the mix of the set of tracks as reference.
[0012] According to a further aspect the disclosure provides a method of training an Al system on existing sound design project files to learn which tracks the sound designer muted and which ones were used for the final mix.
[0013] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0014] BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0016] Fig. 1 shows an example of a sound file that comprises three stereo tracks;
[0017] Fig. 2 schematically shows a process that informs the user about which tracks can be removed because they are perceptually redundant and do not contribute to the impression of the final mix;
[0018] Fig. 3 schematically shows an example of processing that may be performed as part of the process of Fig. 2;
[0019] Fig. 4 describes an embodiment of a process for determining whether tracks of a project are redundant;
[0020] Fig. 5 shows an example of processing performed in a project containing an audio track and a midi track;
[0021] Fig. 6 schematically shows an embodiment of a computation of a perceptual information loss that may be performed in the audio analysis 27 in Fig. 2;
[0022] Fig. 7 shows a process of informing the user that tracks of a project can be deleted without significant perceptual information loss;
[0023] Fig. 8 shows an example of a user interface output for informing the user that a track of a project can be deleted without significant perceptual information loss;
[0024] Fig. 9 shows another example of a user interface output for informing the user that tracks of a project can be deleted without significant perceptual information loss; Fig. 10 shows an example of a user interface output for informing the user about sections of a track where the largest perceptual difference with the original mix is;
[0025] Fig. 11 schematically shows an example of an Al system that is trained to automatically infer whether or not a track of a project can be muted; and
[0026] Fig. 12 schematically illustrates an embodiment of an electronic device.
[0027] DETAILED DESCRIPTION OF EMBODIMENTS
[0028] Before a detailed description of the embodiments under reference of Fig. 5 is given, general explanations are made.
[0029] Tracks are the building blocks of music production. Audio tracks are typically created by capturing live recordings, such as vocals, guitars, or drums. Typically, an audio track is generated for each instrument or vocal part. To record audio, a microphone or electronic instrument is typically connected to an audio interface. The recorded audio appears as a waveform on the track. MIDI tracks don’t contain audio; they carry MIDI data (note information). MIDI tracks are transformed from MIDI data to audio data by applying virtual instruments (synths, pianos, etc.) hosted by the Digital Audio Workstation (DAW). In a Digital Audio Workstation (DAW), multiple individual tracks are layered for a fuller sound experience.
[0030] The embodiments provide an electronic device comprising circuity configured to determine, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
[0031] The set of tracks may for example relate to a session or project a sound designer, music producer or audio mixing engineer is working on.
[0032] The set of tracks may for example comprise audio tracks represented as waveforms or MIDI tracks that are rendered to audio tracks using virtual instruments or the like. Tracks may comprise compressed or uncompressed audio data.
[0033] The perceptual information may for example be an indication of the amount of perceptual information that is lost in the final mix if the selected track (or a selected subset of tracks) is removed.
[0034] Circuitry may include a processor, a memory (RAM, ROM or the like), a storage, input means (mouse, keyboard, camera, etc.), output means (a display, e.g. liquid crystal, (organic) light emitting diode, etc.), loudspeakers, etc., a (wireless) interface, etc., as it is generally known for electronic devices (computers, smartphones, etc.). Moreover, circuitry may include circuitry such as a GPU specialized for implementing a neural network.
[0035] The circuitry may for example be configured to determine, from the set of tracks a track that can be removed from the set of tracks because it is perceptually redundant.
[0036] The circuitry may for example be configured to determine, from the set of tracks a track that can be removed from the set of tracks because it does not significantly contribute to the impression of the final mix of the tracks comprised by the set of tracks.
[0037] The electronic device may for example offer a sound designer, music producer or audio mixing engineer information about the current session or project they are working on, which is helpful in reducing the complexity of the session without impacting the perceptual result of the final mix.
[0038] The embodiments described below may for example alleviate the need to have expensive highgrade audio reproduction systems, such as loudspeakers or headphones, in order to retrieve information about redundant audio tracks.
[0039] Determining the perceptual information loss may for example be based on a perceptual model.
[0040] According to embodiments, determining a perceptually redundant track is based on the perceptual model.
[0041] According to embodiments, the perceptual model only encodes parts of the mix that can be heard. If the two encodings are very close to each other, their perceptual difference will be small and they will be indistinguishable when listened to by a human, even if the actual waveforms are different.
[0042] The perceptual model may for example be configured to simulate human perception and encode only what can be heard in a sound, neglecting imperceptible details. The perceptual model may for example allow to virtually reproduce the perception that a human gets when listening to the sound inside a computer. Any approach that models the human perception of sound can be used as perceptual model. The perceptual model may for example be a MP3 encoder or a PEMO-Q model.
[0043] The circuitry may be configured to compute, as a reference mix, the mix of all the tracks and to encode the reference mix using a perceptual model.
[0044] The circuitry may further be configured to select one track of the set of tracks at a time, to compute a query mix without the selected track, and to encode the query mix with the perceptual model. Selecting one track of the set of tracks at a time may be repeated for all tracks of the mix. The circuitry may be configured to compare the query mix with the reference mix. For example, if the difference between the query mix and the reference mix is lower than a predefined threshold, the selected track can be removed from the mix. If the difference is higher than the threshold, the selected track is kept.
[0045] The perceptual information loss may for example be a value indicative of the difference between the encoded query mix and the encoded reference mix, after the perceptual model has been applied. For example, the perceptual information loss may be a value proportional to the difference between the encoded query mix and the encoded reference mix, after the perceptual model has been applied.
[0046] Determining the query mix may comprise muting the selected track. In this way, the query mix changes by muting a selected track, but the reference mix remains the same.
[0047] Determining the perceptual information loss may be repeated for multiple or all tracks in the set of tracks. For example, determining the perceptual information loss may be iteratively repeated for multiple or all tracks in the set of tracks.
[0048] Determining the perceptual information may be performed for tracks in a subset of tracks comprised in the set of tracks. The circuitry may for example be configured to determine, for a subset of tracks selected from the set of tracks, a perceptual information loss related to removing the tracks of the subset from the set of tracks.
[0049] The circuitry may for example be configured to provide a user interface that, given a set of tracks provided by the user, informs the user about which subset of tracks can be removed because they are perceptually redundant. The user interface may for example be configured to display to the user a list of tracks that can be removed.
[0050] The user interface may for example be configured to allow the user to select which contributions to the perceptual differences should be kept and which should be restored. For example, the user interface may allow the user to manually select parts of tracks on the waveform / spectrogram.
[0051] The user interface is configured to indicate to the user where the largest perceptual loss related to removing the selected track from the reference mix is.
[0052] According to embodiments, the circuitry is configured to perform a perceptual similarity -based retrieval given a query. This may enable the electronic device o to find the most similar recording to a query inside a database. The embodiments also provide a method comprising determining, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
[0053] The method may be a computer-implemented method.
[0054] The embodiments also disclose a computer-readable medium comprising instructions to implement the functionality described herein.
[0055] The embodiments also provide an Al system that is configured to infer whether or not to mute a selected track of a predefined set of tracks, given the mix of the set of tracks as reference.
[0056] The Al system may comprise a neural network that is configured to receive as input a reference mix and a selected individual track belonging to the reference mix. According to embodiments, the reference mix is obtained based on all tracks of a set of tracks and the individual track is a selected individual track belonging to set of tracks comprised by the reference mix.
[0057] The neural network may be configured to output an indication of the probability of muting the selected track.
[0058] According to an embodiment, the Al system is configured to, for a project consisting of multiple tracks, iterate through the tracks and runs inference for each of these tracks, providing the same reference mix as additional input.
[0059] The Al system may be configured to consider muting a track based on perceptual redundancy, noise present in the recording, or clashing of frequencies with other tracks.
[0060] The embodiments also provide a method of training an Al system on existing sound design project files to learn which tracks the sound designer muted and which ones were used for the final mix.
[0061] Fig. 1 shows an example of a sound file that comprises three stereo tracks. A first stereo track (Track 1) comprises a left channel audio signal (L) and a right channel audio signal (R) that relate to a signal of recorded vocals. A second stereo track (Track 2) comprises a left channel audio signal (L) and a right channel audio signal (R) that relate to a signal of recorded drums. A third stereo track (Track 3) comprises a left channel audio signal (L) and a right channel audio signal (R) that relate to a signal of recorded guitars. After recording, the audio tracks can be processed in a digital audio workstation (DAW) with effects like reverb, compression, and equalization to enhance their sound. By sending multiple tracks (such as vocals, drums, guitars above) to a master bus, the audio signals stored in the tracks are layered. When layering the audio signals of the individual tracks, the audio engineer typically controls the mixing of the tracks. When audio engineers perform mixing of recorded tracks, they engage in a process that brings together individual elements (e.g. the vocals, drums, and guitars as shown in Fig. 1) to create a cohesive and balanced final audio mix. For example, the audio engineer adjusts the volume levels of each recorded track (such as vocals, drums, guitars, bass, etc.). This “balancing” ensures that no track dominates the mix, and all elements are audible. Still further, the audio engineer places each track in the stereo field using panning. Tracks can be positioned anywhere between the left and right speakers (or more output channels if 3D audio with multiple output channels is applied).
[0062] In larger recording projects, the number of tracks can vary significantly based on the complexity of the music, the genre, and the producer’s creative choices. In home studios and small projects, there are typically 12 to 24 tracks. This allows for basic arrangements, including drums, bass, guitars, vocals, and a few additional instruments. Professional studios often work with more tracks. Professional studios in pop / rock music typically use 24 to 48 tracks. This range accommodates individual instruments, multiple vocal layers, harmonies, and additional effects. Separate Tracks: Drums (kick, snare, toms, cymbals), bass, guitars (rhythm, lead), keyboards, vocals (lead, backing), and other instruments each get their own track. In orchestral and film scoring, the number of tracks can easily go over hundreds of tracks. Orchestral compositions or film scores involve many individual instrument tracks. Strings, brass, woodwinds, percussion, choir, and more contribute to the rich sound. Each section (violins, cellos, etc.) may have multiple tracks for layering and articulations.
[0063] In larger projects, it is possible for the final set of tracks to be perceptually redundant. Often some of the tracks added during the early stages of sound design or mixing do not contribute anymore to the result, as they have been "covered" by tracks added in a later stage. This is exacerbated by the fact that, for some complex projects, the number of tracks quickly becomes difficult to manage (higher than 800).
[0064] The following embodiments describe a system (accessible through a user interface) that informs the user about which tracks can be removed because they are perceptually redundant and do not contribute to the impression of the final mix.
[0065] The embodiments exploit models that simulate human perception and encode only what can be heard in a sound, neglecting imperceptible details. This allows to virtually reproduce the perception that a human gets when listening to the sound inside a computer. In the following it is referred to these models as "perceptual models". Examples of perceptual models are the MP3 encoder ([1] https: / / www.soundonsound.com / techniques / perceptual-coding-howmp3- compression-works [2] https7 / en. wikipedia. org / wiki / MP3) and PEMO-Q ([3] https: / / ieeexplore.ieee.org / document / 1709880). Any approach that exhaustively models the human perception of sound can be used as perceptual model in the embodiments described below.
[0066] Fig. 2 schematically shows a process that informs the user about which tracks can be removed because they are perceptually redundant and do not contribute to the impression of the final mix. The process comprises a so called “reference” part and a so called “query” part. A set of tracks 21 comprises multiple audio tracks (such as exemplified in Fig. 1) that have been recorded and stored in a project within a digital audio workstation (DAW). This set of tracks 21 may for example comprise all available tracks in the project.
[0067] Within the “reference” part, the set of tracks 21 is processed by track mixing 23. In track mixing 23, the individual tracks comprised in the set of tracks 21 are layered (mixed) to generate reference mix 24R. This reference mix 24R may for example be a single stereo track that stores the mixing result. The reference mix 24R is provided to a perceptual model 25 that simulates human perception and encodes only what can be heard in a sound, neglecting imperceptible details. The perceptual model 25 encodes reference mix 24R to generate encoded reference mix 26R. Encoded reference mix 26R virtually reproduces the perception that a human gets when listening to the reference mix 24R with a sound reproduction system.
[0068] Within the “query” part, the set of tracks 21 is processed by selective track muting 22. Selective track muting 22 selects one ore more of the tracks contained in the set of tracks 21 to generate a query set 21Q which is a modified version of set of tracks 21 in which one or more of the tracks are muted. In track mixing 23, the individual tracks comprised in the query set 21Q are layered (mixed) to generate query mix 24Q. That is, the mix is computed with the remaining tracks that have not been muted by selective track muting 22. This query mix 24Q may for example be a single stereo track that stores the mixing result. The query mix 24Q is provided to perceptual model 25 that simulates human perception and encodes only what can be heard in a sound, neglecting imperceptible details. The perceptual model 25 encodes query mix 24Q to generate encoded query mix 26Q. Encoded query mix 26Q (in the following also called the “query”) virtually reproduces the perception that a human gets when listening to the reference mix 24Q with a sound reproduction system. The encoded reference mix 26R and the encoded query mix 26Q are provided to an audio analysis 27. Audio analysis 27 determines if there is a significant difference between encoded query mix 26Q and encoded reference mix 26Q, or not. Audio analysis 27 determines an analysis result 28 that parametrizes the perceptual difference between encoded query mix 26Q and encoded reference mix 26Q. This analysis result 28 is processed by audio track removal 29. Audio track removal 29 may for example indicate to the user that a specific track selected by selective track muting 22 can be removed from the track without causing a significant perceptual difference to the listener.
[0069] The set of tracks 21 in Fig. 2 may for example relate to audio tracks that comprise audio data in waveform format. However, the same principles as set out in Fig. 2 above also apply to music projects with MIDI tracks or combinations of audio tracks and MIDI tracks. MIDI stands for Musical Instrument Digital Interface. Unlike audio tracks that capture live instruments and microphones, MIDI tracks record sequences of data that relate to electronic (hardware or software) instruments. MIDI data may provide information such as ‘note on / off which indicates which note is pressed and released, ‘velocity’ which represents how hard a key is pressed (ranging from 0 to 127), and other parameters such as ‘aftertouch’ (which captures changes in key pressure), ‘vibrato’ and ‘pitch bend’ which allow expressive control over notes.
[0070] The audio analysis 27 may for example compare two audio signals by quantifying their similarity or difference. This can be done using various techniques known to the skilled person, for example by computing the distance of the data vectors representing the two audio signals in an n-dimensional space, where n is the number of audio samples of the query mix and the reference mix.
[0071] As the reference mix 25R obtained in the “reference” part of the process is the same for all queries, the reference mix 25R can be determined in advance and stored in a memory. For each query, audio analysis may than compare the query mix 26Q obtained by the “query” part of the process with the stored reference mix.
[0072] Fig. 3 schematically shows an example of processing that may be performed as part of the process of Fig. 2. A set of tracks 21 of a music project comprises six audio tracks, namely audio track 1, audio track 2, audio track 3, audio track 4, audio track 5, and audio track 6. In a reference part, at 31R, audio tracks 1 to 6 are processed by mixer 41 which adds the signals contained in the set of tracks 1 to 6. The resulting reference mix is then processed by perceptual model 24 to generate a coded version of the reference track. In a query part, audio track 3 in the set of tracks 21 has been selected for muting (e.g. by selective track muting 22 in Fig. 2). At 3 IQ, audio tracks 1 to 6 (with audio track 3 being muted) are processed by mixer 3 IQ which adds the signals contained in the set of tracks 1-2, and 5 to 6. The resulting query mix is then processed by perceptual model 24 to generate a coded version of the query mix. At 32, the coded query mix and the coded reference mix are subtracted from each other (which is an example of the audio analysis 27 in Fig. 2) to obtain a difference signal (which is an example of the analysis result 28 in Fig. 2). At 33, it is decided, based on the difference signal obtained from subtraction 32, whether track 3 (the track that has been muted and thus is not present in the query mix) can be foreseen for removal from the project, or not.
[0073] The subtraction performed at 32 may for example comprise determining a distance between a data vector representing the query mix and a data vector representing the reference mix. The difference obtained from subtraction 32 of the query mix and the reference mix may for example be compared to an internal threshold, which allows to determine whether the muted track can be removed or not. The intuition behind the comparison is that the perceptual model only encodes parts of the mix that can be heard. If the two encodings are very close to each other, their perceptual difference will be small and they will be indistinguishable when listened to by a human, even if the actual waveforms are different.
[0074] This process is repeated for all available tracks. The query changes by muting a different track, but the reference remains the same.
[0075] At the end of this process, the proposed system may display to the user a list of tracks that can be removed.
[0076] Fig. 4 describes an embodiment of a process for determining whether tracks of a project are redundant. At 41, the process first computes the mix of all the tracks. Then, at 42, the process encodes the mix using a perceptual model. The resulting encoding will be used as global reference. At 43, the process selects one track at a time and computes a new mix without such track. This "incomplete" mix is then encoded by the perceptual model (called "query"). At 44, the query is compared with the reference encoding. If the difference between them is lower than a threshold, at 45, the current tracks can be removed from the mix. If the difference is higher than the threshold, at 46, the track is kept.
[0077] The value of the threshold may for example depend on the perceptual model used and may for example be set manually during development. Steps 43 to 46 may be iteratively performed for all tracks in the project. In this way, the process may inform the user about which subset of tracks can be removed because the tracks of the subset are perceptually redundant.
[0078] The process may for example be implemented in a system (accessible through a user interface) that, given a set of tracks provided by the user that will be mixed together, e.g. by summation, informs the user about which subset of tracks can be removed because they are perceptually redundant.
[0079] In the embodiment described above mixing the tracks to generate the reference mix and the query mix comprises summing the audio waveforms of the tracks. Mixing the tracks to generate the reference mix and the query mix may also comprise additional processing steps that are typical in mixing, such as adapting the gain for each track according according to a setting provided by the user of the system.
[0080] Processing of projects with audio tracks and MIDI tracks
[0081] As tracks, the processes described here may handle not only audio tracks, but also MIDI tracks. According to technology known from Digital Audio Workstations, MIDI tracks are converted to audio tracks by software or hardware instruments before further audio processing.
[0082] Fig. 5 shows an example of processing performed in a project containing an audio track 51 and a midi track 52. Audio track 51 is directly provided to track mixing 23. MIDI track 52 is provided to preprocessing by a virtual instrument 52. Virtual instrument 53 (e.g. a sampler, a rompler, or a synthesizer) generates, based on the MIDI data stored in MIDI track 52, an audio track 53 that represents audio data that resembles data recorded from an instrument modeled by virtual instrument 53. Audio track 53 is then provided to track mixing 23 for further processing together with audio track 51. The virtual instrument 43 may for example be implemented according to the VST (Virtual Studio Technology) standard or Audio Unit (AU) standard, or the like.
[0083] Perceptual information loss
[0084] Fig. 6 schematically shows an embodiment of a computation of a perceptual information loss that may be performed in the audio analysis 27 in Fig. 2. According to this embodiment, a process 61 computes how much perceptual information is lost in the final mix if the current subset of tracks is removed. The process 61 provides this information as perceptual information loss 62. According to an example, the amount of perceptual information loss may for example be determined as a quantity proportional to the difference between the encoded query 26Q and the encoded reference 26R, after the perceptual model has been applied. The perceptual information loss may for example be determined as the square of the difference between the data vector that describes the waveform of reference mix 26R and the data vector that describes the waveform of query mix 26Q.
[0085] This embodiment allows the user to tolerate some loss in the perception of the final mix, gaining through a reduction in complexity of the project (which may be required due to e.g., computational limitations of the hardware used to produce the sound).
[0086] User Interface (UI)
[0087] The following embodiments describe a user interface that informs the user about which tracks of a project can be removed because they are perceptually redundant and do not contribute to the impression of the final mix.
[0088] Fig. 7 shows a process of informing the user that tracks of a project can be deleted without significant perceptual information loss. At 71, for each track selected of a project, the selected track us muted and a perceptual information loss of the resulting query mix as compared to the reference mix is determined. At 72, if the perceptual information loss for a selected track is insignificant, the user is informed that the track can be deleted.
[0089] Fig. 8 shows an example of a user interface output for informing the user that a track of a project can be deleted without significant perceptual information loss. After analysis of a project as set out with regard to the embodiments above, the user interface displays a message to the user informing the user that Track no. 7 titled „light pad“ has a perceptual information loss of 0.002 (which is less than a predefined threshold) and that this track is barely audible. The user interface further asks the user if this track should be removed from the project. The user interface provides the user with two alternative choices, “yes” or “no”. If the user selects “yes”, then the track is removed from the project. If the user selects “no”, then the track remains in the project.
[0090] Fig. 9 shows another example of a user interface output for informing the user that tracks of a project can be deleted without significant perceptual information loss. After analysis of a project as set out with regard to the embodiments above, the user interface displays to the user a track list 91 of the project. In the example of Fig. 8, track nos. 1 to 6 and 8 to 20 with the titles “Violin 1”, “Violin 2”, “Viola”, “Cello”, “Double Bass”, “Flute”, “Oboe”, “Clarinet”, “Bassoon”, “French Horn”, “Trumpet”, “Trombone”, “Tuba”, “Timpani”, “Snare Drum”, “Bass Drum”, “Cymbals”, “Harp”, “Celesta” all have a respective attributed perceptual information loss that is larger than a predefined threshold. These tracks are therefore not marked for removal. Track no. 7 with title “English Horn” has an attributed perceptual information loss of 0.008 which is lower than the predefined threshold. Accordingly, this track is marked as a candidate for removal.
[0091] Based on the value of the perceptual information loss provided by the user interface output of Fig. 8, the user can recognize that some of the tracks in the project are not, or barely (depending on the preset height of the threshold) audible. The user interface output also provides a recommendation which tracks can be removed from the project without significant information loss. Based on this user interface output, the user can thus make a decision on how to proceed. The user can either follow the suggestion to remove the tracks that are marked for removal. Alternatively, the user might decide to change the gain for an inaudible track so that, after mixing, its contribution to the mix becomes larger.
[0092] It should be noted that in the illustrative examples of Figs. 8 and 9, the threshold value has been chosen as 0.01. Tracks with perceptual information loss below this threshold are considered as providing an insignificant contribution to the final mix. This value is, however, illustrative only. The user may set the threshold according to own needs as part of configuration settings of the system. Alternatively, the system might present to the user some preset recommendations for the threshold value, e.g., 0.01 for a decision between “barely audible” and “audible”, or 0.05 for a decision between “inaudible” and “audible”.
[0093] It should be noted that according to modifications of the illustrative example of Fig. 9, the user interface output might only present the respective perceptual information loss for each track to the user, without giving a recommendation for removal.
[0094] According to yet a modification, instead or in addition of giving a recommendation for removal, the user interface might provide some guidance to the user in the form of a translation of the value of each respective perceptual information loss to a human readable attribute, such as “inaudible”, “barely audible”, “audible”, and “clearly audible”, depending on predefined ranges of the perceptual information loss. For example, perceptual information losses smaller than 0.005 could be translated to “inaudible”, perceptual information losses larger than or equal to 0.005 and smaller than 0.01 could be translated to “barely audible”, perceptual information losses larger than or equal to 0.01 and smaller than 0.1 could be translated to “audible”, and perceptual information losses larger than or equal to 0.1 could be translated to “clearly audible”.
[0095] According to another modification, the user interface, once the mix has been simplified, highlights in its waveform or spectrogram where the largest perceptual difference with the original mix is. The user interface also allows the user to select which contributions to the perceptual differences should be kept and which should be restored, by e.g., manually selecting them on the waveform / spectrogram.
[0096] Fig. 10 shows an example of a user interface output for informing the user about sections of a track where the largest perceptual difference with the original mix is. The user interface displays on a timeline the waveform 101 of the track, here a stereo track with left and right channels L, R.
[0097] Sections S2, and S4 highlight in the waveform where the largest perceptual difference with the original mix is. Sections SI, S3, and S5 indicate small perceptual difference with the original mix.
[0098] The system can also operate on longer productions (e.g., movies or music): in such case, the system operates in smaller chunks of the mix. This allows the user to measure in which section of the mix a track is redundant and in which it is not.
[0099] Perceptual similarity-based database search
[0100] The embodiments described above can be linked to a system that performs perceptual similaritybased retrieval given a query. Such a system can find the most similar recording to a query inside a database of audio recordings. Assuming that the database comprises the set of all mixtures where one or more tracks have been removed, the process described in the embodiments above finds the set of those tracks that are most similar to the query, in a perceptually motivated way.
[0101] This has applications other than redundancy reductions, as this information can be useful to a sound designer or music producer e.g., for choosing among multiple perceptually similar recordings.
[0102] Al system
[0103] Besides the perceptual approaches described above, an Al system can be trained to automatically learn whether to mute a track or not, given the mix as reference. According to such an alternative implementation, an Al system can be trained on existing sound design project files to learn which tracks the sound designer muted and which ones were used for the final mix. By this, the system could also learn other characteristics of the tracks that the lead the sound designer to mute a track - for example, there might be the problem of noise in a recording or clashing frequencies with other tracks. This allows the system to consider muting a track for other reasons than perceptual redundancy (noise present in the recording, clashing frequencies with other tracks, etc). Assuming that the system is provided with enough training data representing these instances, the model will learn to consider them. Fig. 11 schematically shows an example of an Al system that is trained to automatically infer whether or not a track of a project can be muted. In the example of Fig. 11, an exemplifying configuration for such an Al system is a neural network 111 that receives as input the reference mix 112 (including all the tracks) and one individual track 113 belonging to it, the individual track 113 being a selected candidate for removal from the mix. The neural network 111 is trained to output a scalar value 114 (here 0.8), indicating the probability of muting the selected track 113.
[0104] For a project consisting of N tracks, the system would iterate through the tracks and run inference for each of them, providing the same reference mix as additional input.
[0105] Existing projects that are used as training data include the ground-truth information of whether a track is muted or not, and this can be used in a supervised setting to perform the learning procedure, using e.g., a binary cross-entropy loss.
[0106] Any network architecture can be used to carry out this task. Convolutional layers can be preferred, as they work well with audio signals, although more powerful layers (e.g., transformers) can also be employed to improve quality.
[0107] Implementation
[0108] Fig. 12 schematically illustrates an embodiment of an electronic device 120. The electronic device 120 may be an electronic device, such as a personal computer, a workstation computer, a tablet computer, or other mobile devices. The electronic device 120 includes a CPU 121 as processor. Additionally, or alternatively, other computation hardware, such as GPU, TPU, DSP etc. may be used.
[0109] The electronic device 120 further includes an audio interface 126 that may connect microphones and loudspeakers to the processor 121. The processor 121 may for example implement a process of determining, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks as described in the embodiments above.
[0110] The electronic device 120 further includes a user interface 129 that is connected to the processor 101. This user interface 129 may act as a man-machine interface and enables a dialogue between a user and the electronic device 120. For example, a user may make configurations to the system using this user interface 129. The user interface 129 may for example comprise a display unit and input means such as a keyboard, a mouse, etc. The user interface may also be provided in part by software running on CPU 121. The electronic device 120 further includes a Bluetooth interface 124, and a WLAN interface 125.
[0111] These units 124, 125 act as I / O interfaces for data communication with external devices. An ethernet interface may also be possible.
[0112] The electronic device 120 further includes a data storage 122 and a data memory 123 (here a RAM). The data memory 123 is arranged to temporarily store or cache data or computer instructions for processing by the processor 121. The data storage 122 is arranged as a long-term storage, which may be accessed via the processor 121. Data storage 122 may store software instructions and data.
[0113] Furthermore, the electronic device 120 includes an artificial intelligence (Al) processor 127. The Al processor 127 may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU). The Al processor 127 may be configured to execute an Al model (e.g., an artificial neural network), for example, the machine learning model.
[0114] Note that the present technology can also be configured as described below.
[0115] (1) An electronic device comprising circuity configured to determine, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
[0116] (2) The electronic device of (1), wherein the circuitry is configured to determine, from the set of tracks a track that can be removed from the set of tracks because it is perceptually redundant.
[0117] (3) The electronic device of (1) or (2), wherein the circuitry is configured to determine, from the set of tracks a track that can be removed from the set of tracks because it does not significantly contribute to the impression of the final mix of the tracks comprised by the set of tracks.
[0118] (4) The electronic device of any one of (1) to (3), wherein determining the perceptual information loss is based on a perceptual model.
[0119] (5) The electronic device of any one of (1) to (4), wherein the perceptual model is configured to simulate human perception and encode only what can be heard in a sound, neglecting imperceptible details.
[0120] (6) The electronic device of any one of (1) to (5), wherein the perceptual model is a MP3 encoder or a PEMO-Q model. (7) The electronic device of any one of (1) to (6), wherein the circuitry is configured to compute, as a reference mix, the mix of all the tracks and to encode the reference mix using a perceptual model.
[0121] (8) The electronic device of any one of (1) to (7), wherein the circuitry is configured to select one track of the set of tracks at a time, to compute a query mix without the selected track, and to encode the query mix with the perceptual model.
[0122] (9) The electronic device of (8), wherein the circuitry is configured to compare the query mix with the reference mix.
[0123] (10) The electronic device of (8) or (9), wherein the perceptual information loss is a value indicative of the difference between the encoded query mix and the encoded reference mix, after the perceptual model has been applied.
[0124] (11) The electronic device of any one of (8) to (10), wherein determining the query mix comprises muting the selected track.
[0125] (12) The electronic device of any one of (1) to (11), wherein determining the perceptual information loss is repeated for multiple or all tracks in the set of tracks.
[0126] (13) The electronic device of any one of (1) to (12), wherein determining the perceptual information is performed for tracks in a subset of tracks comprised in the set of tracks.
[0127] (14) The electronic device of any one of (1) to (13), wherein the circuitry is configured to provide a user interface that, given a set of tracks provided by the user, informs the user about which subset of tracks can be removed because they are perceptually redundant.
[0128] (15) The electronic device of (14), wherein the user interface is configured to allow the user to select which contributions to the perceptual differences should be kept and which should be restored.
[0129] (16) The electronic device of (14) or (15), wherein the user interface is configured to indicate to the user where the largest perceptual loss related to removing the selected track from the reference mix is.
[0130] (17) The electronic device of any one of (1) to (16), wherein the circuitry is configured to perform a perceptual similarity-based retrieval given a query.
[0131] (18) A method comprising determining, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks. (19) An Al system that is configured to infer whether or not to mute a selected track of a predefined set of tracks, given the mix of the set of tracks as reference.
[0132] (20) The Al system of (19), wherein the Al system comprises a neural network that is configured to receive as input a reference mix and a selected individual track belonging to the reference mix.
[0133] (21) The Al system of (20), wherein the neural network is configured to output an indication of the probability of muting the selected track.
[0134] (22) The Al system of any one of (19) to (21), wherein the Al system, for a project consisting of multiple tracks, iterates through the tracks and runs inference for each of these tracks, providing the same reference mix as additional input.
[0135] (23) The Al system of any one of (19) to (22), wherein the Al system is configured to consider muting a track based on perceptual redundancy, noise present in the recording, or clashing of frequencies with other tracks.
[0136] (24) A method of training an Al system on existing sound design project files to learn which tracks the sound designer muted and which ones were used for the final mix.
Claims
CLAIMS1. An electronic device comprising circuity configured to determine, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
2. The electronic device of claim 1, wherein the circuitry is configured to determine, from the set of tracks a track that can be removed from the set of tracks because it is perceptually redundant.
3. The electronic device of claim 1, wherein the circuitry is configured to determine, from the set of tracks a track that can be removed from the set of tracks because it does not significantly contribute to the impression of the final mix of the tracks comprised by the set of tracks.
4. The electronic device of any one of claim 1, wherein determining the perceptual information loss is based on a perceptual model.
5. The electronic device of claim 1, wherein the perceptual model is configured to simulate human perception and encode only what can be heard in a sound, neglecting imperceptible details.
6. The electronic device of claim 1, wherein the perceptual model is a MP3 encoder or a PEMO-Q model.
7. The electronic device of claim 1, wherein the circuitry is configured to compute, as a reference mix, the mix of all the tracks and to encode the reference mix using a perceptual model.
8. The electronic device of claim 1, wherein the circuitry is configured to select one track of the set of tracks at a time, to compute a query mix without the selected track, and to encode the query mix with the perceptual model.
9. The electronic device of claim 8, wherein the circuitry is configured to compare the query mix with the reference mix.
10. The electronic device of claim 8, wherein the perceptual information loss is a value indicative of the difference between the encoded query mix and the encoded reference mix, after the perceptual model has been applied.
11. The electronic device of claim 8, wherein determining the query mix comprises muting the selected track.
12. The electronic device of claim 1, wherein determining the perceptual information loss is repeated for multiple or all tracks in the set of tracks.
13. The electronic device of claim 1, wherein determining the perceptual information is performed for tracks in a subset of tracks comprised in the set of tracks.
14. The electronic device of claim 1, wherein the circuitry is configured to provide a user interface that, given a set of tracks provided by the user, informs the user about which subset of tracks can be removed because they are perceptually redundant.
15. The electronic device of claim 14, wherein the user interface is configured to allow the user to select which contributions to the perceptual differences should be kept and which should be restored.
16. The electronic device of claim 14, wherein the user interface is configured to indicate to the user where the largest perceptual loss related to removing the selected track from the reference mix is.
17. The electronic device of claim 1, wherein the circuitry is configured to perform a perceptual similarity-based retrieval given a query.
18. A method comprising determining, for a track selected from a set of tracks, a perceptual information loss related to removing the selected track from the set of tracks.
19. An Al system that is configured to infer whether or not to mute a selected track of a predefined set of tracks, given the mix of the set of tracks as reference.
20. The Al system of claim 19, wherein the Al system comprises a neural network that is configured to receive as input a reference mix and a selected individual track belonging to the reference mix.
21. The Al system of claim 20, wherein the neural network is configured to output an indication of the probability of muting the selected track.
22. The Al system of claim 19, wherein the Al system, for a project consisting of multiple tracks, iterates through the tracks and runs inference for each of these tracks, providing the same reference mix as additional input.
23. The Al system of claim 19, wherein the Al system is configured to consider muting a track based on perceptual redundancy, noise present in the recording, or clashing of frequencies with other tracks.
24. A method of training an Al system on existing sound design project files to learn which tracks the sound designer muted and which ones were used for the final mix.
Citation Information
Patent Citations
System and method for synchronized multi-track editing
US20100218097A1
Semantic audio track mixer
US20140037111A1
Structures and methods for controlling the playback of music track files
US20160224310A1
Intelligent Crossfade With Separated Instrument Tracks
US20180277076A1
Adaptive audio mixing
US20220193549A1