Audio injection system and method

By separating and processing audio tracks to generate harmonically compatible binaural beats and spatialized tracks, the method enhances neurological effects by synchronizing brainwaves with the music's frequency spectrum, improving EEG synchronization and musical experience.

JP2026509117APending Publication Date: 2026-03-17APPLIED INSIGHTS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-08-23
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing methods for delivering binaural beats as background audio fail to utilize the music material for brainwave synchronization, lacking integration with the musical structure and frequency spectrum, thereby reducing effectiveness.

Method used

The method involves separating a source audio track into multiple tracks, processing each to generate binaural beat tracks, spatialized tracks, and additional sound tracks like infrasound and ultrasonic tracks, ensuring harmonically compatible frequencies and tempo-based spatialization to enhance neurological effects without disturbing the musical experience.

Benefits of technology

This approach increases the probability of successful EEG synchronization by providing multi-layered binaural beats dynamically matched to the source audio, enhancing neurological responses through varied frequencies and spatial awareness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509117000001_ABST
    Figure 2026509117000001_ABST
Patent Text Reader

Abstract

An audio injection system and method are disclosed. A source audio track is separated into multiple audio tracks (e.g., instrumental, vocal, or a mix thereof), and the audio tracks are processed individually to generate multiple binaural beat tracks. At least one spatialized track is also generated by filtering the source audio track to provide a filtered track, generating one or more spatialized trajectories based on specific audio features of the source audio track (e.g., tempo) and a target final state effect, and spatializing the filtered track using the spatialized trajectories. One or more other tracks, such as an ultra-low frequency track, an ultrasonic track, an enhanced low-frequency track, and / or a subharmonic track, may also be generated. The tracks may be played or mixed simultaneously for delivery to an end-user device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 18 / 105,373, filed on February 3, 2023, the entire disclosure of which is incorporated herein by reference.

[0002] Field of the Invention The present invention relates generally to audio systems and methods, and more particularly to systems and methods for digitally processing audio recordings to provide end - users with measurable enhanced effects.

Background Art

[0003] To reach a particular mental state of the human brain, various types of sounds can be used to synchronize the human brain. For example, when two tones with slightly different frequencies are played simultaneously in the left and right ears, the human brain perceives that a third tone has occurred with a frequency equal to the difference in frequencies between the two played tones. This auditory illusion is called a binaural beat and, depending on its frequency, can be used to reduce stress and anxiety, increase concentration and attention, improve sleep quality, or provide other benefits. In some cases, an individual may choose to listen to binaural beats along with music to enhance the listening experience. However, the binaural beats are simply played as the background of the music (or vice versa), that is, the music material itself is not used for the purpose of brainwave synchronization. Therefore, in the art, there remains a need for improved methods of delivering sounds having neuropsychological effects that overcome the drawbacks associated with existing methods and / or provide other advantages compared to existing methods.

Summary of the Invention

Means for Solving the Problems

[0004] The present invention relates to a method and system for injecting sound into a source audio track without disturbance to musical quality or perceived structure in order to provide desired psychological, neurological, and / or physiological results.

[0005] In one embodiment, a source audio track is separated into multiple audio tracks, such as multiple musical stems, and each separated audio track is processed individually to generate a binaural beat track. Each binaural beat track is generated from the audio track by (a) automatically transcribing the audio track to provide a transcription that includes estimated fundamental frequencies and / or estimated amplitude envelopes for each of the multiple notes, and (b) using the transcription to generate the binaural beat track. This embodiment provides multiple binaural beat tracks, one for each separated audio track, providing multiple layers of stimulation. Because the binaural beat tracks are derived from the separated audio tracks, the perception of specific instruments and / or vocals within the audio tracks can be enhanced. Furthermore, the delivery of multi-layered binaural beats dynamically matched to the source audio track results in a polyphonic binaural beat experience that achieves frequencies corresponding to the central tone of the source audio track without distorting the acoustic experience.

[0006] Existing methods for implementing binaural beats typically involve delivering a single drone tone frequency in each of the left and right speaker channels, extracted from a limited frequency spectrum. However, the binaural beat generation method of the present invention extracts a range of pitches that are harmonically and melodically compatible with the source audio track. Since it is an established hypothesis that individuals respond better to certain binaural "carrier" and "offset" frequencies on a subjective level than to other frequencies, this method increases the probability of successful EEG synchronization by increasing the number and diversity of source frequencies that produce the binaural beat effect.

[0007] In another embodiment, a source audio track is processed to generate one or more spatialized tracks. Each spatialized track is generated from one or more spatialized trajectories derived from specific audio features of the source audio track (e.g., tempo, musical event density, harmonic modes, loudness, percussive characteristics, etc.) and the target final state effect. Each spatialized track is configured so that the sound source is perceived as moving in time with the music. That is, each spatialized track provides tempo-based spatialization. Spatialization enhances the neurological effect by creating sustained attention by directing attention to a moving sound source. Furthermore, tracking a moving sound source activates spatial awareness.

[0008] In another embodiment, the source audio track is processed to generate both a binaural beat track and one or more spatialized tracks, as described above. Additional tracks, such as one or more infrasound tracks, ultrasonic tracks, enhanced bass tracks, and / or subharmonic tracks, may also be generated. The various tracks may be played simultaneously or mixed for the delivery of a single enhanced audio track to the end-user device. It should be noted that binaural beat stimulation and tempo-based spatialization are added to infrasound and ultrasonic track synthesis and enhanced bass and subharmonic processing techniques to provide a synergistic effect in enhancing neurological responses.

[0009] Various embodiments of the present invention are described in detail below, or will be apparent to those skilled in the art based on the disclosures provided herein, or can be learned from practicing the present invention. It should be understood that the above summary of the present invention is not intended to identify the main features or essential components of the embodiments of the present invention, nor is it intended to be used as an aid in determining the scope of the claimed subject matter described below. [Brief explanation of the drawing]

[0010] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0011] [Figure 1] This is a block diagram of one embodiment of an audio injection system that can be used to create an enhanced audio track according to the present invention. [Figure 2] Figure 1 is a block diagram of one embodiment of an audio distribution system that can be used to deliver enhanced audio tracks created by the audio injection system to end-user devices. [Figure 3] This is a flowchart of one embodiment of an audio injection process that may be carried out according to the present invention. [Figure 4] Figure 3 is an illustrative graph showing the filtering of outlier notes performed during the "note filtering" process. [Figure 5] This is a flowchart illustrating an exemplary process for generating spatialized tracks. [Figure 6] Figure 5 shows an example of a spatialized orbit generated by the orbital planner. [Figure 7] This is a flowchart illustrating an exemplary process for setting the gain of an audio track. [Figure 8A] Figure 3 shows the source audio track among the exemplary audio tracks at various stages of the audio injection process. [Figure 8B] This shows the individual tracks after the "gain processing" step, including four binaural beat tracks. [Figure 8C] This shows the individual tracks after the "gain processing" step, including four binaural beat tracks. [Figure 8D] This shows the individual tracks after the "gain processing" step, including four binaural beat tracks. [Figure 8E] This shows the individual tracks after the "gain processing" step, including four binaural beat tracks. [Figure 8F] This shows an ultra-low frequency track. [Figure 8G] Shows the ultrasonic track. [Figure 8H] Shows the subharmonic track. [Figure 8I] Shows the enhanced low-frequency track. [Figure 8J] Shows the spatialized track. [Figure 8K] Shows the final enhanced audio track.

Mode for Carrying Out the Invention

[0012] Detailed Description The present invention relates to a method and system for injecting sound into a source audio track to provide desired psychological, neurological, and / or physiological results. The present invention will be described in detail below with reference to various exemplary embodiments, but it should be understood that the present invention is not limited to the specific configurations or techniques of these embodiments. In addition, although the exemplary embodiments are described as embodying several different inventive features, those skilled in the art will understand that in the present invention, any one of these features can be implemented without the others.

[0013] In this disclosure, references to "one embodiment", "an embodiment", "an exemplary embodiment", or "embodiments" mean that one or more of the described features are included in at least one embodiment of the present invention. Separate references to "one embodiment", "an embodiment", "an exemplary embodiment", or "embodiments" in this disclosure do not necessarily refer to the same embodiment, and are not mutually exclusive, unless explicitly stated as such and / or are readily apparent to those skilled in the art from the description. For example, features, structures, functions, etc. described in one embodiment may be included in other embodiments, but not necessarily. Thus, the present invention may encompass various combinations and / or integrations of the embodiments described herein.

[0014] Generally, the methods of the present invention include a variety of digital signal processing steps performed to inject sound into a source audio track to create an enhanced audio track configured to achieve desired psychological, neurological, and / or physiological results.

[0015] The source audio track may include any type of audio recording containing vocals, instrumental music, and / or other types of sound, including a one-channel music track (mono), a two-channel music track (stereo), or a multi-channel music track (surround sound, spatial audio, etc.). In some embodiments, the source audio track is a professionally mixed and mastered stereo music track. Naturally, the present invention is not limited to enhancing commercial audio recordings; that is, any type of audio recording can be enhanced according to the teachings described herein.

[0016] Various types of sounds can be injected into the source audio track, including binaural beats, tempo-based spatialization, infrasound signals, ultrasonic signals, enhanced bass, subharmonics, and so on. It should be understood that in a particular implementation, any combination of these sounds may be used, including (a) binaural beats only, (b) tempo-based spatialization only, (c) binaural beats and tempo-based spatialization, (d) binaural beats and infrasound signals and / or ultrasonic signals, (e) tempo-based spatialization and infrasound signals and / or ultrasonic signals, (f) binaural beats, tempo-based spatialization, and infrasound signals and / or ultrasonic signals, and (g) any combination of the above with enhanced bass and / or subharmonics. Naturally, other combinations of sounds and additional sounds may also be used within the scope of the present invention.

[0017] Sound can be injected into a source audio track in various ways. One injection method involves processing the source audio track to create one or more audio tracks derived from it. Examples of audio tracks that can be created using this injection method include binaural beat tracks, spatialized tracks, enhanced bass tracks, and / or subharmonic tracks. Another injection method involves generating one or more synthesized audio tracks. Examples of audio tracks that can be created using this injection method include infrasound tracks and / or ultrasonic tracks. Naturally, both of the above injection methods may be used, and all audio tracks may be optionally mixed to provide a single enhanced audio track. Therefore, as used herein, “injecting” sound into a source audio track means processing the source audio track to create one or more audio tracks derived from it and / or generating one or more synthesized audio tracks, and optionally mixing audio tracks to provide a single enhanced audio track, all with the aim of obtaining a desired psychological, neurological, and / or physiological result.

[0018] The enhanced audio track, which contains injected sound, may be delivered to the end-user device by any known delivery method, such as streaming or downloading the audio file to the end-user device. The enhanced audio track may include a mixed audio track that can be played back through stereo headphones, earphones, earbuds, one or more speakers, etc. Alternatively, the enhanced audio track may include individual audio tracks that can be played back simultaneously through multiple speakers or multiple headphone drivers.

[0019] The method of the present invention may be carried out in conjunction with additional pre-processing steps performed prior to the above steps. For example, a medical professional or other operator may determine the settings of the input configuration used to create the sound to be injected. These settings may be selected so that the enhanced audio track provides the desired psychological, neurological, and / or physiological results, and in this respect, the settings may differ among different end users. As another example, the source audio track may be separated from other sound elements prior to the above steps; for example, a movie track may include music, voice, and / or sound effects, and the music may be separated from the voice and sound effects for enhancement as described herein. Other pre-processing steps will be obvious to those skilled in the art.

[0020] The method of the present invention may also be carried out in conjunction with additional post-processing steps performed after the above steps. For example, the enhanced audio track may be delivered to an end user via a production speaker system used for health or medical management. The end user's response to the enhanced audio track can then be measured for effectiveness by a healthcare professional. Other post-processing steps will be obvious to those skilled in the art.

[0021] Referring to Figure 1, one embodiment of an audio injection system that may be used to create an enhanced audio track according to the present invention is shown collectively as reference no. 100. The system 100 includes a file server 110 for storing source audio files, an application server 120 for running an audio injection application, a database server 130 for storing input configurations manually entered by an operator, a machine learning server 140 for storing input configurations generated by machine learning, a file server 150 for storing enhanced audio files, and a web server 160 for providing access to the enhanced audio files.

[0022] The file server 110 may include any suitable computer hardware and software known in the art configured for storing source audio files available for enhancement by the application server 120. The source audio files may be provided in any desired audio format. Examples of common audio formats include Advanced Audio Coding (AAC), Free Lossless Audio Codec (FLAC), Windows Media Video (WMV), Waveform Audio File Format (WAV), MPEG-1 Audio Layer 3 (MP3), and Pulse-Code Modulation (PCM). Naturally, other audio formats known in the art may also be used.

[0023] The application server 120 may include any suitable computer hardware and software known in the art, configured to receive source audio files from the file server 110, receive input configurations from the database server 130 and / or the machine learning server 140, run a sound injection application to inject sound into the source audio tracks to create enhanced audio tracks, and store the enhanced audio files in the file server 150. The steps performed by the sound injection application are described in more detail below with reference to Figures 3 to 7.

[0024] The database server 130 may include any suitable computer hardware and software known in the art configured to maintain a database for storing multiple input configurations manually entered by an operator. The data elements of each input configuration may include various different types of settings such as (a) binaural beat frequency (Hz), (b) binaural beat volume (in loudness unit full scale (LUFS) units), (c) infrasound frequency (Hz), (d) infrasound volume (in LUFS units), (e) ultrasound frequency (Hz), (f) ultrasound volume (in LUFS units), (g) subharmonic volume (in LUFS units), (h) low-frequency intensity (in LUFS units), and (i) low-frequency volume (in LUFS units). Naturally, other settings may also be stored in the database 130 in accordance with the present invention. In some embodiments, the operator is prompted to enter values ​​for the settings via a configuration file, but any data entry method known in the art may be used.

[0025] The machine learning server 140 may include any suitable computer hardware and software known in the art that is configured to maintain and provide a set of machine learning models used to generate input configurations. It should be understood that any kind of machine learning approach may be used, such as recommendation engines using supervised learning, reinforcement learning, etc. Generally, the machine learning server 140 monitors the quality of the output from the application server 120 and fine-tunes the configuration settings for optimal performance of the audio injection application.

[0026] It should be understood that input configuration settings for a specific end user or group of end users may be manually entered by an operator, determined by machine learning, or a combination of both. Recommended values ​​for these settings may differ among different end users; that is, the recommended settings for one individual may differ from those for another. These settings may be determined experimentally based on various factors, such as the final state effect of the goal to be achieved, the intended duration of exposure to the enhanced audio track, and whether the treatment plan requires the delivery of the same or different stimuli in different sessions. Therefore, the selection of settings can be understood as enabling the generation of enhanced audio tracks that are individualized or customized for the end user or group of end users, either manually or through a machine learning process.

[0027] The file server 150 may include any suitable computer hardware and software known in the art configured for storing enhanced audio files generated by the application server 120. The enhanced audio files may be provided in any desired audio format, such as AAC, FLAC, WMV, WAV, MP3, and PCM. Naturally, other audio formats known in the art may also be used.

[0028] The web server 160 may include any suitable computer hardware and software known in the art that is configured to handle requests to access (stream, download, etc.) enhanced audio files stored in the file server 150. As shown in Figure 2, the web server 160 connects to multiple end-user devices 1701-170 via the communication network 180. nIt can be used to deliver enhanced audio files to end-user devices 1701-170 via Hypertext Transfer Protocol (HTTP) (e.g., HTTP / 1.0, HTTP / 1.1 Plus, HTTP / 2, or HTTP / 3), Hypertext Transfer Protocol Secure (HTTPS), or any other network protocol used to distribute data and web content. n It can communicate with them.

[0029] End-user devices 1701-170 n Each of these may include a mobile phone, a personal computing tablet, a personal computer, a laptop computer, or any other type of personal computing device capable of communicating (wirelessly or wired) with the web server 160 via the communication network 180. In this embodiment, end-user devices 1701-170 n Each of these communicates with the web server 160 using an internet-enabled application (e.g., a web browser or an installed application). End-user devices 1701-170 n It should be understood that this can be used to play enhanced audio tracks through stereo headphones, earphones, earbuds, one or more speakers, etc. Alternatively, end-user devices 1701-170 n This can be used to store the enhanced audio track in memory (e.g., file copy, USB memory, etc.) for later playback on another device (e.g., MP3 player, iPod®, etc.).

[0030] Communication network 180 consists of web server 160 and end-user devices 1701-170. nIt may include any network or combination of networks that can facilitate data exchange between them. In some embodiments, the communication network 180 enables communication in accordance with one or more cellular standards, such as the Long-Term Evolution (LTE) standard, the Universal Mobile Telecommunications System (UMTS) standard, etc. In other embodiments, the communication network 220 enables communication in accordance with the IEEE 802.3 protocol (e.g., Ethernet) and / or the IEEE 802.11 protocol (e.g., Wi-Fi). Naturally, other types of networks may also be used within the scope of the present invention.

[0031] It should be understood that System 100 is an exemplary embodiment, and other embodiments may not include all the servers shown in Figure 1, or may include additional servers not shown in Figure 1. In particular, the system may be implemented with any number and combination of application servers, database servers, file servers, web servers, etc., which are either located in the same geographical location or in different geographical locations and connected to each other via the communication network 180. Thus, the system may be implemented with any number and combination of servers which are either located in the same location or geographically distributed.

[0032] Furthermore, in another embodiment, the sound injection process used to create the enhanced audio track according to the present invention is performed on a single computing device, such as a mobile phone, personal computing tablet, personal computer, laptop computer, or any other type of personal computing device. In this case, the various servers of System 100 shown in Figure 1 are not required. Thus, it can be understood that the computerized implementation of the present invention can be performed using any type of computing device known in the art, including one or a combination of servers, personal computing devices, etc.

[0033] Referring to Figure 3, one embodiment of a sound injection process that may be performed to create an enhanced audio track according to the present invention is shown collectively as reference no. 300. This process may be performed, for example, by the application server 120 by running the sound injection application described above, or by a personal computing device.

[0034] In step 302, the source audio track is provided as a digital audio track having one or more channels (e.g., mono, stereo, or surround sound). In some embodiments, a compressed or uncompressed audio file (e.g., an audio file provided in AAC, FLAC, WMV, WAV, MP3, or PCM audio format) is read and decoded into a two-channel digital audio track. For example, the application server 120 may read an audio file obtained from the file server 110, decode it, and provide a two-channel digital audio track. In other embodiments, an analog audio signal passes through an analog-to-digital converter to provide a digital audio track.

[0035] In step 304, the source audio track is separated into several audio tracks, commonly referred to as stem tracks. Each stem track may include a vocal track or an instrumental track, which may contain either a single instrument or a combination of individual instruments initially mixed together to create the stem track. In the exemplary embodiment shown in Figure 3, the source audio track is separated into four non-drum stem tracks: a vocal track, a bass track, a keyboard track, and a “other” pitched instrument track. Naturally, the number of stem tracks extracted from the source audio track will vary between implementations, and may include, for example, two stem tracks (vocal and instrumental) or five stem tracks (vocal, drums, bass, keyboard, and other pitched instruments). Other stem models are also possible and may be used within the scope of the present invention. The source audio track can be separated into several stem tracks using any music source separation tool known in the art, such as Spleeter, available from Deezer, located in Paris, France. In other embodiments, the source audio track is provided as a plurality of individual stem tracks, in which case the audio separation performed in step 304 is not necessary.

[0036] When a source audio track is provided as multiple audio tracks, subsequent processing of each audio track, including transcription (step 306), note filtering (step 308), and binaural beat synthesis (step 310), may be performed on a multiprocessor computer or in parallel across multiple computers. Although steps 306, 308, and 310 are described below in relation to the processing of a single stem track, it should be understood that these steps are performed in relation to each stem track that was separated in step 304 or provided as individual stem tracks in step 302.

[0037] In step 306, each stem track is automatically transcribed to provide information about the various notes contained within the stem track. As described below, the transcription may include, for example, an estimated fundamental frequency and / or estimated amplitude envelope for each note in the stem track. If the source audio track contains two or more channels (e.g., a stereo music track), the channels are mixed into mono during the transcription process. Those skilled in the art will understand that this mixing can be performed in various ways depending on the transcription tool used to perform the transcription. In some cases, the channels are mixed into mono, and then a transcription of a single mixed channel is generated. In other cases, each channel is automatically transcribed, and then the channel transcriptions are combined. Of course, the channels of the source audio track may alternatively be mixed into mono before source separation in step 304, but this approach is generally less efficient than performing the mixing using the transcription tool during the transcription process in step 306.

[0038] Depending on the source, either monophonic or polyphonic pitch transcription may be used to estimate the fundamental frequency of each note in the stem track. For example, monophonic pitch transcription may be used for an inherently monophonic source, and polyphonic pitch transcription may be used for an inherently polyphonic source. For example, in the exemplary embodiment shown in Figure 3, the vocal and bass stem tracks are automatically transcribed monophonically, while the keyboard and "other" pitched instrument stem tracks are automatically transcribed polyphonically.

[0039] Pitch transcriptions can be represented in various forms. For example, a pitch transcription may take the form of a note event list, such as a standard Musical Instrument Digital Interface (MIDI) file. In this case, each note is represented as a MIDI note number (i.e., pitch quantized to 12-tone equal temperament note numbers) for the time position of such a note. An example of a MIDI file is shown in Figure 4, which will be discussed below in relation to the note filtering process in step 308. As another example, a pitch transcription may take the form of a time-frequency matrix (i.e., a spectrogram). In both examples, an estimate of the fundamental frequency of each note detected in the stem track is provided.

[0040] The transcription may also include an estimated amplitude envelope for each note in the stem track. As is known in the art, the amplitude envelope for a note can be defined by four parameters: attack time, decay time, sustain level, and release time. During the transcription process, estimates of these four parameters may be obtained for each note in the stem track. Alternatively, a fixed amplitude can be used, where a single time-averaged amplitude value is provided in the input configuration settings. In this case, the transcription will include only frequency information and not amplitude information.

[0041] Each stem track can be automatically transcribed using any transcription tool known in the art, such as the Melodyne transcription tool available from Celemony Software GmbH in Munich, Germany. It should be understood that, of course, other transcription tools may also be used in this invention.

[0042] In step 308, the automatically transcribed notes of each stem track are filtered to remove any outlier notes resulting from errors in the pitch transcription process in step 306. One or more different note filtering methods may be used in this invention.

[0043] One note filtering method involves identifying outlier notes based on a statistical analysis of the stem track's pitch transcription. Specifically, the pitch transcription is analyzed, and the mean and standard deviation of the estimated fundamental frequencies of the auto-transcribed notes are calculated. Outlier notes are then identified as having an estimated fundamental frequency that is either above an upper frequency limit or below a lower frequency limit. The upper frequency limit has a value that includes the mean plus the first multiplier of the standard deviation, and the lower frequency limit has a value that includes the mean minus the second multiplier of the standard deviation. If the pitch transcription is represented as a MIDI file, outlier notes are removed from the note event list. On the other hand, if the pitch transcription is represented as a time-frequency matrix, the energy of outlier notes is set to zero.

[0044] The first and second multipliers used to determine the upper and lower frequency limits may be the same or different, respectively. Figure 4 shows an exemplary MIDI file where both the first and second multipliers are set to 1.5. In this case, notes with estimated fundamental frequencies above the upper frequency limit (i.e., mean + 1.5 × standard deviation) or below the lower frequency limit (i.e., mean - 1.5 × standard deviation) are removed from the music event note list. Naturally, in other implementations, the first and second multipliers may be different. Furthermore, the first and second multipliers may be the same for all stem tracks or may differ between different stem tracks.

[0045] In some embodiments, the first and second multipliers are determined by experimentation using multiple datasets of music selected to represent typical processing. Typically, the first and second multipliers have values ​​ranging from 1.25 (when more filtering is needed) to 2.5 (when less filtering is needed). In most cases, the first and second multipliers are pre-set and do not need to be adjusted when processing different music tracks. Of course, in other embodiments, the first and second multipliers may be part of the input configuration and are therefore input by the operator or determined through machine learning.

[0046] Another note filtering method involves estimating the musical key of the stem track and then filtering out all auto-transcribed notes whose fundamental frequencies fall outside the range of that musical key. The musical key can be estimated using any key estimation algorithm known in the art, such as that described in CLKrumhansl and EJ Kessler, Tracing the Dynamic Changes in Perceived Tonal Organization in a Spatial Representation of Musical Keys, Psychological Review, 89(4):334-368, 1982.

[0047] For example, if the musical key of the stem track is estimated to be A major, and the pitch transcription process in step 306 generates a note list of (A3, C3, C♯3, D3, E3, F3, G♯3), the note filtering process in step 308 may filter that note list to (A3, C♯3, D3, E3, G♯3). In this case, certain notes are filtered out so that the remaining notes match the central note of the stem track. For example, the pitches A3, C♯3, D3, E3, and G♯3 fit the musical key of A major, but the pitches C3 and F3 do not.

[0048] The music key estimation and filtering method can be used on its own or in combination with the statistical analysis method described above. Naturally, other note filtering methods known in the art can also be used within the scope of this invention.

[0049] It should be understood that the note filtering method described above may not be necessary if the error in the pitch transcription process of step 306 is low enough to be within an acceptable range. In this regard, step 308 is optional and may not be performed in all implementations.

[0050] In step 310, the filtered transcriptions of each stem track are used to synthesize the corresponding binaural beat track. The binaural beat track consists of a first binaural channel and a second binaural channel, each containing one or more sinusoidal signals derived from the auto-transcribed notes in the filtered transcription. Various methods may be used to generate these sinusoidal signals.

[0051] In the exemplary embodiment shown in Figure 3, the time-varying frequencies of the sinusoidal signals provided to the first and second binaural channels are determined by identifying a binaural beat frequency parameter, and then using this parameter in combination with the estimated fundamental frequency and temporal duration of each note in the filtered transcription to determine such frequencies.

[0052] The binaural beat frequency parameter is included in the input configuration described above (i.e., the frequency of the binaural beat (Hz)) and represents the magnitude of the frequency deviation from the estimated fundamental frequency of each auto-transcribed note. Specifically, the binaural beat frequency parameter is added to the estimated fundamental frequency to determine the corresponding frequency of the sine wave signal provided to the first binaural channel, and conversely, the binaural beat frequency parameter is subtracted from the estimated fundamental frequency to determine the corresponding frequency of the sine wave signal provided to the second binaural channel. For example, if the binaural beat frequency parameter is 2 Hz and the estimated fundamental frequency of the auto-transcribed note is 440 Hz, then the corresponding frequencies of the sine wave signals provided to the first and second binaural channels are 438 Hz and 442 Hz, respectively. This same approach is used for all auto-transcribed notes. Thus, the binaural effect is synthesized by the low-frequency difference between the sine wave signals provided to the first and second binaural channels.

[0053] For polyphonically processed stems, the filtered transcription may include multiple auto-transcribed notes occurring simultaneously; for example, a “other” pitched instrument stem track may include notes associated with two or more instruments. The estimated fundamental frequencies of these auto-transcribed notes are used to synthesize multiple (i.e., simultaneously emitted) sine wave signals. For monophonically processed stems, only one sine wave signal occurs at a time. Therefore, within the scope of the present invention, it can be understood that each of the first binaural track and the second binaural track may include one to multiple sine wave signals.

[0054] In the exemplary embodiment shown in Figure 3, the amplitudes of the sinusoidal signals provided to the first and second binaural channels can be determined by various methods. In some embodiments, the estimated amplitude envelopes of the notes in the filtered transcription are used to control the relative amplitude of the synthesized sinusoidal signals. For example, if the automatically transcribed amplitude changes over time, the amplitude of the binaural sinusoidal signals will also change over time. In other embodiments, a binary beat volume parameter included in the above input configuration (which defines a single time-averaged amplitude value) is used to determine the amplitude of the binaural sinusoidal signals. This approach may be understood as providing a concise method for determining the amplitude of sinusoidal signals that does not require transcription of the estimated amplitude envelopes of the notes.

[0055] In step 312, the gain of the binaural beat tracks (i.e., one binaural beat track corresponding to each music stem track) is adjusted using a gain setting process described below in relation to Figure 7. Alternatively, any gain adjustment method known in the art, including the use of a fixed gain, may be used to set the gain of the binaural beat tracks.

[0056] In step 314, one or more infrasound tracks and / or one or more ultrasonic tracks are synthesized, each containing a sinusoidal signal having a constant frequency and amplitude across the track. Unlike other tracks described herein, the infrasound tracks and / or ultrasonic tracks are not derived from the source audio track.

[0057] For each infrasound track, the frequency of the sine wave signal is determined from the frequency parameter of the infrasound included in the above input configuration (i.e., the frequency of the infrasound (Hz)), and the amplitude of the sine wave signal is determined from the volume parameter of the infrasound included in the above input configuration (i.e., the volume of the infrasound (LUFS units)). Typically, the infrasound frequencies have values ​​in the range of 0.25 Hz to 20 Hz. The selected frequencies are chosen to be complementary to the frequencies of the source audio track, either by manual selection or by using machine learning. The selected frequencies for each infrasound track may be chosen complementary by combining a reference to the listener's neurological and physiological sensitivity to the probe frequencies, as well as by a reference to the subharmonics of identified prominent frequencies in the source audio track. The balance between these contributions is determined either by manual selection or by using machine learning.

[0058] For each ultrasonic track, the frequency of the sinusoidal signal is determined from the ultrasonic frequency parameter (i.e., ultrasonic frequency (Hz)) included in the above input configuration, and the amplitude of the sinusoidal signal is determined from the ultrasonic volume parameter (i.e., ultrasonic volume (LUFS units)) included in the above input configuration. Typically, the ultrasonic frequency has a value in the range of 20 kHz to the Nyquist frequency (half the audio sample rate). The selected frequency is chosen to be complementary to the frequency of the source audio track, either by manual selection or by using machine learning. The selected frequency for each ultrasonic track may be chosen complementary by combining a reference to the listener's neurological and physiological sensitivity to the probe frequency, as well as a reference to the high-frequency harmonics of identified prominent frequencies in the source audio track. The balance between these contributions is determined either by manual selection or by using machine learning.

[0059] The number of infrasound tracks and / or ultrasonic tracks varies across different contexts; that is, some implementations synthesize one or more infrasound tracks, some synthesize one or more ultrasonic tracks, and some synthesize a combination of one or more infrasound tracks and one or more ultrasonic tracks. In each of these tracks, the input configuration can be understood to include frequency parameters and volume parameters.

[0060] In step 316, the gains of one or more infrasound tracks and / or one or more ultrasonic tracks are adjusted using a gain setting process described below in relation to Figure 7. Alternatively, any gain adjustment method known in the Art, including the use of fixed gains, may be used to set the gains of the infrasound tracks and / or ultrasonic tracks.

[0061] In step 318, the source audio track is filtered to remove non-bass frequencies. In the exemplary embodiment shown in Figure 3, a bandpass filter with a passband of 20 Hz to 120 Hz is used to remove non-bass frequencies. Naturally, other methods may be used instead, for example, source audio isolation may be used to remove non-bass instruments. Thus, the low-frequency components of the source audio track are provided for processing in steps 320 and 324.

[0062] In step 320, a subharmonic track is generated by synthesizing subharmonics of the low-frequency components of the musical audio using a subharmonic generator. Any subharmonic generator known in the art may be used to create subharmonic frequencies below the fundamental frequencies of the low-frequency components, typically using a harmonic division method. The subharmonic track may contain non-sinusoidal signals derived from a combination of frequencies extracted from the source audio track using spectral analysis, such as a Fast Fourier Transform (FFT). The volume (in LUFS units) of the subharmonics provided in the above input configuration is used to set the amplitude of the subharmonic track so that it is in the desired balance with other musical elements in the final mix track (e.g., not excessively strong, not inaudible, etc.).

[0063] In step 322, the gain of the subharmonic track is adjusted using a gain setting process described below in relation to Figure 7. Alternatively, any gain adjustment method known in the art, including the use of a fixed gain, may be used to set the gain of the subharmonic track.

[0064] In step 324, an enhanced bass track is generated by enhancing the low-frequency components of a musical audio using the process described in U.S. Patent No. 5,930,373, which is incorporated herein by reference. In this process, the bass intensity parameter (i.e., bass intensity in LUFS units) included in the above input configuration is used to determine the degree of bass enhancement of the track being processed, and the bass volume parameter (i.e., bass volume in LUFS units) included in the above input configuration is used to determine the loudness of the track being processed. Those skilled in the art will understand that bass enhancement is desirable for the greatest effect (e.g., power, warmth, clarity, etc.).

[0065] In step 326, the gain of the enhanced low-frequency track is adjusted using a gain setting process described below in relation to Figure 7. Alternatively, any gain adjustment method known in the art, including the use of a fixed gain, may be used to set the gain of the enhanced low-frequency track.

[0066] In step 328, the source audio track is filtered to remove low frequencies. In the exemplary embodiment shown in Figure 3, a high-pass filter is used to remove frequencies below 120 Hz. Naturally, other methods may be used as alternatives, for example, source audio isolation may be used to remove bass instruments. A low-pass filter may also be used to remove frequencies above 20 kHz to reduce high frequencies of sibilance sounds (e.g., cymbals) that may cause audible artifacts from the spatialization process. Then, in step 330, the spatialized track is generated by spatializing the non-low frequency components of the source audio track.

[0067] In step 332, the gain of the spatialized track is adjusted using a gain setting process described below in relation to Figure 7. Alternatively, any gain adjustment method known in the art, including the use of a fixed gain, may be used to set the gain of the spatialized track.

[0068] In step 334, the various audio tracks described above—namely, the binaural beat track, the infrasound track and / or the ultrasonic track, the subharmonic track, the enhanced bass track, and the spatialized track—are mixed to create an enhanced audio track. In an exemplary embodiment, mixing is performed by summing and averaging to create a two-channel digital audio track. Then, in step 336, the enhanced audio track is stored as a compressed or uncompressed audio file (e.g., an audio file provided in AAC, FLAC, WMV, WAV, MP3, or PCM audio format). For example, referring to Figure 1, the application server 120 may store the audio file in the file server 150. The audio file can then be streamed or downloaded to an end-user device for playback through stereo headphones, earphones, earbuds, one or more speakers, etc. It is also possible to create individual enhanced tracks (in contrast to the mixed track), and individual enhanced tracks can be played simultaneously by the end user.

[0069] Figures 8A to 8K show exemplary audio tracks at various stages of the sound injection process shown in Figure 3. Specifically, Figure 8A shows the source audio track obtained in step 302, Figure 8B shows the binaural beat track corresponding to the bass stem track after gain adjustment in step 312, Figure 8C shows the binaural beat track corresponding to the vocal stem track after gain adjustment in step 312, Figure 8D shows the binaural beat track corresponding to the keyboard stem track after gain adjustment in step 312, Figure 8E shows the binaural beat track corresponding to the "other" pitched instrument stem track after gain adjustment in step 312, Figure 8F shows the infra-frequency track after gain adjustment in step 316, Figure 8G shows the ultrasonic track after gain adjustment in step 316, Figure 8H shows the subharmonic track after gain adjustment in step 322, Figure 8I shows the enhanced bass track after gain adjustment in step 326, Figure 8J shows the spatialized track after gain adjustment in step 332, and Figure 8K shows the final enhanced audio track after mixing in step 334.

[0070] Referring to Figure 5, one embodiment of the process for generating spatialized tracks according to the present invention is shown overall as reference no. 500. This process can be carried out, for example, as steps 328 and 330 of the method shown in Figure 3.

[0071] In step 502, the source audio track is acquired. In step 504, one or more features are extracted from the source audio track, including (a) the tempo of the source audio track (beats per minute (BPM)), (b) the musical event density of the source audio track (events per second), (c) the harmonic modes of the source audio track (major, minor, etc.), (d) the loudness of the source audio track (in LUFS units), and / or (e) the percussive characteristics of the source audio track (using rising edge transient detection). These features may be extracted using any feature extraction algorithm known in the art and then supplied to the trajectory planner for processing in step 506, as described below.

[0072] Examples of tempo estimation software that can be used to estimate the tempo of a source audio track include Traktor software, available from Native Instruments GmbH in Berlin, Germany; Cubase software, available from Steinberg Media Technologies GmbH in Hamburg, Germany; Ableton Live software, available from Ableton AG in Berlin, Germany; and Hornet Songkey MK4 software, available from Hornet SRL in Italy. Tempo may also be estimated using the approach described in D. Bogdanov et al., Essentia: An audio Analysis Library for Music Information Retrieval, 14th Conference of the International Society for Music Information Retrieval (ISMIR), pages 493-498, 2013.

[0073] The music event density of the source audio track may be estimated using the approach described in D. Bogdanov et al., Essentia: An audio Analysis Library for Music Information Retrieval, 14th Conference of the International Society for Music Information Retrieval (ISMIR), pages 493-498, 2013.

[0074] Examples of harmonic mode estimation software that can be used to estimate the harmonic modes of a source audio track include Traktor software from Native Instruments GmbH in Berlin, Germany; Cubase software from Steinberg Media Technologies GmbH in Hamburg, Germany; Ableton Live software from Ableton AG in Berlin, Germany; Hornet Songkey MK4 software from Hornet SRL in Italy; and Pro Tools software from Avid Technology, Inc. in Burlington, Massachusetts.

[0075] An example of loudness estimation software that can be used to estimate the loudness of a source audio track is the WLM Loudness Meter software, available from Waves Audio Ltd., located in Tel Aviv, Israel.

[0076] The percussive characteristics of the source audio track can be estimated using the approach described in F. Gouyon et al., Evaluating Rhythmic Descriptors for Musical Genre Classification, Proceedings of the AES 25th International Conference, pages 196-204, 2004.

[0077] In step 508, the final state effect intended to be induced in the end user is obtained. Examples of final state effects include, but are not limited to, a calm-stable-controlled state, a motivated-stimulated-energetic state, a relaxed-calm-drowsy state, an undistracted-meditative-focused state, a comfortable-relieved-calm state, a recovered-returned state, and a driven-enhanced-flow state. The final state effect is then supplied to the trajectory planner for processing in step 506, which is described below.

[0078] In step 506, the trajectory planner generates a spatialized trajectory derived from audio features extracted from the source audio track and the target final state effect. The spatialized trajectory can be a two-dimensional (2D) or three-dimensional (3D) trajectory, depending on the implementation and desired effect. In particular, a two-dimensional trajectory allows control of the perceived distance and lateral deviation of the sound, while a three-dimensional trajectory also allows control of the trajectory height. Several different approaches may be used to generate the spatialized trajectory.

[0079] In some embodiments, audio features extracted from the source audio track and the target final state effect are used to index preset spatialization trajectories. Examples of four preset spatialization trajectories, each associated with the tempo of the source audio track and the target final state effect, are provided in Table 1 below.

[0080] [Table 1]

[0081] It should be understood that the preset spatialization trajectories provided in Table 1 are merely examples. Any number of preset spatialization trajectories may be used, each associated with one or more audio features extracted from the source audio track and the target final state effect.

[0082] In other embodiments, the spatialized trajectory is determined using a combined approach that includes a lookup from a set of preset spatialized trajectories and parameter modification of the selected trajectory using one or more audio features extracted from the source audio track. An exemplary algorithm using this approach is provided below. 1. A set of tempo range categories, music genre categories, and target final state effect categories is used to index the trajectory table, with each trajectory instance designated as a list of three-dimensional Vezier spline control points. 2. The percussive characteristic parameter is a binary value normalized to the range of -1.0 to 1.0. Therefore, it is used as a scaling parameter for the trajectory's orientation, controlling the specific lateral deviation of the trajectory relative to the listener's position. 3. The event density parameter is a binary value normalized to the range of -1.0 to 1.0. Therefore, it is used as a scaling parameter for the trajectory height, controlling the specific vertical deviation of the trajectory relative to the listener's position. 4. The harmonic mode parameters are converted to binary values, i.e., a binary value of "0" if the mode is one of the major modes (Ionian, Lydian, Mixolydian), or a binary value of "1" if the mode is one of the minor modes (Aeolian, Dorian, Phrygian, Locrian). If the binary value is "0", the trajectory is not filtered. However, if the binary value is "1", a low-pass filter is applied to each of the three dimensions of the trajectory, scaled by the cutoff frequency using the tempo parameter (measured in beats per minute) and multiplied by the loudness parameter, in order to smooth the trajectory to the extent determined by the musical characteristics. 5. The tempo parameter (measured in beats per minute) is used to calculate the specific speed of passage through the path in seconds.

[0083] An exemplary trajectory loop that can be generated using the above algorithm is shown in Figure 6. As can be seen, the tempo extracted from the source audio track determines the duration of the trajectory loop, which is 16 seconds in this example. The remaining extracted audio features determine the control points of the trajectory loop in terms of the degree of movement, the rate of change of the music's position, the complexity of the trajectory in terms of overlap, or the repetition of the path. Naturally, it should be understood that other algorithms will result in trajectory loops different from the example shown in Figure 6. Typically, multiple trajectory loops are generated for the source audio track, i.e., a trajectory loop for each music section of the source audio track. The trajectory loops generated in step 506 are sent to the spatialization processing unit for further processing in step 512.

[0084] In step 510, the source audio track is filtered to remove low frequencies. It should be understood that this step uses the same process as described above in relation to step 328 of the method shown in Figure 3. Subsequently, in step 512, the spatialization processing unit is used to spatialize the non-low frequency components of the source audio track to produce a spatialized track 514. In exemplary embodiments, the spatialization processing unit uses a head-related transfer function (HRTF) spatialization algorithm such as Dolby Atmos Personalized Rendering, available from Dolby Laboratories, Inc. in San Francisco, California, or dearVR Pro, available from Dear Reality in Düsseldorf, Germany, which uses each trajectory loop generated in step 506 to set the position of the sound in virtual space. It should be noted that the position of the sound may be animated with speed and distance determined from the tempo of the source audio track. Thus, the sound source is perceived to move in time with the music. That is, the spatialized track provides tempo-based spatialization as an added sound.

[0085] In some embodiments, the process may use more than one filter to provide multiple tracks to be spatialized simultaneously. For example, a first filter may remove frequencies below 120 Hz to provide a first track, a second filter may allow frequencies between 20 Hz and 120 Hz to pass through to provide a second track, a third filter may remove frequencies other than between 400 Hz and 1,000 Hz to provide a third track, and a fourth filter may remove frequencies other than between 1,000 Hz and 4,000 Hz to provide a fourth track. Different spatialization trajectories may be used to spatialize each of these tracks, thereby generating multiple spatialized tracks. Other examples will be obvious to those skilled in the art.

[0086] Referring to Figure 7, one embodiment of the process for setting the gain of an audio track according to the present invention is shown collectively as reference no. 700. This process may be performed, for example, instead of steps 312, 316, 322, 326, and 332 of the method shown in Figure 3. In this process, it can be seen that an iterative gain setting algorithm is used to set the gain of the audio track in order to ensure that mixing of the audio track does not result in overdriven signals.

[0087] First, a spatialized track is acquired in step 702, an enhanced low-frequency track is acquired in step 704, a subharmonic track is acquired in step 706, a binaural beat track is acquired in step 708, and an infrasound track and / or ultrasonic track is received in step 710.

[0088] In step 712, the spatialized track, the enhanced bass track, and the subharmonic track are mixed by gain. In step 714, the loudness of the mix is ​​estimated using any loudness estimation algorithm known in the art (e.g., WLM Loudness Meter software available from Waves Audio Ltd., Tel Aviv, Israel). In step 716, it is determined whether the loudness of the mix matches the loudness of the source audio track. If it does not, in step 718, the gains of the spatialized track, the enhanced bass track, and the subharmonic track are iteratively adjusted until their levels substantially match the loudness of the source audio track. Machine learning, including any multi-parameter estimation method such as stochastic gradient descent with a multilayer perceptron, may be used to determine the gain adjustments in each iteration. Of course, simpler methods such as hill-climbing estimation can also be used to adjust the gains.

[0089] If the loudness of the mix matches the loudness of the source audio track, in step 720, the final gain of the mix is ​​used to set the gains of the binaural beat track and the infrasound track and / or ultrasonic track. In an exemplary embodiment, the gains of the binaural beat track and the infrasound track and / or ultrasonic track are determined by multiplying the final gain of the mix by a constant, i.e., a fixed loudness ratio which can be derived by experimentation with an existing dataset or by data analysis. For example, if the final gain of the mix is ​​-6 dB and the constant is 9, the gains of the binaural beat track and the infrasound track and / or ultrasonic track are -54 dB.

[0090] In step 722, various audio tracks, namely binaural beat tracks, infrasound tracks and / or ultrasonic tracks, subharmonic tracks, enhanced bass tracks, and spatialized tracks, are mixed to create an enhanced audio track. Then, in step 724, the enhanced audio track is stored as an audio file. It should be understood that steps 722 and 724 use the same processes described above in relation to steps 334 and 336 of the method shown in Figure 3.

[0091] The above description provides several exemplary embodiments of the subject matter of the present invention. Although each exemplary embodiment represents a single combination of the elements of the invention, the subject matter of the present invention is considered to include all possible combinations of the disclosed elements. Thus, if one embodiment includes elements A, B, and C, and a second embodiment includes elements B and D, the subject matter of the present invention is considered to include other remaining combinations of A, B, C, or D, even if not expressly disclosed.

[0092] Any use of example or illustrative language (e.g., “e.g., “etc.”) provided in reference to a particular embodiment is intended solely to better illustrate the invention and not to limit its scope. Furthermore, the use of relative relational terms such as “first” and “second” is used merely to distinguish one unit or action from the other, and no actual relationship or order is required or suggested between such units or actions. Nothing expressed herein should be construed as indicating that any non-claimed element is essential for the implementation of the invention.

[0093] The use of the terms “comprises,” “comprising,” or other variations thereof is intended to include non-exclusive inclusion, so that a process, method, or system containing a list of elements does not contain only those elements, but may also contain other elements not expressly described in or inherent to such process, method, or system.

[0094] Although the present invention has been described and illustrated with respect to several exemplary embodiments as described above, it should be understood that various modifications can be made to these embodiments without departing from the scope of the invention. Accordingly, the present invention is not limited to any particular configuration or methodology of the exemplary embodiments, unless such limitations are contained within the following claims.

Claims

1. A computer-based method for injecting sound into an audio track, Obtaining the source audio track, Multiple binaural beat tracks, each of which corresponds to one of a plurality of audio tracks separated from the source audio track, and the synthesis of multiple binaural beat tracks, each generated by (a) automatically transcribing the audio track to provide a transcription that includes one or both of the estimated fundamental frequency and estimated amplitude envelope for each of the plurality of notes in the audio track, and (b) using the transcription to generate the binaural beat track. (a) filtering the source audio track to provide a filtered track; (i) extracting one or more audio features from the source audio track; (ii) identifying a target final state effect; (iii) using the audio features extracted from the source audio track and the target final state effect to determine a spatialized trajectory; (b) generating the spatialized trajectory; and (c) spatializing the filtered track using the spatialized trajectory to generate a spatialized track; To generate either or both ultra-low frequency tracks and ultrasonic tracks, A mixed track is generated by mixing the binaural beat track, the spatialized track, and one or both of the ultra-low frequency track and the ultrasonic track. To provide the aforementioned mixed track for distribution to end-user devices, Methods that include...

2. The process involves generating a low-frequency track and a subharmonic track based on the aforementioned source audio track, To generate the aforementioned mixed track, the low-frequency track and the subharmonic track are mixed with the binaural beat track, the spatialized track, and one or both of the infra-frequency track and the ultrasonic track. The method according to claim 1, further comprising:

3. A computer-based method for injecting sound into an audio track, Obtaining the source audio track, The invention provides a method for synthesizing multiple binaural beat tracks, each of which corresponds to one of a plurality of audio tracks separated from the source audio track, and where the binaural beat tracks are generated from the audio track based on either or both of the estimated fundamental frequency and estimated amplitude envelope for each of the plurality of notes automatically transcribed from the audio track. To provide a filtered track, the spatialized track is generated by filtering the source audio track, generating a spatialized trajectory based on one or more audio features and target final state effects extracted from the source audio track, and spatializing the filtered track using the spatialized trajectory to generate the spatialized track. Outputting the binaural beat track and the spatialized track, Methods that include...

4. The method according to claim 3, further comprising generating either or both an ultra-low frequency track and an ultrasonic track.

5. The method according to claim 4, further comprising generating one or both of a low-frequency track and a subharmonic track based on the source audio track.

6. A mixed track is generated by mixing the binaural beat track, the spatialized track, one or both of the ultra-low frequency track and the ultrasonic track, and one or both of the low frequency track and the subharmonic track. To provide the aforementioned mixed track for distribution to end-user devices, The method according to claim 5, further comprising:

7. A computer-based method for synthesizing binaural beat tracks, Multiple binaural beat tracks, each of which corresponds to one of multiple audio tracks separated from a source audio track, and the synthesis of multiple binaural beat tracks generated by (a) automatically transcribing the audio track to provide a transcription that includes one or both of the estimated fundamental frequency and estimated amplitude envelope for each of the multiple notes in the audio track, and (b) using the transcription to generate the binaural beat track. Outputting the aforementioned binaural beat track, Methods that include...

8. The method according to claim 7, further comprising separating the audio track from the source audio track.

9. The method according to claim 7, wherein each of the audio tracks includes one of an instrumental track, a mix of multiple instrumental tracks, a vocal track, or a mix of multiple vocal tracks.

10. The method according to claim 7, wherein the binaural beat track includes (a) one or more sinusoidal signals provided to a first binaural channel, and (b) one or more sinusoidal signals provided to a second binaural channel.

11. Identifying binaural beat frequency parameters that represent the magnitude of frequency deviation, To determine the multiple time-varying frequencies of the one or more sinusoidal signals provided to the first binaural channel, the binaural beat frequency parameter is added to the estimated fundamental frequency for each of the notes, In order to determine the multiple time-varying frequencies of the one or more sinusoidal signals provided to the second binaural channel, the binaural beat frequency parameter is subtracted from the estimated fundamental frequency for each of the notes, The method according to claim 10, further comprising:

12. Identifying binaural beat volume parameters, The binaural beat volume parameter is used to determine the amplitude of the one or more sinusoidal signals provided to the first binaural channel and the one or more sinusoidal signals provided to the second binaural channel, The method according to claim 10, further comprising:

13. The method according to claim 10, further comprising using the estimated amplitude envelope of each of the notes to determine a plurality of time-varying amplitudes of the one or more sinusoidal signals provided to the first binaural channel and the one or more sinusoidal signals provided to the second binaural channel.

14. The method according to claim 7, wherein the transcription includes the estimated fundamental frequency of each of a plurality of outlier notes, and the method is Before generating the binaural beat track, the outlier notes are filtered from the transcription. The method according to claim 7, further comprising:

15. The outlier notes in the above transcription, To determine the average frequency and standard deviation from the average frequency of multiple automatically transcribed notes, the transcription is analyzed, Identifying the automatically transcribed note such that the estimated fundamental frequency is either (a) greater than a first value including the average frequency plus a first multiplier of the standard deviation, or (b) less than a second value including the average frequency minus a second multiplier of the standard deviation, The method according to claim 14, further comprising specifying by

16. In order to provide filtered tracks, the source audio track is filtered, To generate a spatialized track, the filtered track is spatialized using a spatialized trajectory, Outputting the spatialized track, The method according to claim 8, further comprising:

17. The method according to claim 16, further comprising providing the binaural beat track and the spatialized track in order to enable simultaneous playback of the spatialized track and the binaural beat track by a listener.

18. To generate a mixed track, the binaural beat track and the spatialized track are mixed, The method according to claim 16, further comprising providing the mixed track for distribution to an end-user device.

19. The process involves iteratively adjusting a first gain value associated with the spatialized track until the loudness of the spatialized track matches the loudness of the source audio track. The second gain value associated with the binaural beat track is set based on the first gain value and a fixed ratio. To modify the amplitude of the spatialized track and the binaural beat track, respectively, the first gain value and the second gain value are used, The method according to claim 16, further comprising:

20. A computer method for generating spatialized tracks, Obtaining the source audio track, In order to provide filtered tracks, the source audio track is filtered, (a) extracting one or more audio features from the source audio track; (b) identifying the target final state effect; and (c) using the audio features extracted from the source audio track and the target final state effect to determine the spatialized trajectory, thereby generating the spatialized trajectory. To generate a spatialized track, the filtered track is spatialized using the spatialized trajectory, Outputting the spatialized track, Methods that include...

21. The method according to claim 20, wherein a plurality of spatialized trajectories are generated and used to spatialize the filtered track.

22. The method according to claim 20, wherein each of the one or more audio features extracted from the source audio track includes one of tempo, musical event density, harmonic modes, loudness, and percussive characteristic measurement.

23. The method according to claim 20, wherein the spatialized orbit includes a two-dimensional orbit.

24. The method according to claim 20, wherein the spatialized orbit includes a three-dimensional orbit.

25. The aforementioned spatialized orbit is This involves accessing a database that stores multiple preset spatialization trajectories, each associated with one or more audio features and final state effects, From the database, select a preset spatialized trajectory in which one or more audio features extracted from the source audio track match one or more audio features of the preset spatialized trajectory, and the target final state effect matches the final state effect of the preset spatialized trajectory. The method according to claim 20, as determined by...

26. The aforementioned spatialized orbit is further, Modify the selected preset spatialized trajectory based on one or more of the audio features extracted from the source audio track. The method according to claim 25, as determined by...

27. Multiple binaural beat tracks, wherein each of the multiple binaural beat tracks is a composite of multiple binaural beat tracks corresponding to one of multiple audio tracks separated from the source audio track, Outputting the aforementioned binaural beat track, The method according to claim 20, further comprising:

28. The method according to claim 27, further comprising providing the spatialized track and the binaural beat track to enable simultaneous playback of the spatialized track and the binaural beat track by a listener.

29. To generate a mixed track, the spatialized track and the binaural beat track are mixed, The method according to claim 27, further comprising providing the mixed track for distribution to an end-user device.

30. The process involves iteratively adjusting a first gain value associated with the spatialized track until the loudness of the spatialized track matches the loudness of the source audio track. The second gain value associated with the binaural beat track is set based on the first gain value and a fixed ratio. To modify the amplitude of the spatialized track and the binaural beat track, respectively, the first gain value and the second gain value are used, The method according to claim 27, further comprising: