Generation of timbrally harmonic synchronized neural beats for digital audio files
Patent Information
- Application Number
- JP2024524489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-21
- Filing Date
- 2022-10-21
- Publication Date
- 2025-10-23
AI Technical Summary
Existing audio tracks lack the ability to seamlessly integrate neural beats, which can enhance concentration and relaxation, and users often find standalone neural beat tracks boring or distracting, limiting their effectiveness.
A system and method to analyze digital audio files using chromagram features to synchronize neural beats with the audio's pitch and rhythm, adjusting volume and frequency to ensure harmonious integration.
Enables users to enjoy their favorite music while experiencing the benefits of neural entrainment, ensuring the neural beats are harmonious and effective without distracting from the original audio.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] This application relates to the generation of timbrally-harmonic synchronized neural beats for digital audio files. [Background technology]
[0002] In acoustics, a beat is an interference pattern between two sounds of slightly different frequencies, perceived as a periodic variation in volume whose rate is the difference between the two frequencies. In the case of a monaural beat, the listener hears the two different frequencies in the same or both ears. In a binaural beat, the two different frequencies are heard through different ears, for example using headphones or specially placed speakers, and the interference pattern is detected by the listener's brain. More complex interference patterns involving multiple signals can also be used to create a variety of beats.
[0003] Certain types of beats (e.g., mono beats, binaural beats) can be used to promote a desired mental state (e.g., improve an individual's attention or focus). For example, such beats can be used to create neural entrainment in a user who hears the beat, helping the user to better focus or concentrate. Often, these beats can be provided as separate audio tracks, such as an audio track that includes only beats. Alternatively, audio tracks may be prepared with custom additions of mono or binaural beats to the track (i.e., audio tracks that are configured or generated to include mono or binaural beats). In some cases, beats may even be provided that are not within a frequency range that is normally audible. Summary of the Invention
[0004] The present disclosure presents a novel and innovative system and method for generating and adding neural beats to an existing audio track. In a first aspect, the present invention provides a method including receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file, and extracting a plurality of chromagram features of the digital audio file according to a plurality of parameters. The method includes combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file, and extracting from the primary chromagram features dominant pitch classes at a plurality of timestamps in the digital audio file. A plurality of carrier frequencies for the neural beats are selected based on the dominant pitch classes at the plurality of timestamps, and a synchronized neural beat for the digital audio file is synthesized based on the beat frequency and the plurality of carrier frequencies. The method further includes storing at least one of (i) the synchronized neural beats, and (ii) a combined audio track combining the synchronized neural beats and the digital audio file.
[0005] In one embodiment according to the invention in its first aspect, the primary chromagram features include an intensity of each of a plurality of pitch classes at a plurality of timestamps. In one embodiment, a dominant pitch class is selected from among the plurality of pitch classes. In a further embodiment, extracting the dominant pitch classes further comprises generating a probability distribution for each of the plurality of pitch classes at the plurality of timestamps based on the intensities of the plurality of pitch classes using a hidden Markov model. In yet another embodiment, the hidden Markov model is configured to optimize the number and positions of transitions between the dominant pitch classes. In an alternative embodiment, extracting the dominant pitch classes further comprises identifying a sequence of dominant pitch classes within the probability distribution.
[0006] In one embodiment according to the present invention in its first aspect, the multiple timestamps occur every 500 milliseconds or less during the digital audio file.
[0007] In one embodiment according to the present invention in a first aspect, a plurality of chromagram features are linearly combined to form a primary chromagram feature.
[0008] In one embodiment according to the invention in its first aspect, the method further includes adjusting a volume of the synchronized neural beat to track a volume of the audio encoded in the digital audio file over time. In a further embodiment, normalizing the volume of the synchronized neural beat includes generating a loudness profile for a duration of the audio encoded in the digital audio file and forming a volume curve based on the loudness profile. In one embodiment, the method includes adjusting the volume of the synchronized neural beat according to the volume curve.
[0009] In an embodiment according to the present invention in its first aspect, the method further comprises aligning beat frequencies with rhythmic beats in the digital audio file. In an embodiment, aligning beat frequencies comprises estimating positions of rhythmic beats in the digital audio file, estimating a musical tempo in the digital audio file, and adjusting timing of the synchronized neural beats to align peak values in the synchronized neural beats with positions of the rhythmic beats in the digital audio file according to the musical tempo. In an embodiment, the neural beats are at least one of (i) binaural beats and (ii) mono beats.
[0010] In one embodiment according to the present invention in its first aspect, the synchronized neural beats include no more than two audio channels.
[0011] In one embodiment according to the present invention in its first aspect, the synchronized neural beats include three or more audio channels.
[0012] In one embodiment according to the invention in the first aspect, the beat frequency is greater than or equal to 0.5 Hz and less than or equal to 150 Hz.
[0013] In one embodiment according to the invention in its first aspect, the method further includes playing the synchronized neural beats and the digital audio file in parallel via the computing device, hi one embodiment, the method further includes streaming the synchronized neural beats and the digital audio file to the computing device for playback by the computing device.
[0014] The embodiments according to the invention in the first aspect described above are not mutually exclusive, unless the description specifically and explicitly discloses otherwise. In other embodiments according to the invention in the first aspect, features of one embodiment according to the invention in the first aspect are combined with combinations of features of another embodiment according to the first aspect.
[0015] In a second aspect, a system is provided that includes a processor and a memory. The memory may store instructions that, when executed by the processor, cause the processor to perform the method according to the present invention in the first aspect. In one embodiment, the instructions, when executed by the processor, cause the processor to receive a digital audio file and a beat frequency of a neural beat to be added to the digital audio file, and extract a plurality of chromagram features of the digital audio file according to a plurality of parameters. The instructions may also cause the processor to combine the plurality of chromagram features to form a primary chromagram feature of the digital audio file, extract from the primary chromagram features dominant pitch classes at a plurality of timestamps in the digital audio file, and select a plurality of carrier frequencies for the neural beat based on the dominant pitch classes at the plurality of timestamps. The instructions may further cause the processor to synthesize a synchronized neural beat for the digital audio file based on the beat frequency and the plurality of carrier frequencies, and store at least one of (i) the synchronized neural beat and (ii) a combined audio track that combines the synchronized neural beat and the digital audio file.
[0016] In one embodiment according to the present invention in a second aspect, the primary chromagram features include intensities of each of a plurality of pitch classes at a plurality of time stamps. A dominant pitch class may be selected from among the plurality of pitch classes.
[0017] In an embodiment according to the second aspect, the memory stores further instructions that, when executed by the processor while extracting the dominant pitch class, cause the processor to generate, using a hidden Markov model, a probability distribution for each of a plurality of pitch classes at a plurality of timestamps based on the intensities of the plurality of pitch classes.
[0018] The embodiments according to the invention in the second aspect described above are not mutually exclusive, unless the description specifically and explicitly discloses otherwise. In other embodiments according to the invention in the second aspect, features of one embodiment according to the invention in the second aspect are combined with combinations of features of another embodiment according to the second aspect.
[0019] An embodiment according to the present invention in the first aspect can be combined with an embodiment according to the present invention in the second aspect. The features and advantages described herein are not all-inclusive, and in particular many additional features and advantages will be apparent to those skilled in the art in view of the drawings and description. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and explanatory purposes, and is not intended to limit the scope of the disclosed subject matter. [Brief description of the drawings]
[0020] [Figure 1A] FIG. 1 illustrates a system according to an exemplary embodiment of the present disclosure. [Figure 1B] FIG. 1 illustrates a system for audio playback according to an exemplary embodiment of the present disclosure. [Diagram 2] FIG. 1 illustrates a chromagram feature according to an exemplary embodiment of the present disclosure. [Diagram 3] FIG. 2 illustrates a diagram showing dominant pitch classes according to an exemplary embodiment of the present disclosure. [Figure 4] FIG. 2 illustrates a diagram showing selected carrier frequencies according to an exemplary embodiment of the present disclosure. [Diagram 5] FIG. 2 illustrates a volume curve according to an exemplary embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates a method for synthesizing neural beats according to an exemplary embodiment of the present disclosure. [Figure 7A] FIG. 1 illustrates a method according to an exemplary embodiment of the present disclosure. [Figure 7B] FIG. 1 illustrates a method according to an exemplary embodiment of the present disclosure. [Figure 7C]FIG. 1 illustrates a method according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates a computing system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] For purposes of this application, "neural beats" are acoustic beats added to an audio signal. Neural beats do not have to be in the audible frequency range. All references to neural beats in this application are intended to include either or both mono and binaural beats, unless the context clearly limits the discussion to one or the other.
[0022] A "neural beat" may include an audio beat designed to generate or promote a desired mental state of a user. The desired mental state may include neural entrainment, enhanced focus, a calm mood, relaxation, or any other desired mental state. In certain implementations, the neural beat may include a mono or binaural beat that combines a lower beat frequency with a higher carrier frequency. In particular, the "beat frequency" may be selected based on the desired mental state (e.g., where different frequencies promote different types of mental states in an individual). In one embodiment, the beat frequency may range from 0.5 to 150 Hz. In other exemplary embodiments, the frequency may be selected within a range between 1 to 100 Hz, between 1 to 10 Hz, between 10 to 100 Hz, or between 10 to 40 Hz. The "carrier frequency" may be an audio frequency or sound selected to carry or audibly reproduce the beat frequency within an audio track. For example, the beat frequency may be a frequency lower than humans can detect and / or may be in the lower range of human hearing. Thus, to maximize the effectiveness of the neural beats, a carrier frequency of an existing audio signal, such as an existing music or other audio recording, or a particular track in such a recording, can be selected, and the beat frequency can be modulated onto the carrier frequency to form the neural beats. The carrier frequency can range from 207.65 to 392.00 Hz. In various implementations, the neural beats can have a different number of audio channels, such as one audio channel (e.g., a mono beat), two audio channels (e.g., a binaural beat), five audio channels, or more.
[0023] Not all users enjoy listening to audio tracks that contain only neural beats, finding them boring or distracting, limiting the effectiveness of neural entrainment. Furthermore, the limited availability of existing audio tracks with embedded mono beats may not appeal to all users. Certain systems can automatically generate music incorporating mono beats so that users do not have to listen to the same track multiple times. However, even such systems cannot compensate for the possibility that a user may want to listen to a particular track or genre that has not been previously combined with neural beats. Thus, there is a need to automatically add neural beats to existing audio tracks so that users can listen to their favorite tracks or music genres while also experiencing the benefits of neural entrainment, relaxation, and / or increased focus provided by neural beats.
[0024] One solution to this problem is to analyze the pitch characteristics of the digital audio file over time. In particular, chromagram features indicate the intensity of different pitch classes over time of the audio signal. Chromagram features can be generated for a digital audio file and indicate the intensity of different pitch classes over time of the digital audio signal encoded within the digital audio file. This information can then be used to select carrier frequencies for neural beats to be added to the digital audio file. For example, dominant pitch classes may be extracted from the chromagram features at various timestamps within the digital audio file, and the dominant pitch classes may be used to select carrier frequencies for the neural beats at the various timestamps. In certain cases, the dominant pitch classes can be analyzed using a model (e.g., a hidden Markov model) to select carrier frequencies and optimize the number of changes in carrier frequency, e.g., to minimize the number of changes while still achieving a certain accuracy. Neural beats can then be synthesized based on the beat frequencies and the selected carrier frequencies and stored for later use. In certain cases, a combined audio track can be generated that combines the digital audio file with the neural beats. In other cases, the neural beats may be stored in association with the digital audio file. Additionally, in certain cases, the neural beats and / or the combined audio track may be generated in real time as a user device streams the digital audio file, such as by a server from which the digital audio file is streamed or by a user device receiving the streamed digital audio file. The neural beats may then be played along with the digital audio file via the user device (e.g., as separate audio files played simultaneously and / or as a single audio file).
[0025] FIG. 1A illustrates a system 100 according to an exemplary embodiment of the present disclosure. The system 100 may be configured to generate and synchronize neural beats for addition to digital audio files. The system 100 includes a computing device 102 and a server 104. The server 104 stores digital audio files 108, 110 to which neural beats may be added by the computing device 102. For example, the computing device 102 and the server 104 may be part of a digital audio streaming platform configured to stream the digital audio files 106, 108, 110 upon request of a user. Furthermore, the computing device 102 may be configured to add neural beats 168, 174 to the digital audio files 106, 108, 110 upon request of a user. For example, a user may manipulate a preference for adding neural beats 168, 174 to streamed audio files received from the audio streaming platform.
[0026] The computing device 102 can receive the digital audio file 106 from the server 104 and can generate neural beats 168 and / or adjusted neural beats 174 to be added to the digital audio file 106. The computing device 102 can also receive beat frequencies 112 for the neural beats 168, 174. The beat frequencies 112 received from a user, such as via a user-configurable beat frequency setting. The neural beats 168, 174 can be mono beats, binaural beats, or have more audio channels, and the type of neural beats 168, 174 can be selected by the user. Additionally or alternatively, the computing device 102 can select between mono beats and binaural beats based on the audio device from which the user is streaming the digital audio file. For example, if a user is streaming audio from a mono audio device, the computing device 102 may generate mono neural beats, and if a user is streaming audio from a stereo audio device (e.g., stereo speakers, stereo headphones), the computing device 102 may generate binaural neural beats. In yet another implementation, the computing device 102 may select the number of audio channels to be the same as the number of audio channels in the digital audio file 106.
[0027] The computing device 102 can be configured to, among other things, generate neural beats 168, 174 that blend into the digital audio file 106. For example, the computing device 102 can be configured to generate neural beats 168, 174 that are synchronized with the audio pitch in the digital audio file 106 to avoid noticeable and distracting pitch differences that may interfere with a user's neural entrainment. To do so, the computing device 102 can extract a number of chromagram features 116 from the digital audio file 106. The chromagram features 116 can include pitch classes 124, 126 and associated intensities 136, 138 at a number of timestamps 148, 150.
[0028] For example, FIG. 2 illustrates a chromagram feature 200 according to an exemplary embodiment of the present disclosure. The chromagram feature 200 includes intensities (as defined in legend 202) of multiple pitch classes at multiple timestamps T1-T19. The pitch classes include B, A sharp / B flat, A, G sharp / A flat, G, F sharp / G flat, F, E, D sharp / E flat, D, C sharp / D flat, and C, which represent each of the types of notes that may be reproduced within the digital audio file 106. In particular, each pitch class may represent all audible pitches of a song, separated by integer octaves. For example, pitch class C may include middle C, treble C, high C, tenor C, low C, and other octaves of note C. Other pitch classes may be defined to include multiple notes of different octaves as well. In practice, a pitch class may be defined as a collection of frequency bands. For example, pitch class C may be defined as 261.626±0.1 Hz (for middle C), 523.251±0.1 Hz (for tenor C), and similarly for other notes included within the pitch class. As shown, certain sharp or flat notes (e.g., A sharp, B flat, G sharp, A flat, F sharp, G flat, D sharp, E flat, C sharp, D flat) are grouped into a pitch class separate from the pitch class that includes the natural notes A-G. In additional or alternative implementations, pitch classes may be defined to include sharp or flat versions of notes. Similarly, particular implementations may define pitch classes differently (e.g., to include any desired combination of notes). For example, the pitch class of C may include middle C sharp or middle C flat in alternative implementations. Those skilled in the art will appreciate that the chromagram features 200 may be calculated according to any of a number of possible pitch classes, such as equal-tempered tuning (e.g., equal-tempered with 24 tones in 24 pitch classes, equal-tempered with 19 tones in 19 pitch classes, and / or equal-tempered with 7 tones in 7 pitch classes).In practice, the computing device 102 may calculate many more pitch classes than are represented in the chromagram feature 200, and the pitch classes may be combined into the desired pitch classes for the chromagram feature 200. For example, the computing device 102 may calculate 36 pitch classes, which are then combined into the pitch classes indicated for the chromagram feature 200.
[0029] The chromagram feature 200 includes an intensity of each pitch class at each of the timestamps T1-T19. These intensities change over time (e.g., as the music changes within the digital audio file 106). For example, pitch classes A and D both have high intensities at times T1-T5. At times T6-T10, the pitch classes with the highest intensities alternate between C and C sharp / D flat (T8, T12), D (T9-T10, T13, T17-T18), D and D sharp / E flat (T6, T14), E (T7, T15), E and D sharp / E flat (T11, T19), and F (T16). These intensities can be calculated based on an analysis of the frequency domain of the digital audio file 106 at each of the timestamps T1-T19. For example, the computing device 102 can divide the digital audio file 106 into segments for each of the timestamps T1-T19. The computing device 102 may then compute a time-frequency representation (e.g., a frequency distribution at multiple times) for each of the segments (e.g., by performing a Fourier transform, a Fast Fourier Transform (FFT), a Constant-Q transform, a wavelet transform, using a filter bank, etc.). The frequencies of the time-frequency representation may correspond to or be classified into each of the pitch classes (e.g., according to a predetermined frequency band). The intensity of each of the pitch classes may then be computed based on the intensity of the corresponding frequency in the time-frequency representation. This process may be repeated multiple times for the segments corresponding to each of the timestamps T1-T19. In a particular implementation, the timestamps T1-T19 may occur every 50 milliseconds. In additional or alternative implementations, the timestamps T1-T19 may occur more frequently (e.g., every 10 milliseconds, every 5 milliseconds, every 1 millisecond) and / or less frequently (e.g., every 0.5 seconds, every 0.25 seconds, every 0.1 seconds).In certain implementations, rather than performing a frequency domain analysis of the digital audio file 106, the computing device 102 may perform the analysis in the time domain. For example, a filter bank may be used with one or more filters for each pitch class. The intensity of the resulting filtered signal at each timestamp may then be used to determine the intensity of the chromagram feature 200.
[0030] Returning to FIG. 1A , the computing device 102 may compute multiple chromagram features 116 for the digital audio file 106. For example, the multiple chromagram features 116 may be computed to focus on different frequency ranges within the digital audio file 106. As a particular example, a first set of chromagram features may be computed to focus on a lower frequency range (e.g., C4, or below 261.62 Hz) within the digital audio file 106, and a second set of chromagram features may be computed to focus on a higher frequency range (e.g., C1-C8, or 32.70 Hz-4186.01 Hz). In such instances, the computing device 102 may be configured to combine the multiple chromagram features 116 into a set of primary chromagram features 118 for the digital audio file 106. For example, the computing device 102 may linearly combine the chromagram features 116 (e.g., according to predetermined weights) to form the primary chromagram features 118. The data structure of the primary chromagram features 118 may correspond to the data structure of the chromagram features 116. For example, in a particular implementation, the chromagram features 200 may represent a set of primary chromagram features 118 of the digital audio file 106. Additionally, while FIG. 2 illustrates the chromagram features 200 as a plot of data over time, it should be understood that in practice the chromagram features 116 and / or the primary chromagram features 118 may be stored in additional or alternative data structures. For example, the chromagram features 116 and / or the primary chromagram features 118 may be stored as an array including intensity values of pitch classes at timestamps T1-T19.
[0031] The computing device 102 can identify a dominant pitch class 120 based on the primary chromagram features 118. The dominant pitch class has the highest intensity at a particular time or over a time interval of the audio signal. In particular, the computing device 102 can calculate a probability distribution 144, 146 for each of the pitch classes 132, 134 of the dominant pitch class for a particular timestamp 156, 158. For example, FIG. 3 illustrates a dominant pitch class 300 according to an exemplary embodiment of the present disclosure. The dominant pitch class 300 includes a probability (as defined in legend 302) for each of pitch classes B, A-sharp / B-flat, A, G-sharp / A-flat, G, F-sharp / G-flat, F, E, D-sharp / E-flat, D, C-sharp / D-flat, C at each of timestamps T1-T19. In particular, at times T1-T5, pitch classes A and D have a moderately high probability, at times T6 and T14, pitch classes D and D-sharp / E-flat have a moderately high probability, at times T7, T11, and T19, pitch classes E and D-sharp / E-flat have a moderately high probability, at times T8 and T12, pitch classes C and C-sharp / D-flat have a moderately high probability, at times T9, T10, T13, T17, and T18, have a high probability, at time T15, pitch class E has a high probability, and at time T16, pitch class F has a high probability. The probabilities can be calculated to reflect the probability that each pitch class represents a dominant pitch class at a given time point. For example, in a particular case, the probabilities can be calculated by a Hidden Markov Model (HMM). In certain cases, the HMM can be tuned to optimize the number of transitions in dominant pitch classes (e.g., to optimize the number of changes in carrier frequency of neural beats 168, 174) that a user may find distracting and / or that may adversely affect neural entrainment.
[0032] Returning to FIG. 1A, the computing device 102 may determine the carrier frequency 114 based on the dominant pitch class 120. The carrier frequency 114 may include a single selected frequency 160, 162 at each timestamp 164, 166 that serves as the carrier frequency at that time within the neural beat 168, 174. For example, FIG. 4 illustrates a carrier frequency 400 according to an exemplary embodiment of the present disclosure. The carrier frequency 400 includes a single selected pitch class at each timestamp T1-T19. In particular, pitch class D is selected as the carrier frequency for timestamps T1-T14 and T17-T19, and pitch class E is selected for timestamps T15-T16. The carrier frequency may be selected to follow the musical harmony of the digital audio file while avoiding unnecessary changes in the carrier frequency. In particular, the carrier frequency at times T9-T10 and T17-T20 may be selected as pitch class D to align with the dominant pitch class at these times. However, since excessive changes in carrier frequency can be distracting to a user, the selected carrier frequency can be selected to maintain consistency over time in certain cases, such as when selecting between different pitch classes with similar probability or small, short changes in dominant pitch class. For example, in the dominant pitch class 300, pitch classes A and D had similar probability at times T1-T5. However, to avoid a transition from pitch class A to pitch class D at time T6, where pitch class D is dominant, pitch class D may be selected as the carrier frequency at times T1-T5. As another example, at times T7, T11, T19, pitch classes D sharp / E flat and E both have similar probability. However, to reduce the number of carrier frequency changes (e.g., since pitch class D still has a medium probability in the dominant pitch class 300), pitch class D may be selected as the carrier frequency even though it does not have the highest probability at these times. On the other hand, not following musical harmony may have a negative effect on neural entrainment.Thus, at times T17-T19, the carrier frequency switches from E (at time T16) to D (at times T17-T19) to properly track the harmonies in the digital audio file.
[0033] To select the carrier frequency 400, the computing device 102 may be configured to balance maximizing the overall probability of the selected carrier frequency while limiting the number of successive dominant pitch class changes. In a particular implementation, the computing device 102 may perform Viterbi decoding on the dominant pitch classes 300 to find the most likely sequence of individual pitch classes at each timestamp that limits the number of carrier frequency transitions, while ensuring that the carrier frequency 400 is musically aligned with the digital audio file 106.
[0034] 1A , the computing device 102 can synthesize the neural beat 168 based on the beat frequency 112 and the carrier frequency 114. In particular, the computing device 102 can synthesize the neural beat 168 by modulating the beat frequency 112 to the selected carrier frequency 160, 162 at each of the timestamps 164, 166. In certain implementations, the timestamps 164, 166 (e.g., timestamps T1-T19) in the carrier frequency may correspond to timestamps of the audio data in the digital audio file 106. In such instances, the computing device 102 can synthesize the neural beat 168 based directly on the carrier frequency at each of the timestamps 164, 166.
[0035] In certain implementations, the computing device 102 may further adjust one or more aspects of the neural beats 168 based on additional characteristics of the digital audio file 106. For example, the computing device 102 may adjust the volume of the neural beats 168 to match changes in the volume of the digital audio file 106. In particular, if the neural beats 168 are relatively quiet compared to the digital audio file 106, the benefits of the neural beats may be diminished. Additionally or alternatively, if the neural beats 168 are loud relative to the digital audio file 106, the neural beats 168 may prove confusing or distracting to the user, and the benefits provided by the neural beats 168 are interrupted. Thus, the audio mixer 122 may be used to adjust the volume of the neural beats 168 over the course of the digital audio file 106.
[0036] In particular, the audio mixer 122 can determine a loudness profile 170 for the digital audio file 106. The loudness profile 170 is a representation of how loud the audio encoded in the digital audio file 106 is over time (e.g., over the entire duration of the audio encoded in the digital audio file 106). The loudness profile 170 can be calculated as the total intensity (e.g., across audible frequencies) at multiple timestamps in the digital audio file 106. The loudness profile 170 can then be used to generate a volume curve 172 for the neural beat 168. In particular, the loudness profile 170 can be offset (e.g., depending on a maximum desired intensity of the neural beat 168) to generate the volume curve 172. For example, FIG. 5 illustrates a volume curve 500 according to an exemplary embodiment of the present disclosure. The loudness curve 500 illustrates the change in energy (e.g., in dB) over a duration of audio encoded in the digital audio file 106, and the energy of the audio signal in the digital audio file 106 can be used as a proxy for the loudness over time in the digital audio file 106. Returning to FIG. 1A , the loudness curve 172 can be applied to the neural beat 168 to generate the adjusted neural beat 174. In particular, applying the loudness curve 172 to the neural beat 168 can include increasing or decreasing the loudness (e.g., intensity) of the neural beat 168 at different times according to the intensity illustrated in the loudness curve 172 (e.g., such that the adjusted neural beat 174 is louder at times of high intensity of the volume curve 172 and quieter at times of low intensity of the volume curve 172).
[0037] The neural beats 168 and / or adjusted neural beats 174 can then be stored, transmitted, and / or played on the user's device. For example, the computing device 102 can store the neural beats 168 and / or adjusted neural beats 174 in association with the digital audio file 106 (e.g., at the server 104). In certain implementations, the digital audio file 106 and the neural beats 168 and / or adjusted neural beats 174 may be stored separately. In additional or alternative implementations, the computing device 102 can combine the digital audio file 106 with the neural beats 168 and / or adjusted neural beats 174 to generate a combined audio track that can be stored (e.g., at the server 104). As another example, with reference to FIG. 1B and system 190, the digital audio file 106 and the neural beats 168 and / or adjusted neural beats 174 may be transmitted to a user device 192 associated with a user 194. The user device 162 may include a smartphone, a tablet computer, a wearable computing device, a laptop, a personal computer, or any other personal computing device. The user device 192 may also include one or more audio devices for audio playback, such as a speaker, a 3.5 mm audio jack connected to headphones or speakers, wirelessly connected headphones, wirelessly connected speakers, or any other device capable of audio playback. The system 100 can transmit (e.g., stream) the digital audio file 106 and the neural beats 168 and / or the conditioned neural beats 174 to the user device 192. The user device 192 can then receive and play the digital audio file 106 simultaneously with the neural beats 168 and / or the conditioned neural beats 174.Additionally or alternatively, the user device 192 can store the digital audio file 106 and the neural beats 168 and / or the adjusted neural beats 174 for future playback. Additionally or alternatively, the computing device 102 can transmit the combined audio track to the user device 192. In yet another implementation, the neural beats 168 and / or the adjusted neural beats 174 may be generated on the user device 192. In such cases, the neural beats 168 and / or the adjusted neural beats 174 may be played along with the digital audio file 106 on the user device 192 (e.g., as separate audio files, as a combined audio track) and / or may be stored on the user device 192 for future playback at a later time.
[0038] Although not shown, the computing device 102, the server 104, and / or the user device 192 may include at least one processor and / or memory configured to implement one or more aspects of the computing device 102, the server 104, and / or the user device 192. For example, the memory may store instructions that, when executed by the processor, cause the processor to perform one or more operational functions of the computing device 102, the server 104, and / or the user device 192. The processor may be implemented as one or more Central Processing Units (CPUs), Field Programmable Gate Arrays (FPGAs), and / or Graphics Processing Units (GPUs) configured to execute instructions stored in the memory. Additionally, the computing device 102, the server 104, and / or the user device 192 may be configured to communicate using a network. For example, the computing device 102, the server 104, and / or the user device 192 may communicate with a network using one or more wired network interfaces (e.g., an Ethernet interface) and / or wireless network interfaces (e.g., Wi-Fi, Bluetooth, and / or cellular data interfaces). In certain cases, the network may be implemented as a local network (e.g., a local area network), a virtual private network, an L1 and / or a global network (e.g., the Internet).
[0039] In certain implementations, the computing device 102 and the server 104 may be implemented as a single computing device. For example, the computing device 102 may store the digital audio files 106, 108, 110 (e.g., in a local database). In further implementations, the computing device 102 and / or the server 104 may be implemented at least in part by the user device 162. In yet other implementations, the computing device 102, the server 104, and / or the user device 192 may be implemented by multiple computing devices. For example, the computing device 102 may be implemented as multiple software services running in a distributed computing environment (e.g., a cloud computing environment). As another example, the user device 162 may be implemented by multiple personal computing devices (e.g., a wearable computing device such as a smartphone and a smart watch).
[0040] FIG. 6 illustrates a method 600 for synthesizing neural beats according to an exemplary embodiment of the present disclosure. The method 600 may be implemented on a computer system, such as the systems 100, 160. For example, the method 600 may be implemented by the computing device 102 and / or the user device 192. The method 600 may also be implemented by a set of instructions stored on a computer-readable medium that, when executed by a processor, causes the computer system to perform the method 600. For example, all or part of the method 600 may be implemented by a processor and / or memory of the computing device 102 and / or the user device 192. The following example is described with reference to the flowchart illustrated in FIG. 6, although many other ways of performing the operations associated with FIG. 6 may be used. For example, the order of some of the blocks may be changed, certain blocks may be combined with other blocks, one or more of the blocks may be repeated, and some of the described blocks may be optional.
[0041] The method 600 may begin with receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file (block 602). For example, the computing device 102 may receive the digital audio file 106 and a beat frequency 112 of the neural beat to be added to the digital audio file 106. As described above, the computing device 102 may receive the digital audio file 106 from the server 104 and / or may retrieve the digital audio file 106 from local storage. In certain implementations, the digital audio file 106 may be received according to a user request. For example, a user request to play a particular song (e.g., via a music streaming service) may be received from a user device. The computing device 102 may receive the beat frequency 112 from a user (e.g., according to a user request and / or predefined user settings). In certain implementations, the beat frequency 112 may specify a particular frequency (e.g., 3 Hz) for the neural beat to be added to the digital audio file 106. In additional or alternative implementations, the beat frequency 112 can specify a range of frequencies for the neural beats (eg, 4-8 Hz).
[0042] A plurality of chromagram features may be extracted from the digital audio file (block 604). For example, the computing device 102 may extract a plurality of chromagram features 116, 200 from the digital audio file 106. As described above, the chromagram features may include intensity information of a plurality of pitch classes at a plurality of timestamps in the digital audio file 106. In a particular implementation, each of the plurality of chromagram features may be extracted according to different parameters applied to the digital audio file 106 prior to extracting the chromagram feature 116, 200. For example, a first chromagram feature may be extracted focusing on lower frequencies of the digital audio file 106, and a second chromagram feature may be extracted focusing on higher frequencies of the digital audio file 106. As another example, three chromagram features may be extracted from the digital audio file 106: a first chromagram feature focused on lower frequencies (e.g., below 200 Hz), a second chromagram feature focused on mid-level frequencies (e.g., between 200 Hz and 800 Hz), and a third chromagram feature focused on higher frequencies (e.g., above 800 Hz). In practice, the multiple chromagram features 116, 200 may be generated by generating a time-frequency representation as described above, and then selecting octaves and intensities within a desired frequency range for inclusion in the chromagram features 116, 200. In other implementations, the multiple chromagram features may be generated by applying a filter (e.g., a high-pass filter, a low-pass filter, a band-pass filter, etc.) to the digital audio file 106 prior to extracting the chromagram features 116, 200 (e.g., using FFT, Constant-Q transform, filter buckets, and / or other techniques as described above).
[0043] The multiple chromagram features may be combined to form a primary chromagram feature of the digital audio file (block 606). For example, the computing device 102 may combine multiple chromagram features 116, 200 to form a primary chromagram feature 118 of the digital audio file 106. In certain implementations, the multiple chromagram features 116, 200 may be linearly combined (e.g., according to predefined weights) to form the primary chromagram feature 118. In additional or alternative implementations, the multiple chromagram features 116, 200 may be combined according to any other possible combination strategy. For example, the multiple chromagram features 116, 200 may be combined by "stacking" the chromagram features 116, 200 (e.g., combining two chromagram features 116, 200 with 12 pitch classes to form a primary chromagram feature with 24 rows). Generating the primary chromagram features 118 based on multiple chromagram features 116 may better capture the audio frequency characteristics of the digital audio file 106 (e.g., by separately focusing on different frequency ranges, such as different octaves, within the digital audio file 106). In certain implementations, one or both of blocks 604, 606 may be omitted. For example, in certain implementations, rather than extracting multiple chromatogram features and combining them to form the primary chromagram features, a single set of chromagram features may be extracted from the digital audio file 106 and used as the primary chromagram features 118.
[0044] The dominant pitch classes may be extracted at multiple timestamps in the digital audio file (block 608). For example, the computing device 102 may extract the dominant pitch classes 120, 300 at multiple timestamps 156, 158 in the digital audio file 106. The dominant pitch classes 120, 300 may be extracted from the primary chromagram features 118 using a model such as a hidden Markov model. In particular, the dominant pitch classes 120, 300 may be extracted as a probability distribution at multiple timestamps T1-T19. The timestamps T1-T19 may be selected based on the timestamps of the primary chromagram features 118, as described above.
[0045] A plurality of carrier frequencies may be selected for the neural beat (block 610). For example, the computing device 102 may select a plurality of carrier frequencies 114, 400 for the neural beat 168, 174. In particular, the plurality of carrier frequencies 114, 400 may include individual carrier frequencies 160, 162 at a plurality of time stamps 164, 166, T1 through T19. The selected carrier frequency 114, 400 may be selected by a Viterbi process, which may select carrier frequencies such that transitions of carrier frequencies in adjacent time periods are optimized according to transition probabilities, as described further herein. In certain implementations, in addition to selecting a plurality of carrier frequencies 114, 400, a particular beat frequency of the neural beat 168 may be selected. For example, if the beat frequency 112 is received as a range of acceptable frequencies, the computing device 102 may select a beat frequency of the neural beat 168 from within the acceptable range, as described further below.
[0046] A synchronized beat is a beat that changes carrier frequency at different time periods within a digital audio file based on changes in an underlying audio signal, e.g., changes in musical harmony and / or melody at different time periods with the audio signal. A synchronized beat may be synthesized for a digital audio file based on the beat frequency and multiple carrier frequencies (block 612). For example, the computing device 102 may synthesize a neural beat 168 for a digital audio file 106 based on the beat frequency 112 and the carrier frequency 114. In particular, the neural beat 168 may be generated by modulating the beat frequency 112 with two different carrier frequencies 160, 162 at times corresponding to timestamps 164, 166, T1-T19 within the carrier frequency 114, 400. In this manner, the neural beat 168 may be synchronized to changes in musical harmony and / or melody at different time periods within the digital audio file 106. In certain implementations, the neural beats 168 may be synthesized to include a single audio channel (e.g., as a mono beat). In additional or alternative implementations, the neural beats 168 may be synthesized to include two audio channels (e.g., as a binaural beat with two channels, as a mono beat with two channels). In yet other implementations, the neural beats 168 may be synthesized to include more than two audio channels (e.g., three audio channels, four audio channels, five audio channels). In certain implementations, the number of audio channels may be specified by a user or by a predetermined setting. In additional or alternative implementations, the number of audio channels may be selected based on the number of audio channels in the digital audio file 106 (e.g., such that the neural beats 168 have the same number of audio channels as the digital audio file 106).
[0047] At least one of the synchronized neural beats and a combined audio track combining the synchronized neural beats and the digital audio file may be stored (block 614). For example, the computing device 102 may store at least one of the synchronized neural beats 168 or a combined audio track combining the neural beats 168 with the digital audio file 106. For example, as described above, the computing device 102 may store the neural beats 168 and / or the combined audio track on the server 104 and / or local storage within the computing device 102. Additionally or alternatively, the computing device 102 may transmit the neural beats 168 and / or the combined audio track to a user device for storage and playback (e.g., temporary storage for streaming, long-term storage). In implementations in which the computing device 102 is a user device, the computing device 102 may store the neural beats 168 and / or the combined audio track locally for current or future playback. In certain implementations, as further described above, the computing device 102 may be further configured to generate an adjusted neural beat 174 based on the neural beat 168. In such cases, the computing device 102 may be configured to store the equalized neural beat 174 and / or a combined audio track that combines the adjusted neural beat 174 with the digital audio file 106 in a manner similar to that described above.
[0048] In this manner, method 600 enables a computing device to generate neural beats for any digital audio file, allowing increased user choice in the type of music used to generate neural entrainment. Furthermore, the computing device can do this in real-time and ensure that the neural beats blend with the sound quality of the digital audio file and / or the loudness of the digital audio file to minimize user distraction and maximize neural entrainment. Thus, method 600 ensures that the generated neural beats combine constructively with previously created digital audio files.
[0049] 7A-7C illustrate methods 700, 710, 720 according to an exemplary embodiment of the present disclosure. The methods 700, 710, 720 may be performed in combination with at least a portion of the method 600. For example, the method 700 may be performed while performing blocks 608, 610 of the method 600. As another example, the method 710 may be performed between blocks 612 and 614 and / or as part of block 612 of the method 600. As a further example, the method 720 may be performed as part of block 612 of the method 600. The methods 700, 710, 720 may be performed on a computer system such as the systems 100, 190. For example, the methods 700, 710, 720 may be performed by the computing device 102 and / or the user device 192. The methods 700, 710, 720 may also be implemented by a set of instructions stored on a computer-readable medium that, when executed by a processor, causes a computer system to perform the methods 700, 710, 720. For example, all or part of the methods 700, 710, 720 may be implemented by a processor and / or memory of the computing device 102 and / or the user device 192. The following examples are described with reference to the flowcharts shown in Figures 7A-7C, although many other ways of performing the operations associated with Figures 7A-7C may be used. For example, the order of some of the blocks may be changed, certain blocks may be combined with other blocks, one or more of the blocks may be repeated, and some of the described blocks may be optional.
[0050] The method 700 may be performed to select multiple carrier frequencies for neural beats. The method 700 may begin generating a probability distribution for pitch classes at multiple timestamps (block 702). For example, a hidden Markov model may be used to generate a probability distribution for pitch classes (e.g., pitch classes of B, A sharp / B flat, A, G sharp / A flat, G, F sharp / G flat, F, E, D sharp / E flat, D, C sharp / D flat, and C) in the digital audio file 106 at multiple timestamps T1-T19 in the digital audio file 106. The timestamps T1-T19 may be selected based on timestamps in the primary chromagram features 118 (e.g., based on a segment of the digital audio file 106 used to calculate the chromagram features 116 and / or the time-frequency representation of the primary chromagram features 118). The hidden Markov model may be configured by adjusting transition probabilities to select when transitions between different carrier frequencies should occur. In particular, the transition probabilities of the hidden Markov model (e.g., transition probabilities between 0.005 and 0.02) may have been previously received (or may be updated) based on input received from a user, a system administrator, and / or a computational process.
[0051] A sequence of dominant pitch classes may then be identified within the probability distribution (block 704). For example, the computing device 102 may identify a sequence of dominant pitch classes within the probability distribution. In particular, the carrier frequencies 114, 400 may include a sequence of dominant pitch classes to be used as carrier frequencies for the neural beats 168. The sequence of dominant pitch classes may be identified to maximize the joint probability of the selected pitch classes within the probability distribution according to constrained transition probabilities for changes in the selected pitch classes. In particular, the sequence of dominant pitch classes may be selected by a Viterbi process implemented by the computing device 102.
[0052] In this manner, method 700 may be performed to select a sequence of carrier frequencies based on the harmony and melody (e.g., chromagram features) of a received digital audio file. This process thus allows for the application of neural beats 168 to existing digital audio files, while also ensuring that changes in carrier frequency do not confuse or distract a user attempting to trigger neural entrainment using neural beats.
[0053] The method 710 may be performed to adjust the volume of the neural beats 168 based on the volume of the digital audio file 106 at different times within the digital audio file 106. The method 710 may begin with generating a loudness profile for the duration of the digital audio file (block 712). For example, the computing device 102 (e.g., the audio mixer 122) may generate a loudness profile 170 for the duration of the digital audio file 106. The loudness profile 170 may be generated based on the intensity (e.g., audio volume) of the digital audio file 106 at multiple times within the digital audio file 106. For example, the loudness profile 170 may be generated for each data sampling timestamp within the digital audio file 106.
[0054] A volume curve may be formed based on the loudness profile (block 714). For example, the computing device 102 may form a volume curve 172 based on the loudness profile 170. The volume curve 172 may be formed as a percentage of the loudness profile 170 (e.g., 50% of the loudness profile 170). Additionally or alternatively, the volume curve 172 may be formed by normalizing the loudness profile 170 for a maximum volume desired for the neural beat 168). One skilled in the art may similarly recognize one or more additional means for generating the volume curve 172 based on the loudness profile 170 of the digital audio file 106. All such similar implementations are contemplated herein within the scope of this disclosure.
[0055] The volume of the synchronized neural beat may then be adjusted according to the volume curve (block 714). For example, the computing device 102 may adjust the volume of the neural beat 168 based on the volume curve 172 to generate an adjusted neural beat 174. For example, the neural beat 168 may be scaled in intensity to match the desired volume reflected in the volume curve 172.
[0056] In this manner, method 710 may be performed to adjust the neural beats 168, thereby reducing the number of intrusive volume discrepancies between the neural beats in the digital audio file. For example, if the volume of the neural beats in the digital audio file is much lower, the user may not be able to hear the volume of the neural beats, which may reduce their effectiveness in generating neural entrainment. As another example, if the volume of the neural beats is much louder than the digital audio file 106, the user may be distracted or confused by the difference in volume, which may disrupt or reduce the neural entrainment generated by the neural beats.
[0057] Method 720 can be used to synchronize neural beats 168 with rhythmic patterns in a digital audio file 106. Method 720 can begin with estimating the location of rhythmic beats in the digital audio file (block 722). Rhythmic beats refer to a quantitative level of music, which is not the same use of the term as "beat" is used in acoustics. For example, computing device 102 can estimate the location of rhythmic beats in digital audio file 106. The location of rhythmic beats in digital audio file 106 can be estimated using a machine learning model, such as a pre-trained network configured to detect rhythmic beats in audio files. For example, the location of rhythmic beats can be estimated using one or more models similar to models provided by the madmom audio software package, the Essentia audio software package, and the like. In additional or alternative implementations, the location of rhythmic beats can be estimated using one or more algorithmic techniques.
[0058] The timing of the synchronized neural beats may be adjusted based on the location of the rhythmic beats in the digital audio file (box 724). For example, the computing device 102 may adjust the timing of the neural beats 168 based on the location of the rhythmic beats. For example, the computing device 102 may adjust the beat frequency 112 to match (e.g., be a multiple of) the tempo of the digital audio file. For example, if the digital audio file 106 has a tempo of 120 bpm and the beat frequency 112 is 0.6 Hz (e.g., 100 bpm), the computing device 102 may adjust the beat frequency 112 to be an integer multiple of the tempo of 120 beats per minute (e.g., 2 Hz). As a specific example, the computing device 102 may adjust the beat frequency 112 to 0.5 Hz (30 bpm) and / or 1 Hz (60 bpm). In implementations in which a user has specified a desired frequency range for the beat frequency 112, the beat frequency 112 may be selected within the desired frequency range to be an even multiple of the rhythmic frequency and / or as close to a multiple of the rhythmic frequency as possible. Additionally, the timing of the synchronized neural beats may be adjusted so that peak values of the neural beats (e.g., peak values at the beat frequency 112) occur simultaneously with (e.g., align with) the timing of the rhythmic beats in the digital audio file 106.
[0059] In this manner, method 720 can be used to ensure that the rhythmic beats and beat frequencies in a digital audio file are not out of phase. In particular, if the beat frequencies are out of phase with the rhythmic frequencies of the digital audio file, interference between the beat frequencies in the digital audio file can adversely affect sound quality and / or create distracting or confusing interference patterns when the digital audio file and neural beats at the interfering beat frequencies are played simultaneously. Thus, adjusting the beat frequencies based on the rhythmic beats in the digital audio file can reduce these interferences and improve the quality of the neural beats subsequently generated and / or the quality of the neural entrainment produced by the neural beats.
[0060] 8 illustrates an exemplary computer system 800 that may be utilized to implement one or more of the devices and / or components described herein, such as computing device 102. In certain embodiments, one or more computer systems 800 perform one or more steps of one or more methods described or illustrated herein. In certain embodiments, one or more computer systems 800 provide functionality described or illustrated herein. In certain embodiments, software executing on one or more computer systems 800 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Certain embodiments include one or more portions of one or more computer systems 800. As used herein, references to a computer system may encompass a computing device, and vice versa, where appropriate. Additionally, references to a computer system may encompass one or more computer systems, where appropriate.
[0061] This disclosure contemplates any suitable number of computer systems 800. This disclosure contemplates computer system 800 taking any suitable physical form. By way of example, and not limitation, computer system 800 may be an embedded computer system, a System-On-Chip (SOC), a Single-Board Computer (SBC) system (e.g., a Computer-On-Module (COM) or System-On-Module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a Personal Digital Assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computer system 800 may reside in a cloud that may include one or more cloud components that may be single or distributed, across multiple locations, across multiple machines, across multiple data centers, or in one or more networks, including one or more computer systems 800. Where appropriate, one or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not by way of limitation, one or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. One or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations, where appropriate.
[0062] In a particular embodiment, computer system 800 includes a processor 806, memory 804, storage 808, an input / output (I / O) interface 810, and a communications interface 812. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular configuration, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable configuration.
[0063] In particular embodiments, the processor 806 includes hardware for executing instructions, such as those that make up a computer program. By way of example and not limitation, to execute an instruction, the processor 806 may retrieve (or fetch) the instruction from an internal register, an internal cache, memory 804, or storage 808, decode and execute the instruction, and then write one or more results to an internal register, an internal cache, memory 804, or storage 808. In particular embodiments, the processor 806 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates the processor 806 including any suitable number of any suitable internal caches, where appropriate. By way of example and not limitation, the processor 806 may include one or more instruction caches, one or more data caches, and one or more Translation Lookaside Buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 804 or storage 808, and the instruction cache may speed up retrieval of those instructions by the processor 806. The data in the data cache may be a copy of data in memory 804 or storage 808 to be operated on by the computer instruction, a result of a previous instruction executed by the processor 806 accessible to a subsequent instruction or for writing to the memory 804 or storage 808, or any other suitable data. The data cache may speed up read or write operations by the processor 806. The TLB may speed up virtual address translation for the processor 806. In particular embodiments, the processor 806 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates the processor 806 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, the processor 806 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, and may include one or more processors 806.Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0064] In particular embodiments, memory 804 includes a main memory for storing instructions that processor 806 executes or data that processor 806 operates on. By way of example and not limitation, computer system 800 may load instructions into memory 804 from storage 808 or another source (such as another computer system 800). Processor 806 may then load the instructions from memory 804 into an internal register or cache. To execute instructions, processor 806 may retrieve instructions from the internal register or cache and decode them. During or after execution of instructions, processor 806 may write one or more results (which may be intermediate or final results) to an internal register or cache. Processor 806 may then write one or more of those results to memory 804. In particular embodiments, the processor 806 executes only instructions in one or more internal registers or caches or memory 804 (as opposed to storage 808 or elsewhere) and operates only on data in one or more internal registers or caches or memory 804 (as opposed to storage 808 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may couple the processor 806 to the memory 804. The buses may include one or more memory buses, as described in further detail below. In particular embodiments, one or more Memory Management Units (MMUs) reside between the processor 806 and the memory 804 to facilitate accesses to the memory 804 requested by the processor 806. In particular embodiments, the memory 804 includes Random Access Memory (RAM). This RAM may be volatile memory, where appropriate. This RAM may be Dynamic RAM (DRAM) or Static RAM (SRAM), where appropriate. Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM.This disclosure contemplates any suitable RAM. Memory 804 may include one or more memories 804, where appropriate. Although this disclosure describes and illustrates particular memory implementations, this disclosure contemplates any suitable memory implementations.
[0065] In certain embodiments, storage 808 includes mass storage for data or instructions. By way of example, and not limitation, storage 808 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Storage 808 may include removable or non-removable (or fixed) media, where appropriate. Storage 808 may be internal or external to computer system 800, where appropriate. In certain embodiments, storage 808 is non-volatile solid-state memory. In certain embodiments, storage 808 includes read-only memory (ROM). Where appropriate, the ROM may be mask programmable ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage 808 taking any suitable physical form. Storage 808 may include one or more storage control units that facilitate communications between processor 806 and storage 808, where appropriate. Storage 808 may include one or more storages 808, where appropriate. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.
[0066] In particular embodiments, I / O interface 810 includes hardware, software, or both, providing one or more interfaces for communication between computer system 800 and one or more I / O devices. Computer system 800 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person (i.e., a user) and computer system 800. By way of example and not limitation, an I / O device may include a keyboard, a keypad, a microphone, a monitor, a screen, a display panel, a mouse, a printer, a scanner, a speaker, a still camera, a stylus, a tablet, a touch screen, a trackball, a video camera, another suitable I / O device, or a combination of two or more of these. An I / O device may include one or more sensors. Where appropriate, I / O interface 810 may include one or more device or software drivers that enable processor 806 to drive one or more of these I / O devices. I / O interface 810 may, where appropriate, include one or more I / O interfaces 810. Although this disclosure describes and illustrates particular I / O interfaces, this disclosure contemplates any suitable I / O interface or combination of I / O interfaces.
[0067] In certain embodiments, communications interface 812 includes hardware, software, or both to provide one or more interfaces for communications (e.g., packet-based communications, etc.) between computer system 800 and one or more other computer systems 800 or one or more networks 814. By way of example and not limitation, communications interface 812 may include a Network Interface Controller (NIC) or network adapter for communicating with an Ethernet or any other wired-based network, or a Wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a Wi-Fi network. This disclosure contemplates any suitable network 814 and any suitable communications interface 812 for network 814. By way of example, and not limitation, the network 814 may include one or more of an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, the computer system 800 may be in communication with a Wireless PAN (WPAN) (e.g., a Bluetooth WPAN, etc.), a WI-FI network, a WI-MAX network, a cellular network (e.g., a Global System for Mobile Communications (GSM) network, etc.), or any other suitable wireless network, or a combination of two or more of these.Computer system 800 may include any suitable communications interface 812 for any of these networks, where appropriate. Communications interface 812 may include one or more communications interfaces 812, where appropriate. Although this disclosure describes and illustrates particular communications interface implementations, this disclosure contemplates any suitable communications interface implementation.
[0068] The computer system 802 may also include a bus, which may include hardware, software, or both, that may communicatively couple the components of the computer system 800 to one another. By way of example, and not limitation, the bus may include an Accelerated Graphics Port (AGP) or any other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front-Side Bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low-PIN-Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB), or another suitable bus, or a combination of two or more of these buses. A bus may include one or more buses, where appropriate, and although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0069] As used herein, a computer-readable non-transitory storage medium may include, where appropriate, one or more semiconductor-based or other types of integrated circuits (ICs) (e.g., Field Programmable Gate Arrays (FPGAs) or Application-Specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more of these. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.
[0070] As used herein, unless expressly indicated otherwise or indicated otherwise by context, "or" is inclusive and not exclusive. Thus, as used herein, "A or B" means "A, B, or both," unless expressly indicated otherwise or indicated otherwise by context. Moreover, unless expressly indicated otherwise or indicated otherwise by context, "and" is both joint and several. Thus, as used herein, unless expressly indicated otherwise or indicated otherwise by context, "A and B" means "A and B, jointly or severally."
[0071] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that would be understood by a person skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although the present disclosure describes and illustrates each embodiment herein as including certain components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by a person skilled in the art. Furthermore, any reference in the appended claims to an apparatus or system, or an apparatus or system component, adapted, arranged, capable, configured, enabled, enabled, or operable to perform a particular function encompasses that apparatus, system, or component, regardless of whether it or that particular function is activated, turned on, or unlocked, so long as the apparatus, system, or component is so adapted, arranged, capable, configured, enabled, enabled, or operable. Moreover, although the present disclosure has described or illustrated particular embodiments as providing certain advantages, the particular embodiment may provide none, some, or all of these advantages.
[0072] All of the disclosed methods and procedures described in this disclosure can be implemented using one or more computer programs or components. These components may be provided as a series of computer instructions on any conventional computer-readable or machine-readable medium, including volatile and non-volatile memories such as RAM, ROM, flash memory, magnetic or optical disks, optical memory, or other storage media. The instructions may be provided as software or firmware, and may be implemented in whole or in part in hardware components such as ASICs, FPGAs, DSPs, or any other similar devices. The instructions may be configured to be executed by one or more processors, which, when executing the series of computer instructions, perform or facilitate the execution of all or a portion of the disclosed methods and procedures.
[0073] It should be understood that various changes and modifications to the examples described herein will be apparent to those skilled in the art. Such changes and modifications can be made without departing from the scope of the present subject matter and without diminishing its intended advantages. Accordingly, such changes and modifications are intended to be covered by the appended claims.
Claims
1. receiving a digital audio file and beat frequencies of neural beats to be added to the digital audio file; extracting a plurality of chromagram features from the digital audio file according to a plurality of parameters; combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file; extracting from the primary chromagram features dominant pitch classes at multiple timestamps within the digital audio file; selecting a plurality of carrier frequencies for the neural beat based on the dominant pitch class at the plurality of timestamps; synthesizing a synchronized neural beat for the digital audio file based on the beat frequency and the plurality of carrier frequencies; storing at least one of (i) the synchronized neural beats and (ii) a combined audio track that combines the synchronized neural beats and the digital audio file; A method comprising:
2. the primary chromagram features include intensities of each of a plurality of pitch classes at the plurality of timestamps; The method of claim 1 , wherein the dominant pitch class is selected from among the plurality of pitch classes.
3. 3. The method of claim 2, wherein extracting the dominant pitch class further comprises using a hidden Markov model to generate a probability distribution for each of the plurality of pitch classes at the plurality of timestamps based on the intensities of the plurality of pitch classes.
4. The method of claim 3 , wherein the hidden Markov model is configured to optimize the number and locations of transitions between a plurality of dominant pitch classes.
5. The method of claim 3 , wherein extracting the dominant pitch classes further comprises identifying a sequence of dominant pitch classes within the probability distribution.
6. 6. The method of claim 1, wherein the plurality of timestamps occur every 500 milliseconds or less during the audio encoded into the digital audio file.
7. The method of claim 1 , wherein the plurality of chromagram features are linearly combined to form the primary chromagram feature.
8. 6. The method of claim 1, further comprising adjusting a volume of the synchronized neural beats to track a volume of the audio encoded in the digital audio file over time.
9. normalizing the loudness of the synchronized neural beats; generating a loudness profile for a duration of the audio encoded in the digital audio file; generating a volume curve based on the loudness profile; adjusting the volume of the synchronized neural beats according to the volume curve; The method of claim 8, comprising:
10. The method of any of claims 1 to 5, further comprising matching the beat frequency with a rhythmic beat in the digital audio file.
11. said matching said beat frequencies comprising: estimating the location of rhythmic beats within the audio signal encoded within the digital audio file; estimating the musical tempo of the audio signal encoded within the digital audio file; adjusting the timing of the synchronized neural beats to align peak values in the synchronized neural beats with positions of the rhythmic beats in the audio signal encoded in the digital audio file according to the music tempo; The method of claim 10, comprising:
12. 6. The method of claim 1, wherein the neural beats are at least one of (i) binaural beats and (ii) monaural beats.
13. The method of claim 1 , wherein the synchronized neural beats are encoded in no more than two audio channels.
14. The method of claim 1 , wherein the synchronized neural beats are encoded in three or more audio channels.
15. 6. The method according to claim 1, wherein the beat frequency is equal to or greater than 0.5 Hz and equal to or less than 150 Hz.
16. 6. The method of claim 1, further comprising simultaneously playing the synchronized neural beats and the digital audio file via a computing device.
17. 17. The method of claim 16, further comprising streaming the synchronized neural beats and the digital audio file to the computing device for playback by the computing device.
18. A system comprising a processor and a memory, characterized in that the memory is configured to perform the steps of the method according to any one of claims 1 to 5.
19. the primary chromagram features include intensities of each of a plurality of pitch classes at the plurality of timestamps; The system of claim 18 , wherein the dominant pitch class is selected from among the plurality of pitch classes.
20. the memory further comprising instructions; The instruction is: when executed by the processor while extracting the dominant pitch class, 20. The system of claim 19, wherein the processor is configured to use a hidden Markov model to generate a probability distribution for each of the plurality of pitch classes at the plurality of timestamps based on the intensities of the plurality of pitch classes.