Generating pitch-compatible sync neuro beats for digital audio files
By analyzing the pitch characteristics of digital audio files, generating chroma map features, and synthesizing synchronized neural beats, the problem that neural beats cannot be added to existing audio tracks in the prior art is solved, and neural synchronization and attention enhancement are achieved when listening to preferred audio tracks.
Patent Information
- Application Number
- CN202280071048.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-21
- Filing Date
- 2022-10-21
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing technologies struggle to effectively add neural beats to existing audio tracks, preventing users from experiencing the benefits of neural synchronization, relaxation, and enhanced attention while listening to their preferred tracks.
By analyzing the pitch characteristics of digital audio files, generating chroma map features, selecting carrier frequencies, synthesizing synchronous neural beats, and adjusting volume and rhythm alignment, the neural beats are synchronized with existing audio files.
It automatically adds neural beats while the user listens to their preferred audio tracks, enhancing neural synchronization, relaxation, and attention.
Smart Images

Figure CN118176535B_ABST
Abstract
Description
BACKGROUND
[0001] In acoustics, beats are interference patterns between two sounds of slightly different frequencies, viewed as periodic changes in volume, with a rate of change that is the difference between the two frequencies. For monaural beats, a listener hears two different frequencies in the same ear or both ears at the same time. In binaural beats, two different frequencies are heard through different ears (e.g., using headphones or specially placed speakers), and the listener’s brain detects the interference pattern. More complex interference patterns involving multiple signals can also be used to produce various beats.
[0002] Certain types of beats (e.g., monaural beats, binaural beats) can be used to promote a desired mental state (e.g., to improve an individual’s focus or attention). For example, such beats can be used to produce neural synchrony when a user listens to the beats, helping the user to better focus or concentrate. These beats can often be provided as standalone audio tracks (e.g., audio tracks that contain only beats). Alternatively, audio tracks that have been added with monaural or binaural beat customization (i.e., audio tracks that have been composed or generated to contain monaural or binaural beats) can be prepared. In some cases, beats that are not in the frequency range of typical audible sounds can even be provided. SUMMARY
[0003] The present disclosure presents new and innovative systems and methods for generating neural beats and adding them to existing audio tracks. In a first aspect, the present disclosure provides a method comprising receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file, and extracting a plurality of chromagram features of the digital audio file according to a plurality of parameters. The method comprises combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file, and extracting a primary sound level at a plurality of timestamps within the digital audio file from the primary chromagram feature. Selecting a plurality of carrier frequencies for the neural beat based on the primary sound level at the plurality of timestamps, and synthesizing a synchronized neural beat of the digital audio file based on the beat frequency and the plurality of carrier frequencies. The method further comprises storing at least one of (i) the synchronized neural beat and (ii) a combined audio track that combines the synchronized neural beat and the digital audio file.
[0004] In embodiments according to the first aspect of the present disclosure, the primary chromagram feature comprises an intensity of each of the plurality of sound levels at the plurality of timestamps. In one embodiment, the primary sound level is selected from the plurality of sound levels. In another embodiment, extracting the primary sound level further comprises generating a probability distribution for each of the plurality of sound levels at the plurality of timestamps based on the intensities of the plurality of sound levels using a hidden Markov model. In yet another embodiment, the hidden Markov model is configured to optimize a number and a location of transitions between the primary sound levels. In an alternative embodiment, extracting the primary sound level further comprises identifying a sequence of the primary sound levels within the probability distribution.
[0005] In embodiments according to the first aspect of the disclosure, the plurality of timestamps occur once every 500 milliseconds or less during the digital audio file.
[0006] In embodiments according to the first aspect of the disclosure, the plurality of chromagram features are linearly combined to form a primary chromagram feature.
[0007] In embodiments according to the first aspect of the disclosure, the method further comprises adjusting a volume of the synchronized neural beat over time to follow a volume of the encoded audio in the digital audio file. In another embodiment, normalizing the volume of the synchronized neural beat comprises generating a loudness profile of a duration of the encoded audio in the digital audio file, and forming a volume curve based on the loudness profile. In one embodiment, the method comprises adjusting the volume of the synchronized neural beat according to the volume curve.
[0008] In embodiments according to the first aspect of the disclosure, the method further comprises aligning the beat frequency with the rhythmic beat within the digital audio file. In one embodiment, aligning the beat frequency comprises estimating a location of the rhythmic beat within the digital audio file, estimating a musical rhythm within the digital audio file, and adjusting a timing of the synchronized neural beat according to the musical rhythm to align a peak within the synchronized neural beat with the location of the rhythmic beat within the digital audio file. In one embodiment, the neural beat is at least one of: (i) a binaural beat and (ii) a monaural beat.
[0009] In embodiments according to the first aspect of the disclosure, the synchronized neural beat comprises two or fewer audio channels.
[0010] In embodiments according to the first aspect of the disclosure, the synchronized neural beat comprises three or more audio channels.
[0011] In embodiments according to the first aspect of the disclosure, the beat frequency is greater than or equal to 0.5 Hz and less than or equal to 150 Hz.
[0012] In embodiments according to the first aspect of the disclosure, the method further comprises playing the synchronized neural beat and the digital audio file in parallel via the computing device. In one embodiment, the method further comprises streaming the synchronized neural beat and the digital audio file to the computing device for playback by the computing device.
[0013] Unless the disclosure explicitly discloses otherwise, embodiments of the disclosure according to the above first aspect are not mutually exclusive from each other: in other embodiments according to the first aspect of the disclosure, features of one embodiment according to the first aspect of the disclosure are combined with features of another embodiment according to the first aspect.
[0014] In a second aspect, a system comprising a processor and a memory is provided. The memory can store instructions that, when executed by the processor, cause the processor to perform the method according to the first aspect of the disclosure. In one embodiment, the instructions, when executed by the processor, cause the processor to receive a digital audio file and a beat frequency of a neurobeat to be added to the digital audio file, and extract a plurality of chromagram features of the digital audio file according to a plurality of parameters. The instructions can also cause the processor to combine the plurality of chromagram features to form a primary chromagram feature of the digital audio file, extract a master sound level at a plurality of timestamps within the digital audio file from the primary chromagram feature, and select a plurality of carrier frequencies for the neurobeat based on the master sound level at the plurality of timestamps. The instructions can also cause the processor to synthesize a synchronized neurobeat of the digital audio file based on the beat frequency and the plurality of carrier frequencies, and store at least one of (i) the synchronized neurobeat and (ii) a combined audio track that combines the synchronized neurobeat and the digital audio file.
[0015] In one embodiment according to the second aspect of the disclosure, the primary chromagram feature comprises an intensity of each of the plurality of sound levels at the plurality of timestamps. The master sound level can be selected from the plurality of sound levels.
[0016] In one embodiment according to the second aspect, the memory stores further instructions that, when executed by the processor in extracting the master sound level, cause the processor to generate a probability distribution for each of the plurality of sound levels at the plurality of timestamps based on the intensities of the plurality of sound levels using a hidden Markov model.
[0017] Unless the disclosure explicitly discloses otherwise, embodiments of the disclosure according to the above second aspect are not mutually exclusive: in other embodiments according to the second aspect of the disclosure, features according to one embodiment of the second aspect of the disclosure are combined with features according to another embodiment of the second aspect.
[0018] Embodiments according to the first aspect of the disclosure can be combined with embodiments according to the second aspect of the disclosure. Not all of the features and advantages described herein need to be included in every embodiment of the disclosure. Many additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings and descriptions. Moreover, it is to be noted that the language used in the specification has been principally selected for readability and instructional purposes and can not have been selected to expressly delineate or otherwise limit the scope of the disclosed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1A A system according to example embodiments of the disclosure is shown.
[0020] Figure 1B An audio playback system according to example embodiments of the disclosure is shown.
[0021] Figure 2 A chromagram feature is shown in accordance with example embodiments of the present disclosure.
[0022] Figure 3 A master level is shown in accordance with example embodiments of the present disclosure.
[0023] Figure 4 A selected carrier frequency is shown in accordance with example embodiments of the present disclosure.
[0024] Figure 5 A volume curve is shown in accordance with example embodiments of the present disclosure.
[0025] Figure 6 A method of synthesizing a neural beat is shown in accordance with example embodiments of the present disclosure.
[0026] Figures 7A-7C A method is shown in accordance with example embodiments of the present disclosure.
[0027] Figure 8 A computing system is shown in accordance with example embodiments of the present disclosure. DETAILED DESCRIPTION
[0028] For the purposes of this application, a "neural beat" is an acoustic beat added to an audio signal. The neural beat does not necessarily have to be in the audible frequency range. All references to neural beats in this application are intended to include either or both monaural and binaural beats, unless the context clearly limits the discussion to one or the other.
[0029] A "neurobeat" can include an audio beat designed to produce or facilitate a desired mental state in a user. The desired mental state can include neurosynchronization, improved focus, mood calming, relaxation, or any other desired mental state. In certain embodiments, a neurobeat can include a monaural or binaural beat that combines a lower beat frequency with a higher carrier frequency. In particular, the "beat frequency" can be selected based on the desired mental state (e.g., where different frequencies can facilitate an individual to produce different types of mental states). In one embodiment, the beat frequency can range from 0.5 Hz to 150 Hz. In other example embodiments, the frequency can be selected in a range between 1 Hz and 100 Hz, 1 Hz and 10 Hz, 10 Hz and 100 Hz, or 10 Hz and 40 Hz. The "carrier frequency" can be an audio frequency or note that is selected to carry or audibly reproduce the beat frequency within an audio track. For example, the beat frequency can be lower than a frequency that a human can detect, and / or can be in a lower frequency range of human hearing. Thus, to maximize the effectiveness of a neurobeat, a carrier frequency can be selected in an existing audio signal, such as an existing musical or other audio recording, or a particular track in such a recording, and the beat frequency can be modulated onto the carrier frequency to form the neurobeat. The carrier frequency can range from 207.65 Hz to 392.00 Hz. In various embodiments, a neurobeat can have a different number of audio channels, such as one audio channel (e.g., a monaural beat), two audio channels (e.g., a binaural beat), five audio channels, or more.
[0030] Not all users can enjoy listening to audio tracks that contain only neurobeats, and can find it boring or distracting, thereby limiting the effectiveness of neurosynchronization. Furthermore, the availability of existing audio tracks that contain embedded monaural beats can not appeal to all users. Certain systems can automatically generate music that contains monaural beats to prevent users from having to listen to the same track multiple times. However, such systems still cannot make modifications for the possibility that a user wants to listen to a particular track or genre that was not previously combined with a neurobeat. Thus, there is a need to automatically add neurobeats to existing audio tracks so that users can listen to their preferred tracks or music genres while also experiencing the benefits of neurosynchronization, relaxation, and / or improved focus provided by neurobeats.
[0031] One solution to this problem is to analyze the pitch characteristics of a digital audio file over time. Specifically, a chromagram feature indicates the intensity of different pitch levels of an audio signal over time. A chromagram feature can be generated for a digital audio file, indicating the intensity of different pitch levels of a digital audio signal encoded within the digital audio file over time. This information can then be used to select a carrier frequency for a neurobeat to be added to the digital audio file. For example, a dominant pitch level can be extracted from the chromagram feature at different timestamps within the digital audio file, and the dominant pitch level can be used to select a carrier frequency for a neurobeat at the different timestamps. In some cases, the dominant pitch levels can be analyzed with a model (e.g., a hidden Markov model) to select a carrier frequency to optimize the number of changes in carrier frequency, e.g., to minimize the number of changes while still achieving a particular accuracy. The neurobeat can then be synthesized based on the beat frequency and the selected carrier frequency, and stored for later use. In some cases, a combined audio track that combines the digital audio file with the neurobeat can be generated. In other cases, the neurobeat can be stored in association with the digital audio file. Further, in some cases, the neurobeat and / or the combined audio track can be generated in real-time as the digital audio file is streamed by a user device, e.g., by a server that streams the digital audio file or by a user device that receives the streamed digital audio file. The neurobeat can then be played with the digital audio file via the user device (e.g., as a separate audio file played simultaneously and / or as a single audio file).
[0032] Figure 1A A system 100 according to example embodiments of the present disclosure is shown. The system 100 can be configured to generate and synchronize neurobeats for addition to digital audio files. The system 100 includes a computing device 102 and a server 104. The server 104 stores digital audio files 106, 108, 110 to which the computing device 102 can add neurobeats. For example, the computing device 102 and the server 104 can be part of a digital audio streaming platform configured to stream digital audio files 106, 108, 110 upon request by a user. Further, the computing device 102 can be configured to add neurobeats 168, 174 to the digital audio files 106, 108, 110 upon request by a user. For example, a user can have a preference to add neurobeats 168, 174 to streamed audio files received from the audio streaming platform.
[0033] The computing device 102 can receive the digital audio file 106 from the server 104 and can generate the neurobeat 168 and / or the adjusted neurobeat 174 to be added to the digital audio file 106. The computing device 102 can also receive the beat frequency 112 of the neurobeat 168, 174. The beat frequency 112 can be received from a user, for example, via a user-configurable beat frequency setting. The neurobeat 168, 174 can be a monoaurai beat, a binaural beat, or can have more audio channels, and the type of neurobeat 168, 174 can be selected by the user. Additionally or alternatively, the computing device 102 can select between monoaurai and binaural beats based on an audio device from which the user is streaming the digital audio file. For example, if the user is streaming audio from a monoaurai audio device, the computing device 102 can generate a monoaurai neurobeat, and if the user is streaming audio from a stereophonic audio device (e.g., stereo speakers, stereo headphones), the computing device 102 can generate a binaural neurobeat. In yet another implementation, the computing device 102 can select a number of audio channels based on a same number of audio channels in the digital audio file 106.
[0034] In particular, the computing device 102 can be configured to generate the neurobeat 168, 174 to be mixed into the digital audio file 106. For example, the computing device 102 can be configured to generate the neurobeat 168, 174 in synchronization with audio pitches within the digital audio file 106 to avoid apparent and distracting pitch differences that can interfere with the user's neurosynchronization. To this end, the computing device 102 can extract a plurality of chromagram features 116 from the digital audio file 106. The chromagram features 116 can include the sound levels 124, 126 and associated intensities 136, 138 at a plurality of timestamps 148, 150.
[0035] For example, Figure 2A chromagram feature 200 is depicted in accordance with example embodiments of the present disclosure. The chromagram feature 200 includes the intensity of a plurality of tonal levels (as defined in the legend 202) at a plurality of timestamps T1-T19. The tonal levels include B, sharp A / flatted B, A, sharp G / flatted A, G, sharp F / flatted G, F, E, sharp D / flatted E, D, sharp C / flatted D, and C, which represent each type of note that can be reproduced within the digital audio file 106. Specifically, each tonal level can represent all audible pitches separated by an integer number of octaves in a song. For example, the tonal level C can contain the notes C of middle C, treble C, high C, tenor C, low C, and other octaves. Other tonal levels can be similarly defined to contain multiple notes of different octaves. In practice, the tonal levels can be defined as a collection of frequency bands. For example, the tonal level C can be defined as 261.626 ± 0.1 Hz (for middle C), 523.251 ± 0.1 Hz (for tenor C), and similarly for other notes contained within the tonal level. As depicted, certain sharp or flatted notes (e.g., sharp A, flatted B, sharp G, flatted A, sharp F, flatted G, sharp D, flatted E, sharp C, flatted D) are grouped into tonal levels separate from the tonal levels containing natural notes A-G. In additional or alternative implementations, the tonal levels can be defined to contain sharp or flatted versions of the notes. Similarly, certain implementations can define the tonal levels in different ways (e.g., to contain any desired combination of notes). For example, in another implementation, the tonal level for C can contain a central sharp C or a central flatted C. Those skilled in the art will appreciate that the chromagram feature 200 can be computed in accordance with any of a number of conceivable tonal levels, such as equal temperament tuning (e.g., 24 equal temperament with 24 tonal levels, 19 equal temperament with 19 tonal levels, and / or 7 equal temperament with 7 tonal levels). In practice, the computing device 102 can compute more tonal levels than represented in the chromagram feature 200 and can combine the tonal levels into the desired tonal levels of the chromagram feature 200. For example, the computing device 102 can compute 36 tonal levels and then combine the tonal levels into the tonal levels depicted for the chromagram feature 200.
[0036] The chromagram feature 200 includes an intensity of each pitch class at each of the time stamps T1-T19. These intensities vary over time (e.g., as the music in the digital audio file 106 changes). For example, both the pitch class A and the pitch class D have high intensities at times T1-T5. Starting at times T6-T10, the pitch class with the highest intensity alternates between C and C# / D (T8, T12), D (T9-10, T13, T17-18), D and D# / E (T6, T14), E (T7, T15), E and D# / E (T11, T19), and F (T16). These intensities can be computed based on an analysis of the frequency domain of the digital audio file 106 at each of the time stamps T1-T19. For example, the computing device 102 can divide the digital audio file 106 into segments for each of the time stamps T1-T19. The computing device 102 can then compute a time-frequency representation (e.g., a frequency distribution over a plurality of times) for each segment (e.g., by performing a Fourier transform, a Fast Fourier Transform (FFT), a Constant-Q transform, a wavelet transform, using a filter bank, etc.). The frequencies in the time-frequency representation can correspond to or be classified into each pitch class (e.g., according to a predefined frequency band). The intensity of each pitch class can then be computed based on the intensity of the corresponding frequencies in the time-frequency representation. This process can be repeated multiple times for the segments corresponding to each of the time stamps T1-19. In certain implementations, the time stamps T1-19 can occur every 50 milliseconds. In additional or alternative implementations, the time stamps T1-19 can occur more frequently (e.g., every 10 milliseconds, every 5 milliseconds, every millisecond) and / or less frequently (e.g., every 0.5 seconds, every 0.25 seconds, every 0.1 seconds). In certain implementations, the computing device 102 can perform the analysis in the time domain, rather than performing a frequency domain analysis on the digital audio file 106. For example, a filter bank can be used with one or more filters for each pitch class. The intensities of the filtered signals at each time stamp can then be used to determine the intensities of the chromagram feature 200.
[0037] Returning to Figure 1AThe computing device 102 can compute multiple chroma map features 116 of the digital audio file 106. For example, multiple chroma map features 116 can be computed to focus on different frequency ranges within the digital audio file 106. As a specific example, a first set of chroma map features can be computed focusing on lower frequency ranges (e.g., less than C4 or 261.62 Hz) within the digital audio file 106, and a second set of chroma map features can be computed focusing on higher frequency ranges (e.g., C1 to C8 or 32.70 Hz to 4186.01 Hz). In such a case, the computing device 102 can then be configured to combine the multiple chroma map features 116 into a set of principal chroma map features 118 of the digital audio file 106. For example, the computing device 102 can linearly combine the chroma map features 116 (e.g., according to predefined weights) to form the principal chroma map features 118. The data structure of the principal chroma map features 118 can be comparable to the data structure of the chroma map features 116. For example, in some implementations, chroma map feature 200 may represent a set of primary chroma map features 118 of the digital audio file 106. Furthermore, it should be understood that, although... Figure 2 Chroma map feature 200 is depicted as a data graph over time, but in practice, chroma map feature 116 and / or principal chroma map feature 118 can be stored in an additional or alternative data structure. For example, chroma map feature 116 and / or principal chroma map feature 118 can be stored as an array containing the intensity values of the pitch levels at timestamps T1-19.
[0038] The computing device 102 can identify the tonic level 120 based on the principal chroma map feature 118. The tonic level has the highest intensity within a specific time or time interval of the audio signal. Specifically, the computing device 102 can calculate the probability distributions 144 and 146 of each of the tonic levels 132 and 134 at specific timestamps 156 and 158. For example, Figure 3A master key 300 according to example embodiments of the present disclosure is shown. The master key 300 includes probabilities of each of the keys B, sharp A / flattened B, A, sharp G / flattened A, G, sharp F / flattened G, F, E, sharp D / flattened E, D, sharp C / flattened D, C at each of the timestamps T1-T19 (as defined in the legend 302). Specifically, at times T1-T5, the keys A and D have a medium-high probability; at times T6 and T14, the keys D and flattened D / E have a medium-high probability; at times T7, T11, and T19, the keys E and flattened D / E have a medium-high probability; at times T8 and T12, the keys C and flattened C / D have a medium-high probability; at times T9, T10, T13, T17, and T18, the key E has a high probability; at time T15, the key E has a high probability; at time T16, the key F has a high probability. The probabilities can be calculated to reflect the probability of each key representing the master key at a given point in time. For example, in some cases, the probabilities can be calculated by a Hidden Markov Model (HMM). In some cases, the HMM can be adjusted to optimize the number of transitions in the master key (e.g., to optimize the number of changes in the carrier frequency of the neural beat 168, 174), which can cause the user to be distracted and / or can have a detrimental effect on neural synchronization.
[0039] Returning to Figure 1A , the computing device 102 can determine the carrier frequency 114 based on the master key 120. The carrier frequency 114 can include a single selected frequency 160, 162 at each timestamp 164, 166 to use as the carrier frequency at that time within the neural beat 168, 174. For example, Figure 4A carrier frequency 400 is depicted in accordance with example embodiments of the present disclosure. The carrier frequency 400 includes a single selected pitch class at each timestamp T1-19. In particular, pitch class D is selected as the carrier frequency at timestamps T1-14 and T17-19, and pitch class E is selected as the carrier frequency at timestamps T15-16. The carrier frequency can be selected to follow the music and harmony of the digital audio file, while also avoiding unnecessary changes in the carrier frequency. In particular, the carrier frequency at times T9-T10 and T17-20 can be selected to be pitch class D to align with the primary pitch class at these times. However, excessive changes in the carrier frequency can distract the user, so in some cases, such as when selecting between different pitch classes with similar probabilities or making small, transient changes in the primary pitch class, the selected carrier frequency can be selected to remain consistent over time. For example, in the primary pitch class 300, pitch classes A and D have similar probabilities at times T1-5. However, pitch class D can be selected as the carrier frequency from times T1-5 to avoid transitioning from pitch class A to pitch class D at time T6, where pitch class D is the primary pitch class. As another example, at times T7, T11, T19, pitch classes raise D / lower E and E both have similar probabilities. However, pitch class D can be selected as the carrier frequency even though it does not have the highest probability at these times to reduce the number of changes in the carrier frequency (e.g., because pitch class D still has a moderate probability in the primary pitch class 300). On the other hand, failing to follow the music and harmony can also have a negative impact on neural synchronization. Thus, at times T17-T19, the carrier frequency is switched from E (at time T16) to D (at times T17-T19) to properly follow the harmony in the digital audio file.
[0040] To select the carrier frequency 400, the computing device 102 can be configured to balance maximizing the overall probability of the selected carrier frequency while limiting the number of changes in consecutive primary pitch classes. In some implementations, the computing device 102 can perform Viterbi decoding on the primary pitch class 300 to find the most likely sequence of pitch classes at each timestamp that limits the number of carrier frequency transitions while also ensuring that the carrier frequency 400 is musically aligned with the digital audio file 106.
[0041] Returning to Figure 1AThe computing device can synthesize a neural beat 168 based on the beat frequency 112 and the carrier frequency 114. Specifically, the computing device 102 can synthesize the neural beat 168 by modulating the beat frequency 112 onto the selected carrier frequencies 160, 162 at each of timestamps 164, 166. In some embodiments, the timestamps 164, 166 within the carrier frequencies (e.g., timestamps T1-19) can correspond to the timestamps of the audio data within the digital audio file 106. In such cases, the computing device 102 can directly synthesize the neural beat 168 based on the carrier frequency at each of the timestamps 164, 166.
[0042] In some implementations, computing device 102 may further adjust one or more aspects of neural beat 168 based on other characteristics of digital audio file 106. For example, computing device 102 may adjust the volume of neural beat 168 to align with volume changes in digital audio file 106. Specifically, if neural beat 168 is relatively quiet compared to digital audio file 106, the benefits of neural beat may be reduced. Alternatively or additionally, if neural beat 168 is louder relative to digital audio file 106, neural beat 168 may prove to be intrusive or distracting to the user, thereby negating the benefits provided by neural beat 168. Therefore, audio mixer 122 may be used to adjust the volume of neural beat 168 during the processing of digital audio file 106.
[0043] Specifically, the audio mixer 122 can determine a loudness profile 170 of the digital audio file 106. The loudness profile 170 is a representation of how loud the audio encoded in the digital audio file 106 is over time (e.g., within the duration of the audio encoded in the digital audio file 106). The loudness profile 170 can be calculated as a combined intensity (e.g., across audible frequencies) at multiple timestamps within the digital audio file 106. The loudness profile 170 can then be used to generate a volume curve 172 of a neural beat 168. Specifically, the loudness profile 170 can be offset (e.g., based on the maximum expected intensity of the neural beat 168) to generate the volume curve 172. For example, Figure 5 A volume curve 500 is depicted according to an exemplary embodiment of the present disclosure. Volume curve 500 illustrates the change in energy (e.g., in dB) over the duration of audio encoded in the digital audio file 106, wherein the energy of the audio signal within the digital audio file 106 can be used as a proxy for the volume over time within the digital audio file 106. Return Figure 1AThe volume curve 172 can be applied to the neural beat 168 to generate an adjusted neural beat 174. Specifically, applying the volume curve 172 to the neural beat 168 may include increasing or decreasing the volume (e.g., intensity) of the neural beat 168 at different time points according to the intensity indicated in the volume curve 172 (e.g., making the adjusted neural beat 174 louder at high intensity in the volume curve 172 and quieter at low intensity in the volume curve 172).
[0044] Neural beats 168 and / or adjusted neural beats 174 may subsequently be stored, transmitted, and / or played back on a user device. For example, computing device 102 may store neural beats 168 and / or adjusted neural beats 174 in association with a digital audio file 106 (e.g., in server 104). In some embodiments, digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 may be stored separately. In additional or alternative embodiments, computing device 102 may combine digital audio file 106 with neural beats 168 and / or adjusted neural beats 174 to generate a combined audio track that can be stored (e.g., in server 104). As another example, and referring to... Figure 1B With system 190, digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 can be transmitted to user device 192 associated with user 194. User device 192 may include a smartphone, tablet, wearable computing device, laptop, personal computer, or any other personal computing device. User device 192 may also include one or more audio devices for audio playback, such as speakers, a 3.5mm audio jack connected to headphones or speakers, wirelessly connected headphones, wirelessly connected speakers, or any other device capable of playing back audio. System 100 may transmit (e.g., stream) digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 to user device 192. User device 192 can then receive and play back digital audio file 106 simultaneously with neural beats 168 and / or adjusted neural beats 174. Additionally or alternatively, user device 192 may store digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 for future playback. Alternatively, computing device 102 may transmit a combined audio track to user device 192. In yet another embodiment, neural beats 168 and / or adjusted neural beats 174 may be generated on user device 192. In such a case, neural beats 168 and / or adjusted neural beats 174 may be played on user device 192 along with digital audio file 106 (e.g., as separate audio files, as a combined audio track) and / or may be stored on user device 192 for later playback.
[0045] Although not depicted, the computing device 102, the server 104, and / or the user device 192 can include at least one processor and / or memory configured to implement one or more aspects of the computing device 102, the server 104, and / or the user device 192. For example, the memory can store instructions that, when executed by the processor, can cause the processor to perform one or more operational features of the computing device 102, the server 104, and / or the user device 192. The processor can be implemented as one or more central processing units (CPUs), field-programmable gate arrays (FPGAs), and / or graphics processing units (GPUs) configured to execute instructions stored on the memory. Additionally, the computing device 102, the server 104, and / or the user device 192 can be configured to communicate using a network. For example, the computing device 102, the server 104, and / or the user device 192 can communicate with the network using one or more wired network interfaces (e.g., Ethernet interfaces) and / or wireless network interfaces (e.g., Wi-Fi interfaces and / or cellular data interfaces). In some cases, the network can be implemented as a local network (e.g., a local area network), a virtual private network, the L1, and / or a global network (e.g., the Internet). and / or cellular data interfaces). In some cases, the network can be implemented as a local network (e.g., a local area network), a virtual private network, the L1, and / or a global network (e.g., the Internet).
[0046] In some implementations, the computing device 102 and the server 104 can be implemented as a single computing device. For example, the computing device 102 can store the digital audio files 106, 108, 110 (e.g., in a local database). In another implementation, the computing device 102 and / or the server 104 can be implemented, at least in part, by the user device 162. In yet another implementation, the computing device 102, the server 104, and / or the user device 192 can be implemented by multiple computing devices. For example, the computing device 102 can be implemented as multiple software services executing in a distributed computing environment (e.g., a cloud computing environment). As another example, the user device 162 can be implemented by multiple personal computing devices (e.g., a smartphone and a wearable computing device such as a smartwatch).
[0047] Figure 6 A method 600 of synthesizing a neural beat is shown in accordance with example embodiments of the present disclosure. The method 600 can be implemented on a computer system such as the system 100, 160. For example, the method 600 can be implemented by the computing device 102 and / or the user device 192. The method 600 can also be implemented by a set of instructions stored on a computer-readable medium that, when executed by a processor, cause a computer system to perform the method 600. For example, all or part of the method 600 can be implemented by the processor and / or memory of the computing device 102 and / or the user device 192. Although reference is made to the system 100, 160, the method 600 can be implemented by any suitable computing device, such as the computing device 102 and / or the user device 192. Figure 6 The flowchart shown describes the following example, but other examples can be used that perform the same or similar functions asFigure 6 Many other methods of associated actions. For example, the order of some boxes can be changed, some boxes can be combined with other boxes, one or more boxes can be repeated, and some of the boxes described can be optional.
[0048] Method 600 may begin by receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file (block 602). For example, computing device 102 may receive a digital audio file 106 and a beat frequency 112 of a neural beat to be added to the digital audio file 106. As described above, computing device 102 may receive the digital audio file 106 from server 104 and / or retrieve the digital audio file 106 from local storage. In some embodiments, the digital audio file 106 may be received upon a user request. For example, a user request to play a specific song may be received from a user device (e.g., via a music streaming service). Computing device 102 may receive the beat frequency 112 from the user (e.g., based on a user request and / or previously defined user settings). In some embodiments, the beat frequency 112 may specify a specific frequency (e.g., 3Hz) for the neural beat to be added to the digital audio file 106. In additional or alternative embodiments, the beat frequency 112 may specify a frequency range of the neural beat (e.g., 4Hz to 8Hz).
[0049] A plurality of chromagram features can be extracted from the digital audio file (block 604). For example, the computing device 102 can extract a plurality of chromagram features 116, 200 from the digital audio file 106. As discussed above, the chromagram features can include intensity information for a plurality of loudness levels at a plurality of time stamps within the digital audio file 106. In certain implementations, each of the plurality of chromagram features 116, 200 can be extracted according to different parameters applied to the digital audio file 106 prior to extracting the chromagram features 116, 200. For example, a first chromagram feature can be extracted that focuses on lower frequencies of the digital audio file 106, and a second chromagram feature can be extracted that focuses on higher frequencies of the digital audio file 106. As another example, three chromagram features can be extracted from the digital audio file 106: a first chromagram feature extracted that focuses on lower frequencies (e.g., less than 200 Hz), a second chromagram feature extracted that focuses on medium frequencies (e.g., from 200 Hz to 800 Hz), and a third chromagram feature extracted that focuses on higher frequencies (e.g., greater than 800 Hz). In practice, after generating the time-frequency representation as discussed above, the plurality of chromagram features 116, 200 can be generated by selecting octaves and intensities within a desired frequency range to include in the chromagram features 116, 200. In other implementations, the plurality of chromagram features can be generated by applying filters (e.g., high-pass filters, low-pass filters, band-pass filters, etc.) to the digital audio file 106 prior to extracting the chromagram features 116, 200 (e.g., using FFTs, constant-Q transforms, filter banks, and / or other techniques, as discussed above).
[0050] The multiple chromagram features can be combined to form a primary chromagram feature for the digital audio file (block 606). For example, the computing device 102 can combine the multiple chromagram features 116, 200 to form the primary chromagram feature 118 for the digital audio file 106. In certain implementations, the multiple chromagram features 116, 200 can be linearly combined to form the primary chromagram feature 118 (e.g., according to previously defined weights). In additional or alternative implementations, the multiple chromagram features 116, 200 can be combined according to any other conceivable combination strategy. For example, the multiple chromagram features 116, 200 can be combined by "stacking" the chromagram features 116, 200 (e.g., such that combining two chromagram features 116, 200 having 12 octaves results in a primary chromagram feature having 24 rows). Generating the primary chromagram feature 118 based on the multiple chromagram features 116 can better capture the audio frequency characteristics of the digital audio file 106 (e.g., by separately focusing on different frequency ranges within the digital audio file 106, such as different octaves). In certain implementations, one or both of blocks 604, 606 can be omitted. For example, in certain implementations, a single set of chromagram features can be extracted from the digital audio file 106 and used as the primary chromagram feature 118, rather than extracting multiple chromagram features and combining them to form a primary chromagram feature.
[0051] The primary chromagram feature can be extracted at multiple timestamps within the digital audio file (block 608). For example, the computing device 102 can extract the primary chromagram features 120, 300 at multiple timestamps 156, 158 within the digital audio file 106. The primary chromagram features 120, 300 can be extracted from the primary chromagram feature 118 using a model (e.g., a hidden Markov model). In particular, the primary chromagram features 120, 300 can be extracted as probability distributions at multiple timestamps T1-19. As described above, the timestamps T1-19 can be selected based on the timestamps of the primary chromagram feature 118.
[0052] A plurality of carrier frequencies can be selected for the neuro-tempo (block 610). For example, the computing device 102 can select a plurality of carrier frequencies 114, 400 for the neuro-tempo 168, 174. In particular, the plurality of carrier frequencies 114, 400 can include separate carrier frequencies 160, 162 at the plurality of timestamps 164, 166, T1-19. The selected carrier frequencies 114, 400 can be selected by a Viterbi process, which can select carrier frequencies such that transitions of carrier frequencies from adjacent time periods are optimized according to a transition probability, as will be further explained herein. In certain implementations, in addition to selecting the plurality of carrier frequencies 114, 400, a particular tempo frequency for the neuro-tempo 168 can be selected. For example, in cases where the tempo frequency 112 is received as an acceptable frequency range, the computing device 102 can select a tempo frequency for the neuro-tempo 168 from within the acceptable range, as will be further discussed below.
[0053] A sync-tempo refers to a tempo that varies the carrier frequency at different time periods of a digital audio file according to variations in an underlying audio signal, e.g., varying the music and harmony and / or melody at different time periods of the audio signal. A sync-tempo can be synthesized for a digital audio file based on a tempo frequency and a plurality of carrier frequencies (block 612). For example, the computing device 102 can synthesize the neuro-tempo 168 for the digital audio file 106 based on the tempo frequency 112 and the carrier frequencies 114. In particular, the neuro-tempo 168 can be generated by modulating the tempo frequency 112 at times corresponding to the timestamps 164, 166, T1-19 within the carrier frequencies 114, 400 on two different carrier frequencies 160, 162. In this way, the neuro-tempo 168 can be synchronized with variations in the music and harmony and / or melody at different time periods within the digital audio file 106. In certain implementations, the neuro-tempo 168 can be synthesized to include a single audio channel (e.g., as a mono-tempo). In additional or alternative implementations, the neuro-tempo 168 can be synthesized to include two audio channels (e.g., as a binaural tempo with two channels, as a mono-tempo with two channels). In yet further implementations, the neuro-tempo 168 can be synthesized to include more than two audio channels (e.g., three audio channels, four audio channels, five audio channels). In certain implementations, the number of audio channels can be specified by a user or a predetermined setting. In additional or alternative implementations, the number of audio channels can be selected based on the number of audio channels in the digital audio file 106 (e.g., such that the neuro-tempo 168 has the same number of audio channels as the digital audio file 106).
[0054] At least one of the synchronized neurobeat and the combined neurobeat and digital audio file combination track can be stored (block 614). For example, the computing device 102 can store at least one of the synchronized neurobeat 168 or the combined neurobeat 168 and digital audio file 106 combination track. For example, as described above, the computing device 102 can store the neurobeat 168 and / or the combined track on a server 104 and / or local storage within the computing device 102. Additionally or alternatively, the computing device 102 can transmit the neurobeat 168 and / or the combined track to a user device for storage and playback (e.g., temporary storage for streaming, and long-term storage). In implementations where the computing device 102 is a user device, the computing device 102 can store the neurobeat 168 and / or the combined track locally for current or future playback. In certain implementations, as explained further above, the computing device 102 can be further configured to generate an adjusted neurobeat 174 based on the neurobeat 168. In such cases, the computing device 102 can be configured to store the equalized neurobeat 174 and / or combine the adjusted neurobeat 174 with a digital audio file 106 in a combined track in a manner similar to that discussed above.
[0055] In this way, the method 600 enables a computing device to generate a neurobeat for an arbitrary digital audio file, thereby allowing for increased user selection in the types of music used to produce neurosynchronization. Moreover, the computing device is able to do so in real-time and can ensure that the neurobeat is blended with the tonal quality of the digital audio file and / or the loudness of the digital audio file to minimize the user’s distraction and maximize neurosynchronization. Thus, the method 600 ensures that the generated neurobeat constructively combines with a previously created digital audio file.
[0056] Figures 7A-7CMethods 700, 710, and 720 are shown that are in accordance with example embodiments of the present disclosure. Methods 700, 710, 720 can be performed in conjunction with at least a portion of method 600. For example, method 700 can be performed while implementing blocks 608, 610 of method 600. As another example, method 710 can be performed between blocks 612 and 614 and / or a portion of block 612 of method 600. As another example, method 720 can be performed as a portion of block 612 of method 600. Methods 700, 710, 720 can be implemented on a computer system, such as system 100, 190. For example, methods 700, 710, 720 can be implemented by computing device 102 and / or user device 192. Methods 700, 710, 720 can also be implemented by a set of instructions stored on a computer readable medium that, when executed by a processor, cause a computer system to perform methods 700, 710, 720. For example, all or a portion of methods 700, 710, 720 can be implemented by a processor and / or memory of computing device 102 and / or user device 192. Although reference is made to the flowcharts shown, many other methods of performing the actions associated therewith can be used. For example, the order of some blocks can be changed, certain blocks can be combined with other blocks, one or more blocks can be repeated, and some blocks described can be optional. Figures 7A-7C The flowcharts shown describe the following examples, but many other methods of performing the actions associated therewith can be used. For example, the order of some blocks can be changed, certain blocks can be combined with other blocks, one or more blocks can be repeated, and some blocks described can be optional. Figures 7A-7C The flowcharts shown describe the following examples, but many other methods of performing the actions associated therewith can be used. For example, the order of some blocks can be changed, certain blocks can be combined with other blocks, one or more blocks can be repeated, and some blocks described can be optional.
[0057] Method 700 can be performed to select a plurality of carrier frequencies for a neural beat. Method 700 can begin with generating a probability distribution of tonal levels at a plurality of timestamps (block 702). For example, a probability distribution of tonal levels (e.g., B, Bb / A, A, Ab / G, G, F, E, Eb / D, D, C, Db / D) within digital audio file 106 at a plurality of timestamps T1-19 within digital audio file 106 can be generated using a hidden Markov model. Timestamps T1-19 can be selected based on timestamps within primary chromagram features 118 (e.g., based on a segment of digital audio file 106 used to compute time-frequency representations for chromagram features 116 and / or primary chromagram features 118). The hidden Markov model can be configured by adjusting transition probabilities to select times at which transitions between different carrier frequencies should occur. In particular, based on input received from a user, system administrator, and / or computing process, transition probabilities (e.g., transition probabilities of 0.005-0.02) of the hidden Markov model can have been previously received (or can be updated).
[0058] Then, a sequence of tonic levels can be identified within the probability distribution (box 704). For example, computing device 102 can identify a sequence of tonic levels within the probability distribution. Specifically, carrier frequencies 114, 400 can contain a series of tonic levels of carrier frequencies that will be used as neural beats 168. The sequence of tonic levels can be identified based on the constrained transition probabilities of changes in the selected tonic levels to maximize the combination probability of the selected tonic levels within the probability distribution. Specifically, the sequence of tonic levels can be selected via a Viterbi process implemented by computing device 102.
[0059] In this way, method 700 can be executed to select a sequence of carrier frequencies based on the musical harmonies and melodies (e.g., chroma map features) of the received digital audio file. Therefore, this process allows neural beats 168 to be applied to existing digital audio files while ensuring that changes in carrier frequencies do not disturb or distract users attempting to trigger neural synchronization using neural beats.
[0060] Method 710 can be performed to adjust the volume of neural beat 168 based on the volume of the digital audio file 106 at different times within the digital audio file 106. Method 710 can begin by generating a loudness profile for the duration of the digital audio file (box 712). For example, computing device 102 (e.g., audio mixer 122) can generate a loudness profile 170 for the duration of the digital audio file 106. The loudness profile 170 can be generated based on the intensity (e.g., volume) of the digital audio file 106 at multiple times within the digital audio file 106. For example, a loudness profile 170 can be generated for each data sample timestamp within the digital audio file 106.
[0061] A volume curve can be formed based on a loudness profile (box 714). For example, computing device 102 can form a volume curve 172 based on loudness profile 170. Volume curve 172 can be formed as a percentage of loudness profile 170 (e.g., 50% of loudness profile 170). Alternatively or additionally, volume curve 172 can be formed by normalizing loudness profile 170 to the maximum volume desired for neural beat 168. Those skilled in the art will similarly recognize one or more additional means of generating volume curve 172 based on loudness profile 170 of digital audio file 106. Therefore, all such similar embodiments are considered to be within the scope of this disclosure.
[0062] The volume of the synchronized neural beat can then be adjusted based on the volume curve (box 714). For example, computing device 102 can adjust the volume of neural beat 168 based on volume curve 172 to generate an adjusted neural beat 174. For example, the intensity of neural beat 168 can be scaled to match the desired volume reflected in volume curve 172.
[0063] In this way, the method 710 can be performed to adjust the neural beat 168. This can reduce the amount of volume mismatch between the neural beat and the digital audio file. For example, in cases where the volume of the neural beat is much lower than the digital audio file, a user can not be able to hear the volume of the neural beat, thereby reducing its effectiveness in producing neural synchrony. As another example, in cases where the volume of the neural beat is much higher than the digital audio file 106, a user can be distracted or interrupted by the volume difference, thereby interrupting or reducing any neural synchrony produced by the neural beat.
[0064] The method 720 can be used to synchronize the neural beat 168 with the rhythmic pattern in the digital audio file 106. The method 720 can begin by estimating the location of rhythmic beats within the digital audio file (block 722). A rhythmic beat refers to a level of rhythmicity in a piece of music - this is different from the “beat” usage in acoustics. For example, the computing device 102 can estimate the location of rhythmic beats within the digital audio file 106. The location of rhythmic beats within the digital audio file 106 can be estimated using a machine learning model (e.g., a pre-trained network configured to detect rhythmic beats within an audio file). For example, the location of rhythmic beats can be estimated using one or more models provided by the madmom audio package, the Essentia audio package, etc. In additional or alternative implementations, the location of rhythmic beats can be estimated using one or more algorithmic techniques.
[0065] The timing of the synchronized neural beat can be adjusted based on the location of the rhythmic beats within the digital audio file (block 724). For example, the computing device 102 can adjust the timing of the neural beat 168 based on the location of the rhythmic beats. For example, the computing device 102 can adjust the beat frequency 112 to align with (e.g., be a multiple of) the rhythm of the digital audio file. For example, in cases where the digital audio file 106 has a rhythm of 120 bpm and the beat frequency 112 is 0.6 Hz (e.g., 100 bpm), the computing device 102 can adjust the beat frequency 112 to be an integer multiple of the 120 beat per minute (e.g., 2 Hz) rhythm. As a specific example, the computing device 102 can adjust the beat frequency 112 to be 0.5 Hz (30 bpm) and / or 1 Hz (60 bpm). In implementations where the user has specified a desired frequency range for the beat frequency 112, the beat frequency 112 can be selected from within the desired frequency range to be an even multiple of the rhythm frequency and / or as close to a multiple of the rhythm frequency as possible. Furthermore, the timing of the synchronized neural beat can be adjusted such that peaks in the neural beat (e.g., peaks at the beat frequency 112) occur at the same time as (e.g., align with the timing of) the rhythmic beats within the digital audio file 106.
[0066] In this way, the method 720 can be used to ensure that the tempo beats and tempo frequency within the digital audio file are not out of phase. In particular, when the tempo frequency is out of phase with the tempo frequency of the digital audio file, interference between the tempo frequencies in the digital audio file can negatively impact sound quality and / or can create a distracting or interfering interference pattern when the digital audio file and the neural beat at the interfering tempo frequency are played simultaneously. Accordingly, adjusting the tempo frequency based on the tempo beats within the digital audio file can reduce these interferences, thereby improving the quality of the subsequently generated neural beat and / or the quality of the neural synchronization produced by the neural beat.
[0067] Figure 8 An example computer system 800 that can be used to implement one or more devices and / or components discussed herein, such as the computing device 102, is shown. In particular embodiments, one or more computer systems 800 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 800 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 800 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 800. Herein, reference can be made to a computer system including a computing device, where appropriate, and vice-versa, where appropriate. Furthermore, reference can be made to a computer system including one or more computer systems, where appropriate.
[0068] The present disclosure contemplates any suitable number of computer systems 800. The present disclosure contemplates computer system 800 taking any suitable physical form. As example and not by way of limitation, computer system 800 can be an embedded computer system, a system on a chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computer system 800 can include one or more computer systems 800; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which can include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 800 can perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example and not by way of limitation, one or more computer systems 800 can perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 800 can perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.
[0069] In particular embodiments, computer system 800 includes a processor 806, memory 804, storage 808, an input / output (I / O) interface 810, and a communication interface 812. Although this disclosure describes and illustrates a particular computer system having particular components arranged in a particular fashion, this disclosure contemplates any suitable computer system having any suitable components arranged in any suitable fashion.
[0070] In particular embodiments, processor 806 includes hardware for executing instructions (e.g., those making up a computer program). As an example and not by way of limitation, to execute instructions, processor 806 can retrieve (or fetch) the instructions from an internal register, an internal cache, memory 804, or storage 808; decode and execute them; and then write one or more results to the internal register, internal cache, memory 804, or storage 808. In particular embodiments, processor 806 can include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 806 including any suitable number of any suitable internal caches, where appropriate. As an example and not by way of limitation, processor 806 can include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches can be copies of instructions from memory 804 or storage 808, and the instruction caches can speed up retrieval of those instructions by processor 806. Data in the data caches can be copies of data values stored in memory 804 or storage 808 for instructions executing
[0071] In particular embodiments, memory 804 includes main memory for storing instructions executable by processor 806 or data upon which processor 806 operates. By way of example, and not limitation, computer system 800 can load instructions from storage device 808, or another source such as another computer system 800, into memory 804. Processor 806 can then load the instructions from memory 804 into internal registers or internal cache. To execute the instructions, processor 806 can retrieve the instructions from the internal registers or internal cache and decode them. During or after execution of the instructions, processor 806 can write one or more results (which can be intermediate or final results) to the internal registers or internal cache. Processor 806 can then write one or more of these results to memory 804. In particular embodiments, processor 806 executes only instructions in one or more internal registers or internal cache or in memory 804 (as opposed to storage device 808 or elsewhere), and only operates on data in one or more internal registers or internal cache or in memory 804 (as opposed to storage device 808 or elsewhere). One or more memory buses (each of which can include an address bus and a data bus) can couple processor 806 to memory 804. The bus(es) can include one or more memory buses, as will be described in further detail below. In particular embodiments, one or more memory management units (MMUs) reside between processor 806 and memory 804 and facilitate the processor’s 806 requested access to memory 804. In particular embodiments, memory 804 includes random access memory (RAM). Where appropriate, this RAM can be volatile memory or non-volatile memory, or both. Where appropriate, this RAM can be dedicated to at least one of the processor(s) 806. Also, to the extent practical, the RAM can be a
[0072] In particular embodiments, storage 808 includes mass storage for data or instructions. As an example and not by way of limitation, storage 808 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc (e.g., a compact disc or DVD, etc.), a solid-state drive (SSD), a tape drive, a USB drive, a memory card, or any suitable combination of two or more of these. Storage 808 can include removable or non-removable (or fixed) media, where appropriate. Storage 808 can be internal or external to computer system 800, where appropriate. In particular embodiments, storage 808 is nonvolatile, solid-state memory. In particular embodiments, storage 808 includes read-only memory (ROM). Where appropriate, this ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or any suitable combination of two or more of these. This disclosure contemplates mass storage 808 taking any suitable physical form. Storage 808 can include one or more storage control units (SCUs) for controlling a
[0073] In particular embodiments, I / O interface 810 includes hardware, software, or both, providing one or more interfaces for the transfer of information between computer system 800 and one or more I / O devices. Where appropriate, computer system 800 can include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 800. As an example and not by way of limitation, an I / O device can include a keyboard, keypad, microphone, monitor, screen, display, touch screen, touchpad, touchpad, scanner, speaker, still camera, stylus, tablet, dilog, tracker, video camera, another suitable I / O device, or a combination of two or more of these. An I / O device can include one or more sensors. Where appropriate, I / O interface 810 can include one or more device or software drivers enabling processor 806 to drive one or more of these I / O devices. I / O interface 810 can include one or more I / O interfaces 810, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface or combination of I / O interfaces.
[0074] In particular embodiments, communication interface 812 includes hardware, software, or both providing one or more interfaces for communication (such as packet-based communication) between computer system 800 and one or more other computer systems 800 or one or more networks 814. As an example and not by way of limitation, communication interface 812 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or any other wired network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a Wi-Fi network. This disclosure contemplates any suitable network 814 and any suitable communication interface 812 for it. As an example and not by way of limitation, network 814 can include one or more of a direct connection, such as a bus or wire, use of one or more networks 814, such as an Ethernet network, a Fiber Optic network, and / or a
[0075] Computer system 802 can also include a bus. Bus can include hardware, software, or both, and can couple various components of computer system 800 to one another. As an example and not by way of limitation, bus can include an Accelerated Graphics Port (AGP) or any other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) or another suitable bus or interconnect, or a combination of two or more of these. Where appropriate, bus can include one or more buses of the same type or buses of different types. As an example and not by way of limitation, computer system 800 can include buses to communicate with both PCI bus and PCIe bus. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0076] Herein, a computer-readable non-transitory storage medium can include one or more of a semiconductor-based or other type of integrated circuit (IC) (e.g., a field-programmable gate array (FPGA) or an application-specific IC (ASIC)), a hard disk drive (HDD), a hybrid hard drive (HHD), an optical disc, an optical disc drive (ODD), a magneto-optical disc, a magneto-optical drive, a floppy disk, a floppy disk drive (FDD), a magnetic tape, a solid-state drive (SSD), a RAM drive, a SECURE DIGITAL card or drive, any other appropriate computer- readable non-transitory storage medium, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium can be volatile, non-volatile, or a combination of volatile and non-volatile memory, where appropriate. Where appropriate, a computer-readable non-transitory storage medium can include one or more separate components, or multiple components housed within a single physical enclosure.
[0077] Herein, "or" is inclusive and not exclusive, unless expressly indicated or indicated otherwise by context. Therefore, herein, "A or B" means "A, B, or both," unless expressly indicated otherwise or indicated by context. Moreover, "and" is both joint and several, unless expressly indicated otherwise or indicated by context. Therefore, herein, "A and B" means "A and B, jointly or severally," unless expressly indicated otherwise or indicated by context.
[0078] The scope of the disclosure includes all changes, substitutions, alternatives, variations, amendments, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art can contemplate. The scope of the disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although the disclosure describes and illustrates various embodiments herein as including particular components, elements, features, functions, operations, or steps, any one of these embodiments can include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art can contemplate as can be necessary, expedient, or desirable, as will be apparent to those ordinarily skilled in the art. Moreover, although the appended claims
[0079] including the device, system, component, whether the particular function is activated, turned on, or unlocked, as long as the device, system, or component is so adapted, arranged, enabled, configured, enabled, made operable, or operable. Moreover, although the disclosure describes or illustrates particular embodiments as providing particular advantages, an embodiment can not provide the particular advantage, some but not all of the advantages, or even all of the advantages.
[0080] All of the disclosed methods and processes described in this disclosure can be implemented using one or more computers or components. These components can be provided as a series of computer instructions on any conventional computer-readable medium or machine-readable medium, including volatile and non-volatile memory, such as RAM, ROM, flash memory, magnetic or optical disks, optical memory or other storage media. The instructions can be provided as software or firmware, and can be implemented in whole or in part in hardware components such as ASICs, FPGAs, DSPs, or any other similar devices. The instructions can be configured to be executed by one or more processors that perform or facilitate the performance of all or part of the disclosed methods and processes when the series of computer instructions are executed.
[0081] It is to be understood that various alterations and modifications to the examples described herein will be apparent to those skilled in the art. Such alterations and modifications are intended to be within the scope of the subject matter. It is the intent of the claims to cover all such alterations and modifications.
Claims
1. A method for generating a synchronized neurobeat, comprising: receiving, at a server or a user device, a digital audio file and a beat frequency of a neurobeat to be added to the digital audio file; extracting a plurality of chromagram features from the digital audio file; combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file; extracting a primary sound level at a plurality of timestamps within the digital audio file from the primary chromagram feature; selecting a plurality of carrier frequencies for the neurobeat based on the primary sound level at the plurality of timestamps; synthesizing a synchronized neurobeat of the digital audio file based on the beat frequency and the plurality of carrier frequencies; storing, at the server or the user device, at least one of: the synchronized neurobeat, a combined audio track that combines the synchronized neurobeat and the digital audio file; and causing an audio output device to: play the combined audio track, or play the synchronized neurobeat in parallel with the digital audio file.
2. The method of claim 1, wherein, the primary chromagram feature comprises an intensity of each of a plurality of sound levels at the plurality of timestamps, and wherein the primary sound level is selected from the plurality of sound levels.
3. The method of claim 2, wherein, extracting the primary sound level further comprises generating a probability distribution for each of the plurality of sound levels at the plurality of timestamps based on the intensity of the plurality of sound levels using a hidden Markov model.
4. The method of claim 3, wherein, the hidden Markov model is configured to optimize a number and location of transitions between primary sound levels.
5. The method of claim 3, wherein, extracting the primary sound level further comprises identifying a sequence of primary sound levels within the probability distribution.
6. The method of any one of claims 1 to 5, wherein, the plurality of timestamps occur once every 500 milliseconds or less during the digital audio file.
7. The method of any one of claims 1 to 5, wherein, the plurality of chromagram features are linearly combined to form the primary chromagram feature.
8. The method of any one of claims 1 to 5, further comprising: adjusting a volume of the synchronized neurobeat over time to follow a volume of the digital audio file.
9. The method of claim 8, wherein, normalizing the volume of the synchronized neurobeat comprises: generating a loudness profile for a duration of the digital audio file; forming a volume curve based on the loudness profile; and adjusting the volume of the synchronized neurobeat according to the volume curve.
10. The method of any one of claims 1-5, further comprising aligning the beat frequency with a rhythmic beat within the digital audio file.
11. The method of claim 10, wherein, aligning the beat frequency comprises: estimating a location of a rhythmic beat within the digital audio file; estimating a musical tempo within the digital audio file; and adjusting a timing of the synchronized neurobeat according to the musical tempo to align peaks within the synchronized neurobeat with the location of the rhythmic beat within the digital audio file.
12. The method of any one of claims 1 to 5, wherein, the neurobeat is at least one of: (i) a binaural beat, (ii) a monaural beat.
13. The method of any one of claims 1 to 5, wherein, the synchronized neurobeat comprises two or fewer audio channels.
14. The method of any one of claims 1 to 5, wherein, the synchronized neurobeat comprises three or more audio channels.
15. The method of any one of claims 1 to 5, wherein, the beat frequency is greater than or equal to 0.5 Hz and less than or equal to 150 Hz.
16. The method of any one of claims 1-5, further comprising playing the synchronized neurobeat and the digital audio file in parallel via a computing device.
17. The method of claim 16, further comprising streaming the synchronized neuro beat and the digital audio file to the computing device for playback by the computing device.
18. A system for generating a synchronized neuro beat, comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to perform the following operations: receiving a digital audio file and a beat frequency of a neuro beat to be added to the digital audio file; extracting a plurality of chromagram features from the digital audio file; combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file; extracting a primary sound level at a plurality of timestamps within the digital audio file from the primary chromagram feature; selecting a plurality of carrier frequencies for the neuro beat based on the primary sound level at the plurality of timestamps; synthesizing a synchronized neuro beat of the digital audio file based on the beat frequency and the plurality of carrier frequencies; storing at least one of the following at the memory: the synchronized neuro beat, a combined audio track combining the synchronized neuro beat and the digital audio file; and causing an audio output device to: play the combined audio track, or play the synchronized neuro beat in parallel with the digital audio file.
19. The system of claim 18, wherein, the primary chromagram feature comprises an intensity of each of the plurality of sound levels at the plurality of timestamps, and wherein the primary sound level is selected from the plurality of sound levels.
20. The system of claim 19, wherein, the memory stores further instructions that, when executed by the processor in extracting the primary sound level, cause the processor to generate a probability distribution for each of the plurality of sound levels at the plurality of timestamps based on the intensities of the plurality of sound levels using a hidden Markov model.
21. A method for generating a synchronized neuro beat, comprising: receiving, at a server, a digital audio file and a beat frequency of a neuro beat to be added to the digital audio file; extracting a plurality of chromagram features from the digital audio file; combining the plurality of chromagram features to form a primary chromagram feature of the digital audio file; extracting a primary sound level at a plurality of timestamps within the digital audio file from the primary chromagram feature; selecting a plurality of carrier frequencies for the neuro beat based on the primary sound level at the plurality of timestamps; synthesizing a synchronized neuro beat of the digital audio file based on the beat frequency and the plurality of carrier frequencies; creating a combined audio track combining the synchronized neuro beat and the digital audio file; and sending the combined audio track combining the synchronized neuro beat and the digital audio file to a user device.
Citation Information
Patent Citations
Music transcription
CN101652807A
Device, method, and medium for integrating auditory beat stimulation into music
WO2020220140A1