Generating tone-compatible synchronized neural metronomes for digital audio files

By analyzing the pitch characteristics of audio files to generate chroma map features, selecting carrier frequencies, and synthesizing synchronized neural beats, the problem of adding neural beats to existing audio tracks in existing technologies is solved, and the effect of improving neural synchronization and attention in user-preferred audio tracks is achieved.

CN121747504APending Publication Date: 2026-03-27UNIVISER INT MUSIC CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively add neural beats to existing audio tracks, preventing users from experiencing the benefits of neural synchronization, relaxation, and enhanced attention while listening to their preferred tracks.

Method used

By analyzing the pitch characteristics of digital audio files, generating chroma map features, selecting carrier frequencies, synthesizing synchronized neural beats, and adjusting volume and rhythm to align with existing audio files, the automatic addition of neural beats is achieved.

Benefits of technology

It enables the embedding of neural beats into user-preferred audio tracks, enhancing the user's neural synchronization, relaxation, and attention experience, while avoiding distractions caused by pitch differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747504A_ABST
    Figure CN121747504A_ABST
Patent Text Reader

Abstract

The invention relates to generating tone-compatible synchronized neural metronomes for digital audio files. Methods and systems for improved neural rhythm generation of digital audio files are provided. In one embodiment, a method is provided that includes receiving a digital audio file and a beat frequency of a neural beat. Chroma diagram features may be extracted from a digital audio file and may be used to identify dominant sound levels within the digital audio file at a plurality of timestamps. A plurality of carrier frequencies for different time periods within the digital audio file may be selected based on the dominant level. Neural metronomes may be synthesized for a digital audio file based on a metronome frequency of a plurality of carrier frequencies. The neural rhythms may be stored and / or may be combined with digital audio files to generate combined audio tracks that may be stored.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis This application is a divisional application of patent application No. 202280071048.4, filed on October 21, 2022, entitled “Generating Pitch-Compatible Synchronized Neural Beats for Digital Audio Files”, the entire disclosure of which is incorporated herein by reference. Technical Field

[0002] This application generally relates to the field of acoustics, and more specifically, to generating pitch-compatible synchronized neural beats for digital audio files. Background Technology

[0003] In acoustics, a beat is a pattern of interference between two sounds with slightly different frequencies, considered as a periodic change in volume at a rate equal to the difference between the two frequencies. For a monoaural beat, the listener hears two different frequencies simultaneously in the same ear or both ears. In a binaural beat, the two different frequencies are heard separately through different ears (e.g., using headphones or specially positioned speakers), and the listener's brain detects the interference pattern. More complex interference patterns involving multiple signals can also be used to generate various beats.

[0004] Certain types of beats (e.g., monoaural beats, binaural beats) can be used to promote desired mental states (e.g., improving an individual's focus or attention). For example, such beats can be used to create neural synchronization when a user listens to the beat, helping the user to better concentrate or focus. These beats are often provided as standalone tracks (e.g., tracks containing only the beats). Alternatively, tracks can be prepared with custom monoaural or binaural beats added (i.e., tracks that have been created or generated to contain monoaural or binaural beats). In some cases, even beats outside the frequency range of normally audible sounds can be provided. Summary of the Invention

[0005] This disclosure presents novel and innovative systems and methods for generating neural beats and adding them to existing audio tracks. In a first aspect, this disclosure provides a method comprising receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file, and extracting multiple chroma map features of the digital audio file based on multiple parameters. The method includes combining the multiple chroma map features to form a primary chroma map feature of the digital audio file, and extracting a tonic level at multiple timestamps within the digital audio file from the primary chroma map feature. Multiple carrier frequencies are selected for the neural beat based on the tonic levels at the multiple timestamps, and a synchronized neural beat of the digital audio file is synthesized based on the beat frequency and the multiple carrier frequencies. The method also includes storing at least one of the following: (i) the synchronized neural beat and (ii) a combined audio track that combines the synchronized neural beat and the digital audio file.

[0006] In an embodiment according to a first aspect of this disclosure, the primary chroma map features include the intensity of each of a plurality of pitches at a plurality of timestamps. In one embodiment, the tonic pitch is selected from the plurality of pitches. In another embodiment, extracting the tonic pitch further includes generating a probability distribution for each of the plurality of pitches at a plurality of timestamps based on the intensity of the plurality of pitches using a Hidden Markov Model. In yet another embodiment, the Hidden Markov Model is configured to optimize the number and location of transitions between tonic pitches. In an alternative embodiment, extracting the tonic pitch further includes identifying a sequence of tonic pitches within the probability distribution.

[0007] In an embodiment according to a first aspect of this disclosure, multiple timestamps occur every 500 milliseconds or less during a digital audio file.

[0008] In an embodiment according to a first aspect of this disclosure, a plurality of chroma map features are linearly combined to form a principal chroma map feature.

[0009] In an embodiment according to a first aspect of this disclosure, the method further includes adjusting the volume of the synchronized neural beat over time to follow the volume of audio encoded in a digital audio file. In another embodiment, normalizing the volume of the synchronized neural beat includes generating a loudness profile of the duration of the audio encoded in the digital audio file and forming a volume curve based on the loudness profile. In one embodiment, the method includes adjusting the volume of the synchronized neural beat according to the volume curve.

[0010] In an embodiment according to a first aspect of this disclosure, the method further includes aligning a beat frequency with a rhythmic beat within a digital audio file. In one embodiment, aligning the beat frequency includes estimating the position of a rhythmic beat within the digital audio file, estimating a musical rhythm within the digital audio file, and adjusting the timing of a synchronized neural beat according to the musical rhythm to align a peak within the synchronized neural beat with the position of a rhythmic beat within the digital audio file. In one embodiment, the neural beat is at least one of: (i) a binaural beat and (ii) a monoaural beat.

[0011] In an embodiment according to a first aspect of this disclosure, the synchronized neural beat includes two or fewer audio channels.

[0012] In an embodiment according to a first aspect of this disclosure, the synchronized neural beat includes three or more audio channels.

[0013] In an embodiment according to the first aspect of this disclosure, the beat frequency is greater than or equal to 0.5 Hz and less than or equal to 150 Hz.

[0014] In an embodiment according to a first aspect of this disclosure, the method further includes parallel playback of a synchronized neural beat and a digital audio file via a computing device. In one embodiment, the method further includes streaming the synchronized neural beat and the digital audio file to the computing device for playback by the computing device.

[0015] Unless otherwise expressly disclosed in the specification, the embodiments of the present disclosure according to the first aspect are not mutually exclusive: in other embodiments according to the first aspect of the present disclosure, a feature of one embodiment of the first aspect of the present disclosure is combined with a combination of features of another embodiment of the first aspect.

[0016] In a second aspect, a system including a processor and a memory is provided. The memory may store instructions that, when executed by the processor, cause the processor to perform the method according to the first aspect of this disclosure. In one embodiment, when executed by the processor, the instructions cause the processor to receive a digital audio file and a beat frequency of a neural beat to be added to the digital audio file, and to extract multiple chroma map features of the digital audio file based on multiple parameters. These instructions may also cause the processor to combine the multiple chroma map features to form a principal chroma map feature of the digital audio file, extract the tonic at multiple timestamps within the digital audio file from the principal chroma map feature, and select multiple carrier frequencies for the neural beat based on the tonic at the multiple timestamps. These instructions may also cause the processor to synthesize a synchronized neural beat of the digital audio file based on the beat frequency and the multiple carrier frequencies, and to store at least one of: (i) the synchronized neural beat and (ii) a combined audio track that combines the synchronized neural beat and the digital audio file.

[0017] In one embodiment according to a second aspect of this disclosure, the primary chromaticity map features include the intensity of each of a plurality of pitches at a plurality of timestamps. The primary pitch may be selected from the plurality of pitches.

[0018] In one embodiment according to the second aspect, the memory stores additional instructions that, when executed by the processor during the extraction of the tonic, cause the processor to use a hidden Markov model to generate a probability distribution for each of the multiple tonics at multiple timestamps based on the intensity of the multiple tonics.

[0019] Unless otherwise expressly disclosed in the specification, embodiments of the present disclosure according to the second aspect above are not mutually exclusive: in other embodiments according to the second aspect of the present disclosure, a feature of one embodiment of the second aspect of the present disclosure is combined with a combination of features of another embodiment of the second aspect.

[0020] Embodiments according to the first aspect of this disclosure may be combined with embodiments according to the second aspect of this disclosure. Not all features and advantages described herein are exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings and description. Furthermore, it should be noted that the language used in the specification has been chosen primarily for readability and instructional purposes, and not to limit the scope of the disclosed subject matter. Attached Figure Description

[0021] Figure 1A A system according to an exemplary embodiment of the present disclosure is shown.

[0022] Figure 1B An audio playback system according to an exemplary embodiment of the present disclosure is shown.

[0023] Figure 2 A chromaticity map feature according to an exemplary embodiment of the present disclosure is shown.

[0024] Figure 3 The tonic level is shown according to an exemplary embodiment of this disclosure.

[0025] Figure 4 The selected carrier frequency according to an exemplary embodiment of this disclosure is shown.

[0026] Figure 5 A volume curve is shown according to an exemplary embodiment of the present disclosure.

[0027] Figure 6 A method for synthesizing neural beats according to exemplary embodiments of the present disclosure is shown.

[0028] Figures 7A-7C A method according to an exemplary embodiment of the present disclosure is shown.

[0029] Figure 8 A computing system according to an exemplary embodiment of the present disclosure is shown. Detailed Implementation

[0030] For the purposes of this application, "neural beat" is an acoustic beat added to an audio signal. A neural beat does not necessarily have to be in the audible frequency range. All references to neural beats in this application are intended to include either or both monoaural and binaural beats, unless the context clearly limits the discussion to one or the other.

[0031] A “neural beat” can include an audio beat designed to produce or promote a desired mental state for a user. The desired mental state can include neural synchronization, increased attention, calmness, relaxation, or any other desired mental state. In some implementations, a neural beat can include a monoaural or binaural beat that combines a lower beat frequency with a higher carrier frequency. Specifically, the “beat frequency” can be selected based on the desired mental state (e.g., different frequencies may promote different types of mental states in an individual). In one embodiment, the beat frequency can range from 0.5 Hz to 150 Hz. In other example embodiments, the frequency can be selected within the range of 1 Hz and 100 Hz, 1 Hz and 10 Hz, 10 Hz and 100 Hz, or 10 Hz and 40 Hz. The “carrier frequency” can be an audio frequency or note selected to carry or audibly reproduce the beat frequency within a track. For example, the beat frequency can be lower than frequencies detectable by humans, and / or can be within the lower frequency range of human hearing. Therefore, to maximize the effectiveness of neural beats, a carrier frequency in an existing audio signal can be selected, such as existing music or other audio recordings, or a specific track in such a recording, and the beat frequency can be modulated onto the carrier frequency to form a neural beat. The carrier frequency can range from 207.65 Hz to 392.00 Hz. In various implementations, neural beats can have different numbers of audio channels, such as one audio channel (e.g., a monoaural beat), two audio channels (e.g., a binaural beat), five audio channels, or more.

[0032] Not all users enjoy listening to tracks that only contain neural beats and may find them boring or distracting, thus limiting the effectiveness of neural synchronization. Furthermore, the limited availability of existing tracks containing embedded monophonic beats may not appeal to all users. Some systems can automatically generate music with monophonic beats to prevent users from having to listen to the same track multiple times. However, such systems still cannot adapt to the possibility that a user might want to listen to specific tracks or genres that were not previously combined with neural beats. Therefore, there is a need to automatically add neural beats to existing tracks, allowing users to listen to their preferred tracks or genres while also experiencing the neural synchronization, relaxation, and / or attention-enhancing benefits offered by neural beats.

[0033] One approach to this problem is to analyze the pitch characteristics of digital audio files over time. Specifically, chroma map features indicate the intensity of different pitch levels of the audio signal over time. Chroma map features can be generated for digital audio files, indicating the intensity of different pitch levels of the digital audio signal encoded within the file over time. This information can then be used to select the carrier frequency for neural beats to be added to the digital audio file. For example, the tonic pitch can be extracted from the chroma map features at different timestamps within the digital audio file, and the tonic pitch can be used to select carrier frequencies for neural beats at different timestamps. In some cases, a model (e.g., a Hidden Markov Model) can be used to analyze the tonic pitch to select carrier frequencies that optimize the number of carrier frequency variations, for example, minimizing the number of variations while still achieving a certain level of accuracy. Neural beats can then be synthesized based on the beat frequencies and the selected carrier frequencies and stored for later use. In some cases, composite tracks combining digital audio files with neural beats can be generated. In other cases, neural beats can be stored in association with digital audio files. Furthermore, in certain situations, neural beats and / or combined audio tracks can be generated in real time while a user device is streaming a digital audio file, such as via a server streaming the digital audio file or via a user device receiving the streaming digital audio file. The neural beats can then be played along with the digital audio file via the user device (e.g., as separate audio files played simultaneously and / or as a single audio file).

[0034] Figure 1A A system 100 according to an exemplary embodiment of the present disclosure is illustrated. System 100 can be configured to generate and synchronize neural beats for addition to digital audio files. System 100 includes a computing device 102 and a server 104. Server 104 stores digital audio files 108, 110, and computing device 102 can add neural beats to digital audio files 108, 110. For example, computing device 102 and server 104 can be part of a digital audio streaming platform configured to stream digital audio files 106, 108, 110 upon user request. Furthermore, computing device 102 can be configured to add neural beats 168, 174 to digital audio files 106, 108, 110 upon user request. For example, a user can manipulate preferences for adding neural beats 168, 174 to streaming audio files received from an audio streaming platform.

[0035] Computing device 102 can receive digital audio file 106 from server 104 and can generate neural beats 168 and / or adjusted neural beats 174 to be added to digital audio file 106. Computing device 102 can also receive the beat frequency 112 of neural beats 168 and 174. For example, the beat frequency 112 can be received from a user via a user-configurable beat frequency setting. Neural beats 168 and 174 can be monoaural beats, binaural beats, or may have more audio channels, and the type of neural beats 168 and 174 can be selected by the user. Alternatively or additionally, computing device 102 can select between monoaural beats and binaural beats based on the audio device from which the user streams the digital audio file. For example, if a user is streaming audio from a mono audio device, computing device 102 can generate monoaural beats, and if the user is streaming audio from a stereo audio device (e.g., stereo speakers, stereo headphones), computing device 102 can generate binaural beats. In yet another embodiment, computing device 102 can select the number of audio channels based on the same number of audio channels in the digital audio file 106.

[0036] Specifically, computing device 102 can be configured to generate neural beats 168, 174 that are mixed into digital audio file 106. For example, computing device 102 can be configured to generate neural beats 168, 174 synchronized with the audio pitches within digital audio file 106 to avoid obvious and distracting pitch differences that could hinder the user's neural synchronization. To this end, computing device 102 can extract multiple chroma map features 116 from digital audio file 106. Chroma map features 116 may include pitch levels 124, 126 and associated intensities 136, 138 at multiple timestamps 148, 150.

[0037] For example, Figure 2A chroma map feature 200 according to an exemplary embodiment of the present disclosure is depicted. The chroma map feature 200 includes the intensity of multiple pitch classes at multiple timestamps T1-T19 (as defined in Figure 202). Pitch classes include B, A# / Bb, A, G# / Ab, G, F# / Gb, F, E, D# / Eb, D, C# / Db, and C, which represent each type of note that can be reproduced within a digital audio file 106. Specifically, each pitch class can represent all audible pitches separated by an integer number of octaves in a song. For example, pitch class C can contain the middle C, treble C, high C, tenor C, low C, and other octaves of note C. Other pitch classes can be similarly defined as multiple notes containing different octaves. In practice, pitch classes can be defined as a set of frequency bands. For example, the pitch C can be defined as 261.626 ± 0.1 Hz (for middle C), 523.251 ± 0.1 Hz (for tenor C), and similarly for other notes contained within the pitch. As depicted, certain sharps or flats (e.g., A#, Bb, G#, Ab, F#, Gb, D#, Eb, C#, Db) are grouped into pitches separate from those containing the natural note AG. In additional or alternative embodiments, a pitch can be defined as containing the sharp or flat version of a note. Similarly, some embodiments can define pitches in different ways (e.g., containing any desired combination of notes). For example, in another embodiment, the pitch of C can contain middle C# or middle Cb. Those skilled in the art will understand that chromaticity feature 200 can be calculated based on any of a plurality of conceivable pitches, such as equal temperament tuning (e.g., 24-tone equal temperament with 24 pitches, 19-tone equal temperament with 19 pitches, and / or 7-tone equal temperament with 7 pitches). In practice, computing device 102 can calculate more pitches than are represented in chroma map feature 200, and can combine these pitches into the desired pitches of chroma map feature 200. For example, computing device 102 can calculate 36 pitches and then combine these pitches into the pitches depicted for chroma map feature 200.

[0038] The chroma map feature 200 includes the intensity of each pitch at each of the timestamps T1-T19. These intensities vary over time (e.g., as the music in the digital audio file 106 changes). For example, both pitch A and pitch D have high intensities at times T1-T5. Starting from times T6-T10, the pitches with the highest intensities alternate between C and C# / Db (T8, T12), D (T9-10, T13, T17-18), D and D# / Eb (T6, T14), E (T7, T15), E and D# / Eb (T11, T19), and F (T16). These intensities can be calculated based on an analysis of the frequency domain of the digital audio file 106 at each of the timestamps T1-T19. For example, the computing device 102 can divide the digital audio file 106 into segments for each of the timestamps T1-T19. The computing device 102 can then calculate a time-frequency representation of each segment (e.g., a frequency distribution over multiple times) (e.g., by performing a Fourier transform, a Fast Fourier Transform (FFT), a constant Q-transform, a wavelet transform, using a filter bank, etc.). The frequencies in the time-frequency representation can correspond to or be categorized into each pitch level (e.g., according to a predefined frequency band). The intensity of each pitch level can then be calculated based on the intensity of the corresponding frequency in the time-frequency representation. This process can be repeated multiple times for each segment corresponding to a timestamp T1-19. In some embodiments, timestamps T1-19 may occur once every 50 milliseconds. In additional or alternative embodiments, timestamps T1-19 may occur more frequently (e.g., every 10 milliseconds, every 5 milliseconds, every millisecond) and / or less frequently (e.g., every 0.5 seconds, every 0.25 seconds, every 0.1 seconds). In some embodiments, the computing device 102 can perform the analysis in the time domain, rather than performing a frequency domain analysis on the digital audio file 106. For example, a filter bank can be used with one or more filters for each pitch level. The intensity of the chroma map feature 200 can then be determined using the intensity of the filtered signal obtained at each timestamp.

[0039] Back Figure 1AThe computing device 102 can compute multiple chroma map features 116 of the digital audio file 106. For example, multiple chroma map features 116 can be computed to focus on different frequency ranges within the digital audio file 106. As a specific example, a first set of chroma map features can be computed focusing on lower frequency ranges (e.g., less than C4 or 261.62 Hz) within the digital audio file 106, and a second set of chroma map features can be computed focusing on higher frequency ranges (e.g., C1 to C8 or 32.70 Hz to 4186.01 Hz). In such a case, the computing device 102 can then be configured to combine the multiple chroma map features 116 into a set of principal chroma map features 118 of the digital audio file 106. For example, the computing device 102 can linearly combine the chroma map features 116 (e.g., according to predefined weights) to form the principal chroma map features 118. The data structure of the principal chroma map features 118 can be comparable to the data structure of the chroma map features 116. For example, in some implementations, chroma map feature 200 may represent a set of primary chroma map features 118 of the digital audio file 106. Furthermore, it should be understood that, although... Figure 2 Chroma map feature 200 is depicted as a data graph over time, but in practice, chroma map feature 116 and / or principal chroma map feature 118 can be stored in an additional or alternative data structure. For example, chroma map feature 116 and / or principal chroma map feature 118 can be stored as an array containing the intensity values ​​of the pitch levels at timestamps T1-19.

[0040] The computing device 102 can identify the tonic level 120 based on the principal chroma map feature 118. The tonic level has the highest intensity within a specific time or time interval of the audio signal. Specifically, the computing device 102 can calculate the probability distributions 144 and 146 of each of the tonic levels 132 and 134 at specific timestamps 156 and 158. For example, Figure 3A principal pitch 300 according to an exemplary embodiment of the present disclosure is shown. The principal pitch 300 includes the probability (as defined in Figure 302) of each of the pitches B, A# / Bb, A, G# / Ab, G, F# / Gb, F, E, D# / Eb, D, C# / Db, and C at each of the timestamps T1-T5. Specifically, at times T1-T5, pitches A and D have a medium-high probability; at times T6 and T14, pitches D and Db / E have a medium-high probability; at times T7, T11, and T19, pitches E and Db / E have a medium-high probability; at times T8 and T12, pitches C and Cb / D have a medium-high probability; at times T9, T10, T13, T17, and T18, pitch E has a high probability; at time T15, pitch E has a high probability; and at time T16, pitch F has a high probability. Probabilities can be calculated to reflect the probability that each pitch represents the tonic at a given point in time. For example, in some cases, probabilities can be calculated using a Hidden Markov Model (HMM). In some cases, the HMM can be tuned to optimize the number of transitions in the tonic (e.g., optimizing the number of carrier frequency changes in neural beats 168 and 174), transitions in the tonic that may distract the user and / or potentially have an adverse effect on neural synchronization.

[0041] Back Figure 1A The computing device 102 can determine the carrier frequency 114 based on the tonic level 120. The carrier frequency 114 can include a single selected frequency 160, 162 at each timestamp 164, 166 to be used as the carrier frequency for that time within neural beats 168, 174. For example, Figure 4A carrier frequency 400 according to an exemplary embodiment of the present disclosure is depicted. The carrier frequency 400 includes a single selected pitch at each timestamp T1-19. Specifically, pitch D is selected as the carrier frequency for timestamps T1-14 and T17-19, and pitch E is selected as the carrier frequency for timestamp T15-16. The carrier frequency can be selected to follow the music and harmony of a digital audio file while avoiding unnecessary variations in the carrier frequency. Specifically, the carrier frequency at times T9-T10 and T17-20 can be selected as pitch D to align with the tonic at these times. However, excessive variations in the carrier frequency can distract the user; therefore, in certain situations, such as when choosing between different pitches with similar probabilities or when making small, transient changes in the tonic, the selected carrier frequency can be chosen to maintain consistency over time. For example, in tonic 300, pitches A and D have similar probabilities at times T1-5. However, pitch D can be chosen as the carrier frequency from times T1-5 to avoid the transition from pitch A to pitch D at time T6, where pitch D is the tonic. As another example, at times T7, T11, and T19, pitches D# / Eb and E all have similar probabilities. However, pitch D can be chosen as the carrier frequency even if it doesn't have the highest probability at these times, to reduce the number of carrier frequency changes (e.g., because pitch D still has a moderate probability in tonic 300). On the other hand, failure to follow musical harmonies can also have adverse effects on neural synchronization. Therefore, at times T17-T19, the carrier frequency is switched from E (at time T16) to D (at times T17-T19) to properly follow the harmonies in the digital audio file.

[0042] To select carrier frequency 400, computing device 102 can be configured to balance maximizing the overall probability of the selected carrier frequency while limiting the number of consecutive tonic changes. In some embodiments, computing device 102 can perform Viterbi decoding on tonic 300 to find the most probable sequence of individual tonics at each timestamp, a sequence that limits the number of carrier frequency transitions while also ensuring that carrier frequency 400 is musically aligned with digital audio file 106.

[0043] Back Figure 1AThe computing device can synthesize a neural beat 168 based on the beat frequency 112 and the carrier frequency 114. Specifically, the computing device 102 can synthesize the neural beat 168 by modulating the beat frequency 112 onto the selected carrier frequencies 160, 162 at each of timestamps 164, 166. In some embodiments, the timestamps 164, 166 within the carrier frequencies (e.g., timestamps T1-19) can correspond to the timestamps of the audio data within the digital audio file 106. In such cases, the computing device 102 can directly synthesize the neural beat 168 based on the carrier frequency at each of the timestamps 164, 166.

[0044] In some implementations, computing device 102 may further adjust one or more aspects of neural beat 168 based on other characteristics of digital audio file 106. For example, computing device 102 may adjust the volume of neural beat 168 to align with volume changes in digital audio file 106. Specifically, if neural beat 168 is relatively quiet compared to digital audio file 106, the benefits of neural beat may be reduced. Alternatively or additionally, if neural beat 168 is louder relative to digital audio file 106, neural beat 168 may prove to be intrusive or distracting to the user, thereby negating the benefits provided by neural beat 168. Therefore, audio mixer 122 may be used to adjust the volume of neural beat 168 during the processing of digital audio file 106.

[0045] Specifically, the audio mixer 122 can determine a loudness profile 170 of the digital audio file 106. The loudness profile 170 is a representation of how loud the audio encoded in the digital audio file 106 is over time (e.g., within the duration of the audio encoded in the digital audio file 106). The loudness profile 170 can be calculated as a combined intensity (e.g., across audible frequencies) at multiple timestamps within the digital audio file 106. The loudness profile 170 can then be used to generate a volume curve 172 of a neural beat 168. Specifically, the loudness profile 170 can be offset (e.g., based on the maximum expected intensity of the neural beat 168) to generate the volume curve 172. For example, Figure 5 A volume curve 500 is depicted according to an exemplary embodiment of the present disclosure. Volume curve 500 illustrates the change in energy (e.g., in dB) over the duration of audio encoded in the digital audio file 106, wherein the energy of the audio signal within the digital audio file 106 can be used as a proxy for the volume over time within the digital audio file 106. Return Figure 1AThe volume curve 172 can be applied to the neural beat 168 to generate an adjusted neural beat 174. Specifically, applying the volume curve 172 to the neural beat 168 can include increasing or decreasing the volume (e.g., intensity) of the neural beat 168 at different time points according to the intensity indicated in the volume curve 172 (e.g., making the adjusted neural beat 174 louder at high intensity in the volume curve 172 and quieter at low intensity in the volume curve 172).

[0046] Neural beats 168 and / or adjusted neural beats 174 may subsequently be stored, transmitted, and / or played back on a user device. For example, computing device 102 may store neural beats 168 and / or adjusted neural beats 174 in association with a digital audio file 106 (e.g., in server 104). In some embodiments, digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 may be stored separately. In additional or alternative embodiments, computing device 102 may combine digital audio file 106 with neural beats 168 and / or adjusted neural beats 174 to generate a combined audio track that can be stored (e.g., in server 104). As another example, and referring to... Figure 1B With system 190, digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 can be transmitted to user device 192 associated with user 194. User device 192 may include a smartphone, tablet, wearable computing device, laptop, personal computer, or any other personal computing device. User device 192 may also include one or more audio devices for audio playback, such as speakers, a 3.5 mm audio jack connected to headphones or speakers, wirelessly connected headphones, wirelessly connected speakers, or any other device capable of playing back audio. System 100 may transmit (e.g., stream) digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 to user device 192. User device 192 can then receive and play back digital audio file 106 simultaneously with neural beats 168 and / or adjusted neural beats 174. Additionally or alternatively, user device 192 may store digital audio file 106 and neural beats 168 and / or adjusted neural beats 174 for future playback. Alternatively, computing device 102 may transmit a combined audio track to user device 192. In yet another embodiment, neural beats 168 and / or adjusted neural beats 174 may be generated on user device 192. In such a case, neural beats 168 and / or adjusted neural beats 174 may be played on user device 192 along with digital audio file 106 (e.g., as separate audio files, as a combined audio track) and / or may be stored on user device 192 for later playback.

[0047] Although not depicted, computing device 102, server 104, and / or user device 192 may include at least one processor and / or memory configured to implement one or more aspects of computing device 102, server 104, and / or user device 192. For example, the memory may store instructions that, when executed by the processor, cause the processor to perform one or more operational features of computing device 102, server 104, and / or user device 192. The processor may be implemented as one or more central processing units (CPUs), field-programmable gate arrays (FPGAs), and / or graphics processing units (GPUs) configured to execute instructions stored in memory. Additionally, computing device 102, server 104, and / or user device 192 may be configured to communicate using a network. For example, computing device 102, server 104, and / or user device 192 may communicate with a network using one or more wired network interfaces (e.g., Ethernet interfaces) and / or wireless network interfaces (e.g., Wi-Fi®, Bluetooth®, and / or cellular data interfaces). In some cases, a network can be implemented as a local network (e.g., a local area network), a virtual private network, an L1 network, and / or a global network (e.g., the Internet).

[0048] In some implementations, computing device 102 and server 104 may be implemented as a single computing device. For example, computing device 102 may store digital audio files 106, 108, 110 (e.g., in a local database). In another implementation, computing device 102 and / or server 104 may be implemented at least partially by user device 162. In yet another implementation, computing device 102, server 104, and / or user device 192 may be implemented by multiple computing devices. For example, computing device 102 may be implemented as multiple software services executing in a distributed computing environment (e.g., a cloud computing environment). As another example, user device 162 may be implemented by multiple personal computing devices (e.g., smartphones and wearable computing devices such as smartwatches).

[0049] Figure 6 A method 600 for synthesizing neural beats according to exemplary embodiments of the present disclosure is illustrated. Method 600 can be implemented on a computer system such as systems 100, 160. For example, method 600 can be implemented by computing device 102 and / or user device 192. Method 600 can also be implemented by a set of instructions stored on a computer-readable medium, which, when executed by a processor, cause the computer system to perform method 600. For example, all or part of method 600 can be implemented by the processor and / or memory of computing device 102 and / or user device 192. Although referenced... Figure 6 The flowchart shown illustrates the following example, but execution can be used with... Figure 6Many other methods of associated actions. For example, the order of some boxes can be changed, some boxes can be combined with other boxes, one or more boxes can be repeated, and some of the boxes described can be optional.

[0050] Method 600 may begin by receiving a digital audio file and a beat frequency of a neural beat to be added to the digital audio file (block 602). For example, computing device 102 may receive a digital audio file 106 and a beat frequency 112 of a neural beat to be added to the digital audio file 106. As described above, computing device 102 may receive the digital audio file 106 from server 104 and / or retrieve the digital audio file 106 from local storage. In some embodiments, the digital audio file 106 may be received upon a user request. For example, a user request to play a specific song may be received from a user device (e.g., via a music streaming service). Computing device 102 may receive the beat frequency 112 from the user (e.g., based on a user request and / or previously defined user settings). In some embodiments, the beat frequency 112 may specify a specific frequency (e.g., 3 Hz) for the neural beat to be added to the digital audio file 106. In additional or alternative embodiments, the beat frequency 112 may specify a frequency range of the neural beat (e.g., 4 Hz to 8 Hz).

[0051] Multiple chroma map features can be extracted from a digital audio file (box 604). For example, computing device 102 can extract multiple chroma map features 116, 200 from digital audio file 106. As described above, chroma map features can include intensity information of multiple pitch levels at multiple timestamps within digital audio file 106. In some embodiments, each of the multiple chroma map features can be extracted based on different parameters applied to digital audio file 106 before extracting chroma map features 116, 200. For example, a first chroma map feature can be extracted focusing on lower frequencies of digital audio file 106, and a second chroma map feature can be extracted focusing on higher frequencies of digital audio file 106. As another example, three chroma map features can be extracted from digital audio file 106: a first chroma map feature focused on lower frequencies (e.g., less than 200 Hz), a second chroma map feature focused on middle frequencies (e.g., from 200 Hz to 800 Hz), and a third chroma map feature focused on higher frequencies (e.g., greater than 800 Hz). In practice, after generating the time-frequency representation as discussed above, multiple chroma map features 116, 200 can be generated by selecting octaves and intensities within the desired frequency range to be included in the chroma map features 116, 200. In other embodiments, multiple chroma map features can be generated by applying filters (e.g., high-pass filters, low-pass filters, band-pass filters, etc.) to the digital audio file 106 before extracting the chroma map features 116, 200 (e.g., using FFT, constant Q-transform, filter buckets, and / or other techniques as described above).

[0052] Multiple chroma map features can be combined to form the principal chroma map feature of a digital audio file (box 606). For example, computing device 102 can combine multiple chroma map features 116, 200 to form the principal chroma map feature 118 of digital audio file 106. In some embodiments, multiple chroma map features 116, 200 can be linearly combined to form the principal chroma map feature 118 (e.g., according to previously defined weights). In additional or alternative embodiments, multiple chroma map features 116, 200 can be combined according to any other conceivable combination strategy. For example, multiple chroma map features 116, 200 can be combined by “stacking” them (e.g., such that combining two chroma map features 116, 200 with 12 pitch levels forms a principal chroma map feature with 24 rows). Generating the principal chroma map feature 118 based on multiple chroma map features 116 can better capture the audio frequency characteristics of digital audio file 106 (e.g., by focusing on different frequency ranges within digital audio file 106, such as different octaves). In some implementations, one or both of boxes 604 and 606 may be omitted. For example, in some implementations, instead of extracting multiple chroma map features and combining them to form the main chroma map feature, a single set of chroma map features may be extracted from the digital audio file 106 and used as the main chroma map feature 118.

[0053] The tonic level can be extracted at multiple timestamps within a digital audio file (box 608). For example, computing device 102 can extract tonic levels 120 and 300 at multiple timestamps 156 and 158 within a digital audio file 106. Tonic levels 120 and 300 can be extracted from the principal chroma map feature 118 using a model (e.g., a hidden Markov model). Specifically, tonic levels 120 and 300 can be extracted as a probability distribution at multiple timestamps T1-19. As described above, timestamps T1-19 can be selected based on the timestamps of the principal chroma map feature 118.

[0054] Multiple carrier frequencies can be selected for neural beats (box 610). For example, computing device 102 can select multiple carrier frequencies 114, 400 for neural beats 168, 174. Specifically, the multiple carrier frequencies 114, 400 may include individual carrier frequencies 160, 162 at multiple timestamps 164, 166, T1-19. The selected carrier frequencies 114, 400 can be selected by a Viterbi process, which selects carrier frequencies such that the transition of carrier frequencies between adjacent time periods is optimized according to transition probabilities, as will be further explained herein. In some embodiments, in addition to selecting multiple carrier frequencies 114, 400, a specific beat frequency of neural beat 168 can be selected. For example, if a beat frequency 112 is received as an acceptable frequency range, computing device 102 can select a beat frequency of neural beat 168 from the acceptable range, as will be further discussed below.

[0055] Synchronized beats refer to beats that change the carrier frequency at different time periods within a digital audio file based on changes in the underlying audio signal, such as changing musical harmonies and / or melodies at different time periods of the audio signal. Synchronized beats can be synthesized for a digital audio file based on the beat frequency and multiple carrier frequencies (box 612). For example, computing device 102 can synthesize a neural beat 168 for a digital audio file 106 based on the beat frequency 112 and the carrier frequency 114. Specifically, the neural beat 168 can be generated by modulating the beat frequency 112 on two different carrier frequencies 160 and 162 at times corresponding to timestamps 164, 166, and T1-19 within the carrier frequencies 114 and 400. In this way, the neural beat 168 can be synchronized with changes in musical harmonies and / or melodies at different time periods within the digital audio file 106. In some embodiments, the neural beat 168 can be synthesized to include a single audio channel (e.g., as a monoaural beat). In additional or alternative embodiments, the neural beat 168 may be synthesized to include two audio channels (e.g., as a binaural beat with two channels, or as a monoaural beat with two channels). In yet another embodiment, the neural beat 168 may be synthesized to include more than two audio channels (e.g., three audio channels, four audio channels, or five audio channels). In some embodiments, the number of audio channels may be specified by the user or by a predetermined setting. In additional or alternative embodiments, the number of audio channels may be selected based on the number of audio channels in the digital audio file 106 (e.g., such that the neural beat 168 has the same number of audio channels as the digital audio file 106).

[0056] At least one of synchronized neural beats and combined audio tracks that combine synchronized neural beats and digital audio files can be stored (box 614). For example, computing device 102 can store at least one of synchronized neural beat 168 or combined audio tracks that combine synchronized neural beat 168 with digital audio files 106. For example, as described above, computing device 102 can store neural beat 168 and / or combined audio tracks on local storage within server 104 and / or computing device 102. Additionally or alternatively, computing device 102 can transmit neural beat 168 and / or combined audio tracks to user devices for storage and playback (e.g., temporary storage for streaming, and long-term storage). In embodiments where computing device 102 is a user device, computing device 102 can locally store neural beat 168 and / or combined audio tracks for current or future playback. In some embodiments, as further explained above, computing device 102 can be further configured to generate an adjusted neural beat 174 based on neural beat 168. In such a case, computing device 102 can be configured to store equalized neural beats 174 and / or combine audio tracks that combine the adjusted neural beats 174 with digital audio files 106 in a manner similar to that discussed above.

[0057] In this way, method 600 enables a computing device to generate neural beats for any digital audio file, thereby allowing for increased user choice in the types of music used to generate neural synchronization. Furthermore, the computing device is able to do this in real time and can ensure that the neural beats are blended with the pitch quality and / or loudness of the digital audio file to minimize user distraction and maximize neural synchronization. Therefore, method 600 ensures that the generated neural beats are constructively combined with the previously created digital audio file.

[0058] Figures 7A-7CMethods 700, 710, and 720 according to exemplary embodiments of the present disclosure are illustrated. Methods 700, 710, and 720 may be performed in conjunction with at least a portion of method 600. For example, method 700 may be performed concurrently with blocks 608 and 610 of method 600. As another example, method 710 may be performed between blocks 612 and 614 of method 600 and / or between a portion of block 612. As another example, method 720 may be performed as part of block 612 of method 600. Methods 700, 710, and 720 may be implemented on a computer system, such as systems 100 and 190. For example, methods 700, 710, and 720 may be implemented by computing device 102 and / or user device 192. Methods 700, 710, and 720 may also be implemented by a set of instructions stored on a computer-readable medium, which, when executed by a processor, cause the computer system to perform methods 700, 710, and 720. For example, all or part of methods 700, 710, 720 may be implemented by the processor and / or memory of computing device 102 and / or user device 192. Although references Figures 7A-7C The flowchart shown illustrates the following example, but execution can be used with... Figures 7A-7C Many other methods of associated actions. For example, the order of some boxes can be changed, some boxes can be combined with other boxes, one or more boxes can be repeated, and some of the boxes described can be optional.

[0059] Method 700 can be performed to select multiple carrier frequencies for neural beats. Method 700 can begin by generating a probability distribution of pitch levels at multiple timestamps (box 702). For example, a Hidden Markov Model can be used to generate a probability distribution of pitch levels (e.g., B, A# / Bb, A, G# / Ab, G, F# / Gb, F, E, D# / Eb, D, C# / Db, and C) within a digital audio file 106 at multiple timestamps T1-19 within the digital audio file 106. Timestamps T1-19 can be selected based on timestamps within the principal chroma map feature 118 (e.g., based on segments of the digital audio file 106 used to compute time-frequency representations of chroma map feature 116 and / or principal chroma map feature 118). The Hidden Markov Model can be configured to select the timing when transitions between different carrier frequencies should occur by adjusting the transition probabilities. Specifically, based on inputs received from users, system administrators, and / or computation processes, the transition probabilities of the Hidden Markov Model (e.g., transition probabilities of 0.005–0.02) may have been previously received (or may have been updated).

[0060] Then, a sequence of tonic levels can be identified within the probability distribution (box 704). For example, computing device 102 can identify a sequence of tonic levels within the probability distribution. Specifically, carrier frequencies 114, 400 can contain a series of tonic levels of carrier frequencies that will be used as neural beats 168. The sequence of tonic levels can be identified based on the constrained transition probabilities of changes in the selected tonic levels to maximize the combination probability of the selected tonic levels within the probability distribution. Specifically, the sequence of tonic levels can be selected via a Viterbi process implemented by computing device 102.

[0061] In this way, method 700 can be executed to select a sequence of carrier frequencies based on the musical harmonies and melodies (e.g., chroma map features) of the received digital audio file. Therefore, this process allows neural beats 168 to be applied to existing digital audio files while ensuring that changes in carrier frequencies do not disturb or distract users attempting to trigger neural synchronization using neural beats.

[0062] Method 710 can be performed to adjust the volume of the neural beat 168 based on the volume of the digital audio file 106 at different times within the digital audio file 106. Method 710 can begin by generating a loudness profile for the duration of the digital audio file (box 712). For example, computing device 102 (e.g., audio mixer 122) can generate a loudness profile 170 for the duration of the digital audio file 106. The loudness profile 170 can be generated based on the intensity (e.g., volume) of the digital audio file 106 at multiple times within the digital audio file 106. For example, a loudness profile 170 can be generated for each data sample timestamp within the digital audio file 106.

[0063] A volume curve can be formed based on a loudness profile (box 714). For example, computing device 102 can form a volume curve 172 based on loudness profile 170. Volume curve 172 can be formed as a percentage of loudness profile 170 (e.g., 50% of loudness profile 170). Alternatively or additionally, volume curve 172 can be formed by normalizing loudness profile 170 to the maximum volume desired for neural beat 168. Those skilled in the art will similarly recognize one or more additional means of generating volume curve 172 based on loudness profile 170 of digital audio file 106. Therefore, all such similar embodiments are considered to be within the scope of this disclosure.

[0064] The volume of the synchronized neural beat can then be adjusted based on the volume curve (box 714). For example, computing device 102 can adjust the volume of neural beat 168 based on volume curve 172 to generate an adjusted neural beat 174. For example, the intensity of neural beat 168 can be scaled to match the desired volume reflected in volume curve 172.

[0065] In this way, method 710 can be performed to adjust the neural beat 168. This can reduce the amount of interfering volume mismatch between neural beats in the digital audio file. For example, if the volume of the neural beat is much lower than that of the digital audio file, the user may not be able to hear the volume of the neural beat, thus reducing its effectiveness in generating neural synchronization. As another example, if the volume of the neural beat is much higher than that of the digital audio file 106, the user may be distracted or disturbed by the volume difference, thereby interrupting or reducing any neural synchronization generated by the neural beat.

[0066] Method 720 can be used to synchronize neural beats 168 with rhythmic patterns in digital audio file 106. Method 720 can begin by estimating the position of rhythmic beats within the digital audio file (box 722). A rhythmic beat refers to the rhythmic level of a musical piece—this is different from the acoustic usage of "beat." For example, computing device 102 can estimate the position of rhythmic beats within digital audio file 106. The position of rhythmic beats within digital audio file 106 can be estimated using a machine learning model, such as a pre-trained network configured to detect rhythmic beats within an audio file. For example, one or more models provided by software packages such as madmom and Essentia can be used to estimate the position of rhythmic beats. In additional or alternative implementations, one or more algorithmic techniques can be used to estimate the position of rhythmic beats.

[0067] The timing of synchronized neural beats can be adjusted based on the position of rhythmic beats within a digital audio file (box 724). For example, computing device 102 can adjust the timing of neural beat 168 based on the position of rhythmic beats. For example, computing device 102 can adjust the beat frequency 112 to align with the rhythm of the digital audio file (e.g., a multiple thereof). For example, in the case where digital audio file 106 has a rhythm of 120 bpm and the beat frequency 112 is 0.6 Hz (e.g., 100 bpm), computing device 102 can adjust the beat frequency 112 to an integer multiple of 120 beats per minute (e.g., 2 Hz). As a specific example, computing device 102 can adjust the beat frequency 112 to 0.5 Hz (30 bpm) and / or 1 Hz (60 bpm). In embodiments where the user has specified a desired frequency range for the beat frequency 112, the beat frequency 112 can be selected from the desired frequency range as an even multiple of the rhythm frequency and / or as close as possible to a multiple of the rhythm frequency. In addition, the timing of the synchronized neural beats can be adjusted so that the peak in the neural beat (e.g., the peak at beat frequency 112) occurs simultaneously with the rhythm beats in the digital audio file 106 (e.g., aligned with their timing).

[0068] In this way, method 720 can be used to ensure that the rhythm beats and beat frequencies within a digital audio file are not out of phase. Specifically, when the beat frequency is out of phase with the rhythm frequency of the digital audio file, interference between beat frequencies in the digital audio file can negatively affect sound quality and / or may create distracting or interfering patterns when the digital audio file and neural beats at the interfering beat frequencies are played simultaneously. Therefore, adjusting the beat frequency based on the rhythm beats within the digital audio file can reduce these interferences, thereby improving the quality of subsequently generated neural beats and / or the quality of neural synchronization generated by the neural beats.

[0069] Figure 8 An example computer system 800, such as computing device 102, is illustrated that can be used to implement one or more devices and / or components discussed herein. In specific embodiments, one or more computer systems 800 perform one or more steps of one or more methods described or illustrated herein. In specific embodiments, one or more computer systems 800 provide the functionality described or illustrated herein. In specific embodiments, software running on one or more computer systems 800 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. Specific embodiments include one or more portions of one or more computer systems 800. Throughout this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.

[0070] This disclosure contemplates any suitable number of computer systems 800. This disclosure contemplates computer systems 800 employing any suitable physical form. By way of example and not limitation, computer system 800 may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a host, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of the above. Where appropriate, computer system 800 may include one or more computer systems 800, which may be single or distributed, may span multiple locations, span multiple machines, span multiple data centers, or be located in the cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 800 may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.

[0071] In a specific embodiment, computer system 800 includes a processor 806, a memory 804, a storage device 808, an input / output (I / O) interface 810, and a communication interface 812. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0072] In a specific embodiment, processor 806 includes hardware for executing instructions, such as those constituting a computer program. By way of example, and not limitation, to execute instructions, processor 806 may retrieve (or fetch) instructions from internal registers, internal caches, memory 804, or storage device 808, decode and execute the instructions, and then write one or more results to internal registers, internal caches, memory 804, or storage device 808. In a specific embodiment, processor 806 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 806 including any appropriate number of appropriate internal caches where appropriate. By way of example, and not limitation, processor 806 may include one or more instruction caches, one or more data caches, and one or more translation back buffers (TLBs). Instructions in the instruction cache may be copies of instructions in memory 804 or storage device 808, and the instruction cache may accelerate the retrieval of those instructions by processor 806. The data in the data cache may be a copy of data in memory 804 or storage device 808 operated by computer instructions, the result of a previous instruction executed by processor 806 (which may be accessed by subsequent instructions or used to write to memory 804 or storage device 808), or any other suitable data. The data cache can accelerate read or write operations of processor 806. The TLB can accelerate virtual address translation of processor 806. In specific embodiments, processor 806 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates that processor 806 may include any suitable number of suitable internal registers where appropriate. Where appropriate, processor 806 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 806. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.

[0073] In a specific embodiment, memory 804 includes main memory for storing instructions to be executed by processor 806 or data to be operated on by processor 806. By way of example and not limitation, computer system 800 may load instructions from storage device 808 or another source (e.g., another computer system 800) into memory 804. Processor 806 may then load instructions from memory 804 into internal registers or internal caches. To execute instructions, processor 806 may retrieve instructions from internal registers or internal caches and decode them. During or after instruction execution, processor 806 may write one or more results (which may be intermediate or final results) to internal registers or internal caches. Processor 806 may then write one or more of these results to memory 804. In a specific embodiment, processor 806 executes only the instructions in one or more internal registers or internal caches or memory 804 (as opposed to storage device 808 or elsewhere) and operates only on the data in one or more internal registers or internal caches or memory 804 (as opposed to storage device 808 or elsewhere). One or more memory buses (each memory bus may include an address bus and a data bus) couple processor 806 to memory 804. The bus may include one or more memory buses, which will be described in further detail below. In a particular embodiment, one or more memory management units (MMUs) are located between processor 806 and memory 804 and facilitate access to memory 804 requested by processor 806. In a particular embodiment, memory 804 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 804 may include one or more memory 804s. Although this disclosure describes and illustrates specific memory implementations, this disclosure contemplates any suitable memory implementation.

[0074] In a specific embodiment, storage device 808 includes a mass storage device for data or instructions. By way of example, and not limitation, storage device 808 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage device 808 may include removable or non-removable (or fixed) media. Where appropriate, storage device 808 may be internal or external to computer system 800. In a specific embodiment, storage device 808 is a non-volatile solid-state memory. In a specific embodiment, storage device 808 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage device 808 in any suitable physical form. Where appropriate, storage device 808 may include one or more storage control units to facilitate communication between processor 806 and storage device 808. Where appropriate, storage device 808 may include one or more storage devices 808. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.

[0075] In a specific embodiment, I / O interface 810 includes hardware, software, or both, providing one or more interfaces for communication between computer system 800 and one or more I / O devices. Where appropriate, computer system 800 may include one or more of these I / O devices. One or more of these I / O devices enable communication between a person (i.e., a user) and computer system 800. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, screen, display panel, mouse, printer, scanner, speaker, still camera, stylus, input pad, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of the above. I / O devices may include one or more sensors. Where appropriate, I / O interface 810 may include one or more device or software drivers that enable processor 806 to drive one or more of these I / O devices. Where appropriate, I / O interface 810 may include one or more I / O interfaces 810. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure contemplates any suitable I / O interface or combination of I / O interfaces.

[0076] In a specific embodiment, the communication interface 812 includes hardware, software, or both, providing one or more interfaces for communication (e.g., packet-based communication) between computer system 800 and one or more other computer systems 800 or one or more networks 814. By way of example, and not limitation, the communication interface 812 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or any other wired network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as a Wi-Fi network. This disclosure contemplates any suitable network 814 and any suitable communication interface 812 for network 814. By way of example, and not limitation, network 814 may include one or more of ad hoc networks, personal area networks (PANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 800 may communicate with a wireless PAN (WPAN) (e.g., Bluetooth WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), or any other suitable wireless network, or a combination of two or more of the above. Where appropriate, computer system 800 may include any suitable communication interface 812 for any of these networks. Where appropriate, communication interface 812 may include one or more communication interfaces 812. Although this disclosure describes and illustrates specific communication interface implementations, this disclosure contemplates any suitable communication interface implementation.

[0077] Computer system 802 may also include a bus. The bus may include hardware, software, or both, and may communicatively couple components of computer system 800 to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or any other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), an HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB), or another suitable bus, or a combination of two or more of these buses. Where appropriate, the bus may include one or more buses. Although this disclosure describes and illustrates specific buses, this disclosure contemplates any suitable bus or interconnect.

[0078] In this document, a computer-readable non-transitory storage medium may include one or more semiconductor or other types of integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical disk drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more of these media. Where appropriate, a computer-readable non-transitory storage medium may be a volatile storage medium, a non-volatile storage medium, or a combination of volatile and non-volatile storage media.

[0079] In this document, "or" is inclusive rather than exclusive unless otherwise expressly indicated or the context otherwise indicates. Therefore, in this document, "A or B" means "A, B, or both" unless otherwise expressly indicated or the context otherwise indicates. Furthermore, "and" is both joint and separate unless otherwise expressly indicated or the context otherwise indicates. Therefore, in this document, "A and B" means "A and B, jointly or separately" unless otherwise expressly indicated or the context otherwise indicates.

[0080] The scope of this disclosure includes all changes, substitutions, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates various embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein, as will be understood by those skilled in the art. Additionally, the means or components of a device or system adapted, arranged, enabled, configured, enabled, operable, or operable to perform a particular function, as mentioned in the appended claims, are also included. This includes the device, system, or component, whether or not it or the particular function is activated, turned on, or unlocked, provided that the device, system, or component is so adapted, arranged, enabled, configured, or made operable or operational. Furthermore, although this disclosure describes or illustrates specific embodiments to provide particular advantages, these advantages may not be provided, may be provided in part, or may be provided in all of the specific embodiments.

[0081] All the disclosed methods and procedures described herein can be implemented using one or more computer programs or components. These components may be provided as a series of computer instructions on any conventional computer-readable medium or machine-readable medium, including volatile and non-volatile memories such as RAM, ROM, flash memory, magnetic disks or optical disks, optical storage, or other storage media. The instructions may be provided as software or firmware and may be implemented, wholly or partially, in hardware components such as ASICs, FPGAs, DSPs, or any other similar devices. The instructions may be configured to be executed by one or more processors, which, when executed, perform or facilitate the performance of all or part of the disclosed methods and processes.

[0082] It should be understood that various variations and modifications of the examples described herein will be apparent to those skilled in the art. Such changes and modifications can be made without departing from the scope of the subject matter and without diminishing its intended advantages. Therefore, such changes and modifications are intended to be covered by the appended claims.

Claims

1. A method comprising: Receive a digital audio file and the beat frequency of the neural beats to be added to the digital audio file; Multiple chroma map features are extracted from the digital audio file based on multiple parameters; The multiple chroma map features are combined to form the main chroma map features of the digital audio file; Extract the tonic level at multiple timestamps within the digital audio file from the main chroma map features; Based on the tonic level at the multiple timestamps, multiple carrier frequencies are selected for the neural beat; Based on the beat frequency and the multiple carrier frequencies, a synchronized neural beat of the digital audio file is synthesized. as well as Store at least one of the following: (i) the synchronized neural beat, (ii) a combined audio track that combines the synchronized neural beat and the digital audio file.

2. The method according to claim 1, wherein, The primary chromaticity map features include the intensity of each of the multiple pitches at the multiple timestamps, wherein the primary pitch is selected from the multiple pitches.

3. The method according to claim 2, wherein, Extracting the tonic level also includes: using a hidden Markov model to generate a probability distribution for each of the multiple tonic levels at the multiple timestamps based on the intensity of the multiple tonic levels.

4. The method according to claim 3, wherein, The hidden Markov model is configured to optimize the number and location of transitions between tonic levels.

5. The method according to claim 3, wherein, Extracting the tonic also includes identifying the sequence of tonics within a probability distribution.

6. The method according to any one of claims 1 to 5, wherein, During the audio encoded in the digital audio file, the plurality of timestamps occur once every 500 milliseconds or less.

7. The method according to any one of claims 1 to 5, wherein, The multiple chromaticity map features are linearly combined to form the primary chromaticity map feature.

8. The method according to any one of claims 1 to 5, further comprising: The volume of the synchronized neural beat is adjusted over time to follow the volume of the audio encoded in the digital audio file.

9. The method according to claim 8, wherein, Standardizing the volume of the synchronized neural beats includes: A loudness profile is generated for the duration of the encoded audio in the digital audio file; A volume curve is generated based on the loudness profile; and Adjust the volume of the synchronized neural beat according to the volume curve.

10. The method according to any one of claims 1 to 5, further comprising aligning the beat frequency with the rhythm beat within the digital audio file.

11. The method according to claim 10, wherein, Aligning the beat frequencies includes: Estimate the position of the rhythm beat in the audio signal encoded within the digital audio file; Estimate the musical rhythm of the audio signal encoded within the digital audio file; and The timing of the synchronized neural beats is adjusted according to the music rhythm to align the peaks within the synchronized neural beats with the positions of the rhythm beats in the audio signal encoded within the digital audio file.

12. The method according to any one of claims 1 to 5, wherein, The neural beat is at least one of the following: (i) a binaural beat, (ii) a uniaural beat.

13. The method according to any one of claims 1 to 5, wherein, The synchronized neural beats are encoded in two or fewer audio channels.

14. The method according to any one of claims 1 to 5, wherein, The synchronized neural beats are encoded in three or more audio channels.

15. The method according to any one of claims 1 to 5, wherein, The beat frequency is greater than or equal to 0.5 Hz and less than or equal to 150 Hz.

16. The method according to any one of claims 1 to 5, further comprising simultaneously playing the synchronized neural beat and the digital audio file via a computing device.

17. The method of claim 16, further comprising streaming the synchronized neural beat and the digital audio file to the computing device for playback by the computing device.

18. A system comprising a processor and a memory, characterized in that, The memory is configured to perform the steps of the method according to any one of claims 1 to 17.

19. The system according to claim 18, wherein, The primary chromaticity map features include the intensity of each of the multiple pitches at the multiple timestamps, wherein the primary pitch is selected from the multiple pitches.

20. The system according to claim 19, wherein, The memory stores additional instructions that, when executed by the processor during the extraction of the principal pitch, cause the processor to generate a probability distribution for each of the plurality of pitches at the plurality of timestamps based on the intensity of the plurality of pitches using a hidden Markov model.