System and method for identifying musical segments having characteristics suitable for eliciting a physiological response of the autonomic nervous system - Patents.com

JP2024526125A5Inactive Publication Date: 2025-07-01MIIR AUDIO TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577895
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-30
Filing Date
2022-06-15
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing techniques for identifying musical segments that evoke autonomic nervous system responses, such as 'chills', are subjective and time-consuming, lacking objective and quantitative methods to account for diverse musical genres and styles.

Method used

A system and method using Content-Based Music Information Retrieval (CB-MIR) techniques to extract low-level musical features, identify high-level structures associated with chill induction, and combine them into a combinatorial algorithm to objectively detect 'chill moments' in music.

Benefits of technology

Enables the robust identification of musical segments likely to induce autonomic nervous system responses by analyzing multiple acoustic and musical structures, providing an objective and quantitative process for selecting music segments in social media and advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system and method for identifying the most impactful moments or segments of music that are most likely to induce a chill effect in a human listener. A digital music signal is processed with two or more objective processing metrics that measure acoustic features known to be capable of inducing a chill effect. Individual detection events are identified at the output of each metric based on whether the output is above or below a threshold relative to the overall output. A combination algorithm aggregates the matching detection events to generate a successive match data set of the number of matching detection events in the music signal, which may be calculated on a beat-by-beat basis. A phrase detection algorithm may identify the impactful segments of music based on at least one of peaks, peak proximity, and moving averages of the successive match data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and benefit of U.S. Provisional Application No. 63 / 210,863, entitled "SYSTEMS AND METHODS FOR IDENTIFYING SEGMENTS OF MUSIC HAVING CHARACTERISTICS SUITABLE FOR INDUCING AUTONOMIC PHYSIOLOGICAL RESPONSES," filed June 15, 2021, and also claims priority to and benefit of U.S. Provisional Application No. 63 / 227,559, entitled "SYSTEMS AND METHODS FOR IDENTIFYING SEGMENTS OF MUSIC HAVING CHARACTERISTICS SUITABLE FOR INDUCING AUTONOMIC PHYSIOLOGICAL RESPONSES," filed July 30, 2021, the contents of each of which are incorporated by reference in their entireties herein.

[0002] [Field] The present disclosure relates to systems and methods for processing complex audio data such as music, and more particularly to systems and methods for processing music audio data to determine time regions of the audio data that have the strongest characteristics suitable for eliciting a physiological response of the autonomic nervous system in a human listener. [Background technology]

[0003] Recent scientific studies have attempted to better understand the connection between auditory stimuli and physiological responses of the autonomic nervous system, such as chills or goosebumps, a well-known involuntary response to certain sounds or music. In one of the first investigations of physiological responses of the autonomic nervous system to music, researchers collected data on cerebral blood flow, heart rate, respiration, and electrical activity (e.g., electromyogram) originating from skeletal muscles, as well as participants' subjective reports of "chill." The study confirmed that fluctuations in cerebral blood flow in brain regions associated with reward, emotion, and arousal (e.g., ventral striatum, midbrain, amygdala, orbitofrontal cortex, and ventromedial prefrontal cortex) corresponded with participants' self-reports of chill. These regions also activate in response to euphoria-inducing stimuli, such as food, sex, and recreational drugs.

[0004] Thus, a link between music and physiological responses of the autonomic nervous system has been established. However, there is a wide variety of genres, musical styles, and types of acoustic and musical stimuli that can induce a chill response. In order to accurately identify one or more specific segments in a song or musical score that are most likely to induce such an autonomic response, a digital audio processing routine is needed that can detect various individual underlying acoustic / musical structures in a digital recording that are associated with chill induction in a manner that scales well across a wide variety of musical genres / styles, and evaluate the detected chill inducers. Summary of the Invention [Problem to be solved by the invention]

[0005] In the process of creating software applications used in selecting music segments for use in social media and advertising, manually selecting and curating sections of music is a costly and time-consuming task, and efforts have been made to automate this process. One challenge in curating large catalogs and identifying music segments involves varying levels of aesthetic judgments that may be considered subjective. A novel approach to this problem has been to use techniques from the field of Content-Based Music Information Retrieval (herein referred to as "CB-MIR") in combination with academic research from the field of neurological studies, including the idea of ​​the so-called "chill response" in humans (e.g., a physiological response of the autonomic nervous system). This response is considered to be physiological in nature, given the commonality of the human sensory organs and human experience, and is also strongly related to music appreciation, although chill moments are not necessarily subjective.

[0006] Existing techniques for finding such moments require subjective assessment by music experts, or people familiar with any given piece of music. Even so, any individual has a set of biases and uncertainties that inform their assessment of the presence or likelihood of a chill response in the audience as a whole. The embodiments of the present disclosure enable the detection of musical segments associated with the induction of chill as an objective and quantitative process. [Means for solving the problem]

[0007] One aspect that the present disclosure utilizes is the idea that musicians and composers use common tools to affect the emotional state of the listener. Volume contrasts, key changes, chord changes, melodic pitches, and harmonic pitches can all be used in this "musician's toolbox" and are found in curricula wherever music performance and composition are taught. However, these high-level structures do not have clear "sonic signatures," or definitions in terms of signal processing of music recordings. To find these structures, teachings from the field of CB-MIR, with a particular focus on extracting low-level musical information (e.g., feature extraction) from digitally recorded or streaming audio, are leveraged in novel audio processing routines. Using the low-level information provided by traditional CB-MIR techniques as a source, embodiments of the present disclosure include systems and methods for processing and analyzing complex audio data (e.g., music) to identify high-level acoustic and musical structures that neurological studies of music have found to produce chill responses.

[0008] An example of this process begins with the extraction of various CB-MIR data streams (also referred to herein as objective audio processing metrics) from a music recording. Examples of these are loudness, pitch, spectrum, spectral flux, spectral centroid, Mel-frequency cepstral coefficients, etc., which are described in further detail herein. The particular implementation of feature extraction for any given type of feature may have parameterization options that affect the preparation and optimization of the data for subsequent processing steps. For example, the general feature of loudness may be extracted according to several different filters and methodologies.

[0009] The next phrase in this example process involves looking for high level acoustic and musical structures that induce chill. These structures have been described with varying levels of specificity in the academic literature on chill phenomena. Detecting any one of these high level structures from an individual CB-MIR data stream is referred to herein as a "GLIPh", an acronym for Geometric Limbic Impact Phenomenon. More specifically, an embodiment of the present disclosure involves studying chill inducers as described in the academic literature and then designing a GLIPh that represents the inducer as a statistical data pattern. The GLIPh can represent moments of interest within each musical feature, such as pitch, loudness, spectral flux, etc. Once the various GLIPhs that may be included in the extracted feature dataset have been identified, a border can be drawn around a region of interest (ROI) in a graph plot to indicate where the GLIPh is located within the timeline of the digital recording.

[0010] Then, as instances of GLIPh timestamps are accumulated across the various extracted feature datasets, a new dataset can be formed that calculates the amount of coincidence and proximity of GLIPhs within the digital recording. This data processing is referred to herein as a combination algorithm, and the output data is referred to herein as a "chill moment" plot, which can include a moving average of the output to present a continuous, smoother representation of the output of the combination algorithm, which may have large fluctuations in value at the beat-by-beat level (or a minimum time interval is used for one of the input metrics), resulting in "busy" data when visually analyzed, and this moving average of the output can be more useful for visual analysis of the data, especially when trends within a song across multiple beats or tacts are more useful to evaluate. In some embodiments, the GLIPhs are weighted equally, but the combination algorithm can also be configured to generate chill moment data by attributing a weighted value to each GLIPh instance. An example of generating a moving average includes using a convolution of the chill moment plot with a Gaussian filter, which can span, for example, just 2 or 3 beats, or 100 or more beats, and thus be a time-variable, dynamic value based on the length of the beats in the song. Representative example lengths can range from 10 to 50 beats, including 30 beats, the length used for the data presented herein. Basing this smoothing on beats advantageously allows the moving average to adapt to the musical content.

[0011] A tendency observed in artists' songwriting is that chill inducers (e.g., musical features that increase the likelihood of eliciting an autonomic physiological response) may be used simultaneously and consecutively (up to some logical limit), consistent with chill moment plots reflecting the coincidence and proximity of GLIPh. That is, the more frequently a portion of a song (or an entire song) exhibits a pattern of coincidence and proximity in musical features known to be associated with an autonomic physiological response, the more likely it is to elicit chill in the listener. Overall, when two or more of these features are aligned in time, the level of arousal that the musical moment elicits is increased. Thus, certain embodiments of the present disclosure provide a method for processing audio data to identify individual chill inducers and constructing a new data set of one or more peak moments in the audio data that maximize the likelihood of eliciting an autonomic physiological response based at least in part on the rate of coincidence and proximity of the identified chill inducers. Examples include further processing this new data set to identify musical segments and phrases that contain these peak moments and provide them as a new type of metadata that can be used with the original audio data, for example as a timestamp indicating the peak moment or phrase used to create a truncated segment from the original audio data that contains the peak moment or phrase.

[0012] An embodiment of the present disclosure can be used to process digital audio recordings that encode an audio waveform as a series of "sample" values; typically, 44,100 samples per second are used with pulse code modulation, with each sample capturing a complex audio waveform every 22.676 microseconds. Those skilled in the art will appreciate that higher sampling rates are possible and do not significantly impact the data extraction techniques disclosed herein. Exemplary digital audio file formats are MP3, WAV, AIFF. Processing can begin with a digitally recorded audio file, and multiple subsequent processing algorithms are used to extract musical features and identify musical segments with the strongest chill moments. A musical segment can be any subsection of a musical recording, typically between 10 and 60 seconds in length. Exemplary algorithms can be designed to find segments that start and end coincident with the beginning and end of a phrase, such as a chorus or verse.

[0013] The main categories of digital music recording analysis are: (i) Time domain: the analysis of the frequencies contained in a digital recording with respect to time; (ii) Rhythm: a periodic signal that repeats in the time domain and that humans perceive as distinct beats; (iii) Frequency: a periodic signal that repeats in the time domain and that humans perceive as a single sound / note; (iv) Amplitude: The intensity of sound energy at a given moment; (v) Spectral energy: the total amount of amplitude present across all frequencies in a piece of music (or other unit of time) that is perceived as timbre.

[0014] Physiological responses of the autonomic nervous system (e.g., chill) can be elicited by acoustic, musical, and emotional stimulus-driven characteristics. These characteristics include abrupt changes in acoustic properties, high-level structural predictions, and emotional intensity. Recent investigations have attempted to determine what audio characteristics induce chill. In this approach, researchers suggest that the chill experience involves mechanisms based on expectation, peak emotion, and emotion. However, significant shortcomings have been identified in the reviewed literature with regard to study design, validity of experimental variables, chill measures, terminology, and remaining knowledge gaps. The ability to experience chill is also influenced by personality differences, especially "openness to experience." This means that chill-eliciting moments for a given listener may be rare and difficult to predict, due in part to individual dispositional differences. Although the literature has demonstrated some useful associations between acoustic media (music) and physical phenomena (chill), the lack of rigorous definition of numerous musical and acoustic characteristics of chill-inducing musical events makes it difficult to identify specific musical segments that have one or more of these characteristics. Furthermore, many of the identified musical and acoustic characteristics are best understood as complex arrangements of musical and acoustic events that may only have subjectively identifiable characteristics when viewed as a whole. Thus, the existing literature considers the identification of chill-inducing peak moments in complex audio data (e.g., music) to be an unsolved problem.

[0015] Existing research describes chill inducers in aesthetic descriptive terms, not in numerical terms. Complex concepts such as "amazing harmony" currently have no known mathematical description. Although typical CB-MIR feature extraction methods are low-level and objective, they can nevertheless be used as building blocks in the embodiments of the present disclosure to begin to build (and subsequently discover and identify) patterns that can accurately represent high-level complex concepts, as demonstrated by the embodiments of the present disclosure.

[0016] Examples of the present disclosure go beyond subjective identification to enable objective identification of exemplary patterns in an audio signal that correspond to these events (e.g., GLIPh). Several different objective audio processing metrics can be calculated for use in this identification. These include loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, and spectral centroid. However, while known individual objective metrics are not capable of robustly identifying chill moments across a wide variety of music, examples of the present disclosure enable such robust detection by combining multiple metrics in a manner that identifies segments suitable for eliciting a chill response regardless of the overall characteristics of the music (e.g., genre, mood, or instrument placement).

[0017] For example, during the analysis of a given digital recording, once instances of GLIPh timestamps have been accumulated across various extracted feature datasets, a combination algorithm can be used to form a new dataset based on the amount of coincidence and proximity of GLIPhs identified within the digital recording. This dataset is referred to herein as a chill moment plot, and the combination algorithm generates the chill moment plot by attributing a weighted value to each GLIPh instance, for example, determining their coincidence rate, or per unit of time (e.g., per beat or per second). One reason for combining sets of metrics (e.g., metrics that identify individual GLIPhs) is that there are many types of chill inducers. With respect to standard CB-MIR-style feature extraction, there is no single metric that can encode all of the various acoustic and musical patterns known to determine musical segments that have a set of chill moment-inducing characteristics (e.g., chill-inducing characteristics identified by studies such as those by de Fleurian and Pearce). Furthermore, recording artists use many types of tools when composing and recording music, and there is no single tool that is typically used within a given song, and the wide variety of musical styles and genres have many different aesthetic approaches. The extreme diversity of popular music is strong evidence of this. It is common for a feature to have many points in a song. Melody pitch, for example, can potentially have hundreds of points of interest in a song, each of which can correspond to an individual GLIPh in that song. It is only by looking at the co-occurrence of multiple GLIPh features that match across multiple objective metrics that coherent patterns emerge.

[0018] Musical segments can be identified as primary and secondary chill segments, for example, based on GLIPh match, according to embodiments of the present disclosure. These matches, when listened to by experimental participants, result in predictable changes in behavioral and physiological measures, as detailed in the chill literature. Primary chill segments can be those segments in the audio recording with the highest GLIPh match, which can indicate segments most likely to produce chill, and secondary chill segments are those segments identified as inducing chill to a lesser extent based on lower GLIPh match than primary chill segments. Experiments were conducted to validate this predictive capability, the results of which are presented herein. These identified segments can be referred to as "chill phrases" or "chill moments," but because musical chill (e.g., elicitation of an autonomic physiological response in a given listener) is rarely experienced in practice, these segments can also be considered "impactful musical phrases," or musical segments with characteristics suitable for inducing an autonomic physiological response in general.

[0019] As discussed and illustrated in more detail herein, an embodiment of the present disclosure may include a) analyzing synchronization data from five domains (time, pitch, rhythm, loudness, and spectrum), and b) identifying specific acoustic signatures using only a very general musical map as the starting location. An embodiment may output a series of vectors containing the feature data selected for inclusion in the chill moment plot along with a GLIPh meta-analysis of each feature. For example, the loudness-per-beat data output may be stored as a vector of data, after which a threshold (or other detection algorithm) may be applied to determine the GLIPh instances of individual metric data (e.g., the top quartile of loudness-per-beat data), which are stored along with the start and end times for each GLIPh segment of data that falls in the top quartile in two vectors, one storing the start time and the other storing the end time. Each feature can then be analyzed and, for each beat, it can be determined whether the desired start and stop times for that feature fall within this moment in time, and if so, it is added to the value of the chill moment vector according to the particular weighting of that feature.

[0020] The output is thus a collection of numbers, strings, vectors of real numbers, and matrices of real numbers representing the various features under investigation. The chill moment output can be the sum of features (e.g., individual objective audio metrics) that indicate impactful moments for each elicitor (e.g., identified GLIPh or GLIPh matches) at each time step.

[0021] Embodiments of the present disclosure provide the ability to find the most impactful moments from a music recording, and the coincidence of acoustic and musical features that induce chill are predictive of listener arousal.

[0022] One embodiment of the present disclosure is a computer-implemented method for identifying segments in music, the method including receiving digital music data via an input operated by a processor, using the processor to process the digital music data using a first objective audio processing metric to generate a first output, using the processor to process the digital music data using a second objective audio processing metric to generate a second output, using the processor to generate a first plurality of detected segments using a first detection routine based on areas in the first output where a first detection criterion is met, using the processor to generate a second plurality of detected segments using a second detection routine based on areas in the second output where a second detection criterion is met, and using the processor to combine the first and second plurality of detected segments into a single plot representing matches of detected segments in the first and second plurality of detected segments, where the first and second objective audio processing metrics are different. The method may include identifying areas in the single plot that include the greatest number of matches within a predetermined minimum length of time requirement, and outputting a representation of the identified areas. The combining may include calculating a moving average of the single plot. The method may include identifying regions in the single plot where the moving average exceeds an upper limit and outputting an indication of the identified regions. One or both of the first and second objective audio processing metrics may be primary algorithms and / or configured to output primary data. Examples include the first and second objective audio processing metrics selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increase, sustained pitch, harmonic peak ratio, or key change.

[0023] An embodiment of the method may include applying a low pass envelope to the output of either the first or second objective audio processing metric. The first or second detection criteria may include an upper or lower boundary threshold. The method may include applying a length requirement filter to remove detected segments outside a desired length range. The combining may include applying respective weights to the first and second plurality of detections.

[0024] Another embodiment of the present disclosure is a computer system including an input module configured to receive digital music data, an audio processing module configured to receive the digital music data, perform a first objective audio processing metric on the digital music data, and perform a second objective audio processing metric on the digital music data, the first and second metrics generating respective first and second outputs, a detection module configured to receive the first and second outputs as inputs and generate, for each of the first and second outputs, a set of one or more segments for which a detection criterion is satisfied, and a combination module configured to receive the one or more segments detected by the detection module as inputs and aggregate each segment into a single data set including matches of the detection. The system may include a phrase identification module configured to receive the single data set of matches from the combination module as inputs and identify one or more regions where a highest average value of the single data set occurs during a predetermined minimum length of time. The phrase identification module may be configured to identify the one or more regions based on where a moving average of the single data set exceeds an upper limit. The phrase identification module may be configured to apply a length requirement filter to remove regions outside a desired length range. The combination module can be configured to calculate a moving average of the single plot. One or both of the first and second objective audio processing metrics can be primary algorithms and / or are configured to output primary data.

[0025] The system may include first and second objective audio processing metrics selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. The detection module may be configured to apply a low pass envelope to an output of either the first or second objective audio processing metric. The detection criteria may include upper or lower bound thresholds. The detection module may be configured to apply a length requirement filter to remove detected segments outside a desired length range. The combination module may be configured to apply respective weights to the first and second plurality of detections and then aggregate each detected segment based on the respective weights.

[0026] Yet another embodiment of the present disclosure is a computer program product including a tangible, non-transitory computer usable medium having computer readable program code including code configured to instruct a processor to: receive digital music data; process the digital music data using a first objective audio processing metric to generate a first output; process the digital music data using a second objective audio processing metric to generate a second output; generate a first plurality of detected segments using a first detection routine based on areas in the first output where a first detection criterion is met; generate a second plurality of detected segments using a second detection routine based on areas in the second output where a second detection criterion is met; and combine the first plurality of detected segments and the second plurality of detected segments into a single plot based on matches of the detected segments in the first and second plurality of detected segments, where the first and second objective audio processing metrics are different. The first and second objective audio processing metrics may be selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. The computer program product may include instructions for identifying a region in the single plot containing the greatest number of matches during a predetermined minimum length of time requirement and outputting a representation of the identified region. The product may include instructions for identifying one or more regions in which the highest average value of the single data set occurs during a predetermined minimum length of time. The product may include instructions for calculating a moving average of the single plot. The first or second detection criteria may include an upper or lower boundary threshold. The product may include instructions for applying a length requirement to a filter to remove detected segments outside a desired length range.

[0027] Yet another embodiment of the present disclosure is a computer-implemented method for identifying segments in music having characteristics suitable for eliciting an autonomic nervous system psychological response in a human listener, comprising: receiving digital music data via an input operated by a processor; using the processor to process the digital music data using two or more objective audio processing metrics to generate two or more respective outputs; detecting via the processor a plurality of detected segments in each of the two or more outputs based on areas that meet respective detection criteria; and using the processor to combine the plurality of detected segments in each of the two or more outputs into a single chill moment plot based on matches in the plurality of detected segments, wherein the first and second objective audio processing metrics are selected from the group consisting of: loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, anharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. The method may include using the processor to identify one or more regions in the single chill moment plot that include the greatest number of matches within a minimum length requirement; and using the processor to output an indication of the identified one or more regions. Examples include displaying, via a display device, a visual representation of the values ​​of the single chill moment plot relative to the length of the digital music data. Examples can include displaying, via a display device, a visual representation of the digital music data relative to the length of the digital music data overlaid with a visual representation of the values ​​of the single chill moment plot relative to the length of the digital music data. The visual representation of the values ​​of the single chill moment plot can include a curve of a moving average of the values ​​of the single chill moment plot. An embodiment of the method includes identifying an area in the single chill moment plot that includes the greatest number of matches during a predetermined minimum length of time requirement and outputting an indication of the identified area. Outputting can include displaying, via a display device, a visual representation of the identified area.Outputting may include displaying, via a display device, a visual representation of the digital music data relating to a length of the digital music data overlaid with a visual representation of the identified region within the digital music data.

[0028] Yet another embodiment of the present disclosure is a computer-implemented method for providing information identifying impactful moments in music, the method including: receiving, via an input operated by a processor, a request for information related to impactful moments in a digital audio recording, the request including a representation of the digital audio recording; accessing, using the processor, a database storing a plurality of identifications of different digital audio recordings and a corresponding set of information identifying impactful moments in each of the different digital audio recordings, the corresponding set including at least one of: start and stop times of chill phrases, or values ​​of a chill moment plot; matching, using the processor, the received identification of the digital audio recording to one of the plurality of identifications in the database, the matching including finding an exact match or a closest match; and outputting, using the processor, a set of information identifying impactful moments of the matched identification of the plurality of identifications in the database. The corresponding set of information identifying impactful moments in each of the different digital audio recordings may include information created using a single plot of detected matches for each of the different digital audio recordings generated using the method of embodiment 1 for each of the different digital audio recordings. The corresponding set of information identifying impactful moments in each of the different digital audio recordings can include information created using a single chill moment plot for each of the different digital audio recordings generated using the method of Example 29 for each of the different digital audio recordings.

[0029] Another embodiment of the present disclosure is a computer-implemented method for displaying information identifying impactful moments in music, the method including: receiving, via an input operated by a processor, a representation of a digital audio recording; receiving, via a communications interface operated by the processor, information identifying impactful moments in the digital audio recording, the information including at least one of: a start time and a stop time of a chill phrase, or a value of a chill moment plot; displaying, using the processor, the received identification of the digital audio recording to one identification of a plurality of identifications in a database, wherein matching includes finding an exact match or a closest match; and outputting, using a display device, a visual representation of the digital audio recording for a length of time of the digital audio recording overlaid with a visual representation of the chill phrase and / or the value of the chill moment plot for the length of time of the digital audio recording.

[0030] The present disclosure will become more fully understood from the following detailed description taken in conjunction with the accompanying drawings. [Brief description of the drawings]

[0031] [Figure 1A] 4 is a flow chart of an example routine for processing digital music data according to the present disclosure. [Figure 1B] 1B is a detailed flow chart of an example routine for processing the digital music data of FIG. 1A. [Figure 2A] 1 is a graph of amplitude over time of an example waveform of a digital music file. [Figure 2B] 1 is a visual representation of example outputs of a first representative objective audio processing metric along with corresponding plots of identified GLIPhs. [Figure 2C] 13 is a visual representation of example outputs of a second representative objective audio processing metric along with corresponding plots of identified GLIPhs. [Figure 2D] 1 is a visual representation of an example output of an identified GLIPh-based combination algorithm of first and second representative objective audio processing metrics. [Figure 2E] 2C is a visual representation of an example output of a phrase detection algorithm based on the output of the combination algorithm of FIG. 2D. [Figure 3A] FIG. 2 is a visual representation of the waveform of a digital music file. [Figure 3B] 3B is a visual representation of the loudness metric output based on the waveform of FIG. 3A. [Figure 3C] 3B is a visual representation of the output of the loudness band ratio metric in three different loudness bands based on the waveform of FIG. 3A. [Figure 3D] FIG. 4 illustrates an example output of a combination algorithm based on the objective audio processing metrics of FIGS. 3B and 3C overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 3E] FIG. 3D is a visual representation of the waveform of FIG. 3A showing the output of the phrase detection algorithm of FIG. 3D. [Figure 4A] 3B is a visual representation of the output of a dominant pitch melodic metric based on the waveform of FIG. 3A. [Figure 4B] FIG. 4B illustrates an example output of a combination algorithm based on the objective audio processing metrics of FIGS. 3B, 3C, and 4A overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 4C] FIG. 4C is a visual representation of the waveform of FIG. 3A showing the output of the phrase detection algorithm of FIG. 4B, and a comparison with the output of the phrase detection algorithm shown in FIG. 3E. [Figure 5A] FIG. 2 is a visual representation of the waveform of another digital music file. [Figure 5B] 5B is a visual representation of the output of a loudness objective audio processing metric based on the waveform of FIG. 5A. [Figure 5C] 5B is a visual representation of the output of the loudness band ratio algorithm metrics in three different loudness bands based on the waveform of FIG. 5A. [Figure 5D] 5B is a visual representation of the output of a dominant pitch melodic metric performed on the waveform of FIG. 5A. [Figure 5E] FIG. 5B illustrates an example output of a combination algorithm based on the objective audio processing metrics of FIGS. 5B, 5C, and 5D overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 5F] FIG. 5C is a visual representation of the waveform of FIG. 5A showing the output of the phrase detection algorithm of FIG. 5E. [Figure 6A] FIG. 5B is a visual representation of the output of a spectral flux metric based on the waveform of FIG. 5A. [Figure 6B] FIG. 6B illustrates an example output of a combination algorithm based on the objective audio processing metrics of FIGS. 5B, 5C, 5D, and 6A overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 6C] FIG. 5B is a visual representation of the waveform of FIG. 5A showing the output of the phrase detection algorithm of FIG. 6B, compared to the output of the phrase detection algorithm shown in FIG. 5F. [Figure 7] 11 is a set of plots generated using another song's waveform as input, showing detection outputs from a number of objective audio processing metrics based on the song's waveform and the output from a combination algorithm based on the output of the number of objective audio processing metrics, overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 8] 11 is a set of plots generated using yet another song waveform as input showing detection output from a number of objective audio processing metrics based on the song waveform and output from a combination algorithm based on the output of the number of objective audio processing metrics overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. [Figure 9A] 13 is an output plot from the combination algorithm run on the objective audiometric output of a song. [Figure 9B] 13 is an output plot from the combination algorithm run on the objective audiometric output of different songs. [Figure 9C] 13 is an output plot from the combination algorithm run on the objective audiometric output of different songs. [Figure 9D] 13 is an output plot from the combination algorithm run on the objective audiometric output of different songs. [Figure 10A] 1 is a graph of an example of subject data from a behavioral study. [Figure 10B] fMRI data showing a broad network of neural activation associated with increases during peak moments identified by the algorithm in music compared to non-peak moments. [Figure 11] FIG. 1 is an illustration of a mobile device display showing a social media application incorporating an embodiment of the present disclosure. [Figure 12] FIG. 1 is an illustration of a mobile device display showing a music streaming application incorporating an embodiment of the present disclosure. [Figure 13] FIG. 1 is an illustration of a computer display showing a music catalog application incorporating an embodiment of the present disclosure. [Figure 14] FIG. 1 is an illustration of a computer display showing a video production application incorporating an embodiment of the present disclosure. [Figure 15] FIG. 1 is a block diagram of an exemplary embodiment of a computer system for use with the present disclosure. [Figure 16] FIG. 1 is a block diagram of an exemplary embodiment of a cloud-based computer network for use with the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0032] Next, specific exemplary embodiments will be described to provide a general understanding of the principles of structure, function, and use of the devices, systems, and methods disclosed herein. One or more examples of these embodiments are illustrated in the accompanying drawings. Those skilled in the art will understand that the devices, systems, and components associated with or otherwise part of such devices, systems, and methods specifically described herein and illustrated in the accompanying drawings are non-limiting embodiments, and the scope of the disclosure is defined only by the claims. Features illustrated or described in connection with one embodiment can be combined with features of other embodiments. Such modifications and variations are intended to be included within the scope of the disclosure. Some of the embodiments provided herein may be schematic, including some that are not labeled as such, but will be understood by those skilled in the art to be schematic in nature. These may not be to scale or may be rather rough renderings of the disclosed components. Those skilled in the art will understand how to implement these teachings and incorporate them into the working systems, methods, and components provided herein that are associated with each of them.

[0033] To the extent that the present disclosure includes various terms for components and / or processes of the disclosed devices, systems, methods, etc., those skilled in the art will understand, in light of the claims, this disclosure, and the knowledge of one of ordinary skill in the art, that such terms are merely examples of such components and / or processes, and that other components, designs, processes, and / or operations are possible. As a non-limiting example, the present application describes the processing of digital audio data, but alternatively, or in addition, the processing may be performed via similar analog systems and methods, or may include both analog and digital processing steps. In this disclosure, like-numbered and like-lettered components of various embodiments generally have similar characteristics if those components are of a similar nature and / or serve a similar purpose.

[0034] The present disclosure relates to processing complex audio data, such as music, to identify one or more moments in the complex audio data that have the strongest characteristics suitable for eliciting an autonomic nervous system physiological response in a human listener. However, alternative configurations such as the reverse (e.g., moments in the complex audio data that have the weakest characteristics suitable for eliciting an autonomic nervous system physiological response in a human listener) are also disclosed. Thus, one skilled in the art will appreciate that the audio processing routines disclosed herein are not limited to configurations based on characteristics suitable for eliciting an autonomic nervous system physiological response in a human listener, but are broadly capable of identifying a wide range of complex audio characteristics depending on several configuration factors, such as: the individual metrics selected, the thresholds used for each metric to determine positive GLIPh instances, and the weights applied to each metric when combining matching GLIPh instances to generate an output (although referred to herein as a chill moment dataset, this reflects the selection of individual metrics with known associations with the identification of various chill inducers in neuroscience research, and thus in examples where a set of metrics is selected for the identification of different acoustic phenomena, a name reflecting the context for the output would be selected as well). Indeed, for example, while there may be correlations between music and biological responses yet to be discovered in research, embodiments of the present disclosure can be used to identify moments in arbitrarily complex audio data that are most likely to trigger biological activity by combining individual objective acoustic features that are associated with increased likelihood of biological activity.

[0035] Audio Processing 1A is a flow chart of an example routine 11 for processing audio data 101 according to the present disclosure. In FIG. 1A, the routine 11 can begin with audio data 101, which can be digital audio data such as music, which can be received via input 12. In subsequent steps, two or more objective audio processing algorithms 111, 112 (e.g., also referred to herein as metrics, audio metrics, or audio processing metrics) are performed on the audio data 101 to generate outputs representing audio characteristics (e.g., loudness, spectral energy) associated with the metrics 111, 112. For each metric output, a detection algorithm 131, 132 identifies one or more moments in the data where the metric output is relatively elevated (e.g., above a quartile of the data) and outputs these detections as a binary mask indicating positive and null detection regions in the time domain of the originally input audio data 101 (e.g., if the input audio data 101 is 200 seconds long, each binary mask can cover the same 200 seconds).

[0036] The combination algorithm 140 receives the input binary masks and aggregates them into a chill moment plot, which includes values ​​in the time domain of the aggregate matches. For example, if a moment in the audio data 101 returns a positive detection in both metrics, that moment is aggregated for that time in the output of the combination algorithm 140 with a value of "2". Similarly, if only one metric returns a positive detection for a moment, the value is "1". The combination algorithm can normalize the output as well as provide a moving average, or any other typical processing of data known to those skilled in the art. The combination algorithm 140 can be part of or associated with an output 19, which can provide the output of the combination algorithm 140 to, for example, a storage device, or another processor. Additionally, the routine 11 can include a phrase identification algorithm 150, which takes the output data from the combination algorithm 140 as input and detects one or more segments of the audio data that include one or more peaks in the chill moment plot, for example, based on their relative strength and proximity to each other. The phrase identification algorithm 150 may be part of or associated with the output 19, which may provide the output of the combination algorithm 140 to, for example, a storage device or another processor. The phrase identification algorithm 150 may output any data associated with the identified segments, including timestamps, and a detection of a primary segment based on a comparison of all identified segments. The phrase identification algorithm 150 may create and output segments of the original audio data 101 representing the identified segments.

[0037] Figure IB is a detailed flow chart of an example embodiment for processing digital music data using one or more computer processors, showing additional intermediate processing steps not shown in Figure IA. In Figure IB, process 10 can include routine 11 of Figure IA, as well as storage routine 12 and retrieval routine 13. Routine 11' shown in Figure IB can include routine 11 of Figure IA, but is shown here with additional steps that may or may not be included in routine 11 of Figure IA.

[0038] The routine 11' of FIG. 1B may start with audio data 101, which may be an audio waveform that may be encoded using a number of known lossless and lossy techniques, such as MP3, M4A, DSD, or WAV files. The audio data 101 may be received using an input of a computer system or may be obtained from a database, which may be local to the computer system or obtained over the Internet, for example. Once the audio data 101 is obtained, a number of different objective audio metrics 111, 112, 113 are separately executed by a processor of the computer system to extract primary data from the audio data 101, such as loudness per beat, loudness band ratio, pitch melody, etc. In a next optional step, post-processing routines 111', 113' may be implemented using a processor to prepare the data for subsequent detection processing using thresholds. The post-processing routines 111', 113' may include, for example, transforming the loudness per beat data using a low-pass envelope. In a next step, for each metric, upper or lower bound thresholds 121, 122, 123, such as upper or lower quartile functions, may be applied to the output data using a processor based on the distribution of the data. In a next step, based on the application of thresholds 121, 122, 123 in the previous step, a detection algorithm 130 uses a processor to identify segments of data that meet the threshold requirements. The detection algorithm 130 may enforce requirements, such as a requirement dictating that the selected segment must span a defined number of consecutive beats, in some implementations. For example, at least 2 seconds, or between 2 and 10 seconds, or between 1 and 30 seconds, etc. The detection algorithm 130 may output the detection as a binary mask.

[0039] A common need to detect chill-inducing features in a signal involves highlighting regions that represent signal changes, especially sudden or concentrated changes. For example, artists and composers increase loudness to draw attention to certain passages, and generally the more dramatic the loudness change, the more the listener responds. Detecting relevant segments in a signal usually involves identifying the relative highest or lowest regions in a recording. By employing thresholds such as upper or lower quartiles, aspects of the present disclosure detect the regions of greatest change relative to established dynamic ranges in a particular song. The use of quantile-based thresholds relative (e.g., top 25%) is advantageous because there can be a wide diversity of dynamic ranges within different genres and even between individual songs within a genre, and the use of absolute thresholds can result in undesirable over- or under-selection for most music. Furthermore, when the amount of signal variation for a particular recording is small (e.g., constant loudness), the upper quartile of loudness tends to select small, dispersed regions throughout the song that are unlikely to align significantly with other features in subsequent combination routines. However, when the signal peaks are concentrated in specific regions, quantile-based thresholding selects coherent regions that tend to align simultaneously with other features of interest in subsequent combination routines. While most of the feature detection examples in this disclosure employ quantile-based thresholding techniques, there are some features (e.g., key changes) that are not detected by quantile-based thresholding and employ other techniques described elsewhere herein.

[0040] After the individual segments have been identified, their detections are provided to a combination routine 140, which uses a processor to aggregate the segments and determine where selected segments overlap (e.g., match) and a higher numerical "score" is applied. As a result, where there is no overlap between the selections in the data plot, the score will be the lowest and where there is complete overlap between the selections in the data plot, the score will be the highest. The resulting scoring data, referred to herein as a Chill Moment Plot, may itself be output and / or visually displayed as a new data plot at this stage. Routine 11' may include a subsequent step of executing a phrase identification routine 150. In this step 150, the output of the combination routine is analyzed for sections containing high scores and segments using a processor. The segments with the highest overall score value may be considered "primary Chill phrases" and the identified segments with lower scores (but still meeting the criteria to be selected) may be considered "secondary Chill phrases". In a subsequent step, the chill phrases may be output 161 as data in the form of timestamps indicating the start and end of each identified phrase, and / or may be output 161 as an audio file created to contain only the “chill phrase” segments of the original audio data 101.

[0041] The process 10 may include a storage routine 12 that stores any of the data generated during execution of the routines 11, 11'. For example, the chill moment plot data and chill phrases may be stored in a database 170 as either timestamps and / or digital audio files. The database 170 may also store and / or be the source of the original audio data 101.

[0042] Any part of the process may include the operation of a graphical user interface to allow a user to execute any step of the process 10, view output and input data of the process 10, and / or set or change any parameters associated with the execution of the process 10. The process 10 may also include a search routine 13 that includes an interface (e.g., a graphical user interface and / or an interface with another computer system for receiving data) that allows a user to query the accumulation database 170. The user may search 180 the database for songs that ranked highest in the chill scoring as well as several metadata criteria, such as, for example, song title, artist name, song publication year, genre, or song length. The user interface may also allow a user to view details of a selected song, including chill phrase timestamps as well as other standard metadata. The user interface may also interface with an output 190 that, for example, allows playback of chill phrase audio as well as playback of the entire song with markings indicating where the chill phrases are present in the audio (e.g., an overlay on a waveform graphic of the selected song). The output 190 may also allow a user to transfer, download, or view any of the data generated in or associated with the operation of the process 10.

[0043] Figure 2A is a graph 200 of amplitude (y-axis) versus time (x-axis) for an example waveform 201 of a digital music file. The example waveform of Figure 2A, as well as the output audio metrics shown in Figures 2B and 2C, are entirely fictitious and for illustrative purposes only. In operation, an embodiment of the present disclosure includes running two or more objective audio processing metrics (111, 112, 113 in Figure 1B) on the waveform 201 to generate output data, an example of which is shown in Figure 2B.

[0044] FIG. 2B includes a plot 211 of an example output 21 of a first representative objective audio processing metric (e.g., 111 in FIG. 1B) along with a corresponding output mask 221 of the identified GLIPh 204. In FIG. 2B, the output 21 ranges from a minimum to a maximum value, and a threshold 201 can be applied to enable a detection algorithm (e.g., 130 in FIG. 1B) to extract from the output 21 individual acoustic events whose output satisfies the detection criterion (e.g., threshold 201). The detection criterion shown in FIG. 2B is a simple top quintile of the values ​​of the output 21, but other, more complex detection criteria are possible and may require a post-processing 111' step (e.g., performing a differentiation or Fourier transform to detect harmonies between matching notes) before application. Additionally, post-processing 111' can be used to change the time domain from a processing interval (e.g., 0.1 ms) to per-beat. Additionally, post-processing 111' can be used to convert the frequency domain processing to a time domain output. Using a beat-by-beat time frame allows the metric to be adaptive to the basic "atoms" of a song, so that tempo is not a confounding factor. The level of granularity can be deeper, with higher level features encapsulating several features such as pitch, or many other features such as spectral flux or spectral centroids, but this level does not need to be much smaller than the beat level to get effective results.

[0045] In Figure 2B, once a detection criterion (e.g., threshold 201) has been applied, the detection algorithm 130 converts the output 21 into a binary mask 221 of individual detection events 204 (also referred to herein as GLIPh) that are positive (e.g., value 1) in the time domain where a detection occurs and null (e.g., value 0) in the time domain between detections. The output mask 221 is provided as one input to a combination algorithm (e.g., 140 in Figure 1B), and another input mask comes from a second metric processing the same audio waveform (201 in Figure 2A), as shown in Figure 2C.

[0046] FIG. 2C includes a plot 212 of an example output 22 of a second representative objective audio processing metric (e.g., 112 of FIG. 1B) along with a corresponding output mask 222 of the identified GLIPhs 207. In FIG. 2C, the output 22 ranges from a minimum to a maximum value, and a threshold 202 can be applied to enable a detection algorithm (e.g., 130 of FIG. 1B) to extract individual acoustic events from the output 22 whose output satisfies a detection criterion (e.g., threshold 202). The detection criterion shown in FIG. 2C is a simple upper quartile of the value of the output 22, although other, more complex detection criteria are possible and may depend on the nature of the GLIPhs detected in the output 22 of the metric.

[0047] In Figure 2C, once a detection criterion (e.g., threshold 202) has been applied, detection algorithm 130 converts output 22 into a binary mask 222 of individual detection events 207 (also referred to herein as GLIPh) that are positive (e.g., value 1) in the time domain where a detection occurs and null (e.g., value 0) in the time domain between detections. Output mask 222 is provided as an input to combination algorithm 140 along with input mask 221 of Figure 2B, as shown in Figure 2D.

[0048] FIG. 2D includes plots of masks 221, 222 of detections from the two metrics of FIG. 2B and FIG. 2C, and an impact plot 230 of an example output (e.g., chill moment plot) of a combination algorithm 140 used based on the identified GLIPh of the first and second representative objective audio processing metrics. In the impact plot 230 of FIG. 2D, the masks 221, 222 are aggregated and matching detections are added to create a first region 238 where both masks are positive (e.g., match value 2), a second region 239 where only one mask is positive (e.g., match value 1), and a null region in between. In some cases, the input masks 221, 222 have a time domain interval (e.g., beat by beat), but this is not required and the impact plot 230 can be created using any time domain interval (e.g., minimum x-axis interval) to construct the first region 238 and the second region 239. In some cases, and as shown in more detail herein, a moving average of the first region 238 and the second region 239 may be created and included in the impact plot 230. The second region 238, which represents a peak in the chill moment plot, may be used to map individual timestamps back to the audio waveform of FIG. 2A, as shown in FIG. 2E as peak moments 280 of the audio data. Using these peak moments 280, a phrase detection algorithm (e.g., 150 of FIG. 2B) may identify impact regions 290 in the time domain where the peaks 280 exist and, in some cases, are clustered to create output data with timestamps 298, 299 corresponding to the location of the identified phrases 290.

[0049] Audio Processing Examples 3A-3E show processing steps for an example audio file using two objective audio processing metrics according to an embodiment of the present disclosure, and FIGS. 4A-4C show processing of the same audio file with the addition of a third metric.

[0050] 5A-5F show processing steps for different example audio files using three objective audio processing metrics according to an embodiment of the present disclosure, while FIGS. 6A-6C show processing of the same audio file with the addition of a fourth metric.

[0051] 7 and 8 each illustrate an example of eight-metric processing according to an embodiment of the present disclosure using a different example audio file.

[0052] FIG. 3A is a graph 300 of audio data with time in seconds along the x-axis and amplitude along the y-axis. In FIG. 3A, the audio data shown is a visual diagram of a waveform encoded in a digital music file. Audio waveform data can be digitally represented by the amplitude of the frequencies of the audio signal in samples per second. This data can be compressed or uncompressed depending on the file type. FIG. 3A shows the audio data as a vector of amplitudes, with each value representing the frequency value of the original audio file per sample. In the example audio file of FIG. 2, the audio data has a sampling rate of 44.1 kHz and a bit rate of 128-192.

[0053] FIG. 3B is a graph 311 of the output of an objective audio processing metric using the audio data of FIG. 3A as input. In the example of FIG. 3B, the metric is the spectral energy of the beats of the audio signal across the entire spectrum, and the graph 311 is a visual illustration of the output of an embodiment of the first objective audio processing metric 111 of the present disclosure. The data shown in FIG. 3B represents the general loudness of each beat of the audio waveform of FIG. 3A. From this data, an upper envelope and a lower envelope can be generated based on a threshold 301. In FIG. 3B, the threshold 301 is the upper quartile of the amplitude, and segments that fall within this upper quartile are detected and saved as the start and end time points where a beat exists for each detected segment. The upper quartile is a representative threshold, and other values ​​are possible. In general, the threshold 301 can be based on a relative value (e.g., a value based on the value of the data, such as the top 20% of the average or the 20% of the maximum) or an absolute value (e.g., a value that does not change based on the data). Absolute values ​​can be used, for example, if the data is normalized as part of a metric (e.g., the output value of the metric is between 0 and 1) or if the output value is frequency dependent since frequency is a more strict parameter for recording audio data (e.g., the amplitude of a sound can be scaled for a given audio data without changing the nature of the data, such as by making it louder or quieter, whereas absolute frequencies are typically preserved during recording and processing and typically cannot be altered without changing the nature of the data). Since loudness increases are one of the most fundamental chill response elicitors for listeners, the loudness onset and end points can be used as one set of inputs to a combination algorithm that calculates the most impactful moments in the song of the audio waveform of FIG. 3A, as shown in more detail below. The output of the combination algorithm is also referred to herein interchangeably as chill moment data or chill moment plot.

[0054] FIG. 3C is a set of three graphs 312a-c showing the output of an embodiment of the second objective audio processing metric 112 of the present disclosure performed on the waveform of FIG. 3A. Each of the three graphs 312a-c illustrates the spectral energy of beats in one of three different energy bands 312a-c, each represented by a frequency range of the audio signal (e.g., 20-400 Hz in the first graph 312a, 401-1600 Hz in the second graph 312b, and 1601-3200 Hz in the third graph 312c). The amplitude data in FIG. 3C shows the general loudness of each beat of the recording within the three energy bands as a percentage of the total energy. In each energy band 312a-c, a threshold 302 is applied to generate a lower envelope. In FIG. 3C, threshold 302 represents the upper quartile of the envelope data that may be calculated, and a post-processing routine is used to detect moments in the audio data where all bands fall below threshold 302 for all bands 312a-c. These detected moments are where the frequencies are balanced and represent where all the "instruments" in the music are playing at once (e.g., ensemble vs. solo). For example, an instrumental onset may elicit a chill response in the listener, so the detected onsets and end points where the bands are all below threshold for all bands are used to calculate onsets and end points that are combined with the detected segments of the loudness metric processing output of FIG. 3B to be used as input for a combination algorithm, the output of which is shown in FIG. 3D, and represents the most impactful moments of the song based on the objective audio processing metrics of FIG. 3B and FIG. 3C (e.g., spectral energy per beat and the matching spectral energy per beat in three separate energy bands).

[0055] Further, while FIG. 3C shows the same threshold 302 applied to each energy band, in some cases this threshold 302 is related only to the value of the metric in each energy band (e.g., the top 20% of values ​​in the first band 312a rather than the top 20% of values ​​in all bands 312a-c), while in other cases a different threshold may be used in each energy band, varying depending on which band is used and / or the number or size of the individual bands. In some cases, a detection algorithm using a threshold 302 in each energy band 312a-c will return a positive detection if the threshold is met in any one band 312a-c, while in other cases the detection algorithm will return a positive detection if the respective threshold is met in all bands 312a-c, some bands 312a-c, most bands 312a-c, or any other combination thereof. Furthermore, while the threshold has been discussed as being a 20% value relative to the average of the metric, this may alternatively be related to maximum and minimum values. Also, although 20% (eg, top quintile) is used throughout this disclosure, other thresholds are possible, such as top quartile, top half, or above or below.

[0056] In general, since the ultimate objective is to find peak values ​​for a song and across a combination of multiple different metrics, choosing a threshold that is too high (e.g., 0.1%) or too low (e.g., 80%) will effectively negate the contribution of detections from the metrics in the combination by making detections too common or too rare. This is part of the reason why no single individual metric can be robustly correlated with chill-inducing moments in real music. While a balance can be determined between the strength of correlation with any individual metric and the value of the threshold, a simpler approach is to establish that a peak in any one metric is not necessarily the most likely moment to induce chill, since studies have shown that no single acoustic feature alone is strongly predictive of inducing chill.

[0057] Rather, the inventors have discovered and verified that it is the coincidence of relative rises in the individual metrics that are associated with the acoustic moments with the strongest characteristics suitable for eliciting a physiological response of the autonomic nervous system in human listeners, and that detecting these relative rises does not depend strongly on a precise threshold, but rather more simply requires that a portion to a large portion of the rise in each individual metric is detected throughout the entire piece, which can be achieved by a range of thresholds. For example, the threshold is greater than 50% (e.g., the definition of a rise) and reaches 1% (e.g., 1 / 100th of the total moment of the piece), and this upper limit is based on the idea that a chill-inducing moment needs to last more than a few beats of music in order to be impressed and responded to by the listener. Thus, when a very long piece of music, such as an entire symphony, is being processed, 1 / 100th of the piece may represent significantly more than a few beats, and thus a maximum threshold cannot generally be established for all complex audio data (e.g., both pop music and symphonies).

[0058] The detection algorithm 130 is a process that identifies moments in a song where the value of a metric is above a threshold, and outputs these moments as positive detections during these moments into a new data set.

[0059] Figure 3D is an impact graph 330 of the output of a combination algorithm 140 performed using detections (e.g., GLIPh, which are segments in each metric output that exceed a respective threshold) identified by the detection algorithm 130 in the output of the first and second audio processing algorithms of Figures 3B and 3C. Figure 3D also includes the output of a phrase detection algorithm 150 based on the output of the combination algorithms. The example combination algorithm 140 used to generate the chill moment plot 360 of Figure 3D operates by aggregating the coincidences of detections in the objective audio processing metric outputs of Figures 3B and 3C.

[0060] An example combination algorithm can operate as follows: for each beat of a song, if the loudness of that beat rises above the threshold for that feature of the metric (e.g., the detection algorithm returns a positive value for one or more beats or time segments in the loudness metric output of FIG. 3B), the combination algorithm adds a weight of 1* to the aggregate value of each beat or time segment returned by the detection algorithm. Similarly, if the loudness per beat per band ratio value indicates that the feature is below the threshold for that feature, the metric can add a weight of 1* for the loudness per beat per band ratio to the aggregate value. Each beat of a song is considered to be "on" or "off" for the metric, and those binary features are multiplied by the weight of each metric and summed for each beat. This is the general design of the combination algorithm, regardless of the metric that is added. In FIG. 3D4, the y-axis corresponds to values ​​of 0, 1, and 2, and the weight of each metric is simply set to 1. The output of this process is a chill moment plot 360 with a stepped display based on time steps per beat. The combination algorithm can also generate a running average 361 of the chill moment plot 360, which shows the values ​​of the chill moment plot 360 over several beats. Note that in FIG. 3D , the chill moment plot 360 has been normalized to a range of 0 to 1 (from its original values ​​of 0 to 2).

[0061] Using the chill moment plot 360 as an input, the phrase detection algorithm 150 can identify regions 380 in the time domain where both metrics exceed their respective thresholds. In its simplest form, the phrase detection algorithm 150 returns these peak regions 380 as phrases. However, two short moments in music that are only a few beats apart are not processed very independently by a human listener, so from the perspective of identifying impactful moments (or moments that have suitable characteristics to evoke a psychological response of the autonomic nervous system), multiple peak regions 380 clustered together are more accurately considered to be one acoustic "event." Thus, a more robust configuration of the phrase detection algorithm 150 can establish a window around a group of peak regions 380 and attempt to determine where one group of peak regions 380 separates from another.

[0062] In the configuration of the phrase detection algorithm 150 of FIG. 3D, in addition to the moving average 361, an upper limit 371 and a lower limit 372 are considered. The moving average 361 is separately normalized to set the peak 381 to "1". In FIG. 3D, the upper limit 371 is about 0.65 and the lower limit 371 is about 0.40 (relative to the normalized impact rating). In the configuration of the phrase detection algorithm 150 of FIG. 3D, when the moving average 361 is above the upper limit 371, the peak region 380 is considered to be part of the identified phrase 390. The phrase detection algorithm 150 then determines the start and end points of each identified phrase 390 based on the time before and after the peak region 380 where the moving average 361 is below the lower limit 372. In some embodiments, only a single boundary (e.g., upper bound 371) is used, and the values ​​of upper bound 371 and lower bound 372 depend in part on the number of metrics used, the time average length of moving average 361, and also on the thresholds used for the individual metrics, since higher thresholds generally result in shorter duration detection windows.

[0063] Of note, when multiple metrics are used (e.g., eight or more), there may be only one peak region 380, and the value of the peak region 380 may not be the maximum impact rating (e.g., the peak region may correspond to a value of 7 out of 8 possible, assuming eight metrics and equal weighting). Thus, the peak region 380 need not be used by the phrase detection algorithm 150 at all, and the phrase detection algorithm may instead rely entirely on the moving average 361 (or another time-smoothing function of the chill moment plot 360) exceeding the upper limit 371 to establish the moment at which a phrase should be identified. Also, the use of additional metrics does not prevent one or more peak regions 380 from being sufficiently isolated from other elevated regions of the chill moment plot 361 and / or of short enough duration that the moving average 361 does not rise above the upper limit 371, and thus the phrase detection algorithm 150 does not identify phrases around those one or more peak regions 380.

[0064] In some cases, as shown in FIG. 3D, a small lead-in and / or lead-out time buffer may be added to each identified phrase 390, such that the beginning or end of the identified phrase 390 is established only when, for example, the moving average 361 falls below the lower limit 372 by more than the lead-in or lead-out buffer, thereby accounting for imprecision in capturing any musical "build-up" or "let-down" periods before or after the identified phrase 390 by ensuring that at least a few beats before and / or after any impactful moment are captured in the identified phrase 390. Additionally, this may prevent short dips in the moving average 361 that diverge from what may be subjectively considered a single impactful moment to a listener, although as shown in FIG. 3D and described in more detail below, such divergences may still be seen and detected in FIG. 3D, and if close enough and / or short enough, split identified phrases 390 may be merged. 5E, the phrase detection algorithm 150 may also dynamically adjust the length of the lead-in and / or lead-out time buffers based on the length of the identified phrase 390, the strength of or proximity to a peak in the chill moment plot 361 and / or the moving average 361, and / or the inflection of the moving average 361. In some instances, the start and stop instants of the identified phrase 390 may be triggered by the chill moment plot 360 dropping below a threshold or becoming zero.

[0065] The phrase detection algorithm 150 may also identify a single primary phrase, as shown in FIG. 3D as "primary." The phrase detection algorithm 150 may identify a single primary phrase, for example, by comparing, for each identified phrase 390, the moving average 361 or the average of the chill moment plot 360 for each identified phrase 390, and / or the duration of the moving average 361 that is above the upper limit 371, and identifying the identified phrase 390 having the higher value as the primary phrase. Additionally, as shown in FIG. 3D, two identified phrases 390 may be immediately adjacent to each other and may be combined into one identified phrase 390 (as shown in FIG. 3E) at the output of the phrase detection algorithm 150.

[0066] The phrase detection algorithm 150 outputs timestamps for the identified phrases 390, which can be mapped directly onto the original audio waveform, as shown in Figure 3E, which is a graph 340 of the waveform of Figure 3A, showing the identified phrases 390 and their associated timestamps 398, 399.

[0067] 4A-4C show how the chill moment plot 360 and identified phrases 390 of the audio sample of FIG. 3A change when a third objective audio processing metric, dominant pitch melodia, is added. FIG. 4A is a graph 413 of the output of the dominant pitch melodia metric based on the waveform of FIG. 3A, which can be thresholded for use by the detection algorithm 130. FIG. 4A shows the dominant pitch value at each moment as a frequency value, and a confidence value (not shown in FIG. 4A, which represents how clearly the algorithm sees the dominant pitch). This new metric is created by multiplying the frequency value of the pitch by the confidence value. This data is then thresholded using the upper quartile (not shown) in the same way as was done in FIGS. 3A and 3B, and in and out points are saved for the time before and after the data exceeds that threshold. Composers and musicians often make melodies higher during performance as a way to call attention to the melody, and higher pitches are known to elicit chill responses in listeners, so the Dominant Pitch Melodia is designed to find where the melody is "highest" and "strongest". Threshold detection of the Pitch Melodia output is based on the pitch frequency multiplied by a confidence value, which is then normalized and thresholded, for example, using the upper quartile. The start and end points from the detection algorithm 130 are then aggregated in the combination algorithm 140 in the same manner as the metrics in Figures 3A and 3B, and the phrase detection algorithm 150 is re-run to produce the chill moment plot 460, moving average 461, and identified phrases 490 in the impact graph 431 of Figure 4B. In the impact graph 431 of Figure 4B, the y-axis values ​​are normalized from 0, 1, 2, 3 to 0 to 1 to reflect the addition of the third metric. The resulting identified phrase 490 is mapped onto the audio waveform in FIG. 4C, which also shows a comparison of timestamps 498, 499 of the identified phrase 490 with timestamps 398, 399 of the identified phrase 390 using only two metrics (as shown in FIG. 3E).The addition of the third metric did not substantially change the location of the peaks 381, 481 in the moving averages 361, 461, although the duration of the identified phrases 390 both shrunk slightly, which may indicate improved accuracy in detecting the most impactful moments. Furthermore, the highest peak 481 in the moving average 461 in FIG. 4B has a higher prominence over its neighboring peaks than the highest peak 381 in the moving average 361 in FIG. 3D, which may also indicate improved confidence in the temporal location of this particular impactful moment.

[0068] Chill inducers such as relative loudness, instrumental entry and exit, and relative pitch rise have some universality in terms of eliciting physiological responses in humans, so that embodiments of the present disclosure can robustly identify suitable segments across essentially any type and genre of music, possibly using a minimal combination of two metrics. Music is unmediated, and studies have shown it to be an unconscious process. Listeners do not need to understand the language used in the lyrics, nor do they need to be from the culture in which the music originates, to respond to it. The disclosed algorithm focuses primarily on auditory features that have been shown to elicit physiological responses that activate the human reward center, which are nearly universal, acoustically, and the diversity of auditory features identified by the algorithm allows matching of even two of the resulting metrics to identify musical segments with suitable characteristics to elicit physiological responses of the autonomic nervous system across essentially any genre of music.

[0069] Figure 5A is a graph 500 of waveforms of different digital music files. Figure 5B is a graph 511 of the output from a loudness metric on the waveform input of Figure 5A, showing the corresponding thresholds 501 for use in the detection algorithm 130. Figure 5C is a graph 513 of the output from a loudness band ratio metric on the same input waveform of Figure 5A, in three different energy bands 512z, 512b, 512c, and the respective thresholds 502 used in the detection algorithm 130. Figure 5D is a graph of the output from a dominant pitch melodia metric, and the respective thresholds 503 used in the detection algorithm 130.

[0070] FIG. 5E is a graph 530 showing a chill moment plot 560 output from a combination algorithm 140 using the detection of the metrics of FIGS. 5B-5C as inputs, and also shows a moving average 561 of the chill moment plot 560. Similar to the results of FIGS. 3D and 4B, if there is a peak 480 in the chill moment plot 560 and a peak 481 in the moving average 561, and the moving average 561 exceeds the upper limit 571, then the phrase identification algorithm 150 has generated an identified phrase 590. In the configuration of the phrase identification algorithm 150 of FIG. 5E, the beginning and end of each identified phrase 590 is determined by an inflection point 592 in the moving average 561 around the location 591 where the moving average 561 falls below the lower limit 572. FIG. 5E shows timestamps 597, 598, 599 output by the phrase identification algorithm 150 for each identified phrase. The phrase identification algorithm 150 of FIG. 5E also classifies the third phrase as “primary” as a function of the duration of the moving average 561 or chill moment plot 560 above either the upper limit 571 or the lower limit 572, and / or based on the average of the moving average 561 or chill moment plot 560 during the knee 592 and / or the location 591 where the moving average 561 falls below the lower limit 572. In some cases, not shown, the phrase identification algorithm 150 may then enforce a minimum length, such as 30 seconds, on the primary phrase, which may result in the primary phrase overlapping with other phrases, as shown in other examples herein. The phrase identification algorithm 150 may extend the length of the phrase in various ways, for example, equally in both directions or preferentially in the direction of higher values ​​of the moving average 561 or chill moment plot 560.

[0071] In general, the time lengths of these windows 590 may correspond to a number of factors, such as a predetermined minimum or maximum value to capture adjacent detections when they occur within a maximum time characteristic, or other detection characteristics such as an increase in frequency / density where two of the three metrics reach their criteria. Additionally, while FIG. 5E shows an example using three metrics, examples of the present disclosure include dynamically adding (or removing) metrics as inputs to the combination algorithm 140 in response to any of the features of the graph 530, such as the number or length of the identified phrases 590, the value and / or characteristics (e.g., speed change) of the moving average 561 or chill moment plot 560 within the graph 530 and / or within the identified phrases 590. For example, if a three-metric calculation returns three phrases and the addition of one or two more metrics reduces this detection to two phrases, a two-phrase output may be used.

[0072] FIG. 5E illustrates a combination of three metrics based on the respective criteria of each metric, with combinations of two and four metrics (or more) being considered, with some embodiments including adjusting the respective detection criteria of each metric based on the number of metrics used in the combination. For example, when only two metrics are combined, the respective criteria can be tightened (e.g., lowering the threshold percentile for the overall metric output) to allow detections to be more clearly identified in the combination algorithm. Conversely, when more than two metrics are combined, the respective detection criteria can be loosened (e.g., increasing the threshold percentile for the overall metric output) to allow matches of multiple metrics to be more easily identified by the combination algorithm. Alternatively, combining each metric can include assigning a weight to each metric. In the embodiment presented herein, each metric is combined with a weight of 1.0, i.e., detections of each metric are added as 1 in the combination algorithm 150. However, other values ​​are possible and can be assigned based on the individual metrics being combined, or dynamically based on, for example, either the genre of music, or the output of the respective audio processing metric, or the output from other metrics used in the combination algorithm.

[0073] Examples also include running multiple metrics (e.g., 12 or more) and generating a matrix combination of all possible combinations or more. While the configuration of the currently described system and method is configured to make such a matrix unnecessary (e.g., if chill-inducing features are present in the audio signal, it is highly likely that they will be readily identified using any combination of metrics as long as the metrics are correctly associated with the chill-inducing acoustic features), as an academic exercise it may be useful to identify individual peak moments 581 as precisely as possible (e.g., within a beat or two), and the exact location may be sensitive to the number and selection of metrics. Thus, with a matrix combination of all possible combinations, the combinations may themselves be averaged, or trimmed from outliers and then averaged (the results may be substantially identical), to identify individual peak moments. Additionally, the phrase identification algorithm 150 may be run on this matrix output, although the results may not be significantly different from using all metrics in a single combination using the combination algorithm 140, or from using a smaller subset of metrics (e.g., 3, as shown in FIG. 5E).

[0074] Generally, this is considered to be a capacity issue. For example, when processing 1 million songs of a music catalog according to the embodiments of the present disclosure, the choice of using 3 or 12 metrics can make a significant difference in processing time and cost. Dynamically adjusting the number of metrics may therefore be most efficient if, for example, the combination algorithm 140 is first run on a combination of 3 metrics, and then, if certain conditions are met (e.g., the peak 581 does not stand out), a fourth metric can be added, run on demand, to determine whether this achieves the desired reliability at the location of the peak 481. Of course, if capacity is not an issue, running 8 or 12 metrics on all 1 million songs may provide the "best" data, even if the effective results (e.g., timestamps of the identified phrases 590) are not significantly different from the results generated with 3 or 4 metrics. Thus, embodiments of the present disclosure may include a hierarchy or priority list of metrics based on the measured strength of the observed match with the results of the combination with other metrics. This can be established per genre (or any other separation), for example, by running a representative sample of music of a genre across the full set of 12 metrics, and then establishing a hierarchy of those metrics based on their match with the results in a matrix of all possible combinations. This can be established as a subset of metrics less than 12 for use when processing other music of that genre. Alternatively, or additionally, the respective weights of detection from each metric can be adjusted in a similar manner, for example, by maintaining the use of all 12 metrics for all genres, but each with its own set of weights based on identified matches with the matrix results.

[0075] FIG. 5F shows the identified phrases 590 and their associated timestamps 597, 598, 599 from FIG. 5E displayed on the original waveform of FIG. 5A.

[0076] 6A-6C show that adding another suitable audio processing metric (e.g., a metric related to the same phenomenon as the other, in this case, chill-induced acoustic features) may not substantially change the results. FIG. 6A is a plot 614 of the output of another suitable processing metric, spectral flux, using the waveform of FIG. 5A as input and an associated threshold 604. FIG. 6B is a graph 613 of the combination algorithm 140 and phrase identification algorithm 150 re-run on the detections from the metrics of FIG. 5B-5D, adding the detections from the spectral flux metric of FIG. 6A. FIG. 6B shows the resulting chill moment plot 660, the moving average 691, the respective peaks 680, 681, and the indented phrases 690 with their respective timestamps 697, 698, 699, and start / end points 692 (e.g., the bend in the moving average 690 before or after the location 691 where the moving average falls below the lower limit 572).

[0077] FIG. 6C is a plot 640 of the waveform of FIG. 5A and the updated identified phrase of FIG. 6B. FIG. 6C also shows a comparison between the timestamps 697, 698, 699 of the updated phrase and the original timestamps 597, 598, 599 of the three metric output result of FIG. 5F. In FIG. 6C, the identified phrase 690 is generally aligned with the identified phrase 590 of FIG. 5E, as shown by their detection lengths being nearly the same. The length of the primary phrase is shortened due to the introduction of a very slight bend in the moving average 661 (as shown at 692′ in FIG. 6B) that was not present in the three metric result. In general, this is an example of how the addition of a metric can slightly change the length of a phrase by introducing more variability in the data without significantly changing the position of the phrase in capturing the peak event. However, as shown in a comparison of Figures 5E and 6B, the location of the peak 681 of the primary phrase has changed, indicating that while there is a high degree of confidence in the location of the identified phrase 590, additional metrics may be required if the exact location of the exact peak moment of impact 581, 681 is desired. Note, however, that the location of the peaks of the other non-primary phrases did not change significantly between Figures 5E and 6B.

[0078] In some embodiments, the identification of which window is the primary window may be based on several factors, such as the frequency and strength of detection in the identified segment, and the identification of the primary segment may change, for example, if two of the identified windows have substantially similar detection strengths (e.g., detection frequency in the identified windows) and swapping one metric for another subtly changes the balance of detection in each window without changing the detection in the window itself. Furthermore, a metric may increase validity (e.g., robustness) across many songs when adding a metric does not substantially change the results for a particular song. Thus, for example, adding spectral flux may not change the results for one particular song in a particular genre, but may significantly improve the reliability of chill phrase selection in another genre.

[0079] FIG. 7 is a group of plots 730, 711-718, generated using yet another song waveform as input, showing detection output from a plurality of objective audio processing metrics based on the song waveform, and output from a combination algorithm based on the output of the plurality of objective audio processing metrics, overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. In FIG. 8, the audio waveform was from a digital copy of the song "Bad to Me" by Billy J Kramer. Impact graph 730 shows a chill moment plot 760 and associated peaks 780, with primary phrases 790 and secondary phrases 791 identified in the chill moment plot 760 by an embodiment of a phrase identification algorithm. FIG. 7 also shows individual detection plots 711-718 from eight objective audio processing metrics used as input to the combination algorithm to generate impact graph 730. The eight objective audio processing metric plots are loudness 818, spectral flux 712, spectral centroid 713, inharmonicity 714, critical band loudness 815, dominant pitch melody 716, dissonance 717, and loudness band ratio 718. In operation, each of the eight objective audio processing metrics was processed (e.g., using respective thresholds) to generate a GLIPh, which was converted into a binary detection segment as shown in the metrics' corresponding detection plots 711-718. The binary detection segments were aggregated using a combination algorithm to generate the chill moment plot 760 of the impact graph 730.

[0080] Advantageously, the embodiment of the combination algorithm disclosed herein allows for the creation of a combination algorithm in which all combinations of the individual detections from these eight audio processing algorithms are capable of identifying segments or moments in an audio waveform that have audio characteristics suitable for eliciting a physiological response of the autonomic nervous system, as described above. In this embodiment of FIG. 7, a chill moment plot 760 of the impact graph 730 was generated using an equally weighted combination of the detections of each audio processing algorithm (e.g., as shown in plots 711-718), and a peak moment 780 was identified from the combination algorithm that contained the highest summation value in the chill moment plot 760. This peak moment 780 is surrounded by a small inner window 790 depicted within the shaded area representing the identified segment. The length of this segment can be determined in several ways to include one or more regions of maximum detection value, where only a single maximum detection peak 780 is present in the impact plot 730, an inner window 790 extends between adjacent local minima in the chill moment plot 760 to define the identified segment 790, and a larger grey window 791 represents the application of a time-based minimum segment length that extends the inner window to a 30 second window.

[0081] Since each of the audio processing algorithms in FIG. 7 is representative of one or more audio characteristics known to be associated with eliciting a physiological response of the autonomic nervous system, by combining the detection regions 711'-718' from the outputs 711-718 from each audio processing algorithm with equal weighting, as shown in the embodiment of FIG. 7, the combined output 760 (and resulting impact graph 730) can robustly identify the most "impactful" moments in an audio waveform across diverse genres of music, the identified impactful moments having the strongest characteristics suitable for eliciting an autonomic physiological response in a listener, based on which audio characteristics detectable by each audio processing algorithm are equally responsible for eliciting an autonomic physiological response (e.g., applying equal weighting to detected matches). This is based in part on prior and ongoing research, described in more detail below, which a) used embodiments of the present disclosure to determine correlations between brain activity and identified peaks in the combined plot using equal weighting, b) showed that equal weighting produces extremely strong correlations between identified segments and peaks in brain activity in subjects listening to music, and c) is evidence that equal weighting is sufficient to identify moments with the strongest characteristics suitable for eliciting an autonomic physiological response in the listener. Furthermore, a distinct advantage of the present disclosure is that the complexity of music not only allows for the use of a set of audio processing algorithms sufficient to detect a wide range of possible audio characteristics (of the desired type described above), but that equal weighting allows the routine to be useful across the widest range of music genres and types. Conversely, the weighting of the metrics, and adjustments to the individual threshold criteria used to generate the detection regions, can further tune the embodiments of the present disclosure to be more sensitive to specific genres of music.

[0082] Examples of the present disclosure also include making adjustments in each metric to (1) the weighting of detections in the output from each audio processing algorithm, (2) the detection threshold criteria (individually or across all audio processing algorithms), and / or (3) the minimum length of time for detections based on the genre or type of music. These example adjustments are possible without compromising the overall robustness of the output with respect to which audio processing algorithms are more likely to be coordinated with each other (e.g., more likely to produce peaks in the impact plot and cause discrimination) versus non-adjustments where detections in one or more audio processing algorithms are less likely to match detections in other audio processing algorithms due to similarities between music of the same or similar genre. In this example of FIG. 7, the detections 714' of the inharmonicity metric shown in plot 714 are very weakly correlated with any other detections in the output of other audio processing algorithms. If this lack of correlation of these detections 714' is associated with this genre of music, the fidelity of the resulting identifications (e.g., peaks 780 and segments 790) in the impact plot 730 can be increased by increasing the detection criteria of the outlier metric and / or decreasing the weighting of the detected segments 714' in the plot 714.

[0083] FIG. 8 is a group of plots 830, 811-818 generated using yet another song waveform as input, showing detection outputs from multiple objective audio processing metrics based on the song waveform, and output from a combination algorithm based on the output of the multiple objective audio processing metrics, overlaid with the output of a phrase detection algorithm applied to the output of the combination algorithm. In FIG. 8, the audio waveform is from a digital copy of the song "Without You" by Harry Nilsson. Impact graph 830 shows a chill moment plot 860, having a primary phrase 890 and a secondary phrase 890 identified within the chill moment plot 860 by an embodiment of a phrase identification algorithm. FIG. 8 also shows individual detection plots 811-818 from eight objective audio processing metrics used as input to the combination algorithm to generate impact graph 830. The eight objective audio processing metric plots are loudness 818, spectral flux 812, spectral centroid 813, inharmonicity 814, critical band loudness 815, dominant pitch melodia 816, dissonance 817, and loudness band ratio 818. In operation, each of the eight objective audio processing metrics was processed (e.g., using respective thresholds) to generate a GLIPh, which was then converted into a binary detection segment as shown in the metrics' corresponding detection plots 811-818. The binary detection segments were aggregated using a combination algorithm to generate the chill moment plot 860 of the impact graph 830.

[0084] In the impact graph 830, both the primary and secondary phrases 890, 891 have peaks 880 in the chill moment plot 860 of equal maximum values. The primary phrase 890 received a fixed length window of 30 seconds, determined here by the long duration of the chill moment plot 860 at the peak value 880, and the secondary phrase 891 received a window sized accordingly by expanding the window from the identified peak 880 to the local minimum of the chill moment plot 860. Other criteria for expanding the phrase window around the identified moment can be used, such as evaluating the local rate of change in the chill moment plot 860 of the running average change around the identified moment, and / or evaluating the strength of adjacent peaks in the chill moment plot 860 to expand the window to capture nearby regions of the waveform with strong characteristics suitable for eliciting an autonomic nervous system physiological response in the listener. This method produces a window with the highest possible overall average impact within some minimum and maximum time window.

[0085] Impact Curve Classification Method Examples of the present disclosure also include music classifications created using embodiments of the chill moment plot data described herein. The classifications can be based, for example, on where the highest or lowest impact areas occur in a song, or any aspect of the shape of the chill moment plot. Four examples are shown in Figures 9A-9D. Figures 9A-9D show different chill moment plots (stepped lines) 960, 960', 960", 960'" with moving averages (smooth lines) 961, 961', 961", 961"', as well as windows 971-976 showing identified chill moment segments for four different songs. Figure 9A is "Stairway to Heaven" by Lez Zeppelin, Figure 9B is "Every Breath You Take" by The Police, Figure 9C is "Pure Souls" by Kanye West, and Figure 9D is "Creep" by Radiohead. Examples of the disclosure include systems and methods for classifying various examples of chill moment plots, moving averages, and identified phrases to generate a searchable impact curve classification that allows music to be searched based on a song's impact taxonomy. Example searches include peak locations in the chill moment plot or moving average, phrase locations and durations, variability in the chill moment plot or moving average, or other characteristics related to matching chill generating elements. This also allows media producers to match song impact profiles with synchronized media, such as in a video commercial or feature film.

[0086] Objective Audio Processing Metrics The disclosed embodiments provide audio processing routines that combine the output of two or more objective audio metrics into a single audio metric, referred to herein as a Chill Moment Plot. However, the name "Chill Moment Plot" refers to the ability of the disclosed embodiments to detect moments in complex audio data (e.g., music) that have characteristics suitable for eliciting a physiological response of the autonomic nervous system in a human listener, known as "chill". The ability of the disclosed audio processing embodiments to detect moments with these characteristics is a function of both the metrics selected and the processing of the output of those metrics. Thus, some selections of metrics and / or some configurations of the detection and combination algorithms increase or decrease the strength of detection of characteristics suitable for eliciting a physiological response of the autonomic nervous system in a human listener, or detect for other characteristics. The simplest embodiment of detecting other characteristics is obtained by inverting the detection algorithm (e.g., application of a threshold to the output of the objective audio processing metric) or the combination algorithm. By inverting the detection algorithm (e.g., detecting positives as below the bottom 20% threshold, rather than as above the top 20%), the moment that is generally least relevant for eliciting chills is identified in each metric, and the matches of these detections are processed by the combination algorithm to return the peak match of the moment that has the weakest characteristics suitable for eliciting an autonomic physiological response in a human listener. Alternatively, without changing the operation of the detection algorithm, the minimum of the combination algorithm output can generally represent the moment that has the weakest characteristics suitable for eliciting an autonomic physiological response in a human listener, but with potentially less accuracy than if a lower threshold were used for detection in the output of each metric. This inversion is therefore possible when using metrics that individually correspond to acoustic features known to be associated with eliciting an autonomic physiological response in a human listener.

[0087] Alternatively, other metrics with different relevance may be used, such as a set of two or more metrics related to acoustic complexity, or conversely, acoustic simplicity. In these two examples, the combination algorithm may robustly detect peak moments or phrases of acoustic complexity or simplicity. However, overall complexity or simplicity may lack a robust definition that applies to all types and genres of music, which may make the selection of individual metrics difficult. In any case, embodiments of the present disclosure provide a method to utilize multiple different objective audio processing metrics to generate a combination metric that takes into account simultaneous contributions across multiple metrics.

[0088] In contrast to more vague or subjective acoustic descriptions such as complexity or simplicity, the listener's experience of an autonomic physiological response when listening to music is a well-defined test for comprehensive evaluation, even if such an event is not common: the listener either experiences a chill effect while listening to a piece of music or he or she does not. This binary test has enabled research into the phenomenon that establishes a verifiable association between acoustic characteristics and the likelihood that the listener will experience an autonomic physiological response. This research, and the quantifiable acoustic characteristics associated with it, helps establish a set of metrics that are deemed relevant for the present purpose of determining, without human evaluation, the moment or moments in any given piece of music that have the characteristics most suitable for eliciting an autonomic physiological response. Moreover, both the complexity and variety of music make it unlikely that any one objective audio processing metric alone can be reliably and significantly correlated with peak chill-inducing moments in music. The inventors of the present disclosure have discovered that coincidence of relatively elevated (e.g., not necessarily maximal) events in multiple metrics associated with chill-inducing characteristics can solve the problems associated with any single metric and robustly identify individual moments and associated phrases in a complex audio signal (e.g., music) that have the strongest characteristics suitable for inducing an autonomic nervous system physiological response in a human listener. Based on this, a combination algorithm (as described herein) has been developed to combine inputs from two or more individual objective audio processing metrics that can, for example, identify acoustic characteristics associated with a potential listener's experience of chill.

[0089] Examples of the present disclosure include the use of objective audio processing metrics that relate to acoustic features found in a digital recording of a song. The process does not rely on data from an external source, e.g., lyric content from a lyrics database. The underlying objective audio processing metric must be computable and must be specific in that there must be an "effective method" for computing the metric. For example, there are many known effective methods for extracting pitch-melody information from recorded music stored as a .wav file, or any file that can be converted to a .wav file. The method then relies on pitch information and may specifically search for pitch-melody information that is known to induce chill.

[0090] In combination, objective audio processing metrics capable of detecting chill can rely on social consensus to determine inducers known to produce chill. These are currently derived from scientific studies of chill, the expertise of composers and producers, and the expertise of musicians. Many of these are commonly known, such as sudden loudness or pitch melody. If the goal is to identify impactful musical moments, any objective audio processing metric known to exhibit (or empirically found to exhibit through experimentation) associations with positive human responses can be included in the algorithmic approach described herein. Representative example metrics that are well-defined objectively include loudness, loudness band ratio, critical band loudness, melody, inharmonicity, dissonance, spectral centroid, spectral flux, key changes (e.g., modulations), sudden loudness increases (e.g., crescendos), sustained pitch, and harmonic peak ratio. An embodiment of the present disclosure includes any two or more of these example metrics as inputs to the combination algorithm. The use of three or more of these example metrics generally improves detection of the most impactful moments in most music.

[0091] In general, using three or more metrics improves detection across a greater variety of music because some genres of music have common acoustic signatures, and for such genres, a match in two or three metrics may be equivalent to using eight or more metrics. However, other genres, especially those where the acoustic signatures associated with those two or three metrics are less common or less dynamic, may benefit more from adding additional metrics. Although adding additional metrics may dilute or reduce the effectiveness of the combination algorithm for certain types of music, as long as the added metrics measure acoustic properties that are distinct from the other metrics and that are associated with inducing chill effects in listeners, their inclusion improves the overall performance of the combination algorithm across all music types. All of the example metrics listed above meet this criterion when used in any combination, but nothing prevents any one metric from being substituted for another if the criterion is met. Additionally, given the similarities that exist within certain genres of music, embodiments of the present disclosure may include both pre-selecting the use of certain metrics when the genre of music is known and / or applying unequal weighting to the detection of each metric. Embodiments may also include analyzing the output of individual metrics.

[0092] As an extreme example, a solo singer's music may not have the instrumentation to generate meaningful data from certain metrics (e.g., dissonance), so that the detections from these metrics, when present, add a kind of random noise to the output of the combination algorithm. Even if multiple metrics add this kind of noise to the combination algorithm, as long as two or three relevant metrics are used (e.g., measuring acoustic properties that are actually present in the music), matching detections are extremely likely to be found above the noise. However, it is also possible to see when a given metric is providing random or very low strength detections, and the metric's contribution to the combination algorithm can be reduced by lowering its relative weighting based on the likelihood that its output is not meaningful, or alternatively, its contribution can be removed entirely if a sufficiently high confidence can be established that it has no contribution.

[0093] There are also many qualities identified as being associated with chill that do not have a commonly known effective objective method of detection. For example, virtuosity is known to be a musical chill inducer. Although virtuosity is generally considered to be an aesthetic characteristic related to the performer's skill, there is no clearly defined "effective method" for calculating identifiable sections within a musical recording that are suitable for exemplifying a subjective value such as "virtuosity". Also, testing the effectiveness of an algorithm to "identify virtuosity" may prove difficult or impossible.

[0094] The general method of using matching inducers can be applied to any particular use case. Considering the case of identifying irritating or annoying parts of a music recording (e.g., for the use case of avoiding playing music that matches these qualities), as a first step one would need to conceptually identify what irritating or annoying means in aesthetic terms, and then create an effective statistical method to identify those features. Those features can then be aggregated by the methods described herein, and progressively more effective means of identifying the type of part can be constructed through broadening the metrics used, adjusting thresholds for detection, and / or adjusting relative detection weights before being combined according to an embodiment of the combination algorithm.

[0095] Embodiments of the present disclosure may include additional detection metrics not shown in the drawings, such as sudden dynamic increases / crescendos, sustained pitches, harmonic peak ratios, chord changes / modulations, etc.

[0096] Sudden dynamic increases / crescendos: An example would involve first taking the first derivative of loudness as a representation of the change in loudness, and then using a threshold and detection algorithm to identify GLIPhs around areas where the first derivative is greater than the median and where the region of the first derivative peaks above the median plus the standard deviation.

[0097] Sustained pitch: Examples include detection algorithms to identify GLIPh regions, where dominant pitch confidence values ​​and pitch values ​​are analyzed to highlight specific regions where long sustained notes are held in the main melody. The detection metric in this case involves highlighting regions where the pitch frequency has low variance and exceeds a selected duration requirement (e.g., longer than 1 second).

[0098] Harmonic Peak Ratio: Examples include detection algorithms that identify GLIPh regions where the ratio of the base harmonic is compared to the peak harmonic to find sections where the dominant harmonic is not the 1st, 2nd, 3rd, or 4th harmonic. These sections highlight tonal characteristics that correlate with chill-inducing music. The detection metric in this case simply involves selecting regions that fit a certain harmonic ratio in the signal. For example, selecting regions where the 1st harmonic dominates compared to all other harmonics will highlight regions with a certain type of tonal quality. Similarly, selecting regions where upper harmonics dominate will reveal another type of tonal quality.

[0099] Key Changes / Modulations: An example would include using a detection algorithm to identify GLIPh regions where the dominant chords shift dramatically relative to the dominant chords established at the beginning of the piece. This shift indicates a key change or significant chord modulation. The detection metric in this case does not involve a threshold and directly detects the key change in the music.

[0100] Experimental Verification In two separate studies, the chill phenomenon (e.g., the physiological response of the autonomic nervous system associated with the acoustic characteristics analyzed by the embodiments of the present disclosure) was investigated by comparing data from the output of embodiments of the present disclosure with both brain activation and behavioral responses of listeners.

[0101] The implementation configuration of the algorithm was the same in both studies. To generate the prediction data, a combination algorithm was implemented using the GLIPh detection of the eight objective audio processing metrics as input to generate the chill moment plots. The nature of the eight objective audio processing metrics used was described in the previous section. Specifically, for the experimental validation investigated herein, the eight objective audio processing metrics used were loudness, critical band loudness, loudness band ratio, spectral flux, spectral centroid, dominant pitch melodia, inharmonicity, and dissonance, which are the eight metrics shown in Figures 7 and 8.

[0102] In the same manner as described in the previous section, the eight objective audio processing metrics were applied individually to the digital recordings, and the respective thresholds for the output of each metric were used to generate a set of detections (e.g., GLIPh) for each metric. The set of detections was combined using an embodiment of the combination algorithm of the present disclosure to generate a Chill Moment dataset that included running averages of the output of the combination algorithm to present a continuous graph of relative impact within a song to use for comparison. The running averages of the output of the combination algorithm generated for the recordings were compared to temporal data collected from human subjects listening to the same songs in a behavioral study and separately in an fMRI study.

[0103] behavioral research Behavioral studies were conducted to validate the ability of embodiments of the present disclosure to detect moments of peak impact (e.g., the moment with the highest relative likelihood of eliciting a physiological response of the autonomic nervous system) and, in general, to validate the ability of embodiments of the present disclosure to predict listeners' subjective assessments of a song's impactful properties during listening. In the behavioral study, participants listened to self-selected, chill-inducing music recordings (e.g., songs selected by users who were asked to choose songs they knew that had or could have given them a chill) from a list of 100 songs while moving an on-screen slider in real time to indicate their synchronous perception of the song's musical impact (from lowest impact to highest impact). The music selected by participants was generally contemporary popular music, and the length of the selected songs ranged from roughly 3 minutes to 6 minutes. The data for each participant's slider was cross-correlated with the output of each song, which was generated by the output of a combination algorithm run on the output of the eight objective audio processing metrics in which the participant's selected song was used as input.

[0104] The behavioral study was conducted with 1,500 participants. Participants' responses were significantly correlated with the combination algorithm's predictions for each song. Participants showed higher impact during phrases predicted by the combination algorithm to induce chill. In Figure 10A, a graph plotting the results of participants' slider data 1001 (labeled "human") is overlaid on a moving average of the combination algorithm output 1002 (labeled "machine"). In the results in Figure 10A, participant number 8 was listening to the song Fancy by Reba McEntire.

[0105] Using the continuous slider data received from the 1,500 participants while they listened to their chosen songs, we created Pearson correlation coefficients from the moving average of the slider data and the output of the combination algorithm. Table 1 shows the Pearson correlation coefficients for each of the 34 songs chosen by the 1,500 participants (many participants chose the same songs). The sum of the Pearson correlation coefficients for the 1,500 participants was 0.52, with a probability (p-value) of less than 0.001. In other words, we have the strongest possible statistical evidence that the combination algorithm, using detections from the eight objective audio processing metrics, was able to predict impactful moments in music as judged by real human listeners. [Table 1]

[0106] fMRI research We reanalyzed data from a natural music listening task in which participants listened to musical stimuli during a passive listening task. Seventeen participants with no musical training were examined while listening to nine-minute-long segments of symphonies by Baroque composer William Boyce (1711–1779). Using a general linear model, we performed a whole-brain analysis during the listening session, using findings from the same eight objective audio processing metrics used in the behavioral study to determine voxels whose activation levels correlated with higher predicted impact, as predicted by the combination algorithm. Figure 10B is an fMRI snapshot from this study that shows a widespread network of neural activation, as identified by the combination algorithm, and associated with increases during identified peak moments in music compared to non-peak moments.

[0107] Analysis of the fMRI study revealed significant tracking (p<0.01, cluster-corrected at q<0.05; (Cohen's d=0.75)) of the moving average of the output of the combination algorithm in multiple brain regions including dorsolateral and ventrolateral prefrontal cortex, posterior insula, superior temporal sulcus, basal ganglia, hippocampus, and sensorimotor cortex, as shown in FIG. 10B. No brain regions showed a negative correlation with predicted impact. A control analysis with loudness measures showed a significant response only in the sensorimotor cortex, and no brain regions showed a negative correlation with loudness. These results indicate that distributed brain regions involved in perception and cognition are sensitive to musical impact, and that the combination algorithm combined with detection from eight objective audio processing metrics according to embodiments of the present disclosure can identify temporal moments and segments of digital music data that are strongly correlated with peak brain activity in brain regions involved in perception and cognition.

[0108] Moreover, published studies support this. A basic study by Blood and Zatorre concluded that "subjective reports of chill were accompanied by changes in heart rate, electromyogram, and respiration. As the intensity of these chills increased, increases and decreases in cerebral blood flow were observed in brain regions thought to be involved in reward motivation, emotion, and arousal, including the ventral striatum, midbrain, amygdala, orbitofrontal cortex, and ventromedial prefrontal cortex. These brain structures are known to activate in response to other euphoric stimuli, such as food, sex, and addictive drugs." A study by de Fleurian and Pearce concluded that "structures belonging to the basal ganglia have been repeatedly associated with chill. In the dorsal striatum, increased activation was found in the putamen and left caudate nucleus when comparing music listening with and without the experience of pleasant chill."

[0109] Conclusion of the experiment The results of the behavioral and fMRI studies are significant. Clear correlations can be drawn with the academic literature describing the "chill response" in humans and the factors associated with that response. In the self-reported behavioral study, subjects indicated where they experienced high musical impact, which is directly related to the musical arousal required for the chill response. The fMRI study also confirmed that high activation in areas governing memory, pleasure, and reward strongly corresponded with the output of the combined algorithm. Thus, with the strongest statistical significance possible given the nature and scale of the experiment, the behavioral and fMRI studies together validated the ability of embodiments of the present disclosure to predict listeners' neurological activity associated with physiological responses of the autonomic nervous system.

[0110] Industrial Applications and Implementation Several commercial applications of the embodiments of the present disclosure can be employed based on the premise that curating large catalogs and making aesthetic judgments about music recordings is time consuming. For example, time can be saved by automating the ranking and retrieval of recordings for a particular application. The time it takes a human to go through a library of music recordings and select recordings for any application can be prohibitive. Making an aesthetic evaluation usually involves multiple listenings to the recordings. Given that popular music songs are 3-5 minutes long, this evaluation takes 6-10 minutes per song. There is also the aspect of burnout and fatigue: humans can lose objectivity when listening to many songs in a row.

[0111] An example of a representative use case is for a large music catalog holder (e.g., an existing commercial service such as Spotify, Amazon Music, Apple Music, or Tidal). Typically, a large music catalog holder wants to acquire new "paying subscribers" and convert "free users" into paying subscribers. Success can be based, at least in part, on the user's experience when interacting with the free version of the computer application that provides access to the music catalog. Thus, by applying the embodiments of the present disclosure, the music catalog service will have the means to deliver the "most compelling" or "most impactful" music to the user, which in turn is likely to have a direct impact on the user's purchasing decision. In this embodiment, a database of timestamps may be stored with the digital music catalog, the timestamps representing one or more impactful peak moments detected by a combination algorithm previously run on the objective audio processing metrics of each song, and / or one or more impactful musical phrases generated by a phrase detection algorithm previously run on the output of the combination algorithm. In general, metadata in the form of timestamps generated by the embodiments of the present disclosure for all songs in the service's catalog may be provided and used to improve the user's experience. In an example embodiment of the present disclosure, a user may be provided with a sample of the song that includes peak impactful moments and / or the sample may represent one or more identified impactful phrases.

[0112] Another example use case is in the entertainment and television industry. When a director is selecting music for a production, they often have to filter through hundreds of songs to find the right recording and the right part of that recording to use. In an example embodiment of the present disclosure, a software application provides identified impactful phrases and / or chill moment plots to a user (e.g., a film or television editor, producer, director, etc.) and allows the user to filter to impactful music within selected parameters (e.g., genre) to find the right recordings and phrases for the production. This may also include the ability to match impactful moments and phrases in a song to moments in a video.

[0113] In an example embodiment of the present disclosure, a cloud-based system allows a user to search through a large catalog of music recordings stored in the cloud as input, and delivers as output one or more search results that include or identify the most impactful moments of each returned song result. In an example embodiment of the present disclosure, a local or cloud-based computer-implemented service receives digital music recordings as input, which are processed through examples of the present disclosure to create data regarding timestamps of the peak impactful moments and / or most impactful phrases of each song, as well as data regarding any other musical characteristics that are provided as a result of processing using objective audio processing metrics. Examples include using the stored data combined with an organization's existing metadata to use an improved recommendation system using machine learning techniques, or generating actual audio files of the most impactful phrases, depending on the desired output.

[0114] Music therapy has also been shown to improve medical outcomes in a wide variety of settings, including lowering blood pressure, improving surgical outcomes with patient-selected music, pain management, anxiety treatment, depression, post-traumatic stress disorder (PTSD), and autism. Music therapists, like directors and advertisers, have the same problem of curating music and need to find music in a particular genre that patients can relate to and that elicits a positive response from the patient. Thus, embodiments of the present disclosure can be used to provide music therapists with segments of music to improve treatment outcomes by increasing the likelihood of a positive (e.g., chill) response from the patient. Some patients with certain illnesses (e.g., dementia or severe mental illness) may not be able to help therapists select music. If the patient can name a genre but not a specific song or artist, embodiments of the present disclosure can allow the therapist to select impactful music from that genre. Alternatively, if the patient can name an artist and the therapist is not familiar with the artist, embodiments of the present disclosure can be used to sort the most impactful moments from a list of songs and the therapist can play those moments to see if any of them elicit a response from the patient. Another example is a web interface that helps a music therapist search for music based on the patient's age and music that is likely to elicit an emotional response from the patient (e.g., find the most impactful music from the period when the patient was between 19-25 years old). Another example is a web interface that helps a music therapist select the least impactful music from a list of genres used in meditation practices for patients with PTSD.

[0115] Social Media Examples of the present disclosure include social media platforms and applications configured to use the example systems and methods described herein to enable users to find the most impactful chill phrases that can be paired with their video content in hopes of maximizing viewing and engagement time and also reducing the user's search time to find songs and search for sections to use. Examples include controlling the display of a mobile device or computer to display a visual representation of the data of the chill moment plot and / or a visual identification of the identified phrases (e.g., timestamp, waveform, etc.), which can accompany the selection from the respective song. In some examples, the display is interactive to allow the user to play or preview the identified phrases through an audio device. Examples of the present disclosure can provide several advantages to social media systems, including the ability to find impactful music segments to pair with short video content, maximize viewing and engagement time for videos, reduce user input and search time, and reduce licensing costs by diversifying music selections.

[0116] Non-limiting implementations include: a) embodiments of the present disclosure integrated into existing social media platforms; b) systems and methods for previewing multiple chill phrase selections and seeing how they pair with user-generated content; c) user interface and / or UI elements that visually represent chill moments of songs; d) using CB-MIR functionality to help users discover music from different eras and musical genres; e) using CB-MIR functionality to further refine audio selections within social media apps; f) providing users with a way to license songs that are most likely to connect with listeners; g) previewing songs by identified impactful phrases to reduce listening time for music searches; h) providing a way for social media platforms to expand song selection while controlling licensing costs.

[0117] Figure 11 is an illustration of a mobile device display 1100 showing a social media application incorporating an embodiment of the present disclosure. Figure 11 shows a user selection of a photo 1101 and an overlay of audio data 1102 visually indicating the selection of a music track along with a window identifying a chill phrase 1103 and a line 1104 representing the average of the chill moment plot for the selected music track.

[0118] Music Streaming Platform Examples of the present disclosure include integration with music streaming services to help users discover more impactful music and enhance playlists, for example by allowing them to find and add to their playlist music with similar chill moment characteristics and / or tracks predicted by the systems and methods of the present disclosure to have highly positive emotional and physical effects on humans. Examples may also allow users to listen to the most impactful sections during song previews.

[0119] 12 is an illustration of a mobile device display 1200 showing a music streaming application incorporating an embodiment of the present disclosure. FIG. 12 shows an exemplary music streaming application interface 1202, illustrating a user selection of music tracks 1203, 1204, 1205, an overlay of audio data 1206 for each music track with a window 1207 identifying chill phrases, and a line 1208 representing the average chill moment plot for the selected music tracks. Examples include an embodiment of the present disclosure that allows a user of a music streaming platform to search for specific chill plot classifications, which can assist a user, for example, in creating a playlist that includes all songs with impactful endings, beginnings, or middles, as well as a playlist of songs that includes a mixture of song classifications.

[0120] Song Catalog Non-limiting embodiments include systems and methods that assist creators in finding appropriate music for television series and movies, specifically music that matches the timing of the scenes. Using existing technology, this process can be time-consuming, especially from large catalogs. Examples of the present disclosure can assist creators, for example, in filtering song search results by impactful phrases in a song (e.g., phrase length and classification). Examples also enable the creation of new types of metadata related to chill moments (e.g., timestamps indicating chill moment segment locations), which can reduce search time and costs.

[0121] Figure 13 is an illustration of a user interface 1300 presented on a computer display showing a music catalog application incorporating an embodiment of the present disclosure. Figure 13 shows a user selection of a song, presenting audio data 1321 in a window 1320 representing the music track selection, and a separate musical impact window 1310 having an output 1314 from a combination algorithm processing the selected song, and a line 1313 representing the average of the chill moment plot. The musical impact window 1310 also presents a visual display of first and second identified impactful phrases 1311, 1312 for the selected music track.

[0122] Exemplary features include a) the ability to filter a song database by characteristics of a song's chill moment plot, b) identifying predictably impactful songs, c) finding identified chill segments in songs, d) populating a music catalog with new metadata corresponding to any of the data generated using the methods described herein, and e) reducing search time and license costs. Examples of the present disclosure also include a user interface that provides user control over parameters of the combination algorithm and phrase detection algorithm. For example, allowing a user to adjust or remove weights for one or more input metrics to find different types of phrases. This on-the-fly adjustment allows the combination algorithm and phrase detection algorithm to be rerun without reprocessing the individual metrics. This functionality allows, for example, increasing the weights of parameters related to pitch and melody to find songs with larger melodic peaks, or increasing the weights of metrics related to timbre to find moments that feature a similar acoustic profile. Examples include a user interface that allows a user to individually adjust parameters such as metric weights, or a preselected structure that identifies a preselected acoustic profile. Through the use of interactable elements (e.g., toggles, knobs, sliders, or fields), users can instantly and interactively react to the displayed chill moment plot and associated phrase detections.

[0123] Implementations include: a) providing data related to the chill moment plot to a user interface of video editing software, b) providing data related to the chill moment plot to a user interface of a music catalog application to facilitate users previewing tracks using identified phrases and / or seeking to individual tracks based on the chill moment data, c) providing data related to the chill moment plot to a user interface of audio editing software, d) providing data related to the chill moment plot to a user interface of a music selection application on an airliner to assist passengers in selecting music, e) providing data related to the chill moment plot to a user interface of a kiosk in a physical and digital record store, and f) enabling users to preview artists and individual songs using impactful phrases.

[0124] Examples of the present disclosure include systems and methods for: a) providing data related to chill moment plots to social media platforms for instant generation of social media slideshows; b) generating chill moment plots for live music; c) populating existing digital music catalogs with data related to chill moment plots to enable previews with impactful phrases; d) providing data related to chill moment plots to software for previewing multiple chill moment phrases to see how they are paired with visual editing sequences; and e) processing data related to chill moment plots to provide new metadata to catalog holders and new opportunities to license impactful portions of their songs.

[0125] Audio, Film, Television and Advertising Production Film, television, and advertising producers and marketers want to find music that connects with their target audiences. Examples of the present disclosure include systems and methods that use data related to chill moment plots to help users find impactful moments in recorded music and enable them to pair these chill phrases with advertising, television, or movie scenes. One example advantage is the ability to pair identified chill segments of a song with key moments in an advertisement. FIG. 14 is an illustration of a software interface 1400 on a computer display showing a video production application 1401 incorporating an example of the present disclosure. FIG. 14 shows a current video scene 1410 and an audio-video overlay 1420 showing the time alignment of the audio track and the video track 1430. The audio-video overlay 1420 includes two-channel audio data 1421 representing a music track selection along with an adjacent window 1422 identifying an identified chill phrase 1423, as well as a line 1424 representing the average of the chill moment plot 1425 for the selected music track 1421. Implementations in an audio production context include systems and methods that provide visual feedback of chill plots and phrase selection in real time as different mixes of a song's tracks are composed. Examples can also provide a more detailed breakdown of what metrics are being put into the chill plot of the current song being edited / mixed, giving producers insight into how they can improve the music.

[0126] Gaming Examples of the present disclosure include systems and methods for enabling game developers to find and use the most impactful sections of music to enhance the game experience, thereby reducing labor and production costs. Examples of the present disclosure include using the systems and methods disclosed herein to remove subjectivity from game designers, allowing them to identify the most impactful parts of music and synchronize them with the most impactful parts of the game experience. For example, during game design, music that shows cut scenes, level changes, and challenges central to the game experience. Example benefits include increasing user engagement by integrating the most impactful music, providing music discovery for in-app music purchases, matching music segments to game scenarios, and reducing labor and license costs for game makers. Examples include providing music visualizations synchronized with chill plot data, which may include synchronizing in-game visual cues or dynamic lighting systems of the environment in which the music is played. Examples include aiding in the creation of music tempo games that derive timing and interactivity from chill plot peaks. Implementations include cuing chill moment segments of music in real time in sync with a user's gameplay, and using data related to chill moment plots to indicate cut scenes, level changes, and challenges central to the game experience.

[0127] Health and Wellness People often want to find music that may help relieve stress and improve well-being, which can be done by creating a playlist from music recommendations based on data associated with the chill moment plot. Implementations of the disclosed systems and methods include: a) using data associated with the chill moment plot to select music that resonates with Alzheimer's disease or dementia patients; b) using data associated with the chill moment plot as a testing device in a clinical environment to determine music that resonates most with Alzheimer's disease or dementia patients; c) using data associated with the chill moment plot to integrate music into wearable health / wellness products; d) using data associated with the chill moment plot to select music for exercise activities and workouts; e) using data associated with the chill moment plot to help reduce anxiety in patients before surgery; f) using data associated with the chill moment plot in a mobile application where a physician may prescribe a curated playlist to treat pain, depression, and anxiety; g) using data associated with the chill moment plot to select music for meditation, yoga, and other relaxation activities; and h) using data associated with the chill moment plot to help patients with pain, anxiety, and depression.

[0128] Computer Systems and Cloud-Based Implementations FIG. 15 is a block diagram of an exemplary embodiment of a computer system 1500 on which the present disclosure can be constructed, executed, trained, etc. For example, with reference to FIGS. 1A-14, any module or system can be an example of the system 1500 described herein, such as the input 12, the objective audio processing metrics 111, 112, the detection algorithm 130, the combination algorithm 140, and the phrase detection algorithm 150, the output 19, and any of the associated modules or routines described herein. The system 1500 can include a processor 1510, a memory 1520, a storage device 1530, and an input / output device 1540. Each of the components 1510, 1520, 1530, and 1540 can be interconnected, for example, using a system bus 1550. The processor 1510 can process instructions executed within the system 1500. The processor 1510 can be a single-threaded processor, a multi-threaded processor, or a similar device. The processor 1510 may be capable of processing instructions stored in memory 1520 or on the storage device 1530 .The processor 1510 may perform operations such as: a) performing audio processing metrics; b) applying a threshold to the output of one or more audio processing metrics to detect GLIPh; c) performing a combination algorithm based on the detection of two or more audio processing metrics; d) performing a phrase detection algorithm on the output of the combination algorithm; e) storing output data from any of the metrics and algorithms disclosed herein; f) receiving digital music files; g) outputting data from any of the metrics and algorithms disclosed herein; h) generating and / or outputting digital audio segments based on the phrase detection algorithm; i) receiving a user request for data from any of the metrics and algorithms disclosed herein and outputting the results; and j) operating a display device of a computer system, such as a mobile device, to visually present data from any of the metrics and algorithms disclosed herein, among other features described in connection with the present disclosure.

[0129] The memory 1520 can store information within the system 1500. In some implementations, the memory 1520 can be a computer-readable medium. The memory 1520 can be, for example, a volatile or non-volatile memory unit. In some implementations, the memory 1520 can store information related functions for executing objective audio processing metrics and any algorithms disclosed herein. The memory 1520 can also store digital audio data, as well as outputs from the objective audio processing metrics and any algorithms disclosed herein.

[0130] The storage device 1530 can provide mass storage for the system 1500. In some implementations, the storage device 1530 can be a non-transitory computer-readable medium. The storage device 1530 can include, for example, a hard disk device, an optical disk device, a solid state drive, a flash drive, a magnetic tape, and / or some other mass storage device. The storage device 1530 can alternatively be a cloud storage device, for example, a logical storage device including multiple physical storage devices distributed over and accessed using a network. In some implementations, information stored on the memory 1520 can also or instead be stored on the storage device 1530.

[0131] The input / output devices 1540 can provide input / output operations for the system 1500. In some implementations, the input / output devices 1540 can include one or more of the following: a network interface device (e.g., an Ethernet card or an Infiniband interconnect), a serial communication device (e.g., an RS-232 10 port), and / or a wireless interface device (e.g., a short-range wireless communication device, an 802.7 card, a 3G wireless modem, a 4G wireless modem, a 5G wireless modem). In some implementations, the input / output devices 1540 can include a driver device configured to receive input data and send output data to other input / output devices, such as a keyboard, a printer, and / or a display device. In some implementations, mobile computing devices, mobile communication devices, and other devices can be used.

[0132] In some implementations, system 1500 may be a microcontroller. A microcontroller is a device that includes multiple elements of a computer system in a single electronics package. For example, the single electronics package may include a processor 1510, a memory 1520, a storage device 1530, and / or an input / output device 1540.

[0133] 16 is a block diagram of an exemplary embodiment of a cloud-based computer network 1610 for use with the present disclosure. The cloud-based computer network 1610 can include digital storage services 1611 and processing services 1612, each of which can be provided by one or more individual computer processing and storage devices located at one or more physical locations. The cloud-based computer network 1610 can transmit and receive data 1621, 1631 from individual computer systems 1620 (e.g., personal computers or mobile devices) as well as from a network 1630 of individual computer systems 1620 (e.g., servers operating a music streaming service) via the Internet or other digital connection means. The cloud-based computer network 1610 may facilitate or complete the performance of operations such as: a) performing audio processing metrics and applying a threshold to the output of one or more audio processing metrics to detect a GLIPh; b) performing a combination algorithm based on the detection of two or more audio processing metrics; c) performing a phrase detection algorithm based on the output of the combination algorithm; d) storing output data from any of the metrics and algorithms disclosed herein; e) receiving digital music files; f) outputting data from any of the metrics and algorithms disclosed herein; g) generating and / or outputting digital audio segments based on the phrase detection algorithm; h) receiving user requests for data from any of the metrics and algorithms disclosed herein and outputting results; and i) operating a display device of a computer system, such as a mobile device, to visually present data from any of the metrics and algorithms disclosed herein, among other features described in connection with the present disclosure.

[0134] Although an example processing system has been described above, embodiments of the subject matter and functional operations described above can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more of them. Embodiments of the subject matter described herein can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible program carrier, such as a computer-readable medium, for execution by or to control the operation of a processing system. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or a combination of one or more of them.

[0135] Various embodiments of the present disclosure may be implemented, at least in part, in any conventional computer programming language. For example, some embodiments may be implemented in a procedural programming language (e.g., "C" or ForTran95), or an object-oriented programming language (e.g., "C++"). Other embodiments may be implemented as pre-configured, stand-alone hardware elements and / or pre-programmed hardware elements (e.g., application specific integrated circuits, FPGAs, and digital signal processors), or other related components.

[0136] The term "computer system" may encompass all apparatus, devices, and machines for processing data, including, by way of non-limiting example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, a processing system may include code that creates an environment for the execution of the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0137] A computer program (also known as a program, software, software application, script, executable logic, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in several coordinated files (e.g., a file that stores one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer, or on several computers located at one site, or distributed across several sites and interconnected by a communication network.

[0138] Such an embodiment may include a set of computer instructions fixed on any tangible, non-transitory medium such as a computer readable medium. The set of computer instructions may embody all or part of the functionality previously described herein with respect to the system. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile or volatile memory, media and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks or magnetic tapes; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated in special purpose logic circuitry. The components of the system may be interconnected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.

[0139] Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages ​​for use with many computer architectures or operating systems. Furthermore, such instructions can be stored in any memory device, such as semiconductor, magnetic, optical, or other memory devices, and transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technology.

[0140] Among other ways, such computer program products may be distributed as removable media with accompanying printed or electronic documentation (e.g., shrink-wrapped software), preloaded into a computer system (e.g., in a system ROM or fixed disk), or distributed from a server or bulletin board over a network (e.g., the Internet or the World Wide Web). Indeed, some embodiments may be implemented in a Software-as-a-Service model ("SAAS") or a cloud computing model. Of course, some embodiments of the present disclosure may be implemented as a combination of both software (e.g., computer program products) and hardware. Still other embodiments of the present disclosure are implemented entirely as hardware or entirely as software.

[0141] Those skilled in the art will appreciate further features and advantages of the present disclosure based on the description and embodiments provided. Thus, the present invention is not limited by what has been specifically shown and described. For example, the present disclosure provides for processing digital audio data to identify impactful moments and phrases in a song, but the present disclosure can also be applied to other types of audio data, such as speech or environmental noise, to evaluate their acoustic characteristics and their ability to elicit physical responses from human listeners. All publications and references cited herein are expressly incorporated herein by reference in their entirety.

[0142] Examples of the above-described embodiments may include: 1. A computer-implemented method for identifying segments in music comprising: receiving digital music data via an input operated by a processor; using the processor to process the digital music data using a first objective audio processing metric to generate a first output; using the processor to process the digital music data using a second objective audio processing metric to generate a second output; using the processor to generate a first plurality of detected segments using a first detection routine based on areas in the first output where a first detection criterion is met; using the processor to generate a second plurality of detected segments using a second detection routine based on areas in the second output where a second detection criterion is met; and using the processor to combine the first plurality of detected segments and the second plurality of detected segments into a single plot representing matches of detected segments in the first and second plurality of detected segments, wherein the first and second objective audio processing metrics are different. 2. The method of example 1, comprising: identifying an area in the single plot that contains the greatest number of matches within a predetermined minimum length of time requirement; and outputting a representation of the identified area. 3. The method of example 1 or example 2, wherein the combining comprises calculating a moving average of a single plot. 4. The method of example 3, comprising: identifying an area in the single plot where the moving average exceeds an upper limit; and outputting an indication of the identified area. 5. The method of any of Examples 1 to 4, wherein one or both of the first and second objective audio processing metrics are primary algorithms and / or are configured to output primary data. 6. The method of any of Examples 1 to 5, wherein the first and second objective audio processing metrics are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. 7. The method of any of embodiments 1 to 6, further comprising applying a low pass envelope to an output of either the first or second objective audio processing metric. 8. The method of any of Examples 1 to 7, wherein the first or second detection criterion comprises an upper or lower boundary threshold. 9. The method of any of the preceding claims, wherein the detecting includes applying a length requirement filter to remove detected segments outside a desired length range. 10. The method of any of examples 1-9, wherein the combining includes applying respective weights to the first and second plurality of detections. 11. A computer system comprising: an input module configured to receive digital music data; an audio processing module configured to receive the digital music data and perform a first objective audio processing metric on the digital music data and a second objective audio processing metric on the digital music data, the first and second metrics generating respective first and second outputs; a detection module configured to receive as inputs the first and second outputs and to generate, for each of the first and second outputs, a set of one or more segments for which a detection criterion is satisfied; and a combination module configured to receive as inputs one or more segments detected by the detection module and to aggregate each segment into a single dataset comprising detection matches. 12. The computer system of example 11, including a phrase identification module configured to receive as input a single dataset of matches from the combination module and identify one or more regions where the highest average value of the single dataset occurs for a predetermined minimum length of time. 13. The computer system of example 12, wherein the phrase identification module is configured to identify one or more regions based on where a moving average of a single data set exceeds an upper limit. 14. The computer system of any one of examples 12 to 23, wherein the phrase identification module is configured to apply a length requirement filter to remove regions outside a desired length range. 15. The computer system of any of Examples 11 to 14, wherein the combination module is configured to calculate a moving average of a single plot. 16. A computer system according to any of Examples 11 to 15, wherein one or both of the first and second objective audio processing metrics are primary algorithms and / or are configured to output primary data. 17. The computer system of any of Examples 11 to 16, wherein the first and second objective audio processing metrics are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. 18. A computer system according to any of Examples 11 to 17, wherein the detection module is configured to apply a low pass envelope to an output of either the first or second objective audio processing metric. 19. The computer system of any of Examples 11 to 18, wherein the detection criteria include upper or lower boundary thresholds. 20. A computer system according to any of Examples 11 to 1, wherein the detection module is configured to apply a length requirement filter to remove detected segments outside a desired length range. 21. A computer system according to any of embodiments 11 to 20, wherein the combination module is configured to apply respective weights to the first and second plurality of detections and then aggregate each detection segment based on the respective weights. 22. A computer program product comprising a tangible, non-transitory computer usable medium having computer readable program code comprising code configured to instruct a processor to: receive digital music data; process the digital music data using a first objective audio processing metric to generate a first output; process the digital music data using a second objective audio processing metric to generate a second output; generate a first plurality of detected segments using a first detection routine based on areas in the first output where a first detection criterion is met; generate a second plurality of detected segments using a second detection routine based on areas in the second output where a second detection criterion is met; and combine the first plurality of detected segments and the second plurality of detected segments into a single plot based on matches of the detected segments in the first and second plurality of detected segments, wherein the first and second objective audio processing metrics are different. 23. The computer program product of Example 22, wherein the first and second objective audio processing metrics are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. 24. The computer program product of example 22 or 23, comprising instructions for: identifying an area within the single plot that contains the greatest number of matches within a predetermined minimum length of time requirement; and outputting a representation of the identified area. 25. The computer program product of any of Examples 22 to 24, comprising instructions for identifying one or more regions in which the highest average value of a single data set occurs for a predetermined minimum length of time. 26. A computer program product according to any of Examples 22 to 25, comprising instructions for calculating a moving average of a single plot. 27. The computer program product of any of Examples 22 to 26, wherein the first or second detection criterion comprises an upper or lower boundary threshold. 28. A computer program product according to any of examples 22 to 27, comprising instructions for applying a length requirement to a filter to remove detected segments outside a desired length range. 29. A computer-implemented method for identifying segments in music having characteristics suitable for eliciting an autonomic nervous system psychological response in a human listener, comprising: receiving digital music data via an input operated by a processor; processing the digital music data using two or more objective audio processing metrics using the processor to generate respective two or more outputs; detecting via the processor a plurality of detected segments in each of the two or more outputs based on areas where respective detection criteria are met; and combining, using the processor, the plurality of detected segments in each of the two or more outputs into a single chill moment plot based on matches in the plurality of detected segments, wherein the first and second objective audio processing metrics are selected from the group consisting of: loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. 30. The method of example 29, comprising: using a processor to identify one or more regions in the single chill moment plot that contain the greatest number of matches within a minimum length requirement; and using a processor to output a representation of the identified one or more regions. 31. The method of any one of examples 29 to 30, comprising displaying via a display device a visual representation of the value of a single chill moment plot relative to the length of the digital music data. 32. The method of any of Examples 29 to 32, comprising displaying, via a display device, a visual representation of the digital music data relative to the length of the digital music data overlaid with a visual representation of a single chill moment plot value relative to the length of the digital music data. 33. The method of Example 32, wherein the visual representation of the single chill moment plot values ​​comprises a curve of a moving average of the single chill moment plot values. 34. The method of any of examples 29 to 33, comprising: identifying a region within a single chill moment plot that contains the greatest number of matches during a predetermined minimum length of time requirement; and outputting a representation of the identified region. 35. The method of example 33, wherein outputting includes displaying a visual representation of the identified region via a display device. 36. The method of example 33, wherein outputting includes displaying, via a display device, a visual representation of the digital music data relating to the length of the digital music data overlaid with a visual representation of the identified region within the digital music data. 37. A computer-implemented method for providing information identifying impactful moments in music, comprising: receiving, via an input operated by a processor, a request for information related to an impactful moment in a digital audio recording, the request including a representation of the digital audio recording; accessing, using the processor, a database storing a plurality of identifications of different digital audio recordings and a corresponding set of information identifying an impactful moment in each of the different digital audio recordings, the corresponding set including at least one of: a start time and a stop time of a chill phrase, or a value of a chill moment plot; matching, using the processor, the received identification of the digital audio recording to one identification of the plurality of identifications in the database, the matching including finding an exact match or a closest match; and outputting, using the processor, a set of information identifying an impactful moment of the matched identification of the plurality of identifications in the database. 38. The method of example 37, wherein the corresponding set of information identifying impactful moments in each of the different digital audio recordings includes information created using a single plot of detection matches for each of the different digital audio recordings generated using the method of example 1 for each of the different digital audio recordings. 39. The method of example 37, wherein the corresponding set of information identifying impactful moments in each of the different digital audio recordings includes information created using a single chill moment plot for each of the different digital audio recordings, generated using the method of example 29 for each of the different digital audio recordings. 40. A computer-implemented method for displaying information identifying impactful moments in music, comprising: receiving, via an input operated by a processor, a representation of a digital audio recording; receiving, via a communications interface operated by the processor, information identifying impactful moments in the digital audio recording, the information including at least one of: a start time and a stop time of a chill phrase, or a value of a chill moment plot; using the processor, displaying the received identification of the digital audio recording to one identification of a plurality of identifications in a database, wherein matching includes finding an exact match or a closest match; and outputting, using a display device, a visual representation of the digital audio recording for a length of time of the digital audio recording overlaid with a visual representation of the chill phrase and / or a value of the chill moment plot for a length of time of the digital audio recording.

[0143] [Embodiment] (1) A computer-implemented method for identifying segments in music, comprising: receiving digital music data via an input operated by a processor; using a processor to process the digital music data using a first objective audio processing metric to generate a first output; using a processor to process the digital music data using a second objective audio processing metric to generate a second output; using a processor to generate a first plurality of detection segments using a first detection routine based on areas in the first output where a first detection criterion is satisfied; using a processor to generate a second plurality of detection segments using a second detection routine based on areas in the second output where a second detection criterion is satisfied; using a processor to combine the first and second plurality of detection segments into a single plot representing coincidences of detection segments in the first and second plurality of detection segments; The computer-implemented method, wherein the first objective audio processing metric and the second objective audio processing metric are different. (2) identifying an area in said single plot that contains the greatest number of matches within a predetermined minimum length of time requirement; and outputting a representation of the identified regions. (3) The method of claim 1, wherein combining comprises calculating a moving average of the single plot. (4) identifying an area in the single plot where the moving average exceeds an upper limit; and and outputting a representation of the identified regions. (5) The method of embodiment 1, wherein one or both of the first objective audio processing metric and the second objective audio processing metric are primary algorithms and / or are configured to output primary data.

[0144] (6) The method of embodiment 1, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. (7) The method of embodiment 1, further comprising applying a low pass envelope to an output of either the first objective audio processing metric or the second objective audio processing metric. (8) The method of embodiment 1, wherein the first detection criterion or the second detection criterion includes an upper or lower boundary threshold. (9) The method of embodiment 1, wherein detecting includes applying a length requirement filter to remove detected segments outside a desired length range. (10) The method of embodiment 1, wherein the combining includes applying respective weights to the first plurality of detections and the second plurality of detections.

[0145] (11) A computer system comprising: an input module configured to receive digital music data; an audio processing module configured to receive the digital music data, perform a first objective audio processing metric on the digital music data, and perform a second objective audio processing metric on the digital music data, the first metric and the second metric generating respective first and second outputs; a detection module configured to receive as input the first output and the second output and generate, for each of the first output and the second output, a set of one or more segments for which a detection criterion is satisfied; a combination module configured to receive as input the one or more segments detected by the detection module and aggregate each segment into a single dataset comprising matches of the detection. (12) The computer system of claim 11, further comprising a phrase identification module configured to receive as input the single dataset of matches from the combination module and identify one or more regions where a highest average value of the single dataset occurs for a predetermined minimum length of time. (13) The computer system of claim 12, wherein the phrase identification module is configured to identify the one or more regions based on where a moving average of the single data set exceeds an upper limit. (14) The computer system of claim 12, wherein the phrase identification module is configured to apply a length requirement filter to remove regions outside a desired length range. (15) The computer system of claim 11, wherein the combination module is configured to calculate a moving average of the single plot.

[0146] (16) The computer system of embodiment 11, wherein one or both of the first objective audio processing metric and the second objective audio processing metric are primary algorithms and / or are configured to output primary data. (17) The computer system of embodiment 11, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. (18) The computer system of embodiment 11, wherein the detection module is configured to apply a low pass envelope to an output of either the first objective audio processing metric or the second objective audio processing metric. (19) The computer system of embodiment 11, wherein the detection criteria include upper or lower boundary thresholds. (20) The computer system of embodiment 11, wherein the detection module is configured to apply a length requirement filter to remove detected segments outside a desired length range.

[0147] (21) The computer system of embodiment 11, wherein the combination module is configured to apply respective weights to the first plurality of detections and the second plurality of detections and then aggregate each detection segment based on the respective weights. (22) A computer program product, comprising a tangible, non-transitory computer usable medium having computer readable program code thereon, the computer readable program code being configured to cause a processor to: Receiving digital music data; processing the digital music data using a first objective audio processing metric to generate a first output; processing the digital music data with a second objective audio processing metric to generate a second output; generating a first plurality of detection segments using a first detection routine based on areas in the first output where a first detection criterion is satisfied; generating a second plurality of detection segments using a second detection routine based on areas in the second output where a second detection criterion is satisfied; combining the first and second plurality of detection segments into a single plot based on coincidences of detection segments in the first and second plurality of detection segments; code configured to instruct The first objective audio processing metric and the second objective audio processing metric are different. (23) The computer program product of embodiment 22, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. (24) The computer program product of claim 22, further comprising instructions for identifying an area within the single plot that contains the greatest number of matches within a predetermined minimum length of time requirement, and outputting a representation of the identified area. (25) The computer program product of claim 22, further comprising instructions for identifying one or more regions in which a highest average value of the single data set occurs for a predetermined minimum length of time.

[0148] (26) The computer program product of claim 22, further comprising instructions for calculating a moving average of the single plot. (27) The computer program product of embodiment 22, wherein the first detection criterion or the second detection criterion includes an upper or lower boundary threshold. (28) The computer program product of embodiment 22, further comprising instructions for applying a length requirement to a filter to remove detected segments outside a desired length range. (29) A computer-implemented method for identifying segments in music having characteristics suitable for eliciting an autonomic nervous system psychological response in a human listener, comprising: receiving digital music data via an input operated by a processor; using a processor to process the digital music data using two or more objective audio processing metrics to generate two or more respective outputs; detecting, via a processor, a plurality of detection segments in each of the two or more outputs based on an area where a respective detection criterion is satisfied; and combining, using a processor, the plurality of detection segments in each of the two or more outputs into a single chill moment plot based on matches in the plurality of detection segments; 20. The method of claim 19, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodia, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increases, sustained pitch, harmonic peak ratio, or key change. (30) using a processor to identify one or more regions in the single chill moment plot that contain the greatest number of matches within a minimum length requirement; 30. The method of embodiment 29, comprising using a processor to output a representation of the identified one or more regions.

[0149] (31) The method of embodiment 29, comprising displaying via a display device a visual representation of the value of the single chill moment plot relative to the length of the digital music data. (32) The method of embodiment 29, comprising displaying via a display device a visual representation of the digital music data for a length of the digital music data overlaid with a visual representation of the value of the single chill moment plot for the length of the digital music data. (33) The method of embodiment 32, wherein the visual representation of the values ​​of the single chill moment plot comprises a curve of a moving average of the values ​​of the single chill moment plot. (34) identifying a region within said single chill moment plot that contains the greatest number of matches during a predetermined minimum length of time requirement; and outputting a representation of the identified region. (35) The method of embodiment 33, wherein the outputting includes displaying a visual representation of the identified region via a display device.

[0150] (36) The method of embodiment 33, wherein the outputting includes displaying, via a display device, a visual representation of the digital music data relating to a length of the digital music data overlaid with a visual representation of the identified region within the digital music data. (37) A computer-implemented method for providing information identifying impactful moments in music, comprising: receiving, via an input operated by a processor, a request for information related to the impactful moment in a digital audio recording, the request including a representation of the digital audio recording; using a processor to access a database storing a plurality of identifications of different digital audio recordings and a corresponding set of information identifying impactful moments in each of the different digital audio recordings, the corresponding set including at least one of start and stop times of a chill phrase or values ​​of a chill moment plot; using a processor to match a received identification of the digital audio recording to one of the identifications in the database, said matching including finding an exact match or a closest match; and using a processor to output a set of information identifying moments of impact of matched ones of the plurality of identifications in the database. (38) The method of claim 37, wherein the corresponding set of information identifying impactful moments in each of the different digital audio recordings includes information created using a single plot of detection matches for each of the different digital audio recordings generated using the method of claim 1 for each of the different digital audio recordings. 39. The method of claim 37, wherein the corresponding sets of information identifying impactful moments in each of the different digital audio recordings include information created using a single chill moment plot for each of the different digital audio recordings generated using the method of claim 29 for each of the different digital audio recordings. A single plot. (40) A computer-implemented method for displaying information identifying impactful moments in music, comprising: receiving, via an input operated by a processor, an indication of the digital audio recording; receiving, via a communications interface operated by a processor, information identifying impactful moments in the digital audio recording, the information including at least one of start and stop times of chill phrases or values ​​of a chill moment plot; using a processor to display the received identification of the digital audio recording to one of the identifications in the database, said matching including finding an exact match or a closest match; and outputting, using a display device, a visual representation of the digital audio recording for a length of time of the digital audio recording overlaid with a visual representation of the chill phrases and / or the values ​​of the chill moment plot for the length of time of the digital audio recording.

Claims

1. A computer-implemented method for identifying segments in music, comprising: Receiving digital audio data via an input operated by a processor; Using the processor to process the digital audio data using a first objective audio processing metric to generate a first output, the first output including the value of the first objective audio processing metric in each of a first continuous plurality of series of time segments of the digital audio data; Using the processor to process the digital audio data using a second objective audio processing metric to generate a second output, the second output including the value of the second objective audio processing metric in each of a second continuous plurality of series of time segments; Using the processor to generate a first plurality of detection segments using a first detection routine based on a region in the first output where a first detection criterion is met, the first detection criterion defining a first threshold for the first output, each of the first plurality of detection segments defining a display of the respective value of the first objective audio processing metric that meets the first detection criterion in each of the respective time segments of the first plurality of time segments; Using the processor to generate a second plurality of detection segments using a second detection routine based on a region in the second output where a second detection criterion is met, the second detection criterion defining a second threshold for the second output, each of the second plurality of detection segments defining a display of the respective value of the second objective audio processing metric that meets the second detection criterion in each of the respective time segments of the second plurality of time segments; Using the processor to combine the first plurality of detection segments and the second plurality of detection segments into a single plot representing the coincidence of detection segments in the first plurality of detection segments and the second plurality of detection segments; The computer-implemented method, wherein the first objective audio processing metric and the second objective audio processing metric are different.

2. identifying an area in the single plot that contains the most number of matches during a time requirement of a predetermined minimum length; outputting a display of the identified area, the method according to claim 1. **Claim 3** The method according to claim 1, wherein combining includes calculating a moving average of the single plot. **Claim 4** identifying an area in the single plot where the moving average exceeds an upper limit; outputting a display of the identified area, the method according to claim 3. **Claim 5** The method according to claim 1, wherein one or both of the first objective audio processing metric and the second objective audio processing metric is a primary algorithm and / or is configured to output primary data. **Claim 6** The method according to claim 1, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodica, spectral flux, spectral centroid, inharmonicity, dissonance, sudden increase in dynamics, sustained pitch, harmonic peak ratio, or key change. **Claim 7** The method according to claim 1, further comprising applying a low-pass envelope to the output of either the first objective audio processing metric or the second objective audio processing metric. **Claim 8** The method according to claim 1, wherein the first detection criterion or the second detection criterion includes an upper or lower boundary threshold. **Claim 9** The method according to claim 1, wherein detecting includes applying a length requirement filter to remove detection segments outside a desired length range. **Claim 10** The method according to claim 1, wherein combining includes applying respective weights to a first plurality of detections and a second plurality of detections. **Claim 11** A computer system, an input module configured to receive digital audio data; An audio processing module configured to receive the digital audio data, execute a first objective audio processing metric on the digital audio data, and execute a second objective audio processing metric on the digital audio data, wherein the first objective audio processing metric and the second objective audio processing metric each generate a first output and a second output, and each of the first output and the second output includes a value of the respective objective audio processing metric in each of the respective consecutive plurality of series of time segments of the digital audio data. A detection module configured to receive the first output and the second output as inputs and generate a set of one or more detection segments for which respective detection criteria are met for each of the first output and the second output, wherein each of the respective detection criteria defines a respective threshold for the respective output, and each of the respective one or more segments defines a display of the respective value of the respective objective audio processing metric that meets the respective detection criteria in each of the respective time segments of the respective consecutive plurality of series of time segments. A computer system including a combination module configured to receive the one or more segments detected by the detection module as inputs and aggregate each segment into a single data set including the detection segment match.

12. The computer system according to claim 11, further comprising a phrase identification module configured to receive the single data set of matches from the combination module as an input and identify one or more regions in which the highest average value of the single data set occurs during a time of a predetermined minimum length.

13. The computer system according to claim 12, wherein the phrase identification module is configured to identify the one or more regions based on a location where a moving average of the single data set exceeds an upper limit.

14. The computer system according to claim 12, wherein the phrase identification module is configured to apply a length requirement filter to remove regions outside a desired length range.

15. The computer system according to claim 11, wherein the combination module is configured to calculate a moving average of the single plot.

16. The computer system according to claim 11, wherein one or both of the first objective audio processing metric and the second objective audio processing metric are primary algorithms and / or are configured to output primary data.

17. The computer system according to claim 11, wherein the first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodica, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increase, sustained pitch, harmonic peak ratio, or key change.

18. The computer system according to claim 11, wherein the detection module is configured to apply a low-pass envelope to an output of either the first objective audio processing metric or the second objective audio processing metric.

19. The computer system according to claim 11, wherein the detection criteria include an upper or lower boundary threshold.

20. The computer system according to claim 11, wherein the detection module is configured to apply a length requirement filter to remove detection segments outside a desired length range.

21. The computer system according to claim 11, wherein the combination module is configured to apply respective weights to the first detection and the second detection and then aggregate each detection segment based on the respective weights.

22. A computer program product, comprising a tangible non-transitory computer-usable medium having computer-readable program code thereon, the computer-readable program code causing a processor to, receive digital audio data, Processing the digital audio data using a first objective audio processing metric to generate a first output, wherein the first output includes values of the first objective audio processing metric in respective ones of a first plurality of consecutive series of time segments of the digital audio data, Processing the digital audio data using a second objective audio processing metric to generate a second output, wherein the second output includes values of the second objective audio processing metric in respective ones of a second plurality of consecutive series of time segments, Generating a first plurality of detection segments using a first detection routine based on regions in the first output that meet a first detection criterion, wherein the first detection criterion defines a first threshold for the first output, and each of the first plurality of detection segments defines a display of the respective value of the first objective audio processing metric that meets the first detection criterion in each of the respective time segments of the first plurality of consecutive series of time segments, Generating a second plurality of detection segments using a second detection routine based on regions in the second output that meet a second detection criterion, wherein the second detection criterion defines a second threshold for the second output, and each of the second plurality of detection segments defines a display of the respective value of the second objective audio processing metric that meets the second detection criterion in each of the respective time segments of the second plurality of consecutive series of time segments, Combining the first plurality of detection segments and the second plurality of detection segments into a single plot based on a match of detection segments in the first plurality of detection segments and the second plurality of detection segments, including code configured to instruct, A computer program product, wherein the first objective audio processing metric and the second objective audio processing metric are different.

23. The first objective audio processing metric and the second objective audio processing metric are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melody, spectral flux, spectral centroid, inharmonicity, dissonance, sudden increase in dynamics, sustained pitch, harmonic peak ratio, or key change, the computer program product according to claim 22.

24. Identify an area within the single plot that contains the most number of matches during a time requirement of a predetermined minimum length, and output a display of the identified area, the computer program product according to claim 22, comprising instructions.

25. The computer program product according to claim 22, comprising instructions for identifying one or more areas where the highest average value of the single dataset occurs during a time of a predetermined minimum length.

26. The computer program product according to claim 22, comprising instructions for calculating a moving average of the single plot.

27. The first detection criterion or the second detection criterion includes an upper or lower boundary threshold, the computer program product according to claim 22.

28. The computer program product according to claim 22, comprising instructions for applying a length requirement to a filter to remove detection segments outside a desired length range.

29. A computer-implemented method for identifying segments in music that have characteristics suitable for causing a psychological reaction of the autonomic nervous system in a human listener, Receiving digital audio data via an input operated by a processor; Using the processor to process the digital audio data using two or more objective audio processing metrics to generate respective outputs, each of the two or more outputs including respective values of the two or more objective audio processing metrics in respective continuous plural series of time segments of the digital audio data, Detecting, via a processor, a plurality of detection segments in each of the two or more outputs based on regions where respective detection criteria are met, wherein each of the respective detection criteria defines a respective threshold for each of the two or more outputs, and each of the respective plurality of detection segments defines a display of the value of the respective objective audio processing metric that meets the respective detection criteria in each of the respective time segments of the respective continuous plurality of sets of time segments, Using a processor to combine the plurality of detection segments in each of the two or more outputs into a single chirp moment plot based on a match in the plurality of detection segments, The two or more objective audio processing metrics are selected from the group consisting of loudness, loudness band ratio, critical band loudness, dominant pitch melodic, spectral flux, spectral centroid, inharmonicity, dissonance, sudden dynamic increase, sustained pitch, harmonic peak ratio, or key change, a computer-implemented method.

30. Using a processor to identify one or more regions in the single chirp moment plot that contain the most number of matches during a minimum length requirement, Using a processor to output a display of the identified one or more regions, the method according to claim 29.