Audio track analysis techniques to support audio personalization
The method addresses the challenge of inconsistent audio personalization across categories by identifying and adjusting settings based on user input, enhancing the listening experience through efficient selection and configuration of representative audio tracks.
Patent Information
- Application Number
- JP2021088172
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-01
- Filing Date
- 2021-05-26
- Publication Date
- 2025-10-02
- Estimated Expiration
- 2041-05-26
AI Technical Summary
Users face difficulties in efficiently and consistently personalizing audio settings across different audio categories due to varying audio properties, leading to a tedious and suboptimal listening experience.
A computer-implemented method for determining audio personalization settings by identifying audio properties, selecting a representative portion of an audio track, playing it to the user, and adjusting settings based on user input, with the option to suggest alternative tracks.
Enables efficient and effective selection of representative audio tracks for personalized settings, improving the listening experience by allowing users to achieve a balanced audio category-specific configuration.
Smart Images

Figure 0007748203000001 
Figure 0007748203000002 
Figure 0007748203000003
Abstract
Description
[Technical Field]
[0001] FIELD Embodiments of the present disclosure relate generally to audio devices, and more particularly to audio track analysis to support audio personalization. [Background technology]
[0002] Personal entertainment devices may include mobile applications and computer software that allow users to personalize electronic media and audio content. To enhance the user experience while listening to audio content, such as music, videos, video games, and / or online advertisements, such applications may allow users to, for example, select and listen to preferred content or adjust settings. Such applications may also allow users to digitally manipulate audio content to enhance or clarify particular audio qualities.
[0003] However, to achieve a desired audio experience with given audio content, users typically manually adjust various applications and / or settings, which can be tedious, time-consuming, and / or troublesome. For example, to better hear the nuances and artifacts in an audio track and / or achieve other desired goals, users may need to increase or decrease bass or treble levels, adjust frequency band filters, and / or apply compression or equalization based on personal preference.
[0004] Furthermore, personalizing application settings can be difficult for users when switching between different categories of audio content. In particular, audio properties may vary depending on the audio category. For example, audio personalization settings specific to a first audio category (e.g., rock and roll) may be inappropriate for audio content in a second audio category (e.g., classical). As a result, when audio personalization settings for a first audio category are applied to audio content in the second audio category, the audio personalization settings may not be well-suited to the audio content in the second audio category, thereby degrading the listening experience for the audio content in the second audio category. Thus, users may adjust their audio personalization settings each time they switch between categories. This often makes it difficult to consistently achieve a desired listening experience, especially when streaming audio content. Some of these personalization issues can be addressed by storing user personalization settings for each audio category. The user's personalization settings can then be loaded and applied each time an audio track from the corresponding audio category is played to the user.
[0005] However, selecting an audio sample representative of a particular audio category and thereby initially configuring personalization settings for that particular audio category can be difficult. For example, a user may be familiar with a significant amount of audio content within a particular audio category, but may not be able to easily determine a particular audio track to select as a representative sample for creating their own personalization settings. Furthermore, because audio properties typically vary within a single audio content, even if a particular audio track is representative of a particular audio category, not all parts of the particular audio track may be suitable for configuring personalization settings for the particular audio category.
[0006] As a result, users typically go through a tedious, time-consuming, and error-prone personalization process that is likely to result in selecting poor quality representative samples, and configuring personalization settings with the selected representative samples often results in suboptimal personalization settings and a poor listening experience for a large amount of audio content in each audio category.
[0007] Therefore, there is a need for techniques that allow users to better select audio samples to use when configuring personalization settings for various categories of audio content. Summary of the Invention [Means for solving the problem]
[0008] Various embodiments disclose a computer-implemented method for determining audio personalization settings for an audio category, including identifying one or more audio properties of an audio track, selecting a first portion of the audio track representative of the audio category based on the one or more audio properties, playing the first portion of the audio track to a user, and adjusting the user's personalization settings based on user input during playback of the first portion of the audio track.
[0009] Further embodiments provide, among other things, systems and one or more computer-readable storage media configured to implement the above methods.
[0010] At least one technical advantage of the disclosed technology over the prior art is that the disclosed technology enables improved audio personalization by allowing users to more efficiently and effectively select representative audio tracks and representative audio samples of representative audio tracks that contain an appropriate balance of audio characteristics that allows the user to achieve a personalized personalization setting for a particular audio category. Based on the user's selection, the disclosed technology may suggest alternative representative audio tracks to use when creating a personalized setting for a particular audio category. Furthermore, the disclosed technology provides users with a faster and more computationally efficient means for generating portions of audio tracks that contain category-specific balances of audio characteristics that can be used to configure a personalized setting.
[0011] In order to enable the above-listed features of the various embodiments to be understood in detail, a more particular description of the inventive concepts briefly summarized above can be rendered by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical embodiments of the inventive concepts and are therefore not to be construed as limiting the scope in any manner, as there may be other equally effective embodiments. For example, the present application provides the following: (Item 1) 1. A computer-implemented method for determining audio personalization settings for an audio category, comprising: Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of the audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's personalization settings based on the user's input during playback of the first portion of the audio track; The computer-implemented method as described above, (Item 2) creating an audio sample including multiple repetitions of the first portion of the audio track; playing the first portion of the audio track further comprises playing the audio sample. The computer-implemented method of the preceding paragraph. (Item 3) 2. The computer-implemented method of claim 1, wherein creating the audio sample includes shortening or lengthening the duration of the first portion of the audio track such that no tempo discontinuity occurs between the repetitions of the first portion of the audio track in the audio sample. (Item 4) The computer-implemented method of any one of the preceding items, further comprising, before selecting the first portion of the audio track, determining whether the audio track is representative of the audio category based on the one or more audio properties. (Item 5) 20. The computer-implemented method of claim 19, further comprising: suggesting a second audio track representative of the audio category based on the determination. (Item 6) 2. The computer-implemented method of claim 1, wherein the one or more audio properties include at least one of bass level, treble level, frequency spectrum, energy, or tempo. (Item 7) 2. The computer-implemented method of claim 1, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with corresponding audio metrics associated with the audio category. (Item 8) The computer-implemented method of any one of the preceding items, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference. (Item 9) 2. The computer-implemented method of claim 1, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with a range of corresponding audio metrics associated with the audio category. (Item 10) 20. The computer-implemented method of claim 19, further comprising identifying the audio category of the audio track based on metadata associated with the audio track or a user selection. (Item 11) A system comprising a memory and a processor, the memory storing one or more software applications; When the processor executes the one or more software applications, Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's personalization settings based on the user's input during playback of the first portion of the audio track; The system is configured to perform the steps of: (Item 12) The system described in the above item, wherein the processor is further configured to perform the step of determining whether the audio track is representative of the audio category based on the one or more audio properties before selecting the first portion of the audio track. (Item 13) 10. The system of claim 9, wherein the processor is further configured to perform the step of suggesting a second audio track representative of the audio category based on the determination. (Item 14) The system of any one of the preceding items, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with corresponding audio metrics associated with the audio category. (Item 15) The system of any one of the preceding items, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference. (Item 16) The system of any one of the preceding items, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with a range of corresponding audio metrics associated with the audio category. (Item 17) One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's personalization settings based on the user's input during playback of the first portion of the audio track; The one or more non-transitory computer-readable media described above, causing the steps of: (Item 18) The one or more non-transitory computer-readable media described in the preceding item further includes determining whether the audio track is representative of the audio category based on the one or more audio properties before selecting the first portion of the audio track. (Item 19) Associating said personalization settings with said audio categories; and Save the personalization settings, and One or more non-transitory computer-readable media according to any one of the preceding items, further comprising: (Item 20) receiving a selection of a second audio track to play; identifying a second audio category of the second audio track; loading a second personalization setting associated with the second audio category; and generating a customized audio signal by modifying audio of the second audio track according to the second personalization settings; playing the customized audio signal to the user; and One or more non-transitory computer-readable media according to any one of the preceding items, further comprising: (Summary) Various embodiments demonstrate systems and techniques for enabling audio personalization, including identifying audio personalization settings for an audio category, identifying one or more audio properties of an audio track, selecting a first portion of the audio track representative of the audio category based on the one or more audio properties, playing the first portion of the audio track to a user, and adjusting the user's personalization settings based on user input during playback of the first portion of the audio track. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram illustrating an audio personalization system configured to implement one or more aspects of the present disclosure. [Figure 2] FIG. 1 is a conceptual block diagram of a computing system configured to implement one or more aspects of various embodiments of the present disclosure. [Figure 3] 1 is a flowchart of method steps for customizing personalization settings for audio categories, according to various embodiments of the present disclosure. [Figure 4] 1 is a flowchart of method steps for applying audio personalization settings to playback of audio tracks, according to various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] For clarity, where applicable, the same reference numbers have been used to refer to identical elements common to the figures. It is contemplated that features of one embodiment may be incorporated into other embodiments without further recitation.
[0014] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the concepts of the present invention may be practiced without one or more of these specific details.
[0015] 1 is a schematic diagram illustrating an audio personalization system 100 configured to implement one or more aspects of the present disclosure. Audio personalization system 100 includes, but is not limited to, one or more audio environments 110, a user profile database 120, an audio profile database 130, and a computing device 140. Audio personalization system 100 is configured to enable a user to achieve user-preferred personalization settings for corresponding audio categories by enabling the user to more efficiently and effectively select representative audio tracks and representative audio samples of the representative audio tracks. In some embodiments, audio personalization system 100 is configured to enable a user to customize personalization settings for multiple audio categories.
[0016] In some embodiments, audio content for an audio experience is stored locally on computing device 140, while in other embodiments, such audio content is provided by a streaming service 104 implemented on a cloud-based infrastructure 105. Audio content may include music, videos, movies, video games, online advertisements, audiobooks, sounds (ringtones, animal sounds, synthesized sounds), podcasts, sporting events, or any other content that can be acoustically heard or recorded.
[0017] Cloud-based infrastructure 105 may be any technically feasible internet-based computing system, such as a distributed computing system and / or a cloud-based storage system. In some embodiments, cloud-based infrastructure 105 includes, but is not limited to, multiple networks, multiple servers, multiple operating systems, and / or multiple storage devices. The servers may be standalone servers, clusters or "farms" of servers, one or more network appliances, or any other devices suitable for implementing one or more aspects of the present disclosure.
[0018] Each of the one or more audio environments 110 is configured to play audio content for a particular user. For example, the audio environments 110 may include, but are not limited to, one or more smart devices 111, headphones 112, smart speakers 113, and / or other input / output (I / O) devices 119.
[0019] 1 , audio environment 110 plays audio content received from computing device 140 via any technically feasible combination of wireless or wired point-to-point or networked communication links. Networked communication links include any suitable communication links that enable communication between remote or local computer systems and computing devices, such as, but not limited to, Bluetooth® communication channels, wireless and wired LANs (local area networks), Internet-based WANs (wide area networks), and / or cellular networks. As a result, audio environment 110 may include any audio device capable of receiving audio content directly from computing device 140, such as a pair of “dumb” speakers in a home, a vehicle stereo system, and / or a traditional pair of headphones. Furthermore, in the embodiment shown in FIG. 1 , audio environment 110 does not rely on the ability to perform audio signal processing internally or to receive audio content or other information from entities embodied in cloud-based infrastructure 105.
[0020] The smart device 111 may include, but is not limited to, a computing device, which may be a personal computer, a personal digital assistant, a tablet computer, a mobile phone, a smartphone, a media player, a mobile device, or any other device suitable for implementing one or more aspects of the present invention. The smart device 111 may enhance the functionality of the audio personalization system 100 by providing various services, including, but not limited to, telephone services, navigation services, and / or infotainment services. Additionally, the smart device 111 may acquire data from sensors and transmit the data to the audio personalization system 100. The smart device 111 may acquire audio data via an audio input device and transmit the audio data to the audio personalization system 100 for processing. Similarly, the smart device 111 may receive audio data from the audio personalization system 100 and transmit the audio data to an audio output device so that the user can hear audio originating from the audio personalization system 100.
[0021] Headphones 112 may include an audio output device capable of generating sound based on one or more audio signals received from audio personalization system 100 and / or an alternative audio device, such as a power amplifier, associated with audio personalization system 100. More specifically, the audio output device may convert one or more electrical signals into sound waves and transmit the sound waves into the physical environment.
[0022] The smart speaker 113 may include an audio input device that may acquire acoustic data, such as the user's voice, from the surrounding environment and transmit a signal associated with the acoustic data to the audio personalization system 100.
[0023] Each of headphones 112 and smart speaker 113 includes one or more speakers 117 and, in some embodiments, one or more sensors 118. Speaker(s) 117 are audio output devices configured to generate audio output based on customized audio signals received from computing device 140. Sensor(s) 118 are configured to acquire biometric data from a user (e.g., heart rate and / or skin conductance, etc.) and transmit signals associated with the biometric data to computing device 140. The biometric data acquired by sensor(s) 118 may then be processed by personalization application 145 executing on computing device 140 to determine one or more personal audio preferences for a particular user. In various embodiments, sensor(s) 118 may include any type of image sensor, electrical sensor, and / or biometric sensor capable of acquiring biometric data, including, but not limited to, a camera, an electrode, and / or a microphone.
[0024] Other I / O devices 119 include, but are not limited to, input devices, output devices, and devices capable of both receiving input data and generating output data. Other I / O devices 119 may include, but are not limited to, wired and / or wireless communication devices that transmit data to and / or receive data from smart device 111, headphones 112, smart speaker 113, speaker 117, sensor(s) 118, remote databases, and / or other computing devices. Additionally, in some embodiments, other I / O devices 119 may include push-to-talk (PTT) buttons, such as those included in vehicles, mobile devices, and / or smart speakers.
[0025] User profile database 120 stores user-specific information that enables a personalized audio experience to be created for a particular user in any of audio environments 110. As shown, user profile database 120 may be implemented in cloud-based infrastructure 105, such that computing device 140 can access user profile database 120 whenever it has access to a networked communications link. In some embodiments, information associated with a particular user and stored in user profile database 120 is also stored locally on computing device 140 associated with that particular user. In such embodiments, user preference profile(s) 121 and / or personalization setting(s) 122 are stored in local user profile database 143 on computing device 140. The user-specific information stored in user profile database 120 may include one or more of user preference profile(s) 121 and personalization setting(s) 122.
[0026] User preference profile(s) 121 may include user-specific information used to create a personalized audio experience for a particular user. In some embodiments, user preference profile(s) 121 include acoustic filters and / or EQ curves associated with a particular user. In some embodiments, user preference profile(s) 121 include other user-preferred signal processing, such as dynamic range compression, dynamic expansion, audio limiting, and / or spatial processing of the audio signal. In some embodiments, user preference profile(s) 121 may include a preset EQ curve selected by the user while configuring their preferred listening settings. The EQ curve may include one or more individual amplitude adjustments made by the user while configuring their preferred listening settings. The preset EQ curve may be associated with another user, such as a famous musician or celebrity. In some embodiments, the EQ curve may include head-related transfer function (HRTF) information specific to a particular user.
[0027] The personalization setting(s) 122 may include information used to create a personalized audio experience for a particular user during playback of the corresponding audio category. In some embodiments, each personalization setting 122 may be generated based on settings made by a user during playback of an audio track having one or more audio properties representative of a particular audio category. In some embodiments, each personalization setting 122 may be determined from user input received during playback of a portion of an audio track, the portion of the audio track having one or more audio properties representative of a particular audio category.
[0028] In some embodiments, each particular audio category may include any classification of musical or non-musical audio content. For example, an audio category may include a musical genre (such as classical, country, hip hop, and / or rock). An audio category may also include any classification of videos, movies, video games, online advertisements, audiobooks, audio (ringtones, animal sounds, synthesized sounds), podcasts, sporting events, or any other content that can be acoustically heard or recorded. In some embodiments, each particular audio category may include any classification based on a combination of attributes such as rhythm, harmony, instruments, tonality, and / or tempo.
[0029] In some embodiments, audio content selected by a particular user and played in one of the audio environments 110 is modified to conform to that user's personal listening preferences when playing audio tracks of the corresponding audio category. Alternatively, or in addition, in some embodiments, the personalization setting(s) 122 include other user-preferred and category-specific signal processing to apply during playback of the corresponding audio category, such as category-specific dynamic range compression, category-specific dynamic expansion, category-specific audio limiting, and / or category-specific spatial processing of the audio signal. In some embodiments, such category-specific signal processing may also be used by the audio processing application 146 to modify the audio content when the user plays the audio content in one of the audio environments 110.
[0030] Computing device 140 may be any computing device that can be configured to implement at least one aspect of the present disclosure described herein, including a smartphone, an electronic tablet, a laptop computer, a personal computer, a personal digital assistant, a mobile device, or any other device suitable for implementing one or more aspects of the present disclosure. Generally, computing device 140 may be any type of device that can execute application programs, including, but not limited to, instructions associated with personalization application 145 and / or audio processing application 146. In some embodiments, computing device 140 is further configured to store local user profile database 143, which may include one or more user preference profile(s) 121 and / or personalization setting(s) 122. In some embodiments, computing device 140 is further configured to store audio content 144, such as digital recordings of audio content.
[0031] Personalization application 145 is configured to facilitate communication between computing device 140 and user profile database 120, audio profile database 130, and audio environment 110. In some embodiments, personalization application 145 is also configured to present a user interface (not shown) to the user that enables user audio preference tests and / or setting operations, etc., during playback of audio tracks of the corresponding audio category. In some embodiments, personalization application 145 is further configured to generate a customized audio personalization procedure for the audio signal based on the user-specific audio processing information and the category-specific audio processing information.
[0032] The audio processing application 146 may dynamically generate a customized audio signal by processing the initial audio signal with a customized audio personalization procedure generated by the personalization application 145. For example, the audio processing application 146 may generate a customized audio signal by modifying the initial audio signal based on one or more applicable user personalization settings 122 associated with the playback of a particular audio category.
[0033] The audio profile database 130 stores one or more audio metrics 131 for each of a plurality of categories of audio content. Each of the audio metrics 131 associated with a particular audio category is representative of audio samples included in the particular audio category. These one or more audio metrics 131 can be used by the personalization application 145 to assist in selecting representative audio tracks and / or representative audio samples to use when setting the personalization settings 122 for the corresponding audio category. As shown, the audio profile database 130 can be implemented in the cloud-based infrastructure 105, such that the computing device 140 can access the audio profile database 130 whenever the computing device 140 has access to a networked communications link. The audio profile database 130 can store information such as the audio metrics 131.
[0034] In some embodiments, audio metrics 131 may be generated based on an analysis of audio content representative of each of the audio categories. In some embodiments, audio metrics 131 may include data associated with one or more audio properties, such as dynamic properties, bass or treble level, frequency spectrum, energy, and / or tempo.
[0035] In some embodiments, the audio samples used to determine audio metrics 131 for each of the audio categories may be selected from a curated collection of audio samples of pre-labeled and / or categorized audio categories. In some embodiments, the one or more audio categories may be determined using an algorithm that identifies one or more boundaries between various audio properties of the audio samples, which are consistent with the pre-labeling or classification of the audio samples. In some embodiments, the one or more boundaries may be identified using clustering techniques (e.g., k-means cluster analysis), machine learning techniques, etc.
[0036] In some embodiments, audio metrics 131 are stored separately for each audio category. In some embodiments, audio metrics 131 may be generated based on statistical modeling, data mining, and / or other algorithmic analysis of aggregate audio content. In some embodiments, audio metrics 131 may include one or more statistical properties, such as the mean, standard deviation, range of values, and / or median, of one or more audio properties of the audio content of each audio category. As a non-limiting example, audio metrics 131 may include the mean and standard deviation of spectral energy in each of a set of predefined frequency bands, which indicate the typical amount of spectral energy in each of the predefined frequency bands for each audio category. As another non-limiting example, audio metrics 131 may include the mean and standard deviation of the temporal separation between successive tempo pulse signals, energy fluxes, energy spikes, and / or downbeat positions, etc. In some embodiments, audio metrics 131 may include the mean and standard deviation of the frequencies of tempo pulse signals, energy fluxes, energy spikes, and / or downbeat positions, etc. In some embodiments, audio metrics 131 may include the mean and standard deviation of the number of tempo pulse signals, energy fluxes, energy spikes, and / or downbeat positions, etc., over a given period of time.
[0037] In some embodiments, audio metrics 131 may include a tolerance window associated with each audio category. The tolerance window may be a predetermined range of expected values for one or more audio properties of the audio content of the corresponding audio category. In some embodiments, the tolerance window may include a limit of deviation for one or more audio properties.
[0038] In some embodiments, the audio metric may include a relative or absolute weight or score assigned to each of the audio properties in the calculation of a composite or aggregate audio metric that is associated with the degree of fit of the audio sample to the corresponding audio category. In some embodiments, the aggregate audio metric may be associated with a balance of audio properties that can be used to configure preferred personalization settings for the corresponding audio category.
[0039] In some embodiments, audio metrics 131 may be used by personalization application 145 to assist a user in selecting representative audio tracks and representative audio samples for use in customizing audio category personalization settings 122. In some embodiments, a user may select entire audio tracks, portions of audio tracks, or aggregations of one or more portions of one or more audio tracks as potential candidate audio tracks to use when configuring the user's personalization settings 122. In some embodiments, personalization application 145 compares audio properties of an audio track to audio metrics 131 of audio categories associated with the selected audio track. In some embodiments, the audio category of a selected audio track may be identified from classification data and / or other metadata (e.g., genre, subgenre, artist, and / or title) associated with the selected audio track and / or from a user's identification of an audio category. In some embodiments, personalization application 145 may perform a real-time search of classification data and / or other metadata against one or more online databases to identify relevant audio categories. In some embodiments, the personalization application 145 may identify one or more instruments in an audio track and perform one or more audio pattern matching techniques to identify corresponding audio categories.
[0040] In some embodiments, the personalization application 145 identifies one or more audio properties of the selected audio track, such as dynamic properties, bass or treble levels, frequency spectrum, energy, and / or tempo. In some embodiments, the energy of the audio track includes the amplitude (dB level) of various frequency subbands. In some embodiments, the frequency range of the audio track may be divided into frequency subbands. In some embodiments, the subbands are associated with predetermined frequency ranges. In some embodiments, subband coefficients corresponding to the spectral energy in each of the subbands may be identified using time-frequency domain transform techniques, such as a modified discrete cosine transform (MDCT), a fast Fourier transform (FFT), a quadrature mirror filter bank (QMF), and / or a conjugate quadrature mirror filter bank (CQMF).
[0041] In some embodiments, tempo may be determined using bar detection techniques, such as correlating energy flux with impulse signals and / or finding repeated energy spikes, downbeat locations, etc. In some embodiments, tempo may be determined by the average duration between energy spikes and / or downbeat locations, etc. In some embodiments, tempo may be determined by the average frequency of energy spikes and / or downbeat locations, etc. In some embodiments, tempo may be determined by the number of counts of energy spikes and / or downbeat locations, etc. occurring during a predetermined period of time. In some embodiments, personalization application 145 determines energy flux using techniques such as short-time Fourier transform (STFT).
[0042] In some embodiments, the personalization application 145 determines whether a selected audio track is representative of a corresponding audio category by comparing the audio properties of the selected audio track with one or more audio metrics 131 associated with the corresponding audio category. In some embodiments, the personalization application 145 compares the audio properties of the audio track with a combination of one or more statistical properties and / or tolerance windows associated with the corresponding audio category.
[0043] In some embodiments, the personalization application 145 determines whether all or a predetermined percentage (e.g., 90 percent, 80 percent, and / or 75 percent) of the audio properties of the selected audio tracks are within a corresponding range for each audio property in the audio metrics 131. In some embodiments, the range is determined based on a predetermined number of standard deviations from the corresponding mean for each audio metric 131, a tolerance window for each audio metric 131, or the like.
[0044] In some embodiments, the personalization application 145 determines whether the aggregate difference between an audio property and a corresponding audio metric 131 for the corresponding audio category is below a threshold difference. In some embodiments, the difference between an audio property and a corresponding audio metric 131 is based on how much the audio property differs from the mean of the corresponding audio metric 131. In some embodiments, the difference is measured by determining a z-score that indicates the number of standard deviations of the audio property from the mean of the corresponding audio metric. In some embodiments, the differences between the audio property and the corresponding audio metric 131 may be aggregated using a distance function (e.g., Euclidean distance) and / or a weighted sum, etc. In some embodiments, the weights used in the weighted sum may correspond to a weight or score assigned to each audio property, which indicates the importance of the audio property compared to other audio properties in determining the personalization settings associated with the corresponding category.
[0045] In some embodiments, if the personalization application 145 determines that one or more audio properties do not satisfy one or more audio metrics, the personalization application 145 may suggest an alternative audio track. In some embodiments, the personalization application 145 selects an audio track from one or more of the audio samples in the curated library of audio samples used for the audio metrics 131, audio content played via the streaming service 104, audio content 144, web-based programs, programs stored locally on the computing device 140, and / or a playlist, etc. In some embodiments, the personalization application 145 suggests audio samples that have audio properties similar to the audio properties of the corresponding audio category.
[0046] In some embodiments, the personalization application 145 may dynamically generate suggestions for alternative audio tracks for the corresponding audio category. In some embodiments, the personalization application 145 may suggest audio tracks representative of the corresponding audio category based on an analysis of one or more audio samples in the curated library of audio samples used for the audio metrics 131. In some embodiments, the personalization application 145 dynamically generates suggestions for alternative audio tracks by analyzing multiple audio tracks having audio properties similar to those of the corresponding audio category. In some embodiments, the personalization application 145 uses a pre-configured algorithm to automatically select an alternative representative track based on a dynamic analysis of one or more audio properties of one or more audio samples for one or more audio metrics 131 of the corresponding audio category. In some embodiments, the personalization application 145 may suggest an alternative audio track based on historical data of representative track selections by users in the associated audio category, data regarding representative audio tracks for the audio category, and / or demographic data indicative of one or more representative tracks selected by similar users, etc.
[0047] In some embodiments, the personalization application 145 compares audio properties of one or more portions of the audio track with one or more audio metrics 131 to identify portions of the audio track that are representative of a corresponding audio category. In some embodiments, the personalization application 145 divides the selected audio track into one or more frames. In some embodiments, the personalization application 145 compares audio properties of one or more portions of the audio track with a combination of one or more statistical properties and / or tolerance windows associated with the corresponding audio category. In some embodiments, the personalization application 145 identifies portions of the audio track that are most representative of the corresponding audio category using techniques similar to those described above with respect to determining whether the selected audio track is representative of the corresponding audio category.
[0048] In some embodiments, the personalization application 145 creates an audio sample based on a portion of the audio track. In some embodiments, the audio sample may include a predefined length of audio content generated from a portion of the audio track. For example, the audio sample may be a 15-25 second sample selected from a portion of the audio track. In some embodiments, the personalization application 145 pre-selects the audio sample from a portion of the audio track or creates the audio sample based on user input. In some embodiments, the audio sample is a repeating loop generated from a portion of the audio track. In some embodiments, the audio sample includes multiple repetitions of a portion of the audio track.
[0049] In some embodiments, the personalization application 145 creates an audio sample by seamlessly editing together repetitions of a portion of an audio track into an audio sample. In some embodiments, the personalization application shortens or lengthens the length of the portion of the audio track so that there is no tempo discontinuity between the end of a first repetition of the portion of the audio track and the beginning of a second repetition of the audio track. In some embodiments, the shortening or lengthening is selected so that the duration between the last tempo pulse signal, energy spike, and / or downbeat position, etc. in the first repetition and the first tempo pulse signal, energy spike, and / or downbeat position in the second repetition matches the overall tempo of the portion of the audio track. In some embodiments, similar techniques may be used when combining multiple portions of an audio track together to create an audio sample.
[0050] In some embodiments, the personalization application 145 continuously plays one or more particular passages of the audio sample based on a dynamic analysis of one or more audio properties of the audio sample. In some embodiments, the playback of the audio sample is based on comparing the audio properties of the audio sample with one or more audio metrics 131 associated with a corresponding audio category. In some embodiments, the playback of the audio sample redirects the user's focus to one or more particular passages of the audio sample that have a minimum aggregate difference from the one or more audio metrics 131 of the corresponding audio category.
[0051] In some embodiments, the personalization application 145 may then adjust one or more personalization settings of the user based on the user input as the audio sample plays. In some embodiments, the user may increase or decrease bass or treble levels, adjust frequency band filters, apply compression or equalization, perform discrete amplitude adjustments, select or modify pre-set acoustic filters, and / or select preferred signal processing for an audio category (e.g., dynamic range compression, dynamic expansion, audio limiting, spatial processing of the audio signal, etc.). In some embodiments, the user may select a previous personalization setting for the relevant audio category as a starting point and update the personalization setting during playback of the audio sample.
[0052] In some embodiments, the personalization application 145 then saves one or more personalization settings for the audio category, which in some embodiments are saved in the personalization settings 122 in the user profile database 120.
[0053] In some embodiments, audio processing application 146 may apply personalization settings to the playback of an audio track. In some embodiments, a user may select an entire audio track, a portion of an audio track, an aggregation of one or more portions of one or more audio tracks, etc. In some embodiments, audio processing application 146 may identify an audio category for an audio track using techniques similar to those described above with respect to personalization application 145. In some embodiments, audio processing application 146 identifies an audio category for a selected audio track from classification data and / or other metadata associated with the selected audio track, and / or from user input, etc.
[0054] In some embodiments, the audio processing application 146 determines whether personalization settings for a particular audio category are available. In some embodiments, if the audio processing application 146 determines that personalization settings for a particular audio category are not available, the audio processing application 146 provides an option to create personalization settings using the personalization application 145. In some embodiments, if the audio processing application 146 determines that personalization settings for an audio category are available, the audio processing application 146 loads the personalization settings for the audio category. In some embodiments, the audio processing application 146 loads the personalization settings for the audio category from saved personalization settings 122 in the user profile database 120. In some embodiments, the audio processing application 146 applies the personalization settings to the playback of the audio track.
[0055] 2 is a conceptual block diagram of a computing device 200 configured to implement one or more aspects of various embodiments. In some embodiments, computing device 200 corresponds to computing device 140. Computing device 200 may be any type of device capable of executing application programs, including, but not limited to, instructions associated with personalization application 145 and / or audio processing application 146. For example, computing device 200 may be, but is not limited to, an electronic tablet, a smartphone, a laptop computer, an infotainment system integrated into a vehicle, and / or a home entertainment system. Alternatively, computing device 200 may be implemented as a standalone chip, such as a microprocessor, or as part of a more comprehensive solution implemented as an application-specific integrated circuit (ASIC), a system-on-chip (SoC), or the like. It should be noted that the computing systems described herein are exemplary, and that any other technically feasible configurations are within the scope of the present invention.
[0056] As shown, computing device 200 includes, but is not limited to, a processor 250, an input / output (I / O) device interface 260 connected to audio environment 110 of FIG. 1 , an interconnect (bus) 240 connecting memory 210, storage 230, and network interface 270. Processor 250 may be any suitable processor implemented as a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), any other type of processing unit, or a combination of different processing units, such as a CPU configured to interface with a digital signal processor (DSP). For example, in some embodiments, processor 250 includes a CPU and a DSP. In general, processor 250 may be any technically feasible hardware unit capable of processing data and / or executing instructions to facilitate the operation of computing device 200 as described herein. Furthermore, in the context of the present disclosure, the computing elements depicted in computing device 200 may correspond to physical computing systems (e.g., systems in a data center) or may be virtual computing instances running in a computing cloud.
[0057] 1. I / O device interface 260 typically includes the logic necessary to interpret addresses corresponding to audio environment 110 generated by processor 250. I / O device interface 260 may also be configured to perform handshaking between processor 250 and audio environment 110 and / or generate interrupts associated with audio environment 110. I / O device interface 260 may be implemented as any technologically feasible CPU, ASIC, FPGA, or any other type of processing unit or device.
[0058] Network interface 270 is a computer hardware component that connects processor 250 to communications network 205. Network interface 270 may be implemented in computing device 200 as a standalone card, processor, or other hardware device. In some embodiments, network interface 270 may be configured with cellular communication capabilities, satellite telephone communication capabilities, wireless WAN communication capabilities, or other types of communication capabilities that enable communication with communications network 205 and other computing devices outside of computing device 200.
[0059] Memory 210 may include a random access memory (RAM) module, a flash memory unit, or any other type of memory unit, or a combination thereof. Processor 250, I / O device interface 260, and network interface 270 are configured to read and write data from memory 210. Memory 210 includes various software programs executable by processor 250 and application data associated with such software programs, including personalization application 145 and / or audio processing application 146.
[0060] Storage 230 may include a non-transitory computer-readable medium, such as a non-volatile storage device. In some embodiments, storage 230 includes a local user profile database 143.
[0061] 3 is a flowchart of method steps for customizing personalization settings for audio categories according to various embodiments of the present disclosure. Although the method steps are described with respect to the system of FIG. 1, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of the various embodiments. In some embodiments, some or all of the method steps of FIG. 3 may be performed by personalization application 145.
[0062] As shown, method 300 begins at step 301, where a user selects an audio track. In some embodiments, the user may select an entire audio track, a portion of an audio track, an aggregation of one or more portions of one or more audio tracks, etc. In some embodiments, the user may select an audio track from audio content played via streaming service 104 or from locally stored audio content 144 of computing device 140. In some embodiments, the user may select an audio track using a web-based program or a program locally stored on computing device 140. In some embodiments, the audio track may be automatically selected based on data obtained from sensor(s) 118 or sensors located on smart device(s) 111. For example, the audio track may be selected based on sensors capturing user utterances related to the selection, user movements and / or gestures associated with selecting the audio track, and / or user interaction with an input device, etc. In some embodiments, the audio track may be selected from a playlist.
[0063] At step 302, audio properties of the audio track are identified. In some embodiments, one or more audio properties of the selected audio track are identified, such as dynamic properties, bass or treble levels, frequency spectrum, energy, and / or tempo. In some embodiments, the frequency range of the audio track may be divided into frequency sub-bands. In some embodiments, sub-band coefficients corresponding to the spectral energy in each of the sub-bands are identified using frequency domain techniques similar to those described above with respect to personalization application 145 of FIG. 1.
[0064] At step 303, an audio category of the audio track is determined. In some embodiments, the audio category of the selected audio track may be determined from classification data and / or other metadata associated with the selected audio track. In some embodiments, the audio category of the selected audio track may be determined by performing a real-time search of the classification data and / or other metadata against one or more online databases. In some embodiments, the audio category of the selected audio track may be determined by identifying one or more instruments within the audio track and performing one or more audio pattern matching techniques.
[0065] In some embodiments, the audio category is identified based on a user selection. In some embodiments, the audio category may be automatically selected based on data obtained from sensor(s) 118 or sensors disposed on smart device(s) 111. For example, the audio category may be selected based on sensor(s) 118 capturing a voice command identifying the selection of an audio category, a user movement and / or gesture identifying the selection of an audio category, and / or a user interaction with an input device, etc.
[0066] In step 304, the audio properties of the audio track are compared to one or more audio metrics 131 of the audio category to determine whether the selected audio track is representative of the corresponding audio category. In some embodiments, the audio properties of the audio track are compared to a combination of one or more statistical properties and / or tolerance windows associated with the corresponding audio category.
[0067] In some embodiments, audio properties of an audio track are compared to a range or average of corresponding audio metrics 131 to determine what percentage of audio properties are within a corresponding range, within a predetermined number of standard deviations from the corresponding average, and / or within a tolerance window of the corresponding audio metrics 131, etc. In some embodiments, the aggregate difference between an audio property of an audio track and the corresponding audio metrics 131 is compared to a threshold difference. In some embodiments, the aggregate difference is based on a distance function (e.g., Euclidean distance) and / or a weighted sum, etc. In some embodiments, the difference between an audio property and the corresponding audio metrics 131 is measured from the average of the corresponding audio metrics 131, or by determining a z-score indicating the number of standard deviations of the audio property from the average of the corresponding audio metrics.
[0068] If the audio properties do not match the audio metrics 131 of the corresponding audio category (e.g., if too many audio properties are outside the corresponding range and / or the total distance exceeds a threshold distance), an alternative audio track is suggested in step 305. If the audio properties match the audio metrics 131 of the audio category of the audio track, the selected audio track is further processed starting from step 306.
[0069] An alternative audio track is suggested at step 305. In some embodiments, the alternative audio track is suggested based on historical data of representative track selections by the user in the associated audio category, data regarding representative audio tracks for the audio category, and / or demographic data indicative of one or more representative tracks selected by similar users, etc. Steps 301-304 are then repeated to allow the user to select another audio track, and it is determined whether the alternative audio track matches the audio category.
[0070] In step 306, a portion of the audio track is selected that is representative of the audio category. In some embodiments, the audio track is divided into one or more frames or segments. In some embodiments, a technique similar to that used in step 304 is used to identify which frames and / or segments have audio properties that are best representative of the audio category identified in step 303. The best representative frames or segments are then selected as portions of the audio track. In some embodiments, frames and / or segments that have the smallest aggregate difference from one or more audio metrics 131 of the audio category are selected as portions of the audio track.
[0071] At step 307, an audio sample is created based on the portion of the audio track. In some embodiments, the audio sample may include a predefined length of audio content (e.g., a 15-25 second sample) generated from the portion of the audio track. In some embodiments, the audio sample is a repeating loop generated from the portion of the audio track. In some embodiments, the audio sample includes multiple repetitions of a first portion of the audio track. In some embodiments, the audio sample is created by seamlessly editing repetitions of the portion of the audio track together into the audio sample, such that there is no tempo discontinuity between any two repetitions of the first portion of the audio track.
[0072] At step 308, the audio sample is played for the user. The audio sample may be played using any of the devices in the audio environment 110, including, but not limited to, one or more smart devices 111, headphones 112, smart speakers 113, and other input / output (I / O) devices 119. In some embodiments, the audio sample may be played automatically based on data obtained from sensor(s) 118 or sensors disposed on the smart device(s) 111. For example, the audio sample may be played based on the sensor capturing a user uttering a play command, a user movement and / or gesture associated with initiating playback of the audio sample, and / or a user interaction with an input device, etc.
[0073] At step 309, one or more personalization settings of the user are adjusted based on user input as the audio sample is played. In some embodiments, the user may increase or decrease bass or treble levels, adjust frequency band filters, apply compression or equalization, perform discrete amplitude adjustments, select or modify pre-set acoustic filters, and / or select preferred signal processing for an audio category (e.g., dynamic range compression, dynamic expansion, audio limiting, spatial processing of the audio signal, etc.). In some embodiments, the user may select a previous personalization setting for the associated audio category as a starting point and update the personalization setting during playback of the audio sample.
[0074] In some embodiments, the personalization setting(s) are adjusted automatically based on data obtained from sensor(s) 118 or sensors disposed on smart device(s) 111. For example, the personalization setting(s) may be adjusted based on the sensors capturing a user utterance of a command to increase, decrease, select, modify, or adjust a setting. In some embodiments, the personalization setting(s) may be adjusted based on the sensors capturing a user movement and / or gesture associated with adjusting the setting, and / or a user interaction with an input device, etc.
[0075] At step 310, the personalization setting(s) for the audio category are saved. In some embodiments, the user may save the personalization setting(s) as new personalization setting(s) or update previously saved personalization setting(s) for one or more related categories of audio content. In some embodiments, the personalization setting(s) are associated with the audio category. In some embodiments, the personalization setting(s) may be saved automatically based on data obtained from sensor(s) 118 or sensors located on smart device(s) 111. For example, the personalization setting(s) may be saved based on the sensor capturing a user uttering a save or update command, a user movement and / or gesture associated with initiating the saving or updating of the personalization setting, and / or a user interaction with an input device, etc. In some embodiments, the personalization setting(s) are saved in personalization settings 122 in user profile database 120.
[0076] 4 is a flowchart of method steps for applying audio personalization settings to playback of audio tracks. Although the method steps are described with respect to the system of FIG. 1, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of various embodiments. In some embodiments, some or all of the method steps of FIG. 3 may be performed by audio processing application 146.
[0077] As shown, method 400 begins at step 401, where a user selects an audio track to play. In some embodiments, the user may select an entire audio track, a portion of an audio track, an aggregation of one or more portions of one or more audio tracks, etc. The user may select an audio track from audio content played via streaming service 104 or from locally stored audio content 144 of computing device 140. The user may select an audio track using a web-based program or a program locally stored on computing device 140. The audio track may be selected automatically based on data obtained from sensor(s) 118 or sensors located on smart device(s) 111. For example, the audio track may be selected based on sensors capturing a user utterance regarding the selection, a user movement and / or gesture associated with selecting the audio track, and / or a user interaction with an input device, etc.
[0078] At step 402, an audio category of the audio track is determined. In some embodiments, the audio category of the selected audio track may be determined from classification data and / or other metadata associated with the selected audio track. In some embodiments, the audio category of the selected audio track may be determined by performing a real-time search of the classification data and / or other metadata against one or more online databases. In some embodiments, the audio category of the selected audio track may be determined by identifying one or more instruments within the audio track and performing one or more audio pattern matching techniques.
[0079] In some embodiments, the audio category is identified based on a user selection. In some embodiments, the audio category may be automatically selected based on data obtained from sensor(s) 118 or sensors disposed on smart device(s) 111. For example, the audio category may be selected based on sensor(s) 118 capturing a voice command identifying the selection of an audio category, a user movement and / or gesture identifying the selection of an audio category, and / or a user interaction with an input device, etc.
[0080] In some embodiments, the audio category of the selected audio track is identified using techniques similar to those used in step 304. In some embodiments, the audio category is identified by comparing the audio properties of the selected audio track to one or more audio metrics 131 associated with one or more audio categories to find the audio category having one or more audio metrics 131 that best matches the audio properties of the selected track.
[0081] At step 403, a determination is made whether personalization settings for the particular audio category are available. In some embodiments, the software application queries the user profile database 120 to determine whether the stored personalization setting(s) 122 include personalization settings for the particular audio category. In some embodiments, if personalization settings for the particular audio category are not found, an option to create a personalization setting is provided at step 404. In some embodiments, if personalization settings for the particular audio category are available, the selected audio track is further processed starting at step 405.
[0082] At step 404, an option is provided to create a personalization setting. In some embodiments, suggested options for personalization settings for a particular audio category are generated, allowing the user to select a personalization setting for the audio category. In some embodiments, the user is given the option to select a previous personalization setting for the associated audio category and save the personalization setting for the particular audio category. In some embodiments, the user is given the option to initiate the process of customizing the personalization setting for the audio category, such as the method disclosed in FIG. 3.
[0083] The personalization settings for the audio category are loaded at step 405. In some embodiments, the personalization settings for the audio category correspond to the personalization settings saved at step 310.
[0084] The personalization settings are applied to the playback of the audio track in step 406. In some embodiments, a customized audio signal is generated by modifying the audio of the audio track selected in step 401 according to the personalization settings loaded in step 405.
[0085] In summary, various embodiments demonstrate systems and techniques that enable audio personalization by providing an efficient and convenient means for selecting representative audio tracks and representative audio samples. In disclosed embodiments, a software application analyzes an audio track to identify its audio properties and determines whether the audio track is representative of a corresponding audio category by comparing the audio properties of the audio track with one or more audio metrics associated with the corresponding audio category. If the audio track is sufficiently representative of the corresponding audio category, the software application compares the audio properties of one or more portions of the audio track with one or more audio metrics to identify portions of the audio track that are representative of the corresponding audio category. The software application then creates an audio sample based on the portions of the audio track. In some embodiments, the software application may then adjust one or more personalization settings for the user based on user input upon playback of the audio sample. In some embodiments, the one or more personalization settings may be applied to playback of audio tracks of the corresponding audio category.
[0086] At least one technical advantage of the disclosed technology over the prior art is that the disclosed technology enables improved audio personalization by allowing users to more efficiently and effectively select representative audio tracks that contain an appropriate balance of audio properties that allow the user to achieve a personalized personalization setting for a particular audio category. Based on the user's selection, the disclosed technology may suggest alternative representative audio tracks to use when creating a personalized setting for a particular audio category. Furthermore, the disclosed technology provides users with a faster, more computationally efficient means for generating a subset of audio tracks that contain category-specific balances of audio characteristics that can be used to configure a personalized setting.
[0087] 1. In some embodiments, a computer-implemented method for determining audio personalization settings for an audio category, the computer-implemented method including: identifying one or more audio properties of an audio track; selecting a first portion of the audio track that is representative of the audio category based on the one or more audio properties; playing the first portion of the audio track to a user; and adjusting the user's personalization settings based on the user's input during playback of the first portion of the audio track.
[0088] 2. The computer-implemented method of clause 1, further comprising creating an audio sample including multiple repetitions of the first portion of the audio track, and wherein playing the first portion of the audio track further comprises playing the audio sample.
[0089] 3. The computer-implemented method of clause 1 or 2, wherein creating the audio sample includes shortening or lengthening the duration of the first portion of the audio track so that no tempo discontinuity occurs between the repetitions of the first portion of the audio track in the audio sample.
[0090] 4. A computer-implemented method described in any of clauses 1 to 3, further comprising determining whether the audio track is representative of the audio category based on the one or more audio properties before selecting the first portion of the audio track.
[0091] 5. The computer-implemented method of any of clauses 1-4, further comprising: based on said determination, suggesting a second audio track representative of said audio category.
[0092] 6. The computer-implemented method of any of clauses 1-5, wherein the one or more audio properties include at least one of bass level, treble level, frequency spectrum, energy, or tempo.
[0093] 7. The computer-implemented method of any of clauses 1 to 6, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with corresponding audio metrics associated with the audio category.
[0094] 8. A computer-implemented method described in any of clauses 1 to 7, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference.
[0095] 9. A computer-implemented method described in any of clauses 1 to 8, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with a range of corresponding audio metrics associated with the audio category.
[0096] 10. The computer-implemented method of any of clauses 1-9, further comprising identifying the audio category of the audio track based on metadata associated with the audio track or user selection.
[0097] 11. In some embodiments, a system comprising a memory and a processor, wherein the memory stores one or more software applications, and the processor, when executing the one or more software applications, is configured to perform the steps of identifying one or more audio properties of an audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; and adjusting personalization settings of the user based on input from the user upon playback of the first portion of the audio track.
[0098] 12. The system described in clause 11, wherein the processor is further configured to perform the step of determining whether the audio track is representative of the audio category based on the one or more audio properties before selecting the first portion of the audio track.
[0099] 13. The system of clause 11 or 12, wherein the processor is further configured to perform the step of suggesting a second audio track representative of the audio category based on the determination.
[0100] 14. A system described in any of clauses 11 to 13, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with corresponding audio metrics associated with the audio category.
[0101] 15. A system described in any of clauses 11 to 14, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference.
[0102] 16. A system described in any of clauses 11 to 15, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties with a range of corresponding audio metrics associated with the audio category.
[0103] 17. In some embodiments, one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of identifying one or more audio properties of an audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; and adjusting personalization settings of the user based on input from the user during playback of the first portion of the audio track.
[0104] 18. The one or more non-transitory computer-readable media described in clause 17, further comprising, before selecting the first portion of the audio track, determining whether the audio track is representative of the audio category based on the one or more audio properties.
[0105] 19. One or more non-transitory computer-readable media according to clause 17 or 18, further comprising associating the personalization settings with the audio categories and storing the personalization settings.
[0106] 20. One or more non-transitory computer-readable media described in any of clauses 17-19, further comprising: receiving a selection of a second audio track to play; identifying a second audio category of the second audio track; loading second personalization settings associated with the second audio category; generating a customized audio signal by modifying the audio of the second audio track according to the second personalization settings; and playing the customized audio signal to the user.
[0107] Any combination, in any manner, and all combinations of any claim elements recited in any claim and / or any elements described in this application are within the intended scope of the invention and protection.
[0108] The description of various embodiments is presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Numerous modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0109] Aspects of the present embodiments may be embodied as a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "module," a "system," or a "computer." Furthermore, any hardware and / or software technique, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or set of circuits. Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.
[0110] Any combination of one or more computer-readable medium(s) may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media may include an electrical connection having one or more communication lines, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium capable of containing or storing a program for use by or in connection with an instruction execution system, apparatus, or device.
[0111] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. Execution of the instructions via the processor of the computer or other programmable data processing apparatus causes the specific functions / operations of one or more blocks of the flowchart illustrations and / or block diagrams to be performed. Such a processor may be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor, or a field-programmable gate array.
[0112] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, including one or more executable instructions for implementing specific logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0113] While the forgoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which scope is defined by the appended claims.
Claims
1. 1. A computer-implemented method for determining audio personalization settings for an audio category, comprising: Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of the audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's personalization settings based on the user's input during playback of the first portion of the audio track; and 20. A computer-implemented method comprising:
2. creating an audio sample including multiple repetitions of the first portion of the audio track; playing the first portion of the audio track further comprises playing the audio samples. The computer-implemented method of claim 1 .
3. 3. The computer-implemented method of claim 2, wherein creating the audio sample includes shortening or lengthening a duration of the first portion of the audio track such that the audio sample does not introduce a tempo discontinuity between the repetitions of the first portion of the audio track.
4. The computer-implemented method of claim 1 , further comprising, prior to selecting the first portion of the audio track, determining whether the audio track is representative of the audio category based on the one or more audio properties.
5. The computer-implemented method of claim 4 , further comprising: suggesting a second audio track representative of the audio category based on the determination.
6. The computer-implemented method of claim 1 , wherein the one or more audio properties include at least one of bass level, treble level, frequency spectrum, energy, or tempo.
7. The computer-implemented method of claim 1 , wherein selecting the first portion of the audio track comprises comparing each of the one or more audio properties to corresponding audio metrics associated with the audio category.
8. 2. The computer-implemented method of claim 1, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference.
9. 2. The computer-implemented method of claim 1, wherein selecting the first portion of the audio track comprises comparing each of the one or more audio properties to a range of corresponding audio metrics associated with the audio category.
10. The computer-implemented method of claim 1 , further comprising identifying the audio category of the audio track based on metadata associated with the audio track or a user selection.
11. A system comprising a memory and a processor, the memory stores one or more software applications; When the processor executes the one or more software applications, Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's audio personalization settings for the audio category based on the user's input during playback of the first portion of the audio track; a system configured to perform the steps of
12. 12. The system of claim 11, wherein the processor is further configured to perform the step of determining whether the audio track is representative of the audio category based on the one or more audio properties before selecting the first portion of the audio track.
13. The system of claim 12 , wherein the processor is further configured to perform the step of suggesting a second audio track representative of the audio category based on the determination.
14. The system of claim 11 , wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties to a corresponding audio metric associated with the audio category.
15. 12. The system of claim 11, wherein selecting the first portion of the audio track includes determining whether an aggregate difference between each of the one or more audio properties and a corresponding audio metric associated with the audio category is less than a threshold difference.
16. 12. The system of claim 11, wherein selecting the first portion of the audio track includes comparing each of the one or more audio properties to a range of corresponding audio metrics associated with the audio category.
17. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to: Identifying one or more audio properties of the audio track; selecting a first portion of the audio track representative of an audio category based on the one or more audio properties; playing the first portion of the audio track to a user; adjusting the user's audio personalization settings for the audio category based on the user's input during playback of the first portion of the audio track; One or more non-transitory computer-readable media that cause the steps of
18. 20. The one or more non-transitory computer-readable media of claim 17, further comprising, prior to selecting the first portion of the audio track, determining whether the audio track is representative of the audio category based on the one or more audio properties.
19. Associating the personalization settings with the audio categories; saving the personalization settings; 20. The one or more non-transitory computer-readable media of claim 17, further comprising:
20. receiving a selection of a second audio track to play; identifying a second audio category of the second audio track; loading a second personalization setting associated with the second audio category; and generating a customized audio signal by modifying the audio of the second audio track according to the second personalization settings; and playing the customized audio signal to the user; and 20. The one or more non-transitory computer-readable media of claim 17, further comprising:
Citation Information
Patent Citations
FM multiplex receiver
JP1999145860A
DIGITAL AUDIO DATA ABSTRACT METHOD AND APPARATUS, AND COMPUTER PROGRAM PRODUCT
JP2006508390A
Smart Audio Settings
US20200050421A1
Methods and apparatus to adjust audio playback settings based on analysis of audio characteristics
WO2020086771A1