Method and apparatus for adjusting audio playback settings based on analysis of audio characteristics
The system dynamically adjusts audio playback settings using a neural network to optimize equalization based on real-time analysis, addressing the need for manual adjustments in conventional systems and ensuring consistent sound quality across varying media sources and genres.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2026-03-06
AI Technical Summary
Conventional audio equalization systems require manual adjustment by users due to varying audio characteristics across different media sources and genres, leading to inconsistent listening experiences.
A system that dynamically adjusts audio playback settings using a neural network trained on reference media, applying filters and smoothing techniques to optimize equalization based on real-time audio analysis, eliminating the need for user input.
Provides a seamless and optimized listening experience by automatically adapting to changes in audio characteristics, ensuring smooth transitions and consistent sound quality across different media sources and genres.
Smart Images

Figure 0007825694000004 
Figure 0007825694000005 
Figure 0007825694000006
Abstract
Description
[Technical Field]
[0001]
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to audio playback settings, and more particularly to methods and apparatus for adjusting audio playback settings based on an analysis of audio characteristics.
[0002]
[0001] This patent arises from an application claiming the benefit of U.S. Provisional Patent Application No. 62 / 750,113, filed October 24, 2018, U.S. Provisional Patent Application No. 62 / 816,813, filed March 11, 2019, U.S. Provisional Patent Application No. 62 / 816,823, filed March 11, 2019, and U.S. Provisional Patent Application No. 62 / 850,528, filed May 20, 2019. U.S. Provisional Patent Application No. 62 / 750,113, U.S. Provisional Patent Application No. 62 / 816,813, U.S. Provisional Patent Application No. 62 / 816,823, and U.S. Provisional Patent Application No. 62 / 850,528 are incorporated herein by reference in their entireties. Priority is hereby claimed from U.S. Provisional Patent Application Nos. 62 / 750,113, 62 / 816,813, 62 / 816,823, and 62 / 850,528. [Background technology]
[0003]
[0003] In recent years, a large amount of media of various characteristics has been distributed using an increasing number of channels. Media can be received using traditional channels (e.g., radio, cell phone, etc.) or using more recently developed channels, such as using Internet-connected streaming devices. As these channels have developed, systems capable of processing and outputting audio from multiple sources have also been developed. These audio signals may have various characteristics (e.g., dynamic range, volume, etc.). For example, one automobile media system can deliver media from compact discs (CDs), Bluetooth®-connected devices, universal serial bus (USB)-connected devices, Wi-Fi-connected devices, auxiliary inputs, and other sources. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 is a block diagram illustrating an example environment constructed in accordance with the teachings of the present disclosure for dynamic playback setting adjustment based on real-time analysis of media characteristics. [Figure 2] 2 is a block diagram illustrating additional details of the media unit of FIG. 1 for implementing techniques for audio equalization according to at least first, second, and third implementations of the teachings of this disclosure. [Figure 3] 2 is a block diagram illustrating additional details of the content profile engine of FIG. 1 according to a second implementation. [Figure 4] FIG. 2 is a block diagram illustrating additional details of the audio equalization (EQ) engine of FIG. 1. [Figure 5] 3 is a flowchart representing exemplary machine-readable instructions that may be executed to implement the media unit of FIGS. 1 and 2 to dynamically adjust media playback settings based on real-time analysis of media characteristics according to a first implementation. [Figure 6]4 is a flowchart representing example machine-readable instructions that may be executed to implement the media unit 106 of FIGS. 1 and 2 to personalize equalization settings. [Figure 7] 1 is a flowchart representing example machine-readable instructions that may be executed to implement an audio EQ engine to train an EQ neural network according to a first implementation. [Figure 8A] 1 is a first spectrogram of an audio signal that has undergone dynamic audio playback setting adjustment based on real-time analysis of audio characteristics, but without a smoothing filter, according to a first implementation. [Figure 8B] 8B is a plot showing average gain values versus frequency values for the first spectrogram of FIG. 8A. [Figure 9A] 10 is a second spectrogram of an audio signal subjected to dynamic audio playback setting adjustments based on real-time analysis of audio characteristics, including a smoothing filter, according to a first implementation. [Figure 9B] 9B is a plot showing average gain values versus frequency values for the second spectrogram of FIG. 9A. [Figure 10] 10 is a flowchart representing exemplary machine-readable instructions that may be executed to implement the content profile engine of FIGS. 1 and 3 to deliver profile information along with a stream of content to a playback device according to a second implementation. [Figure 11] 10 is a flowchart representing exemplary machine-readable instructions that may be executed to implement the media unit of FIGS. 1 and 2 to play content using modified playback settings according to a second implementation. [Figure 12] 10 is a flowchart representing exemplary machine-readable instructions that may be executed to implement the media unit of FIGS. 1 and 2 to adjust playback settings based on profile information associated with content, according to a second implementation. [Figure 13A] FIG. 1 is a block diagram of an example content profile according to the teachings of the present disclosure. [Figure 13B]FIG. 1 is a block diagram of an example content profile according to the teachings of the present disclosure. [Figure 14] 10 is a flowchart representing machine-readable instructions that may be executed to implement the media unit of FIGS. 1 and 2 to perform real-time audio equalization according to a third implementation. [Figure 15] 10 is a flowchart representing machine-readable instructions that may be executed to implement the media unit of FIGS. 1 and 2 to smooth the equalization curve according to a third implementation. [Figure 16] 10 is a flowchart representing machine-readable instructions that may be executed to implement the audio EQ engine of FIGS. 1 and 4 to collect a data set and train and / or validate a neural network based on a reference audio signal, according to a third implementation. [Figure 17A] 16 is an exemplary representation of an equalized audio signal before implementing the smoothing technique shown and described in connection with FIG. 15. [Figure 17B] 17B is an exemplary representation of the audio signal of FIG. 17A after performing the smoothing technique shown and described in relation to FIG. 15 according to a third implementation. [Figure 18] FIG. 16 is a block diagram of an exemplary first processing platform configured to execute the instructions of FIGS. 5, 6, 11, 12, 14, and 15 to implement the media unit of FIGS. [Figure 19] FIG. 17 is a block diagram of an exemplary second processing platform architecture for executing the instructions of FIGS. 7 and 16 to implement the audio EQ engine of FIGS. 1 and 4. [Figure 20] FIG. 11 is a block diagram of an exemplary second processing platform architecture for executing the instructions of FIG. 10 to implement the content profile engine of FIGS. 1 and 3. DETAILED DESCRIPTION OF THE INVENTION
[0005]
[0027] Generally, the same reference numbers are used throughout the drawings and accompanying specification to refer to the same or like parts.
[0006]
[0028] In conventional media processing implementations, audio signals associated with different media may have different characteristics. For example, different audio tracks may have different frequency profiles (e.g., various volume levels at different frequencies of the audio signal), different overall volumes (e.g., average volume), pitch, timbre, etc. For example, media from one CD may be recorded and / or mastered differently from media from another CD. Similarly, media retrieved from a streaming device may have significantly different audio characteristics from media retrieved from an uncompressed medium such as a CD, and may also differ from media retrieved from the same device via a different application and / or audio compression level. As users increasingly listen to media from a variety of different sources and different genres and types of media, the differences in audio characteristics between sources, and between media from the same source, may become very noticeable and potentially irritating to listeners. Audio equalization is a technique utilized to adjust the volume levels of different frequencies in an audio signal. For example, equalization can be implemented to increase the presence of low, mid, and / or high frequency signals based on preferences associated with the genre of music, the era of music, user preferences, the space in which the audio signal is output, etc. However, optimal or preferred equalization settings may vary depending on the media being presented. Thus, listeners may need to frequently adjust equalization settings to optimize their listening experience based on changes in the media (e.g., changes in genre, changes in era, changes in overall track volume, etc.).
[0007]
[0029] Some conventional approaches allow for the selection of equalization settings associated with a particular genre or type of music. For example, a vehicle media unit may allow a listener to select a "rock" equalizer that is configured to boost frequencies that the user may want to sound loud and cut other frequencies that may be too strong, based on the typical characteristics of rock music. However, such genre-specific, broadly applied equalization settings cannot address the significant differences between different songs and, further, still require the user to manually change the equalization settings when the user starts a new track of a different genre, as often occurs with radio stations and audio streaming applications.
[0008]
[0030] In a first implementation, exemplary methods, apparatus, systems, and articles of manufacture disclosed herein dynamically adjust audio playback settings (e.g., equalization settings, volume settings, etc.) based on real-time characteristics of an audio signal. Examples disclosed herein determine a frequency representation (e.g., a CQT representation) of a sample (e.g., a 3-second sample) of the audio signal and query a neural network to determine equalization settings specific to the audio signal. In some examples disclosed herein, the equalization settings include multiple filters (e.g., a low-shelf filter, a peaking filter, a high-shelf filter, etc.), one or more of which can be selected and applied to the audio signal. In the exemplary methods, apparatus, systems, and articles of manufacture disclosed herein, the neural network that outputs the equalization settings is trained using a library of reference media corresponding to multiple equalization profiles optimized for the media (e.g., determined by an audio engineer).
[0009]
[0031] In a first implementation, the example methods, apparatus, systems, and articles of manufacture disclosed herein periodically (e.g., every second) query an audio sample (e.g., containing 3 seconds of audio) to a neural network to determine equalization settings for a profile and compensate for changes in the audio signal over time (e.g., different parts of a track with different characteristics, song transitions, genre transitions, etc.) The example methods, apparatus, systems, and articles of manufacture disclosed herein utilize a smoothing filter (e.g., an exponential smoothing algorithm, a one-pole recursive smoothing filter, etc.) to transition between filter settings to avoid perceptible changes in the equalization settings.
[0010]
[0032] Additionally, exemplary methods, systems, and articles of manufacture for modifying content playback using preprocessed profile information are described according to a second implementation. The exemplary methods, systems, and articles of manufacture access a stream of content to be delivered to a playback device, identify a piece of content within the stream of content to be delivered to the playback device, determine a profile for the identified piece of content, and deliver the determined profile to the playback device. These operations can be performed automatically (e.g., in real time) on the fly.
[0011]
[0033] In a second implementation, the exemplary method, system, and article of manufacture receives a stream of content at a playback device, accesses profile information associated with the content stream, and modifies playback of the content stream based on the accessed profile information. For example, the exemplary method, system, and article of manufacture receives and / or accesses an audio stream along with profile information that identifies a mood or other characteristic assigned to the audio stream, and modifies playback settings (e.g., equalization settings) of the playback device based on the profile information.
[0012]
[0034] Thus, in a second implementation, the exemplary methods, systems, and articles of manufacture can preprocess a content stream supplied by a content provider to determine a profile for the content stream and deliver the profile to a playback device, which can then play the content stream with, among other things, an adjusted, modified, and / or optimized playback experience.
[0013]
[0035] In a third implementation, the example methods, apparatus, systems, and articles of manufacture disclosed herein analyze and equalize an incoming audio signal (e.g., from a storage device, radio, streaming service, etc.) without requiring user input or adjustment. The techniques disclosed herein analyze the incoming audio signal to determine an average volume value over a buffer period for multiple frequency ranges, a standard deviation value over the buffer period for multiple frequency ranges, and the energy of the incoming audio signal. By utilizing the average frequency value over the buffer period, sudden short-term changes in the incoming audio signal are smoothed when determining the equalization curve to apply, thereby avoiding abrupt changes in the equalization settings.
[0014]
[0036] In a third implementation, the exemplary methods, apparatus, systems, and articles of manufacture disclosed herein generate an input feature set including an average volume value during a buffer period for multiple frequency ranges and / or a standard deviation value during a buffer period for multiple frequency ranges, and input the input feature set into a neural network. The exemplary methods, apparatus, systems, and articles of manufacture disclosed herein utilize a neural network trained on multiple reference audio signals and multiple equalization curves created by audio engineers. In some examples, the reference audio signals and corresponding equalization curves are tagged (e.g., associated with) the instructions of the particular audio engineer who created the equalization curve, allowing the neural network to learn the different equalization styles and preferences of different audio engineers. The exemplary methods, apparatus, systems, and articles of manufacture disclosed herein receive gain / cuts (e.g., volume adjustments) corresponding to specific frequency ranges from the neural network. In some examples, the gain / cuts are applied to a frequency representation of the incoming audio signal, and the equalized frequency representation is then analyzed to determine whether any anomalies (e.g., sharp spikes or dips in volume levels across frequencies) are present.
[0015]
[0037] According to a third implementation, the example methods, apparatus, systems, and articles of manufacture disclosed herein utilize thresholding techniques to remove anomalies in the equalized audio signal before determining an equalization curve (e.g., gain / cut for multiple frequency ranges) to apply to the audio signal. In some examples, the thresholding technique analyzes a set of adjacent frequency values (e.g., three or more adjacent frequency values) and determines whether the difference in volume between these adjacent frequency values (e.g., determined by calculating a second derivative over the frequency range) exceeds a threshold when the EQ gain / cut 241 from the neural network is applied. In some examples, in response to determining that the difference in volume between the adjacent frequency values exceeds the threshold, the volume corresponding to a central one of the frequency values can be adjusted to the midpoint between the volume levels at the adjacent frequency values, thereby removing spikes or dips in the frequency representation of the equalized audio signal. This adjustment has the subjective effect of a more pleasing EQ curve when compared to an EQ curve that has dips and peaks (e.g., local outliers) across the spectral envelope.
[0016]
[0038] In a third implementation, the example methods, apparatus, systems, and articles of manufacture disclosed herein measure energy values (e.g., RMS values) for the incoming audio signal and energy values after an equalization curve is applied to a representation of the incoming audio signal to attempt to normalize the overall volume before and after equalization. For example, if an equalization curve being applied to an audio signal boosts the volume in more frequency ranges than it cuts the volume in, the overall energy of the equalized audio signal may be higher. In some such examples, volume normalization may be performed on the equalized audio signal to remove noticeable volume changes between the incoming audio signal and the equalized audio signal.
[0017]
[0039] In a third implementation, the exemplary methods, apparatus, systems, and articles of manufacture disclosed herein improve audio equalization techniques by dynamically adjusting equalization settings to compensate for changes in the source supplying the incoming audio signal (e.g., radio, media stored on a mobile device, compact disc, etc.) or changes in the characteristics of the media represented by the incoming audio signal (e.g., genre, era, mood, etc.). The exemplary techniques disclosed herein utilize a neural network intelligently trained on audio signals equalized by expert audio engineers, allowing the neural network to learn the preferences and skills from various audio engineers. The exemplary techniques disclosed herein further refine the equalization adjustments provided by the neural network by implementing thresholding techniques to ensure that the final equalization curve is smooth and does not have large loudness differences between adjacent frequency ranges.
[0018]
[0040] 1 is a block diagram illustrating an example environment 100 constructed in accordance with the teachings of this disclosure for dynamic playback setting adjustment based on real-time analysis of media characteristics. The example environment 100 includes media devices 102, 104 that transmit audio signals to a media unit 106. The media unit 106 processes the audio signals (e.g., by implementing the audio equalization techniques disclosed herein) and sends the signals to an audio amplifier 108, which then outputs the amplified audio signals for presentation via an output device 110.
[0019]
[0041] In the example of FIG. 1 , media devices 102, 104 and / or media units 106 communicate over a network 112, such as the Internet, with exemplary content providers 114 or content sources (e.g., broadcast stations, networks, websites, etc.) that supply various types of multimedia content, such as audio and / or video content. Exemplary content providers 114 may include terrestrial or satellite radio stations, online music services, online video services, television broadcasters and / or distributors, networked computing devices (e.g., mobile devices on a network), local audio or music applications, etc. Note that content (e.g., audio and / or video content) may be obtained from any source. For example, the term “content source” is intended to include users and other content owners (artists, labels, movie studios, etc.). In some examples, the content source is a publicly accessible website, such as YouTube™.
[0020]
[0042] In some examples, network 112 may be any network or communication medium that enables communication between content providers 114, media devices 102, media devices 104, media units 106, and / or other networked devices. Exemplary network 112 may be or include a wired network, a wireless network (e.g., a mobile network), a radio or telecommunications network, a satellite network, etc. For example, network 112 may include one or more portions that make up a private network (e.g., a cable television network or a satellite radio network), a public network (e.g., a wireless broadcast channel or the Internet), etc.
[0021]
[0043] The exemplary media device 102 in the illustrated example of FIG. 1 is a portable media player (e.g., an MP3 player). The exemplary media device 102 stores or receives audio and / or video signals corresponding to media from a content provider 114. For example, the media device 102 can receive audio and / or video signals from the content provider 114 via a network 112. The exemplary media device 102 can transmit audio signals to other devices. In the illustrated example of FIG. 1, the media device 102 transmits audio signals to the media unit 106 via an auxiliary cable. In some examples, the media device 102 can transmit audio signals to the media unit 106 via any other interface. In some examples, the media device 102 and the media unit 106 can be the same device (e.g., the media unit 106 can be a mobile device capable of implementing the audio equalization techniques disclosed herein on audio presented to the mobile device).
[0022]
[0044] The exemplary media device 104 in the illustrated example of FIG. 1 is a mobile device (e.g., a cell phone). The exemplary media device 104 can store or receive audio signals corresponding to media and transmit the audio signals to other devices. In the illustrated example of FIG. 1, the media device 104 transmits audio signals wirelessly to the media unit 106. In some examples, the media device 104 can transmit audio signals to the media unit 106 using Wi-Fi, Bluetooth, and / or any other technology. In some examples, the media device 104 can interact with vehicle components or other devices to allow a listener to select media to be presented in the vehicle. The media devices 102, 104 can be any device capable of storing and / or accessing audio signals. In some examples, the media devices 102, 104 are integral to the vehicle (e.g., a CD player, a radio, etc.).
[0023]
[0045] The exemplary media unit 106 in the illustrated example of FIG. 1 can receive and process audio signals. In the illustrated example of FIG. 1, the exemplary media unit 106 receives and processes media signals from media devices 102, 104 and implements the audio equalization techniques disclosed herein. The exemplary media unit 106 can monitor audio being output by output device 110 to determine the average volume level, audio characteristics (e.g., frequency, amplitude, time value, etc.) of audio segments in real time. In some examples, the exemplary media unit 106 is implemented as software and included as part of another device available through a direct connection (e.g., a wired connection) or available over a network (e.g., available on the cloud). In some examples, the exemplary media unit 106 can be incorporated with an audio amplifier 108 and output device 110, and the exemplary media unit 106 itself can output the audio signal after processing the audio signal.
[0024]
[0046] In some examples, media device 102, media device 104, and / or media unit 106 can communicate with content provider 114 and / or content profile engine 116 via network 112. In additional or alternative examples, media device 102 and / or media device 104 can include a tuner configured to receive a stream of audio or video content, process the stream, output information (e.g., digital or analog) usable on a display of media device 102 and / or media device 104, and reproduce the stream of audio or video content by presenting or playing the audio or video content to a user associated with media device 102 and / or media device 104. Media device 102 and / or media device 104 can also include a display or other user interface configured to display the processed stream of content and / or associated metadata. The display may be a flat panel screen, a plasma screen, a light-emitting diode (LED) screen, a cathode ray tube (CRT), a liquid crystal display (LCD), a projector, etc.
[0025]
[0047] In some examples, the content provider 114, the content profile engine 116, the media device 102, the media device 104, and / or the media unit 106 may include one or more fingerprint generators 115 configured to generate identifiers for content being transmitted or broadcast by the content provider 114 and / or content being received or accessed by the media device 102, the media device 104, and / or the media unit 106. For example, the fingerprint generator 115 may include, among other things, a reference fingerprint generator (e.g., a component that calculates a hash value from a portion of the content) configured to generate a reference fingerprint or other identifier for the received content.
[0026]
[0048] In some examples, media unit 106 can be configured to modify the playback experience of content played by media device 102 and / or media device 104. For example, media unit 106 can access a profile associated with a stream of content and utilize the profile to modify, adjust, and / or control various playback settings (e.g., equalization settings) associated with the quality or characteristics for playback of the content. In one example where the content is video or other visual content, the playback settings can include color palette settings, color layout settings, brightness settings, font settings, artwork settings, etc.
[0027]
[0049] 1 is a device that can receive the audio signal being processed (e.g., equalized) by the media unit 106 and can perform appropriate playback setting adjustments (e.g., amplification of specific bands of the audio signal, volume adjustment based on user input, etc.) for output to the output device 110. In some examples, the audio amplifier 108 can be incorporated into the output device 110. In some examples, the audio amplifier 108 amplifies the audio signal based on an amplification output value from the media unit 106. In some examples, the audio amplifier 108 amplifies the audio signal based on input from a listener (e.g., a passenger or driver in a vehicle adjusting a volume selector). In additional or alternative examples, the audio is output directly from the media unit 106 rather than being communicated to an amplifier.
[0028]
[0050] 1 is a speaker. In some examples, output device 110 may be multiple speakers, headphones, or any other device capable of presenting an audio signal to a listener. In some examples, output device 110 may also be capable of outputting visual elements (e.g., a television with speakers). In some examples, output device 110 may be integrated into media unit 106. For example, if media unit 106 is a mobile device, output device 110 may be a speaker integrated into or connected to the mobile device (e.g., via Bluetooth, an auxiliary cable, etc.). In some such examples, output device 110 may be headphones connected to the mobile device.
[0029]
[0051] In some examples, the content profile engine 116 may access a stream of content provided by the content provider 114 via the network 112 and perform various processes to determine, generate, and / or select a profile or profile information for the stream of content. For example, the content profile engine 116 may identify the stream of content (e.g., using audio or video fingerprint comparison) and determine a profile for the identified stream of content. The content profile engine 116 may deliver the profile to the media device 102, the media device 104, and / or the media unit 106, which receives the profile along with the stream of content and plays the stream of content using certain playback settings associated with and / or selected based on, among other things, the information in the received profile.
[0030]
[0052] 1, the environment includes an audio EQ engine 118 that can provide trained models for use by the media unit 106. In some examples, the trained models reside on the audio EQ engine 118, while in some examples, the trained models are exported for direct use on the media unit 106. Machine learning techniques, whether deep learning networks or other experiential / observational learning systems, can be used, for example, to optimize results, locate objects in images, understand speech and convert speech to text, and improve the relevance of search engine results.
[0031]
[0053] While the illustrated example environment 100 of FIG. 1 is described with reference to a playback setting adjustment (e.g., audio equalization) implementation in a vehicle, some or all of the devices included in the example environment 100 can be implemented in any environment and in any combination. For example, the media unit 106 can be implemented (e.g., in whole or in part) in a mobile phone, along with either the audio amplifier 108 and / or the output device 110, and the mobile phone can perform playback setting adjustment (e.g., audio equalization) using the techniques disclosed herein on any media being presented from the mobile device (e.g., streaming music, media stored locally on the mobile device, radio, etc.). In some examples, the environment 100 can be in an entertainment room in a home, and the media devices 102, 104 can be a personal stereo system, one or more televisions, a laptop, other personal computers, tablets, other mobile devices (e.g., smartphones), gaming consoles, virtual reality devices, set-top boxes, or any other devices capable of accessing and / or transmitting media. Furthermore, in some examples, the media can also include visual elements (e.g., television shows, movies).
[0032]
[0054] In some examples, the content profile engine 116 may be part of the content provider 114, the media device 102, the media device 104, and / or the media unit 106. As another example, the media device 102 and / or the media device 104 may include, among other configurations, the content provider 114 (e.g., the media device 102 and / or the media device 104 may be a mobile device with a music playback application and the content provider 114 may be a local store of songs and other audio).
[0033]
[0055] 2 is a block diagram illustrating additional details of the media unit 106 of FIG. 1 for implementing techniques for audio equalization according to at least first, second, and third implementations of the teachings of this disclosure. The exemplary media unit 106 receives an input media signal 202 and processes the signal to determine audio and / or video characteristics. The audio and / or video characteristics are then utilized to determine appropriate audio and / or video playback adjustments based on the characteristics of the input media signal 202. When the input media signal 202 is an audio signal, the media unit 106 sends the output audio signal to an audio amplifier 108 for amplification before being output by an output device 110.
[0034]
[0056] The example media unit 106 includes an example signal converter 204, an example equalization (EQ) model query generator 206, an example EQ filter setting analyzer 208, an example EQ personalization manager 210, an example device parameter analyzer 212, an example historical EQ manager 214, an example user input analyzer 216, an example EQ filter selector 218, an example EQ adjustment implementor 220, an example smoothing filter configurator 222, an example data store 224, and an example update monitor 226. The example media unit 106 further includes an example fingerprint generator 227 and an example synchronizer 228. The exemplary media unit 106 further includes an exemplary buffer manager 230, an exemplary time-to-frequency domain converter 232, an exemplary volume calculator 234, an exemplary energy calculator 236, an exemplary input feature set generator 238, an exemplary EQ curve manager 240, an exemplary volume adjuster 242, an exemplary thresholding controller 244, an exemplary EQ curve generator 246, an exemplary volume normalizer 248, and an exemplary frequency-to-time domain converter 250.
[0035]
[0057] An exemplary media unit 106 is configured to operate according to at least three implementations. In a first implementation, the media unit 106 equalizes media in real time according to filter settings received from a neural network in response to a query including a frequency representation of the input media signal 202. In the first implementation, after processing the filter settings, the media unit 106 can generate an output media signal 252 that is equalized according to at least some of the filter settings. In some examples of the first implementation, the media unit 106 can further apply one or more smoothing filters to the equalized version of the input media signal 202 before outputting the output media signal 252.
[0036]
[0058] In a second implementation, the media unit 106 dynamically equalizes the media according to one or more profiles received from a content profile engine (e.g., the content profile engine 116). In the second implementation, after processing the one or more profiles, the media unit 106 may generate an output media signal 252 that is equalized according to at least some of the one or more profiles. In some examples of the second implementation, the media unit 106 may further apply personalized equalization to the input media signal 202 before outputting the output media signal 252.
[0037]
[0059] In a third implementation, the media unit 106 equalizes the media in real time according to equalization gain and cut values received from the neural network in response to an input feature set that includes features based on the input media signal 202. In the third implementation, after processing the filter settings, the media unit 106 can generate an output media signal 252 that is equalized according to at least some of the gain and cut values. In some examples of the third implementation, the media unit 106 can apply thresholding to the equalized version of the input media signal 202 to remove local outliers in the output media signal 252. First Implementation: Filter-Based Equalization
[0038]
[0060] In a first implementation, the exemplary input media signal 202 may be an audio signal to be processed and output for presentation. The input media signal 202 may be accessed from a wireless signal (e.g., an FM signal, an AM signal, a satellite radio signal, etc.), a compact disc, an auxiliary cable (e.g., connected to a media device), a Bluetooth signal, a Wi-Fi signal, or any other medium. The input media signal 202 is accessed by the signal converter 204, the EQ adjustment implementer 220, and / or the update monitor 226. The input media signal 202 is converted by the EQ adjustment implementer 220 and output by the media unit 106 as the output media signal 252.
[0039]
[0061] 2 converts the input media signal 202 into a frequency and / or characteristic representation of an audio signal. For example, the signal converter 204 may convert the input media signal 202 into a CQT representation. In some examples, the signal converter 204 converts the input media signal 202 using a Fourier transform. In some examples, the signal converter 204 continuously converts the input media signal 202 into a frequency and / or characteristic representation, while in other examples, the signal converter 204 converts the input media signal 202 at regular intervals or in response to a request from one or more other components of the media unit 106 (e.g., whenever needed for dynamic audio playback setting adjustment). In some examples, the signal converter 204 converts the input media signal 202 in response to a signal from the update monitor 226 (e.g., indicating that it is time to update the audio playback settings). The signal converter 204 of the illustrated example communicates a frequency and / or characteristic representation of the input media signal 202 to the EQ model query generator 206 , the fingerprint generator 227 and / or the synchronizer 228 .
[0040]
[0062] The EQ model query generator 206 in the illustrated example of FIG. 2 generates and communicates an EQ query 207 based on a frequency and / or characteristic representation of the input media signal 202. The EQ model query generator 206 selects one or more frequency representation(s) corresponding to a sample time frame (e.g., a 3-second sample) of the input media signal 202 and communicates the frequency representation(s) to a neural network (e.g., the EQ neural network 402 of FIG. 4). The sample time frame corresponds to a duration of the input media signal 202 to be considered when determining audio playback settings. In some examples, an operator (e.g., a listener, an audio engineer, etc.) can configure the sample time frame. In some examples, the EQ model query generator 206 communicates the query 207 (including the frequency representation(s) of the input media signal 202) to the neural network over a network. In some examples, the EQ model query generator 206 queries a model stored in and executed on the media unit 106 (e.g., the data store 224). In some examples, the EQ model query generator 206 generates a new query 207 to determine updated audio playback settings in response to a signal from the update monitor 226 .
[0041]
[0063] The EQ filter setting analyzer 208 in the illustrated example of FIG. 2 accesses EQ filter settings 209 and calculates filter coefficients to be applied to the input media signal 202. The EQ filter setting analyzer 208 accesses EQ filter settings 209 output by an EQ neural network (e.g., EQ neural network 402 of FIG. 4), where the EQ filter settings 209 may include one or more gain values, frequency values, and / or quality factor (Q) values. In some examples, the EQ filter settings 209 include multiple filters (e.g., one low-shelf filter, four peaking filters, one high-shelf filter, etc.). In some such examples, each filter includes multiple adjustment parameters, such as one or more gain values, one or more frequency values, and / or one or more Q values. For example, for an audio signal to which multiple filters are to be applied, the multiple filters may include respective adjustment parameters, including respective gain values, respective frequency values, and respective Q values (e.g., respective quality factor values). In some examples, the EQ filter setting analyzer 208 utilizes different formulas to calculate the filter coefficients based on the filter type. For example, a first equation may be utilized to determine a first filter coefficient for a low-shelf filter, and a second equation may be utilized to determine a second filter coefficient for a high-shelf filter. The EQ filter setting analyzer 208 communicates with the EQ filter selector 218 to determine which of the one or more sets of EQ filter settings 209 received by the EQ filter setting analyzer 208 should be processed (e.g., by calculating filter coefficients) to apply to the input media signal 202.
[0042]
[0064] 2 can generate personalized equalization setting(s) (e.g., personalized EQ settings, curves, filter settings, etc.) and combine the personalized equalization settings with dynamically generated filter settings from a neural network to reflect a listener's personal preferences. The EQ personalization manager 210 includes an exemplary device parameter analyzer 212, an exemplary historical EQ manager 214, and an exemplary user input analyzer 216.
[0043]
[0065] The device parameter analyzer 212 analyzes parameters associated with the media unit 106 and / or the source device providing the input media signal 202. For example, the device parameter analyzer 212 can indicate the app from which the input media signal 202 originated. In some such examples, different apps can be associated with different equalization profiles. For example, an audio signal from an app associated with audiobooks may have a different optimal equalization curve than an audio signal from an app associated with fitness.
[0044]
[0066] In some examples, the device parameter analyzer 212 determines the location of the device. For example, the device parameter analyzer 212 can determine the location of the media unit 106 and / or the location of the device providing the input media signal 202 to the media unit 106. For example, if the media unit 106 is integrated into a mobile device and the location of the mobile device is a gym, a different personalized equalization curve can be generated than if the mobile device is located at the user's home or work. In some examples, the device parameter analyzer 212 determines whether the location of the mobile device is within a geofence of an area (e.g., gym, home, work, library, etc.) for which personalized equalization settings (e.g., personalized EQ settings) are to be determined.
[0045]
[0067] In some examples, the device parameter analyzer 212 determines the user of the media unit 106 and / or the user of the device providing the input media signal 202 to the media unit. For example, if the media unit 106 is integrated into a mobile device, the device parameter analyzer 212 can determine the user of the mobile device based on a login associated with the user device and / or another identifier associated with the user device. In some examples, the user can be asked to select a user profile to indicate who is utilizing the mobile device and / or other devices associated with the media unit 106.
[0046]
[0068] The device parameter analyzer 212 in the illustrated example outputs and / or adjusts a personalized EQ curve based on any parameter to which the device parameter analyzer 212 has access (e.g., location, user identifier, source identifier, etc.).
[0047]
[0069] The historical EQ manager 214 in the illustrated example of FIG. 2 maintains historical data regarding past equalization curves that are utilized to enable subsequent personalized EQ curve adjustments. For example, if a user frequently listens to rock music and frequently utilizes an EQ curve that is best suited to rock music, the historical EQ manager 214 can help adjust and / or generate a personalized EQ curve based on the user's typical music preferences. For example, the historical EQ manager 214 can generate a personalized EQ curve based on a defined historical listening period. For example, the historical EQ manager 214 can generate a personalized EQ curve based on the previous hour of listening, the previous 24 hours of listening, and / or any other time frame. In other words, the historical EQ manager 214 can generate and / or adjust a personalized EQ curve based on EQ settings associated with the previous period. The historical EQ manager 214 takes the EQ curves being generated in real time by the EQ filter setting analyzer 208 and / or neural network and compiles those settings for each band of the EQ (e.g., each of the five bands) into a long-term personalized EQ filter that averages the settings over a historical period. The average curve recognized by the system over the historical period becomes the personalization EQ curve. This curve reflects the average EQ for the type of music the user listens to. For example, if a user has listened to heavy metal for the past 60 minutes, the user will have a different EQ curve stored in their user profile than if the user has listened to Top 40 pop for the past 60 minutes.
[0048]
[0070] The averaging operation can be a rolling average, an IIR filter, an all-pole filter with coefficients set to average over a time frame, or any other averaging technique. This averaging can alleviate the need to maintain buffer information for long periods of time. By utilizing historical EQ data, the EQ settings can have a degree of "stickiness," allowing the system to gradually learn listener preferences over time and create a more useful equalization curve.
[0049]
[0071] In some examples, the historical EQ manager 214 determines a small subset of genres that can be used with a table lookup for a given EQ curve for each genre (rock, country, voice, hip hop, etc.) Based on this subset of genres, an EQ curve can be generated, adjusted, or selected.
[0050]
[0072] The user input analyzer 216 in the illustrated example of FIG. 2 accesses and responds to user inputs corresponding to equalization settings. For example, a user can provide input regarding whether a particular equalization setting is preferred (e.g., by pressing a “like” button, providing a user rating, etc.). These inputs can then be utilized to more heavily weight the equalization settings that the user indicated they prefer when generating the personalized EQ curve. In some examples, user preferences are stored for a defined period of time (e.g., several months, a year, etc.). In some examples, user preferences are stored in association with a particular user account (e.g., a user login identified by the device parameter analyzer 212). In some examples, the user input analyzer 216 receives a “reset” signal from a listener indicating that the user wishes to cancel the automated personalized equalization applied to the audio signal. In some examples, the user input analyzer 216 adjusts the intensity of the equalization based on the intensity input from the listener.
[0051]
[0073] The exemplary EQ filter selector 218 in the illustrated example of FIG. 2 selects one or more of the filters represented by the EQ filter settings received by the EQ filter setting analyzer 208 (e.g., one or more of a low-shelf filter, a peaking filter, a high-shelf filter, etc.) to apply to the input media signal 202. The EQ filter selector 218 in the illustrated example selects one or more filters having the highest magnitude of gain (and therefore likely to have the greatest impact on the input media signal 202). In some examples, one or more additional filters represented by the EQ filter settings may be discarded, such as when a certain number of filters are to be utilized (e.g., five bandpass filters). In some examples, the EQ filter selector 218 determines which filters will have the least perceptible impact to a listener and discards these filters. For example, the EQ filter selector may integrate over the spectral envelope of one or more filters and compare this output between the filters to determine which of the filters represented by the EQ filter settings to discard. In some examples, the EQ filter selector 218 communicates to the EQ filter settings analyzer 208 and / or the EQ adjustment implementer 220 which of the filters should be applied to the input media signal 202 .
[0052]
[0074] 2 applies the filters selected by the EQ filter selector 218 and analyzed by the EQ filter setting analyzer 208. For example, the EQ adjustment implementer 220 may adjust the amplitude, frequency, and / or phase characteristics of the input media signal 202 based on the filter coefficients calculated by the EQ filter setting analyzer 208. In some examples, the EQ adjustment implementer 220 uses a smoothing filter indicated by the smoothing filter configurator 222 to smoothly transition from the previous audio playback setting to the updated audio playback setting (e.g., a new filter configuration). After applying the one or more equalization filter(s), the EQ adjustment implementer 220 outputs the output media signal 252.
[0053]
[0075] In some examples, the EQ adjustment implementer 220 blends between an equalization profile generated based on the EQ filter settings 209 from the neural network and a personalized EQ from the EQ personalization manager 210. For example, a user profile EQ curve can be blended with a real-time curve generated by the neural network. In some examples, weights are used to blend the EQ curves. Multiple weights can also be used. As one example, the final EQ curve that forms the audio that the user ultimately hears can be 0.5 times the current EQ based on the dynamically generated filter settings and 0.5 times the personalized EQ curve. As another example, the initial number can be 0.25 for the current EQ based on the dynamically generated filter settings and 0.75 for the personalized EQ curve.
[0054]
[0076] 2 defines parameters for smoothing between audio playback settings. For example, the smoothing filter configurator 222 may provide an equation and / or parameters for implementing smoothing by the EQ adjustment implementer 220 when applying an audio playback setting (e.g., an exponential smoothing algorithm, a single-pole recursive smoothing filter, etc.). The second spectrogram 900a in FIG. 9A illustrates the benefit of implementing a smoothing filter and shows a spectrogram of an audio signal that has undergone dynamic audio playback setting adjustment using a smoothing filter.
[0055]
[0077] 2 stores the input media signal 202, the output model from the EQ neural network 402 of FIG. 4, one or more profiles 229, EQ filter settings 209, EQ input feature set 239, EQ gain / cut 241, smoothing filter settings, an audio signal buffer, and / or any other data related to the dynamic playback setting adjustment process implemented by the media unit 106. The data store 224 may be implemented with volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory, etc.). Additionally or alternatively, the data store 224 may be implemented with one or more DDR memories, such as double data rate (DDR), DDR2, DDR3, mobile DDR (mDDR), etc. Additionally or alternatively, data store 224 may be implemented by one or more mass storage devices, such as hard disk drive(s), compact disk drive(s), digital versatile disk drive(s), etc. In the illustrated example, data store 224 is shown as a single database, but data store 224 may be implemented by any number and / or type(s) of databases. Furthermore, the data stored in data store 224 may be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, Structured Query Language (SQL) structures, etc.
[0056]
[0078] The exemplary update monitor 226 in the illustrated example monitors the duration between audio playback setting adjustments and determines when an update duration threshold is met. For example, the update monitor 226 can be configured with a 1-second update threshold, such that the EQ model query generator 206 queries the EQ neural network (e.g., EQ neural network 402 of FIG. 4) every 1 second to determine new playback settings. In some examples, the update monitor 226 communicates with the signal converter 204 to simplify samples (e.g., 3-second samples, 5-second samples, etc.) of the input media signal 202 and begin the process of determining updated audio playback settings.
[0057]
[0079] In operation, the signal converter 204 accesses the input media signal 202 and converts the input audio signal into a frequency and / or characteristic form, which is then utilized by the EQ model query generator 206 to query the neural network and determine EQ filter settings 209. The neural network returns the EQ filter settings 209, which are analyzed and processed (e.g., converted into applicable filter coefficients) by the EQ filter setting analyzer 208. The EQ filter selector 218 determines one or more of the filters represented by the EQ settings to apply to the input media signal 202. The EQ adjustment implementer 220 applies the selected filters using smoothing based on parameters from the smoothing filter configurator 222. The update monitor 226 monitors the duration since the previous audio playback settings were applied and updates the audio playback settings when an update duration threshold is met. Second Implementation: Profile-Based Equalization
[0058]
[0080] 2 generates an identifier (e.g., a fingerprint and / or a signature) for an input media signal 202 (e.g., content) received or accessed by media device 102, media device 104, and / or media unit 106. For example, fingerprint generator 227 may include, among other things, a reference fingerprint generator (e.g., a component that calculates a hash value from a portion of the content) configured to generate a reference fingerprint or other identifier for input media signal 202 (e.g., the received content). In some examples, fingerprint generator 227 implements fingerprint generator 115 of FIG. 1.
[0059]
[0081] 2 synchronizes one or more profiles 229 from the content profile engine 116 to the input media signal 202. In some examples, the media unit 106 may include a sequencer for sequencing the playback of media (e.g., songs) (or modifying (e.g., adjusting) the order in which media is played). In additional or alternative examples, the sequencer may be external to the media unit 106.
[0060]
[0082] 2, synchronizer 228 may utilize fingerprint(s) associated with input media signal 202 to synchronize input media signal 202 to one or more profiles 229. For example, one or more profiles 229 may include information relating one or more settings to known fingerprints for input media signal 202, such that synchronizer 228 may align settings to portions of input media signal 202 in order to synchronize one of one or more profiles 229 to input media signal 202 during playback of input media signal 202.
[0061]
[0083] In some examples, the synchronizer 228 may identify various audio or acoustic events (e.g., a snare hit, the start of a guitar solo, the first vocals) within the input media signal 202 and / or alternative representations thereof and align one of the one or more profiles 229 to the events within the input media signal 202 in order to synchronize one of the one or more profiles 229 to the input media signal 202 during playback of the input media signal 202. In additional or alternative examples, the sequencer may organize song sequences as part of adaptive radio, playlist recommendations, playlists of media (e.g., content) in a cloud (music and / or video) specific to the currently rendered media (e.g., content (e.g., using that profile)), a user's profile, device settings known in advance to provide a personalized, optimal experience, etc.
[0062]
[0084] In a second implementation, the exemplary EQ personalization manager 210 of the illustrated example of FIG. 2 can generate personalized equalization setting(s) (e.g., personalized EQ settings, curves, filter settings, etc.) and combine the personalized equalization settings with one or more profiles 229 to reflect the listener's personal preferences.
[0063]
[0085] The device parameter analyzer 212 analyzes parameters associated with the media unit 106 and / or the source device providing the input media signal 202. For example, the device parameter analyzer 212 can indicate the app from which the input media signal 202 originated. In some such examples, different apps can be associated with different equalization profiles. For example, an audio signal from an app associated with audiobooks may have a different optimal equalization curve than an audio signal from an app associated with fitness.
[0064]
[0086] In some examples, the device parameter analyzer 212 determines the location of the device. For example, the device parameter analyzer 212 can determine the location of the media unit 106 and / or the location of the device providing the input media signal 202 to the media unit 106. For example, if the media unit 106 is integrated into a mobile device and the location of the mobile device is a gym, a different personalized equalization curve can be generated than if the mobile device is located at the user's home or work. In some examples, the device parameter analyzer 212 determines whether the location of the mobile device is within a geofence of an area (e.g., gym, home, work, library, etc.) for which personalized equalization settings (e.g., personalized EQ settings) are to be determined.
[0065]
[0087] In some examples, the device parameter analyzer 212 determines the user of the media unit 106 and / or the user of the device providing the input media signal 202 to the media unit. For example, if the media unit 106 is integrated into a mobile device, the device parameter analyzer 212 can determine the user of the mobile device based on a login associated with the user device and / or another identifier associated with the user device. In some examples, the user can be asked to select a user profile to indicate who is utilizing the mobile device and / or other devices associated with the media unit 106.
[0066]
[0088] The device parameter analyzer 212 in the illustrated example outputs and / or adjusts a personalized EQ curve based on any parameter to which the device parameter analyzer 212 has access (e.g., location, user identifier, source identifier, etc.).
[0067]
[0089] The historical EQ manager 214 in the illustrated example of FIG. 2 maintains historical data regarding past equalization curves that are utilized to enable subsequent personalized EQ curve adjustments. For example, if a user frequently listens to rock music and frequently utilizes an EQ curve that is best suited to rock music, the historical EQ manager 214 can help adjust and / or generate a personalized EQ curve based on the user's typical music preferences. For example, the historical EQ manager 214 can generate a personalized EQ curve based on a defined historical listening period. For example, the historical EQ manager 214 can generate a personalized EQ curve based on the previous hour of listening, the previous 24 hours of listening, and / or any other time frame. In other words, the historical EQ manager 214 can generate and / or adjust a personalized EQ curve based on EQ settings associated with a previous period. The historical EQ manager 214 takes one or more profiles 229 being generated in real time and aggregates those settings for each band of the EQ (e.g., each of the five bands) into a long-term personalized EQ filter that averages the EQ settings for the historical period. The average curve that the system recognizes over the historical period becomes the personalization EQ curve. This curve reflects the average EQ for the type of music the user listens to. For example, if a user has listened to heavy metal for the past 60 minutes, that user will have a different EQ curve stored in their user profile than if they have listened to Top 40 pop for the past 60 minutes.
[0068]
[0090] The averaging operation can be a rolling average, an IIR filter, an all-pole filter with coefficients set to average over a time frame, or any other averaging technique. This averaging can alleviate the need to maintain buffer information for long periods of time. By utilizing historical EQ data, the EQ settings can have a degree of "stickiness," allowing the system to gradually learn listener preferences over time and create a more useful equalization curve.
[0069]
[0091] In some examples, the historical EQ manager 214 determines a small subset of genres that can be used with a table lookup for a given EQ curve for each genre (rock, country, voice, hip hop, etc.) Based on this subset of genres, an EQ curve can be generated, adjusted, or selected.
[0070]
[0092] The user input analyzer 216 in the illustrated example of FIG. 2 accesses and responds to user inputs corresponding to equalization settings. For example, a user can provide input regarding whether a particular equalization setting is preferred (e.g., by pressing a “like” button, providing a user rating, etc.). These inputs can then be utilized to more heavily weight the equalization settings that the user indicated they prefer when generating the personalized EQ curve. In some examples, user preferences are stored for a defined period of time (e.g., several months, a year, etc.). In some examples, user preferences are stored in association with a particular user account (e.g., a user login identified by the device parameter analyzer 212). In some examples, the user input analyzer 216 receives a “reset” signal from a listener indicating that the user wishes to cancel the automated personalized equalization applied to the audio signal. In some examples, the user input analyzer 216 adjusts the intensity of the equalization based on the intensity input from the listener.
[0071]
[0093] In a second implementation, the EQ adjustment implementer 220 is configured to modify playback of the input media signal 202 based on one or more profiles 229 for the input media signal 202. In such an additional or alternative example, the EQ adjustment implementer 220 implements an adjustor to modify playback of the input media signal 202 based on the one or more profiles 229. For example, the EQ adjustment implementer 220 can apply information in the one or more profiles 229 to modify or adjust settings of an equalizer and / or dynamic processor of the media unit 106, the media device 102, and / or the media device 104 to adjust and / or tailor equalization during playback of the input media signal 202 (e.g., a stream of content). In other words, the one or more profiles 229 include information that causes the EQ adjustment implementer 220 to adjust equalization of a portion of the input media signal 202. When the media (e.g., content) is video, one or more profiles 229 can be used to adjust video settings such as color temperature, dynamic range, color palette, brightness, sharpness, and any other video-related settings.
[0072]
[0094] In addition to equalization, the EQ adjustment implementer 220 may adjust a variety of different playback settings, such as equalization settings, virtualization settings, spatialization settings, etc. For example, the EQ adjustment implementer 220 may access information identifying a genre assigned to the input media signal 202 (e.g., a content stream) and modify the playback of the input media signal 202 (e.g., a content stream) by matching equalization settings of the playback device to settings associated with the identified genre. As another example, the EQ adjustment implementer 220 may access information identifying signal strength parameters for various frequencies of the content stream and modify the playback of the content stream by matching equalization settings of the playback device to settings that use the signal strength parameters.
[0073]
[0095] In some examples of the second implementation, the EQ adjustment implementer 220 blends between one or more profiles 229 generated by the content profile engine 116 and a personalized EQ from the EQ personalization manager 210. For example, a user profile EQ curve can be blended with a real-time profile. In some examples, weights are used to blend the personalized EQ curve with one or more profiles 229, and multiple weights can also be used. As one example, the final EQ curve that forms the audio that the user ultimately hears can be 0.5 times the current EQ based on dynamically generated filter settings and 0.5 times the personalized EQ curve. As another example, the initial number can be 0.25 for the current EQ based on dynamically generated filter settings and 0.75 for the personalized EQ curve. Third Implementation: Thresholding-Based Equalization
[0074]
[0096] In a third implementation, the exemplary buffer manager 230 in the illustrated example of FIG. 2 receives the input media signal 202 and stores a portion of the input media signal 202 in the data store 224. The buffer manager 230 can configure the buffer (e.g., the portion of the input media signal 202) to be of any duration (e.g., 10 seconds, 30 seconds, 1 minute, etc.). The portion of the input media signal 202 stored in the buffer in the data store 224 is utilized to determine equalization features, thereby allowing the equalization features to represent a longer duration of the input media signal 202 than would be possible if the features were generated based on the instantaneous characteristics of the input media signal 202. The buffer duration can be adjusted based on how responsive the equalization should be. For example, a very short buffer duration may result in a rapid change in the equalization curve when the spectral characteristics of the input media signal 202 change (e.g., between different parts of a song), whereas a longer buffer period will average out these large changes in the input media signal 202 to produce a more consistent equalization profile. The buffer manager 230 can discard portions of the input media signal 202 that are no longer within the buffer period. For example, if the buffer period is 10 seconds, then a portion of the input media 202 is removed after it has been in the buffer for 10 seconds.
[0075]
[0097] In some examples, a neural network is utilized to identify media changes (e.g., track changes, media source changes, etc.), and the output is utilized to adjust equalization in response to the media change. For example, when a new track is detected by the neural network, a short-term instantaneous or average volume (e.g., volume values for a frequency range over a period shorter than a standard buffer period, a standard deviation value for a frequency range over that short period, etc.) can be calculated to trigger a rapid adjustment of the EQ input feature set 239 and therefore the EQ gain / cut 241 received from the EQ neural network 402 of FIG. 4 (e.g., the equalization adjustment output from the EQ neural network 402). In some examples, to avoid rapid fluctuations in the equalization profile across the track during media changes, a longer-term volume averaging technique is utilized (e.g., determining an equalization profile based on a 30-second volume average, determining an equalization profile based on a 45-second volume average, etc.).
[0076]
[0098] In some examples, in addition to or as an alternative to utilizing a neural network to identify media changes, hysteresis-based logic can be implemented to cause faster equalization changes when more abrupt changes in the characteristics of the media represented by the input media signal 202 occur (e.g., a transition from bass-heavy media to treble-heavy media).
[0077]
[0099] In some examples, the media unit 106 can detect a change in the source of the input audio signal and trigger the aforementioned short-term equalization update (e.g., calculating short-term instantaneous or average volume and determining an equalization profile based on these changes) to compensate for differences in media from the new source relative to the previous source.
[0078]
[0100] 2 converts the input media signal 202 from a time-domain representation to a frequency-domain representation. In some examples, the time-frequency domain converter 232 utilizes a fast Fourier transform (FFT). In some examples, the time-frequency domain converter 232 converts the input media signal 202 to a linear spatial and / or logarithmic spatial frequency-domain representation. The time-frequency domain converter 232 can utilize any type of transform (e.g., short-time Fourier transform, constant-Q transform, Hartley transform, etc.) to convert the input media signal 202 from a time-domain representation to a frequency-domain representation. In some examples, the media unit 106 can alternatively perform the audio equalization techniques disclosed herein in the time domain.
[0079]
[0101] The exemplary loudness calculator 234 in the illustrated example of FIG. 2 calculates volume levels in frequency ranges for the input media signal 202. In some examples, the loudness calculator 234 calculates average volume levels (e.g., average loudness representations) over the entire buffer duration (e.g., 10 seconds, 30 seconds, etc.) for frequency bins (e.g., frequency ranges) of the linear spatial frequency representation of the input media signal 202. The loudness calculator 234 in the illustrated example of FIG. 2 generates a frequency representation of the average volume of the portion of the input media signal 202 stored in the buffer. Additionally or alternatively, the loudness calculator 234 in the illustrated example of FIG. 2 calculates a standard deviation over the entire buffer duration for the frequency bins. In some examples, the loudness calculator 234 calculates volume levels for logarithmic spatial frequency bins (e.g., critical frequency bands, bark bands, etc.). In some examples, to calculate the average volume levels for the frequency bins, the loudness calculator 234 converts the frequency representation of the input media signal 202 to real values.
[0080]
[0102] 2 calculates an energy value for a media signal (e.g., an audio signal). In some examples, the energy calculator 236 calculates a root-mean-square (RMS) value of the frequency representation of the audio signal before equalization (e.g., based on the frequency representation of the input media signal 202 stored in a buffer) and after an equalization curve has been applied (e.g., after the EQ curve generator 240 applies an equalization gain / cut to the average frequency representation of the audio signal). In some examples, the energy calculator 236 calculates the energy of a single frequency representation of the input media signal 202 (e.g., based on the volume level at any instant throughout the entire buffer period) and / or calculates the energy of the average frequency representation of the input media signal 202 over the entire buffer period.
[0081]
[0103] In some examples, the energy calculator 236 communicates the pre- and post-equalization energy values to the loudness normalizer 248 to enable loudness normalization and avoid perceptible changes in overall loudness after equalization. The energy calculator 236 in the illustrated example of FIG. 2 calculates the energy of the post-equalization average frequency representation.
[0082]
[0104] The exemplary input feature set generator 238 of the illustrated example of Figure 2 generates features (e.g., audio features) corresponding to the input media signal 202 for input to the EQ neural network 402 of Figure 4. In some examples, the input feature set generator 238 generates a set including average loudness measurements for frequency bins of the frequency representation of the input media signal 202 over the entire buffer period and / or average standard deviation measurements for frequency bins of the frequency representation of the input media signal 202 over the entire buffer period. In some examples, the input feature set generator 238 may include in the set any available metadata that is delivered to the EQ neural network 402 of Figure 4 to assist the EQ neural network 402 of Figure 4 in determining that appropriate equalization settings should be utilized for the input media signal 202.
[0083]
[0105] 2 determines the equalization curve utilized to equalize the input media signal 202. The example EQ curve manager 240 includes an example volume control 242, an example thresholding controller 244, and an example EQ curve generator 246.
[0084]
[0106] 2 receives EQ gain / cut 241 and applies volume adjustments to frequency ranges of an average representation of the input media signal 202. In some examples, the volume controller 242 receives the EQ gain / cut 241 as multiple values (e.g., scalars) to be applied to specific frequency ranges of the audio signal. In other examples, these values may be logarithmic-based gains and cuts (e.g., in decibels). In some such examples, the EQ gain / cut 241 corresponds to multiple logarithmic spatial frequency bins. For example, the EQ gain / cut 241 may correspond to the 25 critical bands used in the Bark Band representation.
[0085]
[0107] In some examples, to apply the EQ gain / cut 241 to the buffered portion of the input media signal 202, the volume adjuster 242 converts the linear spatial-frequency representation of the input media signal 202 (e.g., generated by the time-to-frequency domain converter 232) to a logarithmic spatial-frequency representation of the input media signal 202. In some such examples, the volume adjuster 242 may apply the EQ gain / cut 241 in decibels to the volume level of the logarithmic spatial-frequency representation to generate an equalized logarithmic spatial-frequency version of the buffered portion of the input media signal 202. The volume adjuster 242 in the illustrated example communicates the equalized logarithmic spatial-frequency version of the buffered portion of the input media signal 202 to the thresholding controller 244. In some examples, the EQ gain / cut 241 may be implemented in a linear spatial-frequency representation and / or other representation and applied to a common (i.e., linear spatial) representation of the buffered portion of the input media signal 202.
[0086]
[0108] In some examples, the volume controller 242 has access to information about the technical limitations of the source of the input media signal 202 and / or other technical characteristics related to the input media signal 202, and uses these technical prototypes or characteristics to refine which frequency ranges are subject to volume changes. For example, the volume controller 242 may have access to information about the type of encoding of the input media signal 202 (e.g., as determined by a decoder of the media unit 106, by analyzing the input media signal 202 for encoding artifacts, etc.). In some such examples, the volume controller 242 may prevent volume adjustments that may negatively affect the quality of the audio signal (e.g., adjustments that boost the volume in frequency ranges that include encoding artifacts).
[0087]
[0109] 2 implements a technique for smoothing an equalized version of the buffered portion of the input media signal 202 (e.g., from the volume controller 242). In some examples, after the volume controller 242 applies the EQ gain / cut 241 to the buffered portion of the input media signal 202, the frequency representation of the equalized audio signal may have local outliers (e.g., irregularities that appear as short-term peaks or dips on a frequency-loudness plot of the equalized audio signal) that may result in perceptible artifacts in the equalized audio signal. As used herein, the term local outlier refers to irregularities on a frequency-loudness plot of the equalized audio signal, such as large loudness differences between adjacent frequency values. In some examples, local outliers are detected by determining whether the second derivative of the loudness over a frequency range exceeds a threshold.
[0088]
[0110] 2 selects multiple frequency values at which to begin the thresholding technique. The thresholding controller 244 determines the volume levels at multiple frequency values and then calculates a measure of the difference between these frequency values. In some examples, the thresholding controller 244 calculates a second derivative of the volume values across multiple frequency values. As an example, if three frequency values are being analyzed to determine whether the center value of the three frequency values corresponds to a local outlier (e.g., an irregularity), the second derivative can be calculated using the following equation, where array val[] contains the volume values and index "i" corresponds to the frequency value index: |(val[i-2]-(2(val[i-1])+val[i])| ···(Formula 1)
[0089]
[0111] The thresholding controller 244 may compare the output of Equation 1 with a threshold value. In some examples, if the output of Equation 1, or any other equation used to calculate the relative difference in loudness at one of the frequency values relative to the loudness at adjacent frequency values, meets (e.g., exceeds) a threshold, a smoothing calculation may be used to remove the irregularity. In some examples, the thresholding controller 244 adjusts the volume level of the detected irregularity by changing the volume to the midpoint between the volume levels at adjacent frequency values. FIG. 17B shows an example of using this midpoint volume adjustment with respect to a local outlier shown in the equalized audio signal depicted in FIG. 17A. In some examples, the thresholding controller 244 may use any other technique to change the volume of the detected local outlier. For example, the thresholding controller 244 may set the volume of the detected local outlier equal to the volume at the adjacent frequency value or some other volume to attempt to remove the local outlier.
[0090]
[0112] In some examples, the thresholding controller 244 iteratively moves across the frequency range of the equalized audio signal to identify volume levels indicative of irregularities. In some examples, after analyzing all of the frequency values / ranges of the equalized audio signal, the thresholding controller 244 may perform one or more additional iterations across the equalized audio signal to determine whether local outliers remain after the initial adjustment stage (e.g., after volume levels for detected local outliers have been changed). In some examples, the thresholding controller 244 is a neural network and / or other artificial intelligence trained for irregularity (e.g., anomaly) detection. In some such examples, the thresholding controller 244 may eliminate irregularities with a single adjustment, without requiring additional iterations.
[0091]
[0113] After the thresholding controller 244 removes local outliers from the equalized frequency representation of the audio signal, or after another stopping condition is reached (e.g., 10 iterations of local outlier detection and adjustment across the entire frequency range), the thresholding controller 244 may communicate the final equalized representation of the audio signal to the EQ curve generator 246, which may then determine the equalization curve to apply to the input media signal 202.
[0092]
[0114] The EQ curve generator 246 of the illustrated example of FIG. 2 determines a final equalization curve to apply to the buffered portion of the input media signal 202. In some examples, the EQ curve generator 246 of the illustrated example of FIG. 2 subtracts the original average logarithmic spatial-frequency representation of the buffered portion of the input media signal 202 from the equalized version output from the thresholding controller 244 to determine the final equalization curve to utilize for equalization. In some such examples, after this subtraction, the EQ curve generator 246 converts the final equalization curve into a form (e.g., a linear spatial-frequency form) that can be applied to a frequency-domain representation of the buffered audio signal. In some such examples, the EQ curve generator 246 of the illustrated example then applies the final EQ curve (e.g., the linear spatial-frequency representation of the final EQ curve) to a corresponding representation (e.g., the linear spatial-frequency representation of the buffered audio signal). The EQ curve generator 246 may communicate the resulting equalized audio signal to the energy calculator 236, the volume normalizer 248, and / or the frequency-to-time domain converter 250. As used herein, an EQ curve includes gain / cut and / or other volume adjustments that correspond to a frequency range of the audio signal.
[0093]
[0115] The exemplary volume normalizer 248 in the illustrated example of FIG. 2 accesses an indication of the change in energy level of the input media signal 202 before and after equalization. The volume normalizer 248 in the illustrated example of FIG. 2 performs volume normalization to compensate for the overall change in the audio signal before and after equalization. In some examples, if the change in energy level of the input media signal 202 before and after equalization exceeds a threshold, the volume normalizer 248 applies a scalar volume adjustment to compensate for the change in energy level. In some examples, the volume normalizer 248 may utilize a dynamic range compressor. In some examples, the energy calculator 236 may calculate a ratio of the energy before and after the equalization process, and the volume normalizer 248 may utilize this ratio to counteract this overall volume change. For example, if the overall energy of the audio portion of the input media signal 202 doubles, the volume normalizer 248 may apply an overall volume cut to reduce the volume by half. In some examples, the loudness normalizer 248 may determine that the change in energy is insufficient to justify loudness normalization. The loudness normalizer 248 in the illustrated example communicates the final equalized audio signal (the equalized audio signal after volume adjustment, if applicable) to the frequency-to-time domain converter 250.
[0094]
[0116] The frequency-to-time domain converter 250 in the illustrated example of FIG. 2 converts the final equalized audio signal from the frequency domain to the time domain, for final output from the media unit 106 .
[0095]
[0117] 1 is illustrated in Figure 2, although one or more of the elements, processes, and / or devices illustrated in Figure 2 may be combined, divided, rearranged, omitted, removed, and / or implemented in any other manner. Additionally, an example signal converter 204, an example EQ model query generator 206, an example EQ filter setting analyzer 208, an example EQ personalization manager 210, an example device parameter analyzer 212, an example history EQ manager 214, an example user input analyzer 216, an example EQ filter selector 218, an example EQ adjustment implementer 220, an example smoothing filter configurator 222, an example data store 224, an example update monitor 226, an example fingerprint generator 227, an example synchronizer 228, an example buffer manager 230, an example The time-to-frequency domain converter 232, the exemplary volume calculator 234, the exemplary energy calculator 236, the exemplary input feature set generator 238, the exemplary EQ manager 240, the exemplary volume adjuster 242, the exemplary thresholding controller 244, the exemplary EQ curve generator 246, the exemplary volume normalizer 248 and / or the exemplary frequency-to-time domain converter 250, and / or more generally the exemplary media unit 106 of FIG. 2, may be implemented by hardware, software, firmware, and / or any combination of hardware, software and / or firmware.Thus, for example, an exemplary signal converter 204, an exemplary EQ model query generator 206, an exemplary EQ filter setting analyzer 208, an exemplary EQ personalization manager 210, an exemplary device parameter analyzer 212, an exemplary history EQ manager 214, an exemplary user input analyzer 216, an exemplary EQ filter selector 218, an exemplary EQ adjustment implementer 220, an exemplary smoothing filter configurator 222, an exemplary data store 224, an exemplary update monitor 226, an exemplary fingerprint generator 227, an exemplary synchronizer 228, an exemplary buffer manager 230, an exemplary time-to-frequency domain converter 232, an exemplary volume calculator 234, an exemplary energy calculator 236, an exemplary input feature set generator 238, an exemplary EQ manager 240 ... 0, the exemplary volume control 242, the exemplary thresholding controller 244, the exemplary EQ curve generator 246, the exemplary volume normalizer 248 and / or the exemplary frequency-to-time domain converter 250, and / or more generally any of the exemplary media unit 106 of FIG. 2 may be implemented by one or more analog or digital circuit(s), logic circuit(s), programmable processor(s), programmable controller(s), graphics processing unit(s) (GPU(s)), digital signal processor(s) (DSP(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)).When reading any of the apparatus or system claims of this patent that encompass purely software and / or firmware implementations, the exemplary signal converter 204, the exemplary EQ model query generator 206, the exemplary EQ filter setting analyzer 208, the exemplary EQ personalization manager 210, the exemplary device parameter analyzer 212, the exemplary history EQ manager 214, the exemplary user input analyzer 216, the exemplary EQ filter selector 218, the exemplary EQ adjustment implementer 220, the exemplary smoothing filter configurator 222, the exemplary data store 224, the exemplary update monitor 226, the exemplary fingerprint generator 227, the exemplary synchronizer 228, the exemplary buffer manager 230, the exemplary time-frequency domain controller 232, the exemplary time-frequency domain controller 234, the exemplary time-frequency domain controller 236, the exemplary time-frequency domain controller 238, the exemplary time-frequency domain controller 239, the exemplary time-frequency domain controller 240, the exemplary time-frequency domain controller 241, the exemplary time-frequency domain controller 242, the exemplary time-frequency domain controller 243, the exemplary time-frequency domain controller 244, the exemplary time-frequency domain controller 245, the exemplary time-frequency domain controller 246, the exemplary time-frequency domain controller 247, the exemplary time-frequency domain controller 248, the exemplary time-frequency domain controller 249, the exemplary time-frequency domain controller 250, the exemplary time-frequency domain controller 251, the exemplary time-frequency domain controller 252, the exemplary time-frequency domain controller 253, the exemplary time-frequency domain controller 254, the exemplary time-frequency domain controller 255, the exemplary time-frequency domain controller 256, the exemplary time-frequency domain controller 257, the exemplary time-frequency domain controller 258, the exemplary time-frequency domain controller 259, the exemplary time-frequency domain controller 260, the exemplary time-frequency domain controller 261, the 2, the exemplary input feature set generator 238, the exemplary EQ manager 240, the exemplary volume adjuster 242, the exemplary thresholding controller 244, the exemplary EQ curve generator 246, the exemplary volume normalizer 248, and / or the exemplary frequency-to-time domain converter 250, and / or more generally at least one of the exemplary media unit 106 of FIG. 2 is hereby expressly defined to include a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disc (DVD), a compact disc (CD), a Blu-ray disc, etc., that includes software and / or firmware. Furthermore, the exemplary media unit 106 of FIG. 1 can include one or more elements, processes, and / or devices in addition to or instead of those shown in FIG. 2 and / or can include any or all of the illustrated elements, processes, and devices. As used herein, the phrase "in communication" (including variations thereof) encompasses direct communication and / or indirect communication via one or more intermediate components and does not require direct physical (e.g., wired) communication and / or constant communication, but rather further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0096]
[0118] 3 is a block diagram illustrating additional details of the content profile engine 116 of FIG. 1 according to a second implementation. The exemplary content profile engine 116 includes an exemplary content retriever 302, an exemplary fingerprint generator 304, an exemplary content identifier 306, an exemplary profiler 308, and an exemplary profile data store 310. As described herein, in some examples, systems and methods generate and / or determine profiles for delivery to the media device 102, media device 104, and / or media unit 106 that identify media (e.g., content) to be streamed or transmitted to the media device 102, media device 104, and / or media unit 106 and provide information associated with the mood, style, or other attributes of the content. In some examples, the profile may be an identifier that identifies a content type. For example, the profile may identify the media (e.g., content) as news, action movies, sporting events, etc. Based on the profile, different settings for the TV may be adjusted in real time (e.g., on the fly). Similarly, a profile may identify a radio talk show, a song, a jingle, a song genre, etc. Thus, audio settings can be adjusted in real time (e.g., on the fly) to improve the audio delivered to the listener.
[0097]
[0119] In the example of FIG. 3, the content retriever 302 accesses and / or retrieves the input media signal 202 (e.g., a stream of content to be delivered to a playback device (e.g., media device 102, media device 104, media unit 106, etc.)) before delivering it to the media unit 106. For example, the content retriever 302 may access the input media signal 202 from a content provider 114 that is supplying the input media signal 202 (e.g., a stream of content) to the playback device (e.g., media device 102, media device 104, media unit 106, etc.) via the network 112. As another example, the content retriever 302 may access the input media signal 202 (e.g., a stream of content) from the content provider 114 that is stored locally by the playback device (e.g., media device 102, media device 104, media unit 106, etc.).
[0098]
[0120] 3, content retriever 302 may have access to various types of media (e.g., various types of content streams), such as audio content streams, video streams, etc. For example, content retriever 302 may have access to streams of songs or other music, streams of voice content, podcasts, YouTube™ videos and clips, etc.
[0099]
[0121] 3 generates an identifier (e.g., a fingerprint and / or a signature) for the input media signal 202 (e.g., content) received or accessed by the content profile engine 116. For example, the fingerprint generator 304 may include, among other things, a reference fingerprint generator (e.g., a component that calculates a hash value from a portion of the content) configured to generate a reference fingerprint or other identifier for the input media signal 202 (e.g., the received content). In some examples, the fingerprint generator 304 implements the fingerprint generator 115 of FIG. 1.
[0100]
[0122] 3, the content identifier 306 identifies a portion of media (e.g., a portion of content) within the input media signal 202 (e.g., a stream of content) to be delivered to a playback device (e.g., media device 102, media device 104, media unit 106, etc.). The content identifier 306 may identify the portion of media (e.g., a portion of content) through various processes, including comparing a fingerprint of the input media signal 202 (e.g., the content) to a reference fingerprint of known media (e.g., the content), such as a reference fingerprint generated by the fingerprint generator 304. For example, the content identifier 306 may generate and / or access a query fingerprint for a portion of the input media signal 202 or a frame or block of frames of the input media signal 202, and perform a comparison of the query fingerprint against the reference fingerprint to identify a piece of content or stream of content associated with the input media signal 202.
[0101]
[0123] 3, the profiler 308 determines one or more profiles 229 for an identified segment or portion (e.g., stream content) within the input media signal 202 and delivers the one or more profiles 229 to a playback device (e.g., media device 102, media device 104, media unit 106, etc.). For example, the profiler 308 may determine one or more characteristics for the input media signal 202 and / or may determine one or more characteristics for portions of the input media signal 202, such as frames or blocks of frames of the input media signal 202. In some examples, the profiler 308 stores the one or more profiles 229 in a profile data store 310.
[0102]
[0124] The example profiler 308 can render, generate, and / or determine one or more profiles 229 for an input media signal 202, such as audio content, having a variety of different characteristics. For example, the one or more profiles 229 can include characteristics associated with EQ settings, such as various audible frequencies within the audio content. The one or more profiles 229 can include various types of information. Exemplary profile information may include: (1) information identifying a category associated with a song, such as a category for the style of music (e.g., rock, classical, hip hop, instrumental, voice, jingle, etc.); (2) information identifying a category associated with a video segment, such as a style of video (e.g., drama, science fiction, horror, romance, news, TV show, documentary, advertisement, etc.); (3) information identifying a mood associated with a song or video clip, such as an upbeat mood, a relaxed mood, a soft mood, etc.; (4) information identifying signal strength parameters for various frequencies within the content, such as low frequencies for bass and other similar tones, high frequencies for voice or singing tones, and / or (5) information identifying color palette, brightness, sharpness, motion, blur, the presence of text and / or subtitles or subtitles, particular content with said text or subtitles, scene cuts, black frames, the presence of display format adjustment bars / pillars, the presence or absence of faces, landscape, or other objects, the presence of particular companies, networks, or broadcast logos, etc.
[0103]
[0125] Thus, the one or more profiles 229 may represent playback attributes (e.g., "DNA") of the input media signal 202, which may be used by the media unit 106 to control a playback device (e.g., media device 102, media device 104, media unit 106, etc.) to, among other things, optimize or enhance the experience during playback of the input media signal 202. As shown in Figure 3, the content profile engine 116 may generate and deliver one or more profiles 229 to the media unit 106, which in turn may adjust playback settings of the playback device (e.g., media device 102, media device 104, media unit 106, etc.) during playback of the input media signal 202 (e.g., a stream of content).
[0104]
[0126] 3 , the profile data store 310 stores one or more profiles, one or more reference fingerprints, and / or any other data related to the dynamic playback setting adjustment process implemented by the media unit 106 via one or more profiles 229. The profile data store 310 may be implemented with volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus dynamic random access memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory, etc.). Additionally or alternatively, the profile data store 310 may be implemented with one or more DDR memories, such as double data rate (DDR), DDR2, DDR3, mobile DDR (mDDR), etc. Additionally or alternatively, the profile data store 310 may be implemented with one or more mass storage devices, such as hard disk drive(s), compact disk drive(s), digital versatile disk drive(s), etc. In the illustrated example, the profile data store 310 is shown as a single database, but any number and / or type(s) of databases may implement the profile data store 310. Furthermore, the data stored in the profile data store 310 may be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, Structured Query Language (SQL) structures, etc.
[0105]
[0127] An example manner of implementing the content profile engine 116 of Figure 1 is illustrated in Figure 3, although one or more of the elements, processes, and / or devices illustrated in Figure 3 may be combined, divided, rearranged, omitted, removed, and / or implemented in any other manner. Furthermore, the example content retriever 302, the example fingerprint generator 304, the example content identifier 306, the example profiler 308, the example profile data store 310, and / or more generally the example content profile engine 116 of Figure 3 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the exemplary content retriever 302, the exemplary fingerprint generator 304, the exemplary content identifier 306, the exemplary profiler 308, the exemplary profile data store 310, and / or, more generally, any of the exemplary content profile engines 116 of FIG. 3 may be implemented by one or more analog or digital circuit(s), logic circuit(s), programmable processor(s), programmable controller(s), graphics processing unit(s) (GPU(s)), digital signal processor(s) (DSP(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)).When reading any of the apparatus or system claims of this patent that encompass a purely software and / or firmware implementation, at least one of the example content retriever 302, the example fingerprint generator 304, the example content identifier 306, the example profiler 308, the example profile data store 310, and / or, more generally, the example content profile engine 116 of Figure 3 is expressly defined hereby to include a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, etc., that includes software and / or firmware. Furthermore, the example content profile engine 116 of Figure 3 can include one or more elements, processes, and / or devices in addition to or instead of those shown in Figure 3 and / or can include any or all of the illustrated elements, processes, and devices. As used herein, the phrase "communicating" (including variations thereof) encompasses direct communication and / or indirect communication via one or more intermediate components and does not require direct physical (e.g., wired) communication and / or constant communication, but rather further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0106]
[0128] Figure 4 is a block diagram illustrating additional details of the audio EQ engine 118 of Figure 1. The example audio EQ engine 118 is configured to operate according to at least two implementations: in some examples, the trained model resides on the audio EQ engine 118 (e.g., within the EQ neural network 402), while in some examples, the trained model is exported for direct use on the media unit 106.
[0107]
[0129] Machine learning techniques, whether deep learning networks or other experiential / observational learning systems, can be used, for example, to optimize results, locate objects in images, understand speech and convert speech to text, and improve the relevance of search engine results. While many machine learning systems are supplied with initial features and / or network weights that are modified through learning and updating of the machine learning network, deep learning networks train themselves to identify "good" features for analysis. Using multi-layer architectures, machines that utilize deep learning techniques can process raw data better than machines that use traditional machine learning techniques. Using various layers of evaluation or abstraction, it becomes easier to examine the data for groups of highly correlated or distinctive themes.
[0108]
[0130] Machine learning techniques, whether neural networks, deep learning networks, and / or other experiential / observational learning systems, can be used, for example, to generate optimal results, locate objects in images, understand speech and convert speech to text, and improve the relevance of search engine results. Deep learning is a subset of machine learning that uses a set of algorithms to model high-level abstractions in data using deep graphs with multiple processing layers, including linear and nonlinear transformations. While many machine learning systems are supplied with initial features and / or network weights that are modified through learning and updating of the machine learning network, deep learning networks train themselves to identify "good" features for analysis. Using a multi-layer architecture, machines utilizing deep learning techniques can process raw data better than machines using traditional machine learning techniques. Using various layers of evaluation or abstraction, it becomes easier to examine data for groups of highly correlated or distinctive themes.
[0109]
[0131] For example, deep learning utilizing convolutional neural networks (CNNs) uses convolutional filters to segment data and locate and identify learned observable features within the data. Each filter, or layer, in a CNN architecture transforms the input data in a way that improves the selectivity and invariance of the data. This abstraction of the data allows the machine to focus on the features within the data it is attempting to classify and ignore irrelevant background information.
[0110]
[0132] Deep learning operates on the understanding that many datasets contain high-level features that contain low-level features. For example, while inspecting an image, rather than looking for objects, it is more efficient to look for edges that form motifs that form each part, edges that form the object you are looking for. These feature hierarchies can be found in many different forms of data.
[0111]
[0133] Learned observable features include objects and quantifiable regularities learned by machines during supervised learning. Machines with large sets of well-classified data are better equipped to distinguish and extract features relevant to successful classification of new data.
[0112]
[0134] A deep learning machine that utilizes transfer learning can successfully map data features to certain classifications supported by human experts. Conversely, the same machine can update parameters for a classification when notified of an incorrect classification by a human expert. For example, settings and / or other configuration information can be guided by the use of learned settings and / or other configuration information, and as the system is further used (e.g., repeatedly and / or by multiple users), some variation and / or other probability of settings and / or other configuration information can be reduced for a given situation.
[0113]
[0135] An exemplary deep learning neural network can be trained, for example, on a set of expert-classified data. This set of data establishes the first parameters for the neural network, which constitutes the supervised learning stage. During the supervised learning stage, the neural network can be tested to determine whether the desired behavior has been achieved. An exemplary flowchart representing machine-readable instructions for training the EQ neural network 402 is shown and described in connection with FIGS. 7 and 16. First Implementation: Filter-Based Equalization
[0114]
[0136] In a first implementation, the exemplary EQ neural network 402 of the illustrated example can be trained using a library of reference audio signals for which audio playback settings are specifically tuned and optimized (e.g., by audio engineering). In some examples, the EQ neural network 402 is trained by associating a sample of one of the reference audio signals (e.g., training data 408) with known audio playback settings for the reference audio signal. For example, gains, frequencies, and / or Q-factors for one or more filters recommended to apply to a track can be associated with individual audio signal samples of the track, thus training the EQ neural network 402 to associate similar audio samples with optimized playback settings (e.g., gains, frequencies, and / or Q-factors for one or more recommended filters). In some examples, different biases associated with different playback settings can also be indicated. For example, if the first 10 tracks are used for training and the audio playback settings (e.g., EQ parameters corresponding to the audio playback settings) for the first 10 tracks are determined by a first engineer, and the second 10 tracks are used for training and the audio playback settings for the second 10 tracks are determined by a second engineer, the EQ neural network 402 can be further trained to learn different preferences and / or biases associated with the first and second audio engineers and mitigate the first and second audio engineers to generate a more objective model.
[0115]
[0137] In some examples, a loss function may be utilized to train the EQ neural network 402. For example, Equation 2 represents one exemplary loss function that may be utilized, where f corresponds to frequency in Hertz, g corresponds to gain in decibels, and q corresponds to a unitless Q factor.
[0116]
number
[0138] After a desired neural network behavior is achieved (e.g., the machine has been trained to operate according to a particular threshold), the neural network can be deployed for use (e.g., to test the machine on “real” data). During operation, neural network classifications can be confirmed or rejected (e.g., by an expert user, an expert system, a reference database, etc.) to continue improving the neural network behavior. The exemplary neural network then enters a state of transfer learning as parameters for the classifications that determine neural network behavior are updated based on ongoing interactions. In some examples, a neural network such as EQ neural network 402 can provide direct feedback to another process, such as audio EQ scoring engine 404. In some examples, EQ neural network 402 outputs data that is buffered (e.g., via the cloud, etc.) and confirmed (e.g., via EQ confirmation data 410) before being fed to another process.
[0117]
[0139] In the example of FIG. 4 , the EQ neural network 402 receives input from previous results data associated with audio playback setting training data and outputs an algorithm for predicting audio playback settings to be associated with an audio signal. The EQ neural network 402 can be provided with some initial correlations, and then the EQ neural network 402 can learn from ongoing experience. In some examples, the EQ neural network 402 continuously receives feedback from at least one audio playback setting training data. In the example of FIG. 4 , throughout the operational life of the audio EQ engine 118, the EQ neural network 402 is continuously trained via feedback, and the example audio EQ engine validator 406 can be updated as needed based on the EQ neural network 402 and / or additional audio playback setting training data 408. The EQ neural network 402 can learn and evolve based on role, position, situation, etc.
[0118]
[0140] In some examples, the level of accuracy of the model generated by the EQ neural network 402 can be determined by the exemplary audio EQ engine validator 406. In such examples, at least one of the audio EQ scoring engine 404 and the audio EQ engine validator 406 receives a set of audio playback setting verification data 410. In such examples, the audio EQ scoring engine 404 further receives inputs (e.g., CQT data) associated with the audio playback setting verification data 410 and predicts one or more audio playback settings associated with the inputs. The predicted results are distributed to the audio EQ engine validator 406. The audio EQ engine validator 406 further receives known audio playback settings associated with the inputs and compares the predicted audio playback settings received from the audio EQ scoring engine 404 with the known audio playback settings. In some examples, the comparison produces a level of accuracy of the model generated by the EQ neural network 402 (e.g., if 95 comparisons produce matches and 5 produce errors, the model is 95% accurate, etc.). After the EQ neural network 402 reaches a desired level of accuracy (e.g., the EQ neural network 402 is trained and ready for deployment), the audio EQ engine validator 406 may output the model (e.g., output 414) to the data store 224 of FIG. 2 for use by the media unit 106 to determine audio playback settings. In some examples, after being trained, the EQ neural network 402 outputs sufficiently accurate EQ filter settings (e.g., EQ filter settings 209) to the media unit 106. Third Implementation: Thresholding-Based Equalization
[0119]
[0141] In a third implementation, the example EQ neural network 402 of the illustrated example can be trained using a library of reference audio signals for which audio equalization profiles (e.g., gains, cuts, etc.) have been determined (e.g., by an audio engineer). In the illustrated example of FIG. 4, the EQ neural network 402 receives example training data 408 (e.g., reference audio signals, EQ curves, and engineer tags). The engineer tags indicate, for a particular track, which of multiple audio engineers created the equalization profile for the track. In some examples, the engineer tags can be represented by a one-hot vector, with each entry in the one-hot vector corresponding to an engineer tag. In some examples, the EQ neural network 402 can ultimately average out relative style differences among various audio engineers without informing the EQ neural network 402 which engineer created the equalization profile for the track. For example, if a first set of reference audio signals has an EQ curve created by an audio engineer that generally emphasizes the bass frequency range, and a second set of reference audio signals has an EQ curve created by an audio engineer that generally emphasizes the mid-frequency range, the EQ neural network 402 can cancel out these relative differences during training if it is unaware of which audio engineer created the EQ curve. By providing an engineer tag associated with one of the multiple reference audio signals and the corresponding EQ curve, the EQ neural network 402 intelligently learns to recognize various equalization styles and effectively utilizes such styles when providing output 414 (e.g., EQ gain / cut 241) in response to EQ input feature set 239. In some examples, the EQ neural network 402 is trained by associating a sample of one of the reference audio signals in training data 408 with a known EQ curve for the reference audio signal.
[0120]
[0142] In some examples, a reference audio signal can be generated by taking a professionally engineered track and degrading the audio by applying an equalization curve that targets matching the spectral envelope of a non-professionally engineered track (e.g., of a lesser-known artist). The EQ neural network 402 can then be trained to undo the degradation by applying the equalization curve to restore the track to its original quality level. Thus, professionally engineered tracks can be utilized with this degradation technique to enable high-loudness training.
[0121]
[0143] In some examples, a loss function may be utilized to train the EQ neural network 402. For example, Equation 3 represents one exemplary loss function that may be utilized, where g i is the ground truth gain value in bin "i",
[0122]
number
[0123]
number
[0144] After a desired neural network behavior is achieved (e.g., the machine has been trained to operate according to a particular threshold), the neural network can be deployed for use (e.g., to test the machine on "real" data). In some examples, the neural network can then be used without further modifications or updates to the neural network parameters (e.g., weights).
[0124]
[0145] In some examples, during operation, neural network classifications can be confirmed or rejected (e.g., by an expert user, an expert system, a reference database, etc.) to continue improving neural network behavior. The example neural network then enters a state of transfer learning as parameters for the classifications that determine neural network behavior are updated based on ongoing interactions. In some examples, a neural network such as EQ neural network 402 can provide direct feedback to another process, such as audio EQ scoring engine 404. In some examples, EQ neural network 402 outputs data that is buffered (e.g., via the cloud, etc.) and confirmed before being fed to another process.
[0125]
[0146] In some examples, the EQ neural network 402 may be fed some initial correlations, and then the EQ neural network 402 may learn from ongoing experience. In some examples, throughout the operational life of the audio EQ engine 118, the EQ neural network 402 may be continuously trained via feedback, updating the EQ neural network 402 and / or the example audio EQ engine validator 406 as needed based on the EQ neural network 402 and / or additional audio playback setting training data 408. In some examples, the EQ neural network 402 may learn and evolve based on role, position, situation, etc.
[0126]
[0147] In some examples, the level of accuracy of the model generated by the EQ neural network 402 can be determined by an example audio EQ engine validator 406. In such examples, at least one of the audio EQ scoring engine 404 and the audio EQ engine validator 406 receives a set of audio playback setting training data (e.g., training data 408). The audio EQ scoring engine 404 of the illustrated example of FIG. 4 can determine the validity of the output 414 (e.g., EQ gain / cut 241) output by the EQ neural network 402 in response to the input 412 (e.g., EQ input feature set 239). In some examples, the audio EQ scoring engine 404 communicates with the audio EQ engine validator 406 during a validation procedure to determine how closely the output of the EQ neural network 402 in response to the input feature set corresponds to a known EQ curve for the input 412 (e.g., EQ input feature set 239). For example, the EQ input feature set 239 may be an audio sample for which an audio engineer has generated an EQ curve, and the audio EQ engine validator 406 may compare the output (e.g., EQ gain / cut 241) output by the EQ neural network 402 with the EQ curve (e.g., gain / cut) generated by the audio engineer.
[0127]
[0148] After being trained, the EQ neural network 402 of the illustrated example of FIG. 4 responds to the input 412 (e.g., the EQ input feature set 239) by providing an output 414 (e.g., the EQ gain / cut 241) to the media unit 106. For example, the EQ neural network 402 can determine multiple equalization adjustments (e.g., the EQ gain / cut 241) based on inferences associated with at least a reference audio signal, an EQ curve, and engineer tags. In some examples, the EQ gain / cut 241 includes multiple volume adjustment values (e.g., gain / cut) corresponding to multiple frequency ranges. In some examples, the EQ gain / cut 241 includes multiple volume adjustment values corresponding to multiple frequency ranges. For example, the EQ gain / cut 241 output by the EQ neural network 402 can include 24 gain or cut values corresponding to 24 critical bands of hearing.
[0128]
[0149] In some examples, the EQ neural network 402 can learn equalization settings based on user input(s). For example, if a user continually adjusts the equalization in a particular way (e.g., increasing the volume of bass frequencies and decreasing the volume of treble frequencies), the EQ neural network 402 learns these adjustments and outputs EQ gain / cut 241 that reflects the user preferences.
[0129]
[0150] In some examples, the comparisons produce a level of accuracy for the model generated by the EQ neural network 402 (e.g., if 95 comparisons produce matches and 5 produce errors, the model is 95% accurate, etc.). In some examples, after the EQ neural network 402 reaches a desired level of accuracy (e.g., the EQ neural network 402 is trained and ready for deployment), the audio EQ engine validator 406 can output the model to the data store 224 of FIG. 2 for use by the media unit 106 to determine audio playback settings.
[0130]
[0151] An example manner of implementing the audio EQ engine 118 of Figure 1 is illustrated in Figure 4, although one or more of the elements, processes, and / or devices illustrated in Figure 4 may be combined, divided, rearranged, omitted, removed, and / or implemented in any other manner. Furthermore, the example EQ neural network 402, the example audio EQ scoring engine 404, the example audio EQ engine validator 406, and / or more generally, the example audio EQ engine 118 of Figure 4 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the example EQ neural network 402, the example audio EQ scoring engine 404, the example audio EQ engine validator 406, and / or more generally any of the example audio EQ engine 118 of FIG. 4 may be implemented by one or more analog or digital circuit(s), logic circuit(s), programmable processor(s), programmable controller(s), graphics processing unit(s) (GPU(s)), digital signal processor(s) (DSP(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)). When reading any of the apparatus or system claims of this patent that encompass a purely software and / or firmware implementation, at least one of the example EQ neural network 402, the example audio EQ scoring engine 404, the example audio EQ engine validator 406, and / or, more generally, the example audio EQ engine 118 of FIG. 4 is expressly defined hereby to include a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, or the like, that includes software and / or firmware.Additionally, the example audio EQ engine 118 of Figure 4 may include one or more elements, processes, and / or devices in addition to or instead of those shown in Figure 4 and / or may include any more than one or all of the illustrated elements, processes, and devices. As used herein, the phrase "in communication" (including variations thereof) encompasses direct communication and / or indirect communication via one or more intermediate components and does not require direct physical (e.g., wired) communication and / or constant communication, but rather further includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0131]
[0152] Flowcharts depicting exemplary hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the media unit 106 of FIGS. 1 and 2 are shown in FIGS. 5, 6, 11, 12, 14, and 15. The machine-readable instructions may be an executable program or portions of an executable program for execution by a computer processor, such as the processor 1812 shown in the exemplary processor platform 1800 discussed below in connection with FIG. 18. The program may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, memory associated with the processor 1812, or the like, although alternatively, the entire program and / or portions thereof may be executed by a device other than the processor 1812 and / or may be embodied in firmware or dedicated hardware. Additionally, although the exemplary program is described with reference to the flowcharts shown in FIGS. 5, 6, 11, 12, 14, and 15, many other ways of implementing the exemplary media unit 106 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0132]
[0153] Flowcharts representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the audio EQ engine 118 of FIGS. 1 and 2 are shown in FIGS. 7 and 16. The machine-readable instructions may be an executable program or portions of an executable program for execution by a computer processor, such as the processor 1912 shown in the example processor platform 1900 discussed below in connection with FIG. 19. The program may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, memory associated with the processor 1912, or the like; alternatively, the entire program and / or portions thereof may be executed by a device other than the processor 1912 and / or may be embodied in firmware or dedicated hardware. Furthermore, although the example program is described with reference to the flowcharts shown in FIGS. 7 and 16, many other ways of implementing the example audio EQ engine 118 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0133]
[0154] A flowchart representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the content profile engine 116 of FIGS. 1 and 3 is shown in FIG. 10. The machine-readable instructions may be an executable program or portions of an executable program for execution by a computer processor, such as the processor 2012 shown in the example processor platform 2000 discussed below in connection with FIG. 20. The program may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, memory associated with the processor 2012, or the like; alternatively, the entire program and / or portions thereof may be executed by a device other than the processor 2012 and / or may be embodied in firmware or dedicated hardware. Furthermore, although the example program is described with reference to the flowchart shown in FIG. 10, many other ways of implementing the example content profile engine 116 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be changed, removed, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0134]
[0155] 5, 6, 7, 10, 11, 12, 14, 15, and 16 can be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium, such as a hard disk drive, flash memory, read-only memory, compact disk, digital versatile disk, cache, random access memory, and / or any other storage device or storage disk, on which information is stored for any duration (e.g., over an extended time frame, permanently, for short instances, while temporarily buffering, and / or while caching information). As used herein, the term non-transitory computer readable medium is expressly defined to include any type of computer readable storage device and / or storage disk, to exclude propagated signals, and to exclude transmission media.
[0135]
[0156] The terms "including" and "comprising" (and all forms and tenses thereof) are used herein to be open-ended terms. Thus, whenever a claim utilizes any form of "include" or "comprise" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within any type of claim recitation, it is understood that additional elements, terms, etc. may be present without departing from the scope of the corresponding claim or recitation. The term "at least," when used herein as a transitional term, for example, in the preamble of a claim, is open-ended in the same way that "comprising" and "including" are open-ended. The term "and / or," when used in the form A, B, and / or C, for example, refers to any combination or subset of A, B, and C, such as (1) A only, (2) B only, (3) C only, (4) A and B, (5) A and C, (6) B and C, and (7) A, B, and C. In the context of describing structures, components, items, objects, and / or things herein, the phrase "at least one of A and B" shall refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, in the context of describing structures, components, items, objects, and / or things herein, the phrase "at least one of A or B" shall refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. In the context of describing the implementation or performance of processes, instructions, operations, activities, and / or steps herein, the phrase "at least one of A and B" shall refer to an implementation that includes either (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.Similarly, in the context of describing the implementation or performance of processes, instructions, operations, activities and / or steps herein, the phrase "at least one of A or B" shall refer to implementations that include any of: (1) at least one A; (2) at least one B; and (3) at least one A and at least one B. First Implementation: Filter-Based Equalization
[0136]
[0157] 5 is a flowchart depicting exemplary machine-readable instructions 500 that may be executed to implement the media unit 106 of FIGS. 1 and 2 to dynamically adjust media playback settings based on real-time analysis of media characteristics, according to a first implementation. With reference to the preceding figures and associated description, the exemplary machine-readable instructions 500 begin with the exemplary media unit 106 accessing an audio signal (block 502). In some examples, the signal converter 204 accesses the input media signal 202.
[0137]
[0158] At block 504, the exemplary media unit 106 converts the audio signal into a frequency representation. In some examples, the signal converter 204 converts the input media signal 202 into a frequency and / or characteristic representation (e.g., a CQT representation, an FFT representation, etc.).
[0138]
[0159] At block 506, the exemplary media unit 106 inputs the frequency representation into the EQ neural network. In some examples, the EQ model query generator 206 inputs the frequency representation of the input media signal 202 into the EQ neural network 402. In some examples, the EQ model query generator 206 inputs the input media signal 202 into the model output by the EQ neural network 402.
[0139]
[0160] At block 508, the example media unit 106 accesses a plurality of filter settings including gain, frequency, and Q-factor. In some examples, the EQ filter setting analyzer 208 accesses the plurality of filter settings including gain, frequency, and Q-factor. In some examples, the EQ filter setting analyzer 208 accesses a plurality of filter settings (e.g., a set of filter settings) including gain, frequency, and Q-factor output by the EQ neural network 402. In some examples, the EQ filter setting analyzer 208 accesses one or more high-shelf filters, one or more low-shelf filters, and / or one or more peaking filters output by the EQ neural network 402.
[0140]
[0161] At block 510, the exemplary media unit 106 selects one or more filters to apply to the input media signal 202. In some examples, the EQ filter selector 218 selects the one or more filters to apply to the input media signal 202. For example, to implement a five-band filter, the EQ filter selector 218 may select one low-shelf filter, one high-shelf filter, and three peaking filters from the set of filters output by the EQ neural network 402.
[0141]
[0162] At block 512, the example media unit 106 calculates filter coefficients based on the settings of the selected filter(s). In some examples, the EQ filter setting analyzer 208 calculates the filter coefficients based on the filter settings of the selected filter(s) to enable application of one or more filter(s) to the input media signal 202.
[0142]
[0163] At block 514, the exemplary media unit personalizes the equalization settings. In some examples, the EQ personalization manager 210 personalized the equalization settings (e.g., personalized the EQ settings). Detailed exemplary machine-readable instructions for personalizing the equalization settings are shown and described in connection with FIG. 6.
[0143]
[0164] At block 516, the exemplary media unit 106 applies the selected filter(s) with smoothing to transition from the previous filter setting (e.g., the previous audio playback setting). In some examples, the EQ adjustment implementer 220 applies the selected filter(s) and transitions to the new playback setting based on the smoothing filter indicated by the smoothing filter configurator 222. In some examples, the EQ adjustment implementer 220 can implement the EQ filter (e.g., the audio playback setting) without the smoothing filter.
[0144]
[0165] At block 518, the exemplary media unit 106 determines whether the update duration threshold is met. In some examples, the update monitor 226 determines whether the update duration threshold is met. For example, if the update duration threshold is set to 1 second, the update monitor 226 determines whether 1 second has elapsed since the previous audio playback settings were determined and implemented. In response to the update duration threshold being met, processing proceeds to block 502. Conversely, in response to the update duration threshold not being met, processing proceeds to block 520.
[0145]
[0166] At block 520, the exemplary media unit 106 determines whether dynamic audio playback setting adjustment is enabled. In response to dynamic audio playback setting adjustment being enabled, processing proceeds to block 518. Conversely, in response to dynamic audio playback setting adjustment not being enabled, processing ends.
[0146]
[0167] 6 is a flowchart depicting example machine-readable instructions 514 and / or example machine-readable instructions 1106 that may be executed to implement the media unit 106 of FIGS. 1 and 2 to personalize equalization settings. With reference to the preceding figures and associated description, the example machine-readable instructions 514 and / or example machine-readable instructions 1106 begin with the example media unit 106 accessing past personalization settings (block 602).
[0147]
[0168] At block 604, the example media unit 106 generates a personalized EQ curve or initiates a new personalized EQ curve based on past personalization settings. In some examples, the historical EQ manager 214 generates a personalized EQ curve or initiates a new personalized EQ curve based on past personalization settings.
[0148]
[0169] At block 606, the exemplary media unit 106 determines whether historical EQ is enabled. In some examples, the historical EQ manager 214 determines whether historical EQ is enabled (e.g., historical equalization is enabled). In response to historical EQ being enabled, processing proceeds to block 608. Conversely, in response to historical EQ not being enabled, processing proceeds to block 610.
[0149]
[0170] At block 608, the example media unit 106 adjusts the personalized EQ curve based on the EQ curve from the historical period. In some examples, the historical EQ manager 214 adjusts the personalized EQ curve based on the EQ curve from the historical period (e.g., past hours, past days, etc.).
[0150]
[0171] At block 610, the exemplary media unit 106 determines whether user preference data (e.g., data indicative of a user's preferences) is available. In some examples, the user input analyzer 216 determines whether the user preference data is available. For example, the user input analyzer 216 may determine user EQ preferences based on when the user presses a "like" button while listening to music. In response to the user preference data being available (e.g., availability of user preference data), processing proceeds to block 612. Conversely, in response to the user preference data not being available, processing proceeds to block 616.
[0151]
[0172] At block 612, the media example unit 106 determines the EQ parameters based on past user preference input (e.g., "like" ratings, etc.). In some examples, the user input analyzer 216 determines the EQ parameters based on past user preference input.
[0152]
[0173] At block 614, the example media unit 106 adjusts the personalized EQ curve based on past user preference input. In some examples, the user input analyzer 216 adjusts the personalized EQ curve based on EQ curves from a historical period.
[0153]
[0174] At block 616, the exemplary media unit 106 determines whether location data is available. In some examples, the device parameter analyzer 212 determines whether location data is available. In response to location data being available (e.g., location data availability), processing proceeds to block 618. Conversely, in response to location data not being available, processing proceeds to block 620.
[0154]
[0175] At block 618, the example media unit 106 adjusts the personalized EQ curve based on the location of the device. In some examples, the device parameter analyzer 212 adjusts the personalized EQ curve based on the location of the device. For example, if the device is at the gym, a different personalized EQ curve may be generated than if the device is at work.
[0155]
[0176] At block 620, the exemplary media unit 106 determines whether user identification is available. In some examples, the device parameter analyzer 212 determines whether user identification is available. In response to user identification being available (e.g., availability of user identification), processing proceeds to block 622. Conversely, in response to user identification not being available, processing proceeds to block 624.
[0156]
[0177] At block 622, the example media unit 106 adjusts the personalized EQ curve based on the user identification. For example, the device parameter analyzer 212 may determine that the media unit 106 is being used by a first user who has a historical profile (e.g., via the historical EQ manager 214) that mostly listens to rock music. In such an example, the device parameter analyzer 212 may adjust the personalized EQ curve to better suit rock music. Thus, the data stored in the historical EQ manager 214 may be filterable based on a particular user, location, app providing the audio, etc.
[0157]
[0178] At block 624, the exemplary media unit 106 determines whether the source information is available. In some examples, the device parameter analyzer 212 determines whether the source information is available. In response to the source information being available (e.g., availability of the source information), processing proceeds to block 626. Conversely, in response to the source information not being available, processing proceeds to block 628.
[0158]
[0179] At block 626, the example media unit 106 adjusts the personalized EQ curve based on the source information. In some examples, the device parameter analyzer 212 adjusts the personalized EQ curve based on the source information. For example, the source information may indicate a particular app of the mobile device (e.g., a music app, a fitness app, an audiobook app, etc.). The personalized EQ curve may be adjusted based on the source of the input media signal 202.
[0159]
[0180] At block 628, the example media unit 106 adjusts the selected EQ filter(s) to be applied to the input media signal 202 by blending the dynamically generated filter output with the personalized EQ curve. For example, weights may be applied to each of the dynamically generated curves and the personalized EQ curve (e.g., based on output from a query submitted to the EQ neural network), and an average curve may be generated and applied to the input media signal 202. This average curve thus reflects both variations between tracks as well as personal preferences. After block 628, the machine-readable instructions 514 and / or the machine-readable instructions 1106 return to block 516 of the machine-readable instructions 500 and block 1108 of the machine-readable instructions 1100, respectively.
[0160]
[0181] 7 is a flowchart representing example machine-readable instructions 700 that may be executed to implement the audio EQ engine 118 of FIG. 4 to train the EQ neural network 402 according to a first implementation. With reference to the preceding figures and associated description, the example machine-readable instructions 700 begin with the example audio EQ engine 118 accessing a library of reference audio signals (block 702). In some examples, the EQ neural network 402 accesses the library of reference audio signals. The library of reference audio signals includes audio signals for which audio playback settings have been determined (e.g., by an expert).
[0161]
[0182] At block 704, the exemplary audio EQ engine 118 accesses EQ parameters associated with the reference audio signal. In some examples, the EQ neural network 402 accesses the EQ parameters (e.g., audio playback settings) associated with the reference audio signal. For example, the EQ neural network 402 may access one or more filters, one or more gain values, frequency values, Q values, etc.
[0162]
[0183] The example audio EQ engine 118 selects a reference audio signal from the plurality of reference audio signals at block 706. In some examples, the EQ neural network 402 selects a reference audio signal from the plurality of reference audio signals.
[0163]
[0184] The example audio EQ engine 118 samples the reference audio signal at block 708. In some examples, the EQ neural network 402 samples the reference audio signal by creating a predetermined number of samples (e.g., 300, 500, etc.) from the audio signal.
[0164]
[0185] At block 710, the example audio EQ engine 118 associates samples of the reference audio signal with EQ parameters (e.g., audio playback settings) corresponding to the reference audio signal. In some examples, the EQ neural network 402 associates samples of the reference audio signal with EQ parameters corresponding to the reference audio signal.
[0165]
[0186] At block 712, the example audio EQ engine 118 determines whether there are additional reference audio signals to use for training. In some examples, the EQ neural network 402 determines whether there are additional reference audio signals to use for training. In response to there being additional reference audio signals to use for training, processing proceeds to block 706. Conversely, in response to there not being additional reference audio signals to use for training, processing ends.
[0166]
[0187] 8A illustrates a first spectrogram 800a of an audio signal that has undergone dynamic audio playback setting adjustment based on real-time analysis of audio characteristics, but without a smoothing filter. The first spectrogram 800a illustrates frequency values in Hertz on a horizontal axis 802 (e.g., x-axis) and time values of the audio signal in seconds on a vertical axis 804 (e.g., y-axis). The shading in the first spectrogram 800a represents the amplitude of the audio signal at particular frequencies and times for the audio signal. The shading in the first spectrogram 800a illustrates sharp transitions between audio signal amplitudes at multiple frequencies. For example, the shading in the first spectrogram 800a transitions sharply between light and dark shading within individual frequency bands, due, at least in part, to transitions between audio playback settings implemented by the dynamic audio playback setting adjustment techniques discussed herein, which are implemented without a smoothing filter.
[0167]
[0188] 8B is a first plot 800b showing average gain values versus frequency values for the first spectrogram 800a of FIG. 8A. The first plot 800b includes frequency values in Hertz on a horizontal axis 806 (e.g., x-axis) and average gain values in decibels on a vertical axis 808 (e.g., y-axis). A comparison of the first plot 800b, which represents average gain values for audio signals whose audio playback settings have been adjusted without smoothing, and the second plot 900b, which represents average gain values for audio signals whose audio playback settings have been adjusted with smoothing, illustrates the benefit of applying a smoothing filter when transitioning between audio playback settings.
[0168]
[0189] FIG. 9A is a second spectrogram 900a of an audio signal that has undergone dynamic audio playback setting adjustments based on real-time analysis of audio characteristics, including a smoothing filter. The second spectrogram 900a includes frequency values in Hertz on a horizontal axis 902 (e.g., x-axis) and time values in seconds on a vertical axis 904 (e.g., y-axis). The second spectrogram 900a corresponds to the original input audio signal of the first spectrogram 800a (FIG. 8A), but a smoothing filter was utilized when applying audio playback settings across the entire track. Compared to the first spectrogram 800a, the second spectrogram 900a exhibits smoother (e.g., gradual) transitions between audio signal amplitudes at multiple frequencies. For example, the shading of the first spectrogram 800a smoothly transitions between lighter and darker shading within individual frequency bands, rather than the relatively abrupt transitions exhibited in the first spectrogram 800a of FIG. 8A.
[0169]
[0190] FIG. 9B is a second plot 900b illustrating average gain values versus frequency values in the second spectrogram 900a of FIG. 9A. The second plot 900b includes frequency values in Hertz on a horizontal axis 906 (e.g., x-axis) and average gain values in decibels on a vertical axis 908 (e.g., y-axis). Compared to the first plot 800b of FIG. 8B, the second plot 900b shows smoother transitions between average gain values across multiple frequency bands. For example, the multiple abrupt transitions in average gain values visible around 77 Hz in the first plot 800b are absent in the second plot 900b; instead, the second plot 900b shows a gradual, smooth decrease in average gain values around 77 Hz. Second Implementation: Profile-Based Equalization
[0170]
[0191] 1 and 3 to deliver profile information (e.g., one or more profiles 229) along with a stream of content (e.g., input media signal 202) to a playback device. As described herein, in some examples, the content profile engine 116 determines and / or generates one or more profiles 229 for the input media signal 202 to be delivered to, among other things, the media device 102, the media device 104, and / or the media unit 106. With reference to the preceding figures and associated description, the example machine-readable instructions 1000 begin when the content profile engine 116 accesses a stream of content to be delivered to the playback device (block 1002). For example, the content retriever 302 may access the input media signal 202 from a content provider 114, which is supplying the input media signal 202 to the playback device via the network 112. As another example, the content retriever 302 may access an input media signal 202 (e.g., a stream of content) from a content provider 114 that is stored locally by the playback device. As described herein, the content retriever 302 may access various types of content streams, such as an audio content stream, a video stream, etc. For example, the content retriever 302 may access a stream of songs or other music, a stream of voice content, a podcast, etc.
[0171]
[0192] At block 1004, the content profile engine 116 identifies a portion of the input media signal 202 (e.g., a portion of content within a stream of content) to be delivered to a playback device. For example, the content identifier 306 can identify the portion of the input media signal 202 using a variety of processes, including comparing a fingerprint for the content against a set of reference fingerprints associated with known content, such as reference fingerprints generated by the reference fingerprint generator 227. Of course, the content identifier 306 can also identify a piece of content using other information, such as metadata associated with the piece of content (e.g., information identifying an associated title, artist, genre, etc.), information associated with the content provider 114, etc.
[0172]
[0193] In some examples, the content identifier 306 may identify a certain category type or genre to be associated with a portion (e.g., a piece of content) of the input media signal 202. For example, instead of identifying the input media signal 202 as a particular piece of content (e.g., a particular song, YouTube™ video / clip, TV show, movie, podcast, etc.), the content identifier 306 may use techniques described herein to identify a genre or category that applies to the portion (e.g., a piece of content) of the input media signal 202.
[0173]
[0194] At block 1006, the content profile engine 116 determines a profile for the identified piece of content. For example, the profiler 308 may determine one or more characteristics for an entire piece of content and / or may determine one or more characteristics for multiple portions of a portion of the input media signal 202 (e.g., a piece of content), such as a frame or block of frames of content. For example, the one or more profiles 229 may include a first set of one or more characteristics for a first portion of the input media signal 202 (e.g., a piece of content), a second set of one or more characteristics for a second portion of the input media signal 202 (e.g., a piece of content), etc.
[0174]
[0195] In some examples, the profiler 308 may render, generate, create, and / or determine one or more profiles 229 for an input media signal 202 (e.g., a piece of content), such as audio content, having a variety of different characteristics. For example, the determined or generated one or more profiles 229 may include characteristics associated with equalization (EQ) settings, spatialization settings, virtualization settings, video settings, etc.
[0175]
[0196] At block 1008, the content profile engine 116 delivers the one or more profiles 229 to the playback device. For example, the profiler 308 can deliver the one or more profiles 229 to the playback device over the network 112 or over other communication channels.
[0176]
[0197] For example, the content profile engine 116 can access a piece of content that is a song to be streamed to a playback device that is a car stereo, identify the song as a particular song associated with the genre of "classical music," determine a profile that includes a set of equalization settings (e.g., signal strength indicators for various frequencies in the song, speaker spatialization settings, etc.) to use when playing the song through the car stereo, and deliver the profile to the car stereo to be consumed by a network associated with the car stereo, such as a car area network (CAN), which controls the operation of the car stereo.
[0177]
[0198] In another example, the content profile engine 116 can access a piece of content, a movie, to be streamed over a broadcast network or the Internet to a playback device that is a TV set or set-top box, identify the movie as a particular movie associated with the genre of "action" and having many fast-paced action sequences, determine a profile that includes a set of image processing settings (e.g., color palette settings, frame rate upscaling settings, contrast enhancement settings for low-contrast scenes, etc.) to use when playing the movie through the TV set or other device, and deliver the profile to the TV set or other device to adjust the rendering and therefore the content experience by the user.
[0178]
[0199] 1 and 2 to play content using modified playback settings. As described herein, in some examples, the media unit 106 modifies or adjusts playback of content by, among other things, a playback device (e.g., media device 102, media device 104, and / or media unit 106). With reference to the preceding figures and associated description, the example machine-readable instructions 1100 begin when the media unit 106 receives and / or accesses a stream of content for a playback device or a stream of content associated with a playback device (block 1102). For example, the media unit 106, and / or more specifically the synchronizer 228, may access an input media signal 202 (e.g., a content stream) to be played by the playback device.
[0179]
[0200] At block 1104, the media unit 106 accesses profile information associated with the stream of content. For example, the media unit 106, and more specifically the synchronizer 228, may receive a profile or profile information generated by the content profile engine 116. As described herein, the content profile engine 116 may determine a profile by identifying the stream of content based on a comparison of a fingerprint associated with the stream of content to a set of fingerprints associated with known content, and may select or determine one or more profiles 229 to be associated with the identified input media signal 202 (e.g., the stream of content).
[0180]
[0201] The one or more profiles 229 may include various types of information, such as information identifying a category or genre associated with the song, information identifying a mood associated with the song, such as an upbeat mood, a relaxed mood, a soft mood, etc., information identifying signal strength parameters for various frequencies within the content, such as low frequencies for bass and other similar tones, high frequencies for voice or singing tones, prosodic information and / or linguistic information obtained from the audio content.
[0181]
[0202] Additionally or alternatively, one or more profiles 229 may include information identifying a category or genre associated with a video or segment of a video clip, information identifying a mood associated with the video, information identifying brightness, color palette, color contrast, brightness range, blur, display format, video scene information, information obtained from visual object detection and / or recognition, face detection and / or recognition, or broadcast logo detection and / or recognition algorithms, the presence and / or content of text or subtitles, the presence and / or content of watermarks, etc.
[0182]
[0203] In block 1106, the media unit 106 personalizes the equalization settings. In some examples, the EQ personalization manager 210 personalizes the equalization settings. Detailed instructions for personalizing the equalization settings are shown and described in connection with FIG. 6.
[0183]
[0204] At block 1108, the media unit 106 modifies playback of the input media signal 202 (e.g., a stream of content) based on the accessed profile information and / or the personalized EQ profile generated at block 1106. For example, the EQ adjustment implementer 220 may modify playback of the input media signal 202 on a playback device based on the one or more profiles 229 and a blended equalization generated based on the personalized EQ profile. In another example, the EQ adjustment implementer 220 may apply information in the one or more profiles 229 to modify or adjust equalizer settings of a playback device to adjust and / or tailor equalization during playback of the input media signal 202 (e.g., a stream of content). In addition to equalization, the EQ adjustment implementer 220 may adjust various different playback settings, such as virtualization settings, spatialization settings, etc.
[0184]
[0205] In some examples, the media unit 106 can access a profile that includes multiple settings related to different portions of the content. For example, a song may include portions with different tempos, and a corresponding profile generated for the song may include, among other things, a first portion with a "slow" setting, a second portion with a "fast" setting, and a third portion with a "slow" setting. The media unit 106 can receive the profile from a different platform than the playback device and can synchronize the profile to the song in order to precisely adjust playback settings using the multiple settings included in the profile.
[0185]
[0206] 1 and 2 to adjust playback settings based on profile information associated with content. For example, the media unit 106 may adjust playback settings based on profile information associated with content, according to some examples. With reference to the preceding figures and associated description, the example machine-readable instructions 1200 begin when the media unit 106 accesses one or more profiles 229 for the input media signal 202 (e.g., a piece of content) (block 1202). For example, the media unit 106, and / or more specifically the synchronizer 228, may have access to various types of profiles, such as a single settings profile, multiple settings profiles, etc.
[0186]
[0207] At block 1204, the media unit 106 synchronizes one or more profiles 229 to the input media signal 202 (e.g., a piece of content). For example, the synchronizer 228 may utilize fingerprint(s) associated with the input media signal 202 (e.g., a piece of content) to synchronize the input media signal 202 (e.g., a piece of content) to one or more profiles 229. The one or more profiles 229 may include information relating one or more settings to known fingerprints for the piece of content and aligning the settings to a portion of the input media signal 202 (e.g., a piece of content) in order to synchronize the one or more profiles 229 to the piece of content during playback of the input media signal 202. As another example, the synchronizer 228 can identify various audio events within a piece of content (e.g., a snare hit, the start of a guitar solo, the first vocals) and align one or more profiles 229 to the events within the input media signal 202 in order to synchronize the one or more profiles 229 to the piece of content during playback of the input media signal 202.
[0187]
[0208] At block 1206, the media unit 106 modifies playback of the input media signal 202 using a playback device (e.g., media device 102, media device 104, media unit 106, etc.) based on the synchronized profile for the input media signal 202. For example, the EQ adjustment implementer 220 may apply information in one or more profiles 229 to modify or adjust equalizer settings of the playback device to adjust and / or tailor equalization during playback of the input media signal 202 (e.g., a stream of content). Similarly, when the content is video, the one or more profiles 229 may be used to adjust video-related settings.
[0188]
[0209] 13A-13B are block diagrams of example content profiles according to the teachings of the present disclosure. FIG. 13A illustrates a content profile 1300a that includes a single setting 1302, i.e., "Mood #1," for the entire content. Meanwhile, FIG. 13B illustrates a content profile 1300b that includes multiple different settings for a piece of content. For example, content profile 1300b includes, among other settings, a first setting 1304 (e.g., "Mood #1"), a second setting 1306 (e.g., "Mood #2"), a third setting 1308 (e.g., "Mood #3"), and a fourth setting 1310 (e.g., "Mood #4"). Thus, in some examples, a media unit 106 can utilize a composite or multi-tiered profile that includes different settings to apply to different portions of content in order to, among other things, dynamically adjust the content playback experience at various times during playback of the content.
[0189]
[0210] Thus, the systems and methods described herein can provide a platform that facilitates real-time or near-real-time processing and delivery of profile information (e.g., content profiles) to playback devices, which utilize the content profiles to, among other things, adjust the playback experience (e.g., video and / or audio experience) associated with playing the content to a user. This may involve buffering the content before rendering it until the profile can be retrieved or predicted. As an example, a particular profile may be applied based on usage history (e.g., a user has consumed a particular content type associated with a particular profile at this time of day / week over the past few days / weeks, and therefore the same profile is applied again after determining the usage pattern). In another example, a user may have previously established a preference for a particular profile with a particular type of content (e.g., video clips categorized as TV dramas), and therefore, moving forward with the content, the profile is automatically applied for content of the same or similar types. Another method of predicting a profile for a user may be by applying collaborative filtering methods, where other users' profiles are inferred for a particular user based on usage patterns, demographic information, or any other information about the user or user group. Yet another example is to include content source settings, e.g., device settings such as the selected input on a TV set, such as between the input that connects to a set-top box and the input that connects to a DVD player or game console, to determine or influence profile selection.
[0190]
[0211] Many playback devices can utilize such a platform, including: (1) car stereo systems that receive and play content from online, satellite, or terrestrial radio stations, and / or from locally stored content players (e.g., CD players, MP3 players, etc.); (2) home stereo systems that receive and play content from online, satellite, or terrestrial radio stations, and / or from locally stored content players (e.g., CD players, MP3 players, TV sets, set-top boxes (STBs), game consoles, etc.); and (3) mobile devices (e.g., smartphones or tablets) that receive and play content (e.g., video and / or audio) from online, satellite, or terrestrial radio stations, and / or from locally stored content players (e.g., MP3 players).
[0191]
[0212] In some examples, systems and methods can improve and / or optimize low-quality or low-volume recordings and other content. For example, the content profile engine 116 can identify a stream of content (e.g., a homemade podcast) as having low audio quality and generate a profile for the low-quality stream of content that includes instructions for boosting playback of the content. The media unit 106 can then adjust playback settings of the playback device (e.g., mobile device, media device 102, media device 104, media unit 106) to, among other things, boost the fidelity of playback of the low-quality content.
[0192]
[0213] In some examples, the systems and methods can reduce the quality of certain types of content, such as advertisements, in a content stream. For example, the content profile engine 116 can identify that the content stream includes commercial breaks and generate a profile for the content stream that reduces playback quality during the commercial breaks. The media unit 106 can then adjust playback settings of the playback device (e.g., mobile device, media device 102, media device 104, media unit 106) to, among other things, reduce the fidelity of the content playback during the commercial breaks. Of course, other scenarios may be possible. Third Implementation: Thresholding-Based Equalization
[0193]
[0214] 1 and 2 to perform audio equalization according to a third implementation. With reference to the preceding figures and associated description, the exemplary machine-readable instructions 1400 begin with the exemplary media unit 106 storing the input media signal 202 in a buffer (block 1402). In some examples, the exemplary buffer manager 230 stores the input media signal 202 in the data store 224. In some examples, the buffer manager 230 removes any portion of the input media signal 202 that exceeds its storage duration in the buffer (e.g., 10 seconds, 30 seconds, etc.).
[0194]
[0215] At block 1404, the exemplary media unit 106 performs a frequency transform on the buffered audio. In some examples, the time-to-frequency domain converter 232 performs a frequency transform (e.g., an FFT) on the portion of the input media signal 202 in the buffer.
[0195]
[0216] At block 1406, the exemplary media unit 106 calculates mean and standard deviation values for the linear spatial frequency bins across the duration of the buffer. In some examples, the loudness calculator 234 calculates mean and standard deviation values for the linear spatial frequency bins across the duration of the buffer. In some examples, the mean loudness values can be calculated in different domains (e.g., the time domain) or at different unit intervals (e.g., logarithmic intervals).
[0196]
[0217] At block 1408, the exemplary media unit 106 calculates a pre-equalization RMS value based on the frequency representation of the input media signal 202. In some examples, the energy calculator 236 calculates the pre-equalization RMS value based on the frequency representation of the input media signal 202. In some examples, the energy calculator 236 utilizes different types of calculations to determine the energy value of the input media signal 202.
[0197]
[0218] At block 1410, the exemplary media unit 106 inputs the mean and standard deviation values for the linear spatial bins along with a representation of the engineer tag to the EQ neural network 402. In some examples, the input feature set generator 238 inputs the mean and standard deviation values for the linear spatial frequency bins across the entire duration of the buffer to the EQ neural network 402. For non-reference or unidentified audio, the engineer tag is set to a particular value within a set of possible values. For example, when the audio is unidentified, the input feature set generator 238 may be configured so that the engineer tag is always set to a particular engineer designation. In some examples, the engineer tag is represented as a vector, with one of the vector elements set to “1” for the selected engineer and the remaining vector elements set to “0.” In some examples, the mean and / or standard deviation values for the input media signal 202 may be input to the EQ neural network 402 in another format (e.g., time-domain format, instantaneous volume rather than average, etc.).
[0198]
[0219] At block 1412, the example media unit 106 receives gain / cut values for the logarithmic spatial frequency bins from the EQ neural network 402. In some examples, the volume control 242 receives the gain / cut values for the logarithmic spatial frequency bins from the EQ neural network 402. In some examples, the gain / cut values may be in a linear spatial frequency representation and / or another domain.
[0199]
[0220] At block 1414, the exemplary media unit 106 converts the linear spatial average frequency representation of the input media signal 202 to a logarithmic spatial average frequency representation. In some examples, the volume adjuster 242 converts the linear spatial average frequency representation of the input media signal 202 to a logarithmic spatial average frequency representation in order to apply the EQ gain / cut 241 received in a logarithmic spatial format. In some examples, if the EQ gain / cut 241 is received in a different format, the volume adjuster 242 adjusts the average frequency representation of the input media signal 202 to correspond to the same format as the EQ gain / cut 241.
[0200]
[0221] At block 1416, the exemplary media unit 106 applies a gain / cut to the log-spatial average frequency representation to determine an equalized log-spatial average frequency representation. In some examples, the volume control 242 applies a gain / cut to the log-spatial average frequency representation to determine an equalized log-spatial average frequency representation of the input media signal 202. As with all steps of the machine-readable instructions 1400, in some examples, applying a gain / cut to the average representation of the incoming audio signal may occur in different regions and / or at different unit intervals.
[0201]
[0222] At block 1418, the exemplary media unit 106 performs thresholding to smooth the equalization curve. In some examples, the thresholding controller 244 performs thresholding to smooth the equalization curve. Detailed instructions for performing thresholding to smooth the equalization curve are shown and described in connection with FIG. 15.
[0202]
[0223] At block 1420, the exemplary media unit 106 calculates the equalized RMS value. In some examples, the energy calculator 236 calculates the equalized RMS value based on the equalized audio signal after the thresholding controller 244 finishes smoothing the equalization curve (e.g., after reducing irregularities). In some examples, the energy calculator 236 calculates another measure of the energy of the equalized audio signal. In some examples, the energy calculator 236 calculates the equalized RMS value after the EQ curve generator 246 generates and applies a final equalization curve (e.g., in a linear spatial frequency representation) to the input media signal 202.
[0203]
[0224] At block 1422, the exemplary media unit 106 determines a loudness normalization based on the calculation of the pre-equalization RMS and the post-equalization RMS. In some examples, the energy calculator 236 calculates a ratio (or other comparison metric) of the post-equalization RMS to the pre-equalization RMS, and the loudness normalizer 248 determines whether this ratio exceeds a threshold associated with a maximum allowable change in the energy of the audio signal (e.g., associated with a tolerable change). In some such examples, in response to the ratio exceeding the threshold, the loudness normalizer 248 applies a normalized overall gain to the post-equalization audio signal. For example, if the total energy of the post-equalization audio signal is twice the total energy before equalization, the loudness normalizer 248 may apply an overall gain of ½ to normalize the overall loudness of the audio signal.
[0204]
[0225] In block 1424, the exemplary media unit 106 subtracts the average frequency representation from the equalized log-spatial frequency representation of the audio signal to determine a final equalization curve. In some examples, the EQ curve generator 246 subtracts the average frequency representation from the equalized log-spatial frequency representation of the audio signal to determine a final equalization curve.
[0205]
[0226] At block 1426, the exemplary media unit 106 applies a final equalization curve to the linear spatial frequency representation of the input media signal 202. In some examples, the EQ curve generator 246 applies the final equalization curve and also makes any overall gain adjustments indicated by the volume normalizer 248. In some examples, the volume normalizer 248 can perform volume normalization before or after the EQ curve generator 246 applies the final equalization curve.
[0206]
[0227] At block 1428, the exemplary media unit 106 performs an inverse frequency transform on the equalized frequency representation of the input media signal 202. In some examples, the frequency-to-time domain converter 250 performs an inverse frequency transform on the equalized frequency representation of the input media signal 202 to generate the output media signal 252.
[0207]
[0228] At block 1430, the exemplary media unit 106 determines whether to continue equalization. In response to continuing equalization, processing proceeds to block 1402. Conversely, in response to not continuing equalization, processing ends.
[0208]
[0229] 1 and 2 to smooth the equalization curve according to a third implementation. With reference to the preceding figures and associated description, the exemplary machine-readable instructions 1500 begin with the exemplary media unit 106 selecting multiple frequency values (block 1502). In some examples, the thresholding controller 244 selects multiple frequency values to analyze for irregular changes in volume (e.g., local outliers). In some examples, the thresholding controller 244 selects a set of adjacent frequency values (e.g., three distinct, consequential frequency values) to analyze at a time.
[0209]
[0230] The exemplary media unit 106 determines the volume at multiple frequency values at block 1504. In some examples, the thresholding controller 244 determines the volume at multiple frequency values.
[0210]
[0231] At block 1506, the exemplary media unit 106 determines a second derivative of the volume across the plurality of frequency values. In some examples, the thresholding controller 244 determines the second derivative of the volume across the plurality of frequency values. In some examples, the thresholding controller 244 utilizes another technique to determine the amount of change in volume across the plurality of frequency values. One exemplary technique for determining the second derivative of the volume across the plurality of frequency values includes utilizing Equation 1, discussed in connection with FIG. 2 earlier in this description.
[0211]
[0232] At block 1508, the exemplary media unit 106 determines whether the absolute value of the second derivative exceeds a threshold. In some examples, the thresholding controller 244 determines whether the absolute value of the second derivative exceeds a threshold. In some examples, the thresholding controller 244 compares another calculation of the amount of change in volume across multiple frequency values to a threshold. In response to the absolute value of the second derivative exceeding the threshold, processing proceeds to block 1510. Conversely, in response to the absolute value of the second derivative not exceeding the threshold, processing proceeds to block 1512.
[0212]
[0233] At block 1510, the exemplary media unit 106 adjusts the volume level of the center value of the plurality of values to be the midpoint between the volume levels at adjacent frequency values. In some examples, the thresholding controller 244 adjusts the volume level of the center value of the plurality of values to be the midpoint between the volume levels at adjacent frequency values. In some examples, the thresholding controller 244 uses another method to adjust the center value of the plurality of values to be more similar to the volume at adjacent frequency values, thereby reducing irregularities in the equalization curve.
[0213]
[0234] At block 1512, the exemplary media unit 106 determines whether there are any additional frequency values to analyze. In some examples, the thresholding controller 244 determines whether there are any additional frequency values to analyze. In some examples, the thresholding controller 244 iterates analyzing all of the frequency values one or more times. In some examples, the thresholding controller 244 iterates until all irregularities are removed or until only a threshold number of irregularities remain. In response to there being additional frequency values to analyze, processing proceeds to block 1502. Conversely, in response to there not being additional frequency values to analyze, processing returns to the machine-readable instructions of FIG. 14 and proceeds to block 1420.
[0214]
[0235] 16 is a flowchart representing example machine-readable instructions 1600 that may be executed to implement the audio EQ engine 118 of FIG. 4 to collect a data set and train and / or validate a neural network based on reference audio signals, according to a third implementation. With reference to the preceding figures and associated description, the example machine-readable instructions 1600 begin with the example audio EQ engine 118 accessing a library of reference audio signals (block 1602). In some examples, the EQ neural network 402 accesses the library of reference audio signals.
[0215]
[0236] At block 1604, the example audio EQ engine 118 accesses an equalization curve associated with the reference audio signal. In some examples, the EQ neural network 402 accesses an equalization curve associated with the reference audio signal.
[0216]
[0237] At block 1606, the example audio EQ engine 118 accesses engineer tags and / or other metadata associated with the reference audio signal. In some examples, the EQ neural network 402 accesses engineer tags and / or other metadata associated with the reference audio signal.
[0217]
[0238] At block 1608, the example audio EQ engine 118 associates the samples of the reference audio signal with corresponding EQ curves and engineer tag(s). In some examples, the EQ neural network 402 associates the samples of the reference audio signal with corresponding EQ curves and engineer tag(s).
[0218]
[0239] At block 1610, the example audio EQ engine 118 determines whether there are additional reference audio signals to use for training. In some examples, the EQ neural network 402 determines whether there are additional reference audio signals, EQ curves, or engineer tags to utilize for training. In response to there being additional reference audio signals to train, processing proceeds to block 1602. Conversely, in response to there not being additional reference audio signals to use for training, processing ends.
[0219]
[0240] FIG. 17A is an exemplary first plot 1700a of an equalized audio signal before implementing the smoothing technique shown and described in connection with FIG.
[0220]
[0241] The exemplary first plot 1700a includes an exemplary frequency axis 1702 that indicates frequency values that increase from left to right (e.g., along the x-axis). The first plot 1700a includes an exemplary volume axis 1704 that indicates volume values that increase from bottom to top (e.g., along the y-axis). In general, the first plot 1700a shows that the audio signal has higher volume levels at lower frequency values, and that the volume generally increases as the frequency value increases. However, the first plot 1700a includes an exemplary irregularity 1706.
[0221]
[0242] The first plot 1700a includes an exemplary first frequency value 1708, an exemplary second frequency value 1710, and an exemplary third frequency value 1712. The first frequency value 1708 corresponds to an exemplary first volume 1714, the second frequency value 1710 corresponds to an exemplary second volume 1716, and the third frequency value 1712 corresponds to an exemplary third volume 1718. When the media unit 106 attempts to perform a thresholding procedure on the signal shown in the first plot 1700a (e.g., via the thresholding controller 244), the media unit 106 can detect irregularities 1706 (e.g., local outliers) because the volume changes significantly between the first frequency value 1708 and the second frequency value 1710, and between the second frequency value 1710 and the third frequency value 1712. If the thresholding controller 244 calculates the second derivative of the volume (or other measure of the volume change) between the volume levels at the first frequency value 1708, the second frequency value 1710, and the third frequency value 1712, the thresholding controller 244 can determine that the second derivative exceeds a threshold and corresponds to the irregularity 1706.
[0222]
[0243] FIG. 17B is an exemplary second plot 1700b of the audio signal of FIG. 17A after performing the smoothing technique shown and described in connection with FIG. 15. In the illustrated example of FIG. 17B, after detecting the irregularity 1706, the thresholding controller 244 adjusts the volume level associated with a second frequency value 1710 (e.g., the center value of the three frequency values being analyzed). The second plot 1700b of FIG. 17B is substantially identical to the first plot 1700a, except that the second frequency value 1710 corresponds to an exemplary fourth volume 1720 rather than the previous second volume 1716. In the illustrated example, the thresholding controller 244 adjusted the second volume 1716 to be the fourth volume 1720 by setting the volume at the second frequency value 1710 to the midpoint between the first volume 1714 and the third volume 1718. In the illustrated example, the remaining portion of the equalization curve between these frequency values is then generated as a smooth line. In the illustrated example of Figure 17B, the adjusted portion of the equalization curve is shown as a dashed line connecting the first volume 1714, the fourth volume 1720, and the third volume 1718. The thresholding controller 244 may utilize any other technique for adjusting the volume levels at the detected irregularities.
[0223]
[0244] Figure 18 is a block diagram of an exemplary processor platform 1800 configured to execute the instructions of Figures 5, 6, 11, 12, 14, and 15 to implement the media unit 106 of Figures 1 and 2. The processor platform 1800 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smartphone, a tablet such as an iPad®), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a gaming console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0224]
[0245] The illustrated example processor platform 1800 includes a processor 1812. The illustrated example processor 1812 is hardware. For example, the processor 1812 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from a desired family or manufacturer. The hardware processor 1812 may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1812 includes an example signal converter 204, an example EQ model query generator 206, an example EQ filter setting analyzer 208, an example EQ personalization manager 210, an example device parameter analyzer 212, an example history EQ manager 214, an example user input analyzer 216, an example EQ filter selector 218, an example EQ adjustment implementer 220, an example smoothing filter configurator 222, an example data store 224, an example update monitor 226 ... 26, an exemplary fingerprint generator 227, an exemplary synchronizer 228, an exemplary buffer manager 230, an exemplary time-to-frequency domain converter 232, an exemplary volume calculator 234, an exemplary energy calculator 236, an exemplary input feature set generator 238, an exemplary EQ manager 240, an exemplary volume adjuster 242, an exemplary thresholding controller 244, an exemplary EQ curve generator 246, an exemplary volume normalizer 248, and / or an exemplary frequency-to-time domain converter 250.
[0225]
[0246] The processor 1812 of the illustrated example includes a local memory 1813 (e.g., a cache). The processor 1812 of the illustrated example communicates with a main memory, including a volatile memory 1814 and a nonvolatile memory 1816, via a bus 1818. The volatile memory 1814 may be implemented with synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The nonvolatile memory 1816 may be implemented with flash memory and / or any other desired type of memory device. Access to the main memory 1814, 1816 is controlled by a memory controller.
[0226]
[0247] The processor platform 1800 of the illustrated example also includes an interface circuit 1820. The interface circuit 1820 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI express interface.
[0227]
[0248] In the illustrated example, one or more input devices 1822 are connected to the interface circuit 1820. The input device(s) 1822 allow a user to input data and / or commands to the processor 1812. For example, the input device(s) may be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0228]
[0249] The interface circuit 1820 of the illustrated example is also connected to one or more output devices 1824. For example, the output device(s) 1824 may be implemented by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Thus, the interface circuit 1820 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0229]
[0250] The interface circuitry 1820 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate the exchange of data with external machines (e.g., computing devices of any type) over a network 1826. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, etc.
[0230]
[0251] The processor platform 1800 of the illustrated example also includes one or more mass storage devices 1828 for storing software and / or data. Examples of such mass storage devices 1828 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0231]
[0252] The machine-readable instructions 1832 of FIG. 18, the machine-readable instructions 500 of FIG. 5, the machine-readable instructions 514 of FIG. 6, the machine-readable instructions 1100 of FIG. 11, the machine-readable instructions 1106 of FIG. 6, the machine-readable instructions 1200 of FIG. 12, the machine-readable instructions 1400 of FIG. 14 and / or the machine-readable instructions 1418 of FIG. 15 may be stored on a mass storage device 1828, a volatile memory 1814, a non-volatile memory 1816, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0232]
[0253] Figure 19 is a block diagram of an exemplary processor platform 1900 configured to execute the instructions of Figures 7 and 16 to implement the audio EQ engine 118 of Figures 1 and 4. The processor platform 1900 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., neural network), a mobile device (e.g., a cell phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a gaming console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0233]
[0254] The processor platform 1900 of the illustrated example includes a processor 1912. The processor 1912 of the illustrated example is hardware. For example, the processor 1912 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from a desired family or manufacturer. The hardware processor 1912 may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1912 implements the example EQ neural network 402, the example audio EQ scoring engine 404, and / or the example audio EQ engine validator 406.
[0234]
[0255] The processor 1912 of the illustrated example includes a local memory 1913 (e.g., a cache). The processor 1912 of the illustrated example communicates with a main memory, including a volatile memory 1914 and a non-volatile memory 1916, via a bus 1918. The volatile memory 1914 may be implemented with synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 1916 may be implemented with flash memory and / or any other desired type of memory device. Access to the main memory 1914, 1916 is controlled by a memory controller.
[0235]
[0256] The processor platform 1900 of the illustrated example also includes an interface circuit 1920. The interface circuit 1920 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI express interface.
[0236]
[0257] In the illustrated example, one or more input devices 1922 are connected to the interface circuit 1920. The input device(s) 1922 allow a user to input data and / or commands to the processor 1912. For example, the input device(s) may be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0237]
[0258] The interface circuit 1920 of the illustrated example is also coupled to one or more output devices 1924. For example, the output device(s) 1924 may be implemented by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Thus, the interface circuit 1920 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0238]
[0259] The interface circuitry 1920 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate the exchange of data with external machines (e.g., computing devices of any type) over a network 1926. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, etc.
[0239]
[0260] The processor platform 1900 of the illustrated example also includes one or more mass storage devices 1928 for storing software and / or data. Examples of such mass storage devices 1928 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0240]
[0261] The machine-readable instructions 1932 of FIG. 19, the machine-readable instructions 700 of FIG. 7, and / or the machine-readable instructions 1600 of FIG. 16 may be stored on mass storage device 1928, volatile memory 1914, non-volatile memory 1916, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0241]
[0262] Figure 20 is a block diagram of an exemplary processor platform 2000 configured to execute the instructions of Figure 10 to implement the content profile engine 116 of Figures 1 and 3. The processor platform 2000 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a gaming console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0242]
[0263] The illustrated example processor platform 2000 includes a processor 2012. The illustrated example processor 2012 is hardware. For example, the processor 2012 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from a desired family or manufacturer. The hardware processor 2012 can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 2012 implements an example content retriever 302, an example fingerprint generator 304, an example content identifier 306, an example profiler 308, and / or an example profile data store 310.
[0243]
[0264] The processor 2012 of the illustrated example includes a local memory 2013 (e.g., a cache). The processor 2012 of the illustrated example communicates with a main memory, including a volatile memory 2014 and a non-volatile memory 2016, via a bus 2018. The volatile memory 2014 may be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus® Dynamic Random Access Memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 2016 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 2014, 2016 is controlled by a memory controller.
[0244]
[0265] The processor platform 2000 of the illustrated example also includes an interface circuit 2020. The interface circuit 2020 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI express interface.
[0245]
[0266] In the illustrated example, one or more input devices 2022 are connected to the interface circuit 2020. The input device(s) 2022 allow a user to input data and / or commands to the processor 2012. For example, the input device(s) may be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0246]
[0267] Also connected to the interface circuitry 2020 of the illustrated example are one or more output devices 2024. For example, the output device(s) 2024 may be implemented by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Thus, the interface circuitry 2020 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0247]
[0268] The interface circuitry 2020 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate the exchange of data with external machines (e.g., computing devices of any type) over a network 2026. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, etc.
[0248]
[0269] The processor platform 2000 of the illustrated example also includes one or more mass storage devices 2028 for storing software and / or data. Examples of such mass storage devices 2028 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0249]
[0270] The machine-readable instructions 2032 of FIG. 20 and / or the machine-readable instructions 1000 of FIG. 10 may be stored on mass storage device 2028, volatile memory 2014, non-volatile memory 2016, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0250]
[0271] From the foregoing, it should be appreciated that exemplary methods, apparatus, and articles of manufacture have been disclosed for dynamically adjusting audio playback settings to adapt to individual track changes, changes between tracks, changes in genre, and / or any other changes in the audio signal by analyzing the audio signal and utilizing a neural network to determine optimal audio playback settings. Additionally, exemplary methods, apparatus, and articles of manufacture have been disclosed for utilizing a smoothing filter to intelligently adjust audio playback settings without perceptible sharp shifts in volume level or equalization settings. Furthermore, the techniques disclosed herein enable dynamic adjustment between tracks as well as equalization techniques that synchronize user preferences (represented by personalized EQ profiles).
[0251]
[0272] Additionally, exemplary methods, apparatus, and articles of manufacture disclosed herein intelligently equalize audio signals to compensate for differences in source and / or other characteristics of the audio signals (e.g., genre, instruments present, etc.). The exemplary techniques disclosed herein utilize a neural network trained on a reference audio signal that has been equalized by an audio engineer and input to the neural network along with the instructions of the particular audio engineer who equalized the reference audio signal. The use of such training allows the neural network to provide an expert equalization output, allowing for subtle adjustments both between different tracks and even within the same track. Furthermore, the exemplary techniques disclosed herein refine the equalization output of the neural network by implementing thresholding techniques to ensure that the final equalization curve applied to the incoming audio signal is smooth and has minimal irregularities perceptible to a listener.
[0252]
[0273]
[0013] Exemplary methods, apparatus, systems, and articles of manufacture are disclosed herein for adjusting audio playback settings based on an analysis of audio characteristics. Further embodiments and combinations thereof include:
[0253]
[0274] Example 1 includes an apparatus including: an equalization (EQ) model query generator for generating a query to a neural network, the query including a representation of samples of an audio signal; an EQ filter setting analyzer for accessing a plurality of audio playback settings determined by the neural network based on the query and determining filter coefficients to apply to the audio signal based on the plurality of audio playback settings; and an EQ adjustment implementer for applying the filter coefficients to the audio signal for a first duration.
[0254]
[0275] Example 2 includes the apparatus of example 1, wherein the representation of the samples of the audio signal corresponds to a frequency representation of the samples of the audio signal.
[0255]
[0276] Example 3 includes the device of example 1, wherein the plurality of audio playback settings include one or more filters, each of the one or more filters including one or more respective gain values, respective frequency values, and respective quality factor values associated with samples of the audio signal.
[0256]
[0277] Example 4 includes the apparatus of example 1, wherein the EQ filter setting analyzer is for determining filter coefficients to apply to the audio signal based on a type of filter associated with the filter coefficients to be applied to the audio signal.
[0257]
[0278] Example 5 includes the apparatus of example 1, wherein the EQ adjustment implementer is to apply a smoothing filter to the audio signal to reduce sharp transitions in an average gain value of the audio signal between the first duration and the second duration.
[0258]
[0279] Example 6 includes the apparatus of example 1, further including a signal converter for converting the audio signal into a frequency representation of samples of the audio signal.
[0259]
[0280] Example 7 includes the apparatus of example 1, wherein the EQ adjustment implementer is for adjusting at least one of an amplitude characteristic, a frequency characteristic, and a phase characteristic of the audio signal based on the filter coefficients.
[0260]
[0281] Example 8 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least generate a query to a neural network, the query including a representation of samples of an audio signal; access a plurality of audio playback settings determined by the neural network based on the query; determine filter coefficients to apply to the audio signal based on the plurality of audio playback settings; and apply the filter coefficients to the audio signal for a first duration.
[0261]
[0282] Example 9 includes the non-transitory computer-readable storage medium of example 8, wherein the representation of the samples of the audio signal corresponds to a frequency representation of the samples of the audio signal.
[0262]
[0283] Example 10 includes the non-transitory computer-readable storage medium of Example 8, wherein the plurality of audio playback settings include one or more filters, each of the one or more filters including one or more of a respective gain value, a respective frequency value, and a respective quality factor value associated with samples of the audio signal.
[0263]
[0284] Example 11 includes the non-transitory computer-readable storage medium of example 8, wherein the instructions, when executed, cause one or more processors to determine filter coefficients to apply to the audio signal based on a type of filter associated with the filter coefficients to apply to the audio signal.
[0264]
[0285] Example 12 includes the non-transitory computer-readable storage medium of example 8, wherein the instructions, when executed, cause the one or more processors to apply a smoothing filter to the audio signal to reduce sharp transitions in the average gain value of the audio signal between the first duration and the second duration.
[0265]
[0286] Example 13 includes the non-transitory computer-readable storage medium of example 8, wherein the instructions, when executed, cause one or more processors to convert an audio signal into a frequency representation of samples of the audio signal.
[0266]
[0287] Example 14 includes the non-transitory computer-readable storage medium of example 8, wherein the instructions, when executed, cause one or more processors to adjust at least one of amplitude characteristics, frequency characteristics, and phase characteristics of the audio signal based on the filter coefficients.
[0267]
[0288] Example 15 includes a method including generating a query to a neural network, the query including a representation of samples of an audio signal; accessing a plurality of audio playback settings determined by the neural network based on the query; determining filter coefficients to apply to the audio signal based on the plurality of audio playback settings; and applying the filter coefficients to the audio signal for a first duration.
[0268]
[0289] Example 16 includes the method of example 15, wherein the representation of the samples of the audio signal corresponds to a frequency representation of the samples of the audio signal.
[0269]
[0290] Example 17 includes the method of example 15, wherein the plurality of audio playback settings include one or more filters, each of the one or more filters including one or more of a respective gain value, a respective frequency value, and a respective quality factor value associated with a sample of the audio signal.
[0270]
[0291] Example 18 includes the method of example 15, further including determining filter coefficients to apply to the audio signal based on a type of filter associated with the filter coefficients to apply to the audio signal.
[0271]
[0292] Example 19 includes the method of example 15, further including applying a smoothing filter to the audio signal to reduce sharp transitions in the average gain value of the audio signal between the first duration and the second duration.
[0272]
[0293] Example 20 includes the method of example 15, further including converting the audio signal into a frequency representation of samples of the audio signal.
[0273]
[0294] Example 21 includes an apparatus comprising: an equalization (EQ) model query generator for generating a query to a neural network, the query including a representation of a sample of the audio signal; an EQ filter setting analyzer for accessing a plurality of audio playback settings determined by the neural network based on the query and determining filter coefficients to apply to the audio signal based on the plurality of audio playback settings; an EQ personalization manager for generating personalized EQ settings; and an EQ adjustment implementer for blending the personalized EQ settings and the filter coefficients to generate blended equalization and applying the blended equalization to the audio signal for a first duration.
[0274]
[0295] Example 22 includes the device of example 21, further including a historical EQ manager for generating personalized EQ settings based on past personalization settings and adjusting the personalized EQ settings based on EQ settings associated with a previous period in response to historical equalization being enabled.
[0275]
[0296] Example 23 includes the device of example 21, further including a user input analyzer for determining EQ parameters based on the data indicative of the user's preferences in response to availability of data indicative of the user's preferences, the EQ parameters corresponding to audio playback settings, and adjusting the personalized EQ settings based on the EQ parameters determined based on the data indicative of the user's preferences.
[0276]
[0297] Example 24 includes the apparatus of example 21, further including a device parameter analyzer for adjusting the personalized EQ settings based on the playback device location data in response to availability of the playback device location data, adjusting the personalized EQ settings based on a profile associated with the user in response to availability of an identification of the user, and adjusting the personalized EQ settings based on a source of the audio signal in response to availability of information associated with the source of the audio signal.
[0277]
[0298] Example 25 includes the apparatus of example 21, wherein the EQ adjustment implementer is for applying weights to the first personalized EQ setting, the second personalized EQ setting, and the filter coefficients to generate a hybrid equalization.
[0278]
[0299] Example 26 includes the device of example 21, wherein the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicative of a user's preferences, location data of the playback device, a profile associated with the user, or a source of the audio signal.
[0279]
[0300] Example 27 includes the apparatus of example 21, wherein the EQ adjustment implementer applies a smoothing filter to the audio signal to reduce sharp transitions in the average gain value of the audio signal between the first duration and the second duration.
[0280]
[0301] Example 28 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least generate a query to a neural network, the query including a representation of samples of the audio signal; access a plurality of audio playback settings determined by the neural network based on the query; determine filter coefficients to apply to the audio signal based on the plurality of audio playback settings; generate personalized EQ settings; blend the personalized EQ settings and the filter coefficients to generate a blended equalization; and apply the blended equalization to the audio signal for a first duration.
[0281]
[0302] Example 29 includes the non-transitory computer-readable storage medium of example 28, wherein the instructions, when executed, cause the one or more processors to generate personalized EQ settings based on past personalization settings and, in response to historical equalization being enabled, adjust the personalized EQ settings based on EQ settings associated with a previous period of time.
[0282]
[0303] Example 30 includes the non-transitory computer-readable storage medium of Example 28, wherein the instructions, when executed, cause one or more processors to, in response to availability of data indicative of user preferences, determine EQ parameters based on the data indicative of the user preferences, where the EQ parameters correspond to audio playback settings, and adjust the personalized EQ settings based on the EQ parameters determined based on the data indicative of the user preferences.
[0283]
[0304] Example 31 includes the non-transitory computer-readable storage medium of Example 28, wherein the instructions, when executed, cause the one or more processors to adjust the personalized EQ settings based on the location data of the playback device in response to availability of the location data of the playback device, adjust the personalized EQ settings based on a profile associated with the user in response to availability of an identification of the user, and adjust the personalized EQ settings based on the source of the audio signal in response to availability of information associated with the source of the audio signal.
[0284]
[0305] Example 32 includes the non-transitory computer-readable storage medium of example 28, wherein the instructions, when executed, cause one or more processors to apply weights to the first personalized EQ setting, the second personalized EQ setting, and the filter coefficients to generate a hybrid equalization.
[0285]
[0306] Example 33 includes the non-transitory computer-readable storage medium of Example 28, wherein the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicative of a user's preferences, location data of the playback device, a profile associated with the user, or a source of the audio signal.
[0286]
[0307] Example 34 includes the non-transitory computer-readable storage medium of example 28, wherein the instructions, when executed, cause one or more processors to apply a smoothing filter to the audio signal to reduce sharp transitions in the average gain value of the audio signal between the first duration and the second duration.
[0287]
[0308] Example 35 includes a method including generating a query to a neural network, the query including a representation of a sample of an audio signal; accessing a plurality of audio playback settings determined by the neural network based on the query; determining filter coefficients to apply to the audio signal based on the plurality of audio playback settings; generating personalized EQ settings; blending the personalized EQ settings and the filter coefficients to generate blended equalization; and applying the blended equalization to the audio signal for a first duration.
[0288]
[0309] Example 36 includes the method of example 35, further including generating personalized EQ settings based on past personalization settings, and adjusting the personalized EQ settings based on EQ settings associated with a previous period in response to historical equalization being enabled.
[0289]
[0310] Example 37 includes the method of Example 35, further including, in response to availability of data indicative of user preferences, determining EQ parameters based on the data indicative of the user preferences, wherein the EQ parameters correspond to audio playback settings, and adjusting the personalized EQ settings based on the EQ parameters determined based on the data indicative of the user preferences.
[0290]
[0311] Example 38 includes the method of Example 35, further including: adjusting the personalized EQ settings based on location data of the playback device in response to availability of location data of the playback device; adjusting the personalized EQ settings based on a profile associated with the user in response to availability of an identification of the user; and adjusting the personalized EQ settings based on a source of the audio signal in response to availability of information associated with the source of the audio signal.
[0291]
[0312] Example 39 includes the method of example 35, further including applying weights to the first personalized EQ setting, the second personalized EQ setting, and the filter coefficients to generate a hybrid equalization.
[0292]
[0313] Example 40 includes the method of example 35, wherein the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicative of a user's preferences, location data of the playback device, a profile associated with the user, or a source of the audio signal.
[0293]
[0314] Example 41 includes an apparatus comprising: a synchronizer for accessing an equalization (EQ) profile corresponding to the media signal in response to receiving the media signal to be played on a playback device; an EQ personalization manager for generating personalized EQ settings; and an EQ adjustment implementer for modifying playback of the media signal on the playback device based on a blended equalization generated based on the EQ profile and the personalized EQ settings.
[0294]
[0315] Example 42 includes the device of example 41, further including a historical EQ manager for generating personalized EQ settings based on past personalization settings and adjusting the personalized EQ settings based on EQ settings associated with previous time periods in response to historical equalization being enabled.
[0295]
[0316] Example 43 includes the device of example 41, further including a user input analyzer for determining EQ parameters based on the data indicative of the user's preferences in response to availability of data indicative of the user's preferences, the EQ parameters corresponding to audio playback settings, and adjusting the personalized EQ settings based on the EQ parameters determined based on the data indicative of the user's preferences.
[0296]
[0317] Example 44 includes the apparatus of example 41, further including a device parameter analyzer for adjusting the personalized EQ settings based on the playback device location data in response to availability of playback device location data, adjusting the personalized EQ settings based on a user profile in response to availability of a user identification, and adjusting the personalized EQ settings based on a source of the media signal in response to availability of information associated with the source of the media signal.
[0297]
[0318] Example 45 includes the apparatus of Example 41, wherein the EQ adjustment implementer is for applying weights to the first personalized EQ setting, the second personalized EQ setting, and the EQ profile to generate a mixed equalization.
[0298]
[0319] Example 46 includes the device of example 41, wherein the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicative of user preferences, location data of the playback device, a user profile, or a source of the media signal.
[0299]
[0320] Example 47 includes the device of example 41, wherein the EQ profile includes playback attributes corresponding to at least one of (1) information identifying a category associated with the song, (2) information identifying a category associated with the video segment, (3) information identifying a mood associated with the song or video segment, or (4) information identifying signal strength parameters for various frequencies along with a portion of the media signal.
[0300]
[0321] Example 48 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to, at least, in response to receiving a media signal to be played on a playback device, access an equalization (EQ) profile corresponding to the media signal, generate personalized EQ settings, and modify playback of the media signal on the playback device based on a blended equalization generated based on the EQ profile and the personalized EQ settings.
[0301]
[0322] Example 49 includes the non-transitory computer-readable storage medium of example 48, wherein the instructions, when executed, cause one or more processors to generate personalized EQ settings based on past personalization settings and, in response to historical equalization being enabled, adjust the personalized EQ settings based on EQ settings associated with a previous period of time.
[0302]
[0323] Example 50 includes the non-transitory computer-readable storage medium of Example 48, wherein the instructions, when executed, cause one or more processors to, in response to availability of data indicative of user preferences, determine EQ parameters based on the data indicative of the user preferences, where the EQ parameters correspond to audio playback settings, and adjust the personalized EQ settings based on the EQ parameters determined based on the data indicative of the user preferences.
[0303]
[0324] Example 51 includes the non-transitory computer-readable storage medium of Example 48, wherein the instructions, when executed, cause one or more processors to adjust the personalized EQ settings based on the playback device location data in response to availability of the playback device location data, adjust the personalized EQ settings based on a user profile in response to availability of a user identification, and adjust the personalized EQ settings based on the source of the media signal in response to availability of information associated with the source of the media signal.
[0304]
[0325] Example 52 includes the non-transitory computer-readable storage medium of example 48, wherein the instructions, when executed, cause one or more processors to apply weights to the first personalized EQ setting, the second personalized EQ setting, and the EQ profile to generate a blended equalization.
[0305]
[0326] Example 53 includes the non-transitory computer-readable storage medium of Example 48, in which the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicative of user preferences, location data of the playback device, a user profile, or a source of the media signal.
[0306]
[0327] Example 54 includes the non-transitory computer-readable storage medium of Example 48, in which the EQ profile includes playback attributes corresponding to at least one of (1) information identifying a category associated with the song, (2) information identifying a category associated with the video segment, (3) information identifying a mood associated with the song or video segment, or (4) information identifying signal strength parameters for various frequencies along with a portion of the media signal.
[0307]
[0328] Example 55 includes a method that includes, in response to receiving a media signal to be played on a playback device, accessing an equalization (EQ) profile corresponding to the media signal, generating a personalized EQ setting, and modifying playback of the media signal on the playback device based on a blended equalization generated based on the EQ profile and the personalized EQ setting.
[0308]
[0329] Example 56 includes the method of example 55, further including generating personalized EQ settings based on past personalization settings, and adjusting the personalized EQ settings based on EQ settings associated with previous periods in response to historical equalization being enabled.
[0309]
[0330] Example 57 includes the method of Example 55, further including, in response to availability of data indicative of user preferences, determining EQ parameters based on the data indicative of the user preferences, wherein the EQ parameters correspond to audio playback settings, and adjusting personalized EQ settings based on the EQ parameters determined based on the data indicative of the user preferences.
[0310]
[0331] Example 58 includes the method of example 55, further including adjusting the personalized EQ settings based on location data of the playback device in response to availability of location data of the playback device, adjusting the personalized EQ settings based on a user profile in response to availability of an identification of the user, and adjusting the personalized EQ settings based on a source of the media signal in response to availability of information associated with the source of the media signal.
[0311]
[0332] Example 59 includes the method of example 55, further including applying weights to the first personalized EQ setting, the second personalized EQ setting, and the EQ profile to generate a blended equalization.
[0312]
[0333] Example 60 includes the method of example 55, wherein the personalized EQ settings are based on at least one of EQ settings associated with a previous period, data indicating user preferences, location data of the playback device, a user profile, or a source of the media signal.
[0313]
[0334] Example 61 includes an apparatus comprising: a volume controller for applying multiple equalization adjustments to an audio signal to generate an equalized audio signal, the multiple equalization adjustments being output from a neural network in response to an input feature set including an average loudness representation of the audio signal; a thresholding controller for detecting irregularities in a frequency representation of the audio signal after application of the multiple equalization adjustments, where the irregularities correspond to changes in loudness between adjacent frequency values where the irregularities exceed a threshold, and adjusting the loudness at a first one of the adjacent frequency values to reduce the irregularities; an EQ curve generator for generating an equalization (EQ) curve to apply to the audio signal when the irregularities are reduced; and a frequency-to-time domain converter for outputting the equalized audio signal in the time domain based on the EQ curve.
[0314]
[0335] Example 62 is the apparatus of Example 61, further including an energy calculator for determining a first root mean square (RMS) value of the frequency representation of the audio signal before application of the plurality of equalization adjustments, determining a second RMS value of the frequency representation of the audio signal after reduction of irregularities, and determining a ratio between the second RMS value and the first RMS value.
[0315]
[0336] Example 63 includes the apparatus of Example 61, further including a volume normalizer for determining whether a ratio between (1) a first RMS value of the frequency representation of the audio signal after irregularity reduction and (2) a second RMS value of the frequency representation of the audio signal before application of the plurality of equalization adjustments exceeds a threshold associated with an allowable energy change in the audio signal, and applying gain normalization of the frequency representation of the audio signal in response to the ratio exceeding the threshold.
[0316]
[0337] Example 64 includes the device of example 61, wherein the plurality of equalization adjustments includes a plurality of volume adjustment values corresponding to a plurality of frequency ranges.
[0317]
[0338] Example 65 includes the apparatus of example 61, wherein the thresholding controller is for selecting a plurality of frequency values in the frequency representation of the audio signal, determining a plurality of volume values associated with the plurality of frequency values, determining a second derivative of the volume across the plurality of frequency values, and adjusting the volume at a first frequency value of adjacent frequency values in response to an absolute value of the second derivative exceeding a threshold to reduce irregularities.
[0318]
[0339] Example 66 includes the apparatus of example 61, wherein the plurality of equalization adjustments are based on at least a reference audio signal, an EQ curve, and tags associated with the plurality of audio engineers who created the EQ curve, and the neural network determines the plurality of equalization adjustments based on inferences associated with at least the reference audio signal, the EQ curve, and tags associated with the plurality of audio engineers.
[0319]
[0340] Example 67 includes the apparatus of example 66, wherein the input feature set includes a mean loudness representation of the audio signal and a mean standard deviation measure for frequency bins of the frequency representation of the audio signal.
[0320]
[0341] Example 68 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least: apply a plurality of equalization adjustments to an audio signal to generate an equalized audio signal, wherein the plurality of equalization adjustments are output from a neural network in response to an input feature set including an average loudness representation of the audio signal; detect irregularities in a frequency representation of the audio signal after application of the plurality of equalization adjustments, wherein the irregularities correspond to changes in loudness between adjacent frequency values that exceed a threshold; adjust the loudness at a first one of the adjacent frequency values to reduce the irregularities; generate an equalization (EQ) curve to apply to the audio signal when the irregularities are reduced; and output the equalized audio signal in the time domain based on the EQ curve.
[0321]
[0342] Example 69 includes the non-transitory computer-readable storage medium of Example 68, wherein the instructions, when executed, cause one or more processors to determine a first root mean square (RMS) value of the frequency representation of the audio signal before application of the plurality of equalization adjustments, determine a second RMS value of the frequency representation of the audio signal after reduction of irregularities, and determine a ratio between the second RMS value and the first RMS value.
[0322]
[0343] Example 70 includes the non-transitory computer-readable storage medium of example 68, wherein the instructions, when executed, cause one or more processors to determine whether a ratio between (1) a first RMS value of the frequency representation of the audio signal after reducing irregularities and (2) a second RMS value of the frequency representation of the audio signal before applying a plurality of equalization adjustments exceeds a threshold associated with an allowable energy change in the audio signal, and in response to the ratio exceeding the threshold, apply gain normalization of the frequency representation of the audio signal.
[0323]
[0344] Example 71 includes the non-transitory computer-readable storage medium of example 68, wherein the plurality of equalization adjustments includes a plurality of volume adjustment values corresponding to a plurality of frequency ranges.
[0324]
[0345] Example 72 includes the non-transitory computer-readable storage medium of example 68, wherein the instructions, when executed, cause one or more processors to select a plurality of frequency values in a frequency representation of the audio signal, determine a plurality of volume values associated with the plurality of frequency values, determine a second derivative of the volume across the plurality of frequency values, and, in response to an absolute value of the second derivative exceeding a threshold, adjust the volume at a first frequency value of an adjacent frequency value to reduce irregularities.
[0325]
[0346] Example 73 includes the non-transitory computer-readable storage medium of example 68, wherein the plurality of equalization adjustments are based on at least a reference audio signal, an EQ curve, and tags associated with the plurality of audio engineers who created the EQ curve, and the neural network determines the plurality of equalization adjustments based on inferences associated with at least the reference audio signal, the EQ curve, and tags associated with the plurality of audio engineers.
[0326]
[0347] Example 74 includes the non-transitory computer-readable storage medium of example 73, wherein the input feature set includes a mean loudness representation of the audio signal and a mean standard deviation measure for frequency bins of the frequency representation of the audio signal.
[0327]
[0348] Example 75 includes a method including applying a plurality of equalization adjustments to an audio signal to generate an equalized audio signal, the plurality of equalization adjustments being output from a neural network in response to an input feature set including an average loudness representation of the audio signal; detecting irregularities in a frequency representation of the audio signal after application of the plurality of equalization adjustments, the irregularities corresponding to changes in loudness between adjacent frequency values where the irregularities exceed a threshold; adjusting the loudness at a first one of the adjacent frequency values to reduce the irregularities; generating an equalization (EQ) curve for application to the audio signal when the irregularities are reduced; and outputting the equalized audio signal in the time domain based on the EQ curve.
[0328]
[0349] Example 76 includes the method of Example 75, further including determining a first root mean square (RMS) value of the frequency representation of the audio signal before applying the plurality of equalization adjustments, determining a second RMS value of the frequency representation of the audio signal after reducing irregularities, and determining a ratio between the second RMS value and the first RMS value.
[0329]
[0350] Example 77 includes the method of Example 75, further including determining whether a ratio between (1) a first RMS value of the frequency representation of the audio signal after reducing irregularities and (2) a second RMS value of the frequency representation of the audio signal before applying the plurality of equalization adjustments exceeds a threshold associated with an allowable energy change in the audio signal, and applying gain normalization of the frequency representation of the audio signal in response to the ratio exceeding the threshold.
[0330]
[0351] Example 78 includes the method of example 75, wherein the plurality of equalization adjustments includes a plurality of volume adjustment values corresponding to a plurality of frequency ranges.
[0331]
[0352] Example 79 includes the method of Example 75, further including selecting a plurality of frequency values in a frequency representation of the audio signal, determining a plurality of volume values associated with the plurality of frequency values, determining a second derivative of the volume across the plurality of frequency values, and adjusting the volume at a first frequency value of adjacent frequency values in response to the absolute value of the second derivative exceeding a threshold to reduce irregularities.
[0332]
[0353] Example 80 includes the method of example 75, wherein the plurality of equalization adjustments are based on at least a reference audio signal, an EQ curve, and tags associated with the plurality of audio engineers who created the EQ curve, and the neural network determines the plurality of equalization adjustments based on inferences associated with at least the reference audio signal, the EQ curve, and tags associated with the plurality of audio engineers.
[0333]
[0354] Although certain exemplary methods, apparatus, and articles of manufacture have been disclosed herein, the scope of protection of this patent is not limited thereto, but rather encompasses all methods, apparatus, and articles of manufacture that expressly fall within the scope of the claims of this patent.
Claims
1. 1. A computing system comprising: a processor; a non-transitory computer-readable storage medium having stored thereon program instructions that, when executed by the processor, cause the execution of a set of operations; The set of actions may include: applying an equalization adjustment to the audio signal to generate an equalized audio signal; detecting irregularities in a frequency representation of the equalized audio signal, the irregularities corresponding to loudness changes between a set of frequency values that exceed a threshold, the set of frequency values including adjacent frequency values; adjusting a volume at a first frequency value of the set of frequency values to reduce the irregularity; Computing system.
2. The computing system of claim 1, wherein the irregularity is detected by determining local outliers based on whether the second derivative of the volume over a frequency range exceeds a threshold.
3. The set of actions may include: selecting the adjacent frequency values in the set of frequency representations of the audio signal; determining volume values associated with the adjacent frequency values; determining the change in volume between the adjacent frequency values; and adjusting a volume at the first one of the neighboring frequency values to reduce the irregularity in response to determining that the absolute value of the change in volume exceeds the threshold. The computing system of claim 1 .
4. the change in volume is represented by a value of a second derivative of volume across the set of frequency values; The computing system of claim 1 .
5. The first frequency value is a center value of the set of frequency values, and the set of operations comprises: adjusting the volume at the first frequency value to be a midpoint between volume levels at other frequency values in the set of frequency values. The computing system of claim 1 .
6. the irregularities correspond to at least one of short-term peaks or short-term dips in volume values across the set of frequency values, which may result in perceptible artifacts in the audio signal. The computing system of claim 1 .
7. The threshold is a first threshold, and the set of actions is: detecting and reducing at least one additional irregularity in the frequency representation of the audio signal until a number of remaining irregularities in the frequency representation of the audio signal meets a second threshold. The computing system of claim 1 .
8. A non-transitory computer-readable storage medium having stored thereon program instructions that, when executed by a processor, cause the execution of a set of operations, comprising: The set of actions may include: applying an equalization adjustment to the audio signal to generate an equalized audio signal; detecting irregularities in a frequency representation of the equalized audio signal, the irregularities corresponding to loudness variations between frequency values exceeding a threshold, the set of frequency values including adjacent frequency values; adjusting a volume at a first one of the frequency values to reduce the irregularity. A non-transitory computer-readable storage medium.
9. A non-transitory computer-readable storage medium as described in claim 8, wherein the irregularity is detected by determining a local outlier based on whether the second derivative of the volume over a frequency range exceeds a threshold.
10. When the instructions are executed, the set of actions are: selecting the adjacent frequency values in the frequency representation of the audio signal; determining volume values associated with the adjacent frequency values; determining the change in volume; and adjusting a volume at the first one of the neighboring frequency values to reduce the irregularity in response to determining that the absolute value of the change in volume exceeds the threshold. The non-transitory computer-readable storage medium of claim 8.
11. the change in volume is represented by a value of a second derivative of volume across the set of frequency values; The non-transitory computer-readable storage medium of claim 8.
12. the first frequency value is a center value of the set of frequency values; The set of actions may include: adjusting the volume at the first frequency value to be a midpoint between volume levels at other frequency values in the set of frequency values. The non-transitory computer-readable storage medium of claim 8.
13. the irregularities correspond to at least one of short-term peaks or short-term dips in volume values across the set of frequency values, which may result in perceptible artifacts in the audio signal. The non-transitory computer-readable storage medium of claim 8.
14. the threshold is a first threshold, The set of actions may include: detecting and reducing at least one additional irregularity in the frequency representation of the audio signal until a number of remaining irregularities in the frequency representation of the audio signal meets a second threshold. The non-transitory computer-readable storage medium of claim 8.
15. applying an equalization adjustment to the audio signal to generate an equalized audio signal; detecting irregularities in a frequency representation of the equalized audio signal, the irregularities corresponding to loudness changes between a set of frequency values that exceed a threshold, the set of frequency values including adjacent frequency values; adjusting a volume at a first frequency value of the set of frequency values to reduce the irregularity. method.
16. The method described in claim 15, wherein the irregularity is detected by determining local outliers based on whether the second derivative of the volume over a frequency range exceeds a threshold.
17. selecting said adjacent frequency values in said frequency representation of said audio signal; determining volume values associated with said adjacent frequency values; determining the change in volume; and adjusting a volume at the first one of the neighboring frequency values in response to determining that the absolute value of the change in volume exceeds the threshold to reduce the irregularity.
16. The method of claim 15.
18. the change in volume is represented by a value of a second derivative of volume across the set of frequency values; 16. The method of claim 15.
19. the first frequency value is a center value of the set of frequency values; adjusting the volume at the frequency value to be midpoint between volume levels at other frequency values in the set of frequency values; 16. The method of claim 15.
20. the irregularities correspond to at least one of short-term peaks or short-term dips in volume values across the set of frequency values, which may result in perceptible artifacts in the audio signal.
16. The method of claim 15.
Citation Information
Patent Citations
Channel gain correction system and noise reduction method in voice communication
JP2003526109A
Sound signal processing method, sound signal processing device and computer program
JP2008076676A
receiver
JP2010272997A
In-vehicle radio noise reducing system
JP2013009066A
Sound adjusting device and sound adjusting method
JP2017163447A