Computer readable storage medium, method to be performed by computer, and computing device for dynamic volume adjustment via audio classification
Audio classification-based volume adjustment in media systems addresses volume inconsistencies by applying targeted gain values, enhancing audio quality and user experience.
Patent Information
- Application Number
- JP2025107517
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-10-12
- Filing Date
- 2025-06-25
- Publication Date
- 2025-11-10
- Estimated Expiration
- 2039-09-06
AI Technical Summary
Conventional audio systems face issues with noticeable volume level differences between various media sources, leading to listener irritation and degradation of audio quality due to continuous dynamic range compression.
Implementing audio classification to identify signal categories and apply targeted gain adjustments, minimizing the need for dynamic range compression by using a combination of classification and real-time input measurements to achieve a consistent volume range.
Reduces the amount of compression required, maintaining audio quality by ensuring volume consistency across different media sources, allowing for seamless transitions and user control over volume levels.
Smart Images

Figure 2025168612000001_ABST
Abstract
Description
Related Applications
[0001]
[0001] This patent claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 728,677, filed September 7, 2018, and U.S. Provisional Patent Application No. 62 / 745,148, filed October 12, 2018. U.S. Provisional Patent Application No. 62 / 702,734 and U.S. Provisional Patent Application No. 62 / 745,148 are incorporated herein by reference in their entireties. FIELD OF THE DISCLOSURE
[0002] This disclosure relates generally to volume control, and more particularly to a method and apparatus for dynamic volume control via audio classification.
[0003]
[0003] In recent years, a large amount of media of various characteristics is being distributed using an ever-increasing number of channels. This media can be received using more traditional channels (e.g., radio) or more recently developed channels, such as using Internet-connected streaming devices. As these channels evolve, systems capable of processing and outputting audio from multiple sources are also being developed. For example, some automobile media systems are capable of delivering media from compact discs (CDs), Bluetooth-connected devices, Universal Serial Bus (USB)-connected devices, Wi-Fi-connected devices, auxiliary inputs, and other sources. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system constructed in accordance with the teachings of this disclosure for dynamic volume adjustment via audio classification.
[0005] [Figure 2] FIG. 2 is a block diagram illustrating further details of the media unit of FIG.
[0006] [Figure 3] FIG. 3 is a block diagram illustrating an audio classification engine capable of providing trained models for use by the media units of FIGS.
[0007] [Figure 4] FIG. 4 is a flowchart representing exemplary machine-readable instructions that may be used to implement the media unit 106 of FIGS. 1 and 2 to perform dynamic volume adjustment via audio classification. [Figure 5] FIG. 5 is a flowchart representing exemplary machine-readable instructions that may be used to implement the media unit 106 of FIGS. 1 and 2 to perform dynamic volume adjustment via audio classification.
[0008] [Figure 6] FIG. 6 is a schematic diagram of an exemplary processor platform capable of executing the instructions of FIGS. 4 and 5 to implement the exemplary media unit 106 of FIGS.
[0009]
[0009] The drawings are not to scale. Wherever possible, the same reference numbers are used throughout the drawings and the accompanying specification to refer to the same or like parts. DETAILED DESCRIPTION
[0010]
[0010] In conventional audio media implementations, audio signals associated with different media may have different volume levels. For example, media on one CD may be recorded and / or mastered at a significantly different volume level than media on another CD. Similarly, media obtained from a streaming device may have a significantly different volume level than media obtained from a different device or from the same device through a different application. As users increasingly listen to media from a variety of different sources, the differences in volume levels between sources and between media from the same source may become very noticeable and irritate listeners.
[0011]
[0011] Some conventional approaches to volume leveling utilize a dynamic range compressor to compress the entire dynamic range of an audio signal to meet a volume threshold. In some conventional implementations, such dynamic range compression continuously monitors and adjusts the volume of an audio signal to meet the audio signal's volume threshold. Such continuous adjustments have a noticeable effect on a listener's perception of the audio signal, as the original dynamics of the track are significantly altered. In some instances, dynamic range compression significantly degrades the perceived quality of the audio signal (e.g., by introducing artifacts into the audio).
[0012] Exemplary methods, devices, systems, and articles of manufacture disclosed herein use audio classification to identify a category of an audio signal and then perform a volume adjustment to minimize the amount of dynamic range compression required to bring the audio signal within a target volume range. The exemplary methods, devices, systems, and articles of manufacture disclosed herein utilize a combination of an audio signal's classification and real-time input audio measurements to identify a target gain value applicable to the audio signal. For example, after identifying a classification group associated with the audio signal, a classification gain value can be obtained (e.g., from a lookup table that associates volume gain adjustment values with classification groups). Furthermore, an input volume for the audio signal can be identified. A target gain value can then be identified based on the input volume and the recommended classification gain value. The target gain value is a volume adjustment applied to the input audio signal to bring the volume closer to the target volume range (e.g., within ±1 dbFS of −21 dbFS), such that when the gain-adjusted signal is provided to a compressor, the amount of compression required to bring the gain-adjusted signal within the target volume range is reduced.
[0013]
[0013] Exemplary methods, apparatus, systems, and articles of manufacture disclosed herein calculate a target gain value based on a classification of an input audio signal and an input volume of the audio signal to reduce the amount of compression required to bring the volume of the audio signal within a target volume range. In some examples, when the input audio signal is initially detected, the dynamic range of the audio signal is initially compressed to bring the volume of the audio signal within a target volume range by the time the input audio signal is classified and the volume of the input audio signal is determined. In some examples, if compression alone is used to adjust the audio signal when the audio signal is initially detected, a listener may easily notice the compression as a reduction in audio level without any manual volume adjustment. However, once the initial volume of the audio signal and the classification of the audio signal are determined, a target gain value is calculated to reduce the amount of compression required to bring the volume of the audio signal within the target volume range. In some examples, the classification and initial volume determination can occur quickly enough (e.g., within 5 seconds, within 1 second, etc.) that the initial use of compression is not noticeable to a listener.
[0014] Some example methods, devices, systems, and articles of manufacture disclosed herein identify and accommodate changes in the source of an audio signal. In some examples, an initial volume adjustment is performed in addition to or instead of using compression. For example, in response to a change in audio signal input (e.g., from no audio signal to a presented audio signal, from one audio signal input source to another, etc.), an initial volume level can be identified (e.g., based on previous volume adjustment settings specific to the source of the audio signal) and an initial volume level adjustment can be performed. In some examples, the initial volume level adjustment is performed using a “fade-in” technique, which gradually increases the audio volume level after a change in the input signal. In some examples, the initial volume level adjustment can be based on stored settings associated with the type of audio input signal (e.g., FM radio, AM radio, CD, auxiliary audio source, etc.).
[0015]
[0015] Exemplary methods, apparatus, systems, and articles of manufacture disclosed herein classify audio signals into one or more of a plurality of classification groups. In identifying classification groups, characteristics of the classification groups (e.g., amount of available headroom, typical dynamic range, etc.) can be used to adjust the volume of the audio signal with minimal loss (e.g., using minimal dynamic range compression). In some examples, pattern recognition in training data can be used to identify classification groups. For example, audio signals can be grouped based on factors such as the instruments represented in the signal, the year the audio signal was produced, the musical genre, etc. Once the training data is grouped, characteristics such as a distribution of dynamic range values, a distribution of volume values, or any other audio characteristic can be associated with the classification groups and stored (e.g., in a lookup table). In some examples, a probability distribution can be determined when classifying an audio signal (e.g., as opposed to outputting one particular classification group to which the audio signal belongs). For example, the classification process may output that the audio signal has a 50% chance of belonging to a group representing music without drums from 1976-1995, a 30% chance of belonging to a group representing music without drums from 1996 to the present, an 18% chance of belonging to a group representing music with synthetic drums from 1976-1995, or a 2% chance of belonging to some other group. In some such examples, selecting gain values associated with the classification groups to perform the volume adjustment may include averaging techniques (e.g., identifying gain values associated with each group and weighting each value according to the probability that the audio signal belongs to the respective group).
[0016] In some example methods, apparatus, systems, and articles of manufacture disclosed herein, a large corpus of loudness profiles of representative audio signals (e.g., representing multiple genres, multiple time periods, etc.) is utilized to train an audio signal classifier to perform classification of audio signals. For example, the loudness profile includes loudness values at multiple times within a song. In some examples, other profiles and / or representations of audio signals can be utilized in addition to or instead of the loudness profile to train the audio signal classifier. In some examples, clustering is performed on the loudness profiles to train the audio signal classifier. In some examples, the audio signal classifier is trained to identify clusters of loudness profiles based on the loudness, dynamic range, and / or any other property of the loudness profile. The audio signal classifier can cluster the loudness profiles into dynamic range groups, and then the audio signal classifier can assign incoming audio (e.g., an input audio signal) to one or more of the classification groups.
[0017]
[0017] In the exemplary methods, apparatus, systems, and articles of manufacture disclosed herein, after identifying a classification group for an audio signal, a volume level of the audio signal can be adjusted by applying a gain value to the audio signal. The gain value can be specific to the classification group. For example, if the classification group is associated with an audio signal having a relatively narrow normalized dynamic range (such as some pop music), a significant volume adjustment can be made to bring the volume level of the audio signal closer to the target volume range (e.g., because the approximate volume deviation across the track can be determined). Conversely, if the classification group is associated with an audio signal having a relatively wide dynamic range, a lesser volume adjustment can be made to keep the audio signal within an audible level.
[0018] Following application of a gain value based on the classification group associated with the audio signal, compression can be used to bring the volume of the audio signal within a target volume range. Because dynamic range compression can result in a reduction in overall audio quality (e.g., some loss of the audio signal), the exemplary methods, apparatus, systems, and articles of manufacture disclosed herein improve volume adjustment techniques by first applying a gain value specific to the type of audio being presented (e.g., specific to the classification group), thus reducing the amount of dynamic range compression required to fit the volume level of the audio signal within a target volume range.
[0019]
[0019] In some exemplary methods, devices, systems, and articles of manufacture disclosed herein, once an audio signal is classified in a dynamic volume adjustment situation, characteristics of the audio signal are inferred from its classification group, and these characteristics are used to identify a target gain value that will bring the volume of the audio signal closer to a target volume threshold with minimal or no compression.
[0020] In some example methods, devices, systems, and articles of manufacture disclosed herein, input loudness measurements are taken into account when determining target gain values. For example, if the input loudness is identified as -15 dbFS and the target loudness range is within ±1 dbFS of -21 dbFS (e.g., -20 dbFS to -22 dbFS), the target gain value needs to be a less negative gain value than if the input loudness was identified as -10 dbFS, even if the classification group remains unchanged. In some such examples, input loudness measurements are weighted more heavily than classification gain values when determining target gain values because, ultimately, the actual input loudness level of a particular audio signal is more indicative of the amount by which the loudness can be adjusted than a class-based prediction (e.g., real-time measurements may be more accurate than a prediction associated with the class of the audio signal). In some examples, the classification gain values and the input loudness are averaged to calculate the target gain value. For example, if the input volume is identified as -15 dbFS and the classification gain value (e.g., determined based on the average dynamic range of the audio signal for the classification group) indicates that the volume can be adjusted by -6 dbFS, but the target volume range is -21 dbFS to ±1 dbFS, relying solely on the classification gain value would leave an extremely small margin of error (e.g., if the dynamic range is wider than expected, the volume is likely to often fall outside the -0.20 dbFS to 22 dbFS target volume range). If instead the target gain value is calculated as the midpoint (e.g., the average) between the input volume and the classification gain value, the target gain value will move the volume of the audio signal closer to the target gain value while still leaving some margin of error.
[0021] In some example methods, apparatus, systems, and articles of manufacture disclosed herein, input volume levels are measured at regular intervals (e.g., every 3 seconds, every 10 seconds, etc.) and classification is performed at regular intervals. New target gain values can be identified in response to changes in input volume (e.g., changes in the average input volume for the interval, changes in the variance of the input volume for the interval) and / or in response to changes in classification group. In some examples, when transitioning between target gain values, a smoothing filter can be utilized to smoothly transition between two gain values to avoid noticeable fluctuations in volume at each interval. In some examples, large changes in target gain value are ramped at a slower rate than relatively small changes in target gain value.
[0022]
[0022] Exemplary methods, devices, systems, and articles of manufacture disclosed herein adjust the volume level of an audio signal within a target volume range. In some examples, a listener can then manually adjust the volume level (e.g., by turning a volume knob, providing a voice command to change the volume level, etc.), which adjustment is then performed by applying a gain value to the volume-adjusted audio signal. In this way, the listener can still choose the volume at which to listen to the audio signal, but can choose the volume from a consistent, standard volume level (e.g., from a target volume range) rather than adjusting for variations between different sources, variations between tracks, etc. Thus, the techniques disclosed herein enable input audio to be adjusted to lock within a consistent volume range. In some exemplary methods, devices, systems, and articles of manufacture disclosed herein, dynamic volume adjustments can be discontinued upon manual volume adjustment. For example, if a user manually adjusts the volume level (e.g., by turning a volume knob, providing a voice command to change the volume level, etc.), automatic adjustment of the audio level (e.g., by classifying the audio, selecting a gain value based on the classification, monitoring the audio level, etc.) can be discontinued, allowing the user full control over the audio level.
[0023] Some exemplary methods, devices, systems, and articles of manufacture disclosed herein can identify audio signals to further improve volume control. For example, some exemplary techniques disclosed herein utilize audio fingerprinting to identify media to obtain metadata associated with the audio signal. Audio fingerprinting is a technique used to identify media, such as television broadcasts, radio broadcasts, advertisements (television and / or radio), downloaded media, streaming media, and prepackaged media. Existing audio watermarking techniques identify media by embedding one or more audio codes (e.g., one or more fingerprints), such as media identification information and / or identifiers mappable to media identification information, into the audio and / or video components. In some examples, the audio or video components are selected to have sufficient signal characteristics to conceal the watermark. As used herein, the terms "fingerprint," "code," "signature," or "watermark" are used interchangeably and defined to mean any identification (e.g., identifier) that can be inserted or embedded into audio or video of media (e.g., a program or advertisement) for purposes of identifying the media or for other purposes such as tuning (e.g., packet identification headers). As used herein, "media" refers to audio and / or visual (still or moving) content and / or advertisements. To identify the fingerprinted media, the fingerprint(s) are extracted and used to access a table of reference fingerprints, which are mapped to media identification information.
[0024] In examples disclosed herein, volume adjustments can be performed by components of a vehicle's audio system or components in communication with the audio system. In some examples, a vehicle's head unit can include a media unit that includes a dynamic volume adjuster or other component capable of dynamic volume adjustment. In such examples, the vehicle head unit can receive audio signals from an auxiliary input, a CD input, a wireless signal receiver input, an external stream from a smart device, a Bluetooth input, a network connection (e.g., a connection to the Internet), or via any other source. For example, dynamic volume adjustments can be performed on a media system in a home entertainment system, where multiple sources (e.g., DVD players, set-top boxes, etc.) can transmit audio signals, and the audio signals are dynamically adjusted to attempt to normalize volume levels between sources and media. In other examples, dynamic volume adjustments can be performed in any situation or for any media device(s).
[0025]
[0025] In an exemplary procedure for dynamic volume adjustment via audio classification, an audio signal corresponding to normalized loud pop music is accessed. After detecting a change in audio signal input associated with the audio signal, a dynamic range compressor compresses the audio to a target volume range (e.g., -21 dBFS). In parallel with this compression, an audio signal classifier identifies a classification group corresponding to the audio signal. For example, the classification group may correspond to music containing synthetic drums and bass from the period 1996 to the present. This classification group may be associated with a particular volume adjustment level (e.g., -15 dBFS). In some examples, this volume adjustment level associated with the classification group may be considered in addition to or instead of a volume level adjustment value identified based on the current audio volume level. Following the volume adjustment associated with this volume adjustment level, only a small amount of audio compression may be performed to reach the target volume range. For example, if a volume adjustment step reduces the volume to a first value (e.g., −17.50 dbFS) and the target volume range is near a second value (e.g., −21 dbFS) that is greater than the first value, then a small amount of audio compression can be performed to bring the audio signal to the second value (e.g., near −21 dbFS but within the target volume range). Thus, only dynamic range compression is performed that reduces the signal by a small amount (e.g., 3.5 dbFS), and the audio quality is significantly better than reducing the signal that needs compression from the original audio input to the target volume range (e.g., compressing the audio signal by −21 dbFS).
[0026] 1 is a schematic diagram of an exemplary system 100 constructed in accordance with the teachings of this disclosure for dynamic volume adjustment. The exemplary system 100 includes media devices 102, 104 that transmit audio signals to a media unit 106. The media unit 106 processes the audio signals and transmits the signals to an audio amplifier 108, which subsequently outputs an amplified audio signal for presentation via an output device 110.
[0027] 1 is a portable media player (e.g., an MP3 player). The exemplary media device 102 stores or receives audio signals corresponding to media and is capable of transmitting the audio signals to other devices. In the example shown in FIG. 1, the media device 102 transmits the audio signals to the media unit 106 via an auxiliary cable. In some examples, the media device 102 can transmit the audio signals to the media unit 106 via any other interface.
[0028] The exemplary media device 104 in the example shown in FIG. 1 is a mobile device (e.g., a mobile phone). The exemplary media device 104 can store or receive audio signals corresponding to media and transmit the audio signals to other devices. In the example shown in FIG. 1, the media device 104 transmits the audio signals to the media unit 106 wirelessly. In some examples, the media device 104 can transmit the audio signals to the media unit 106 using Wi-Fi, Bluetooth, and / or any other technology. In some examples, the media device 104 can interact with vehicle components or other devices to allow a listener to select media to be presented in the vehicle. The media devices 102, 104 can be any device capable of storing and / or accessing audio signals. In some examples, the media devices 102, 104 can be integrated into a vehicle (e.g., a CD player, a radio, etc.).
[0029] The exemplary media unit 106 in the example shown in FIG. 1 is capable of receiving and processing an audio signal. In the example shown in FIG. 1, the exemplary media unit 106 receives a media signal from the media devices 102, 104 and processes the media signal to perform dynamic volume adjustment. The exemplary media unit 106 is capable of identifying the audio signal based on an identifier (e.g., a fingerprint, a watermark, a signature, etc.) embedded in the media. The exemplary media unit 106 is further capable of accessing metadata corresponding to the media associated with the audio signal. In some examples, the metadata is stored on a storage device of the media unit 106. In some examples, the metadata is accessed from another location (e.g., from a server over a network). Furthermore, the exemplary media unit 106 is capable of performing dynamic volume adjustment by identifying and applying an average gain value based on the metadata to adjust the average volume of the audio signal to meet a volume threshold. The exemplary media unit 106 is also capable of monitoring the audio being output by the output device 110 to determine the average volume level of the audio segments in real time. If the audio signal is not identified as corresponding to media and / or metadata containing volume information is not available for the audio signal, the exemplary media unit 106 is capable of dynamic range compression, which provides compression of the audio signal to achieve a desired volume level. In some examples, the exemplary media unit 106 is included as part of another device in a vehicle (e.g., a car radio head unit). In some examples, the exemplary media unit 106 is implemented as software and included as part of another device available either via a direct connection (e.g., a wired connection) or a network (e.g., available over the cloud). In some examples, the exemplary media unit 106 can be incorporated into an audio amplifier 108 and an output device 110 and can output the audio signal independently following processing of the audio signal.
[0030] 1 is a device capable of receiving the audio signal processed by the media unit 106 and performing appropriate amplification of the signal for output by the output device 110. In some examples, the audio amplifier 108 may be integrated with the output device 110. In some examples, the audio amplifier 108 amplifies the audio signal based on an amplified output value from the media unit 106. In some examples, the audio amplifier 108 amplifies the audio signal based on input from a listener (e.g., adjustment of a volume selector by a vehicle passenger or driver).
[0031] 1 is a speaker. In some examples, output device 110 can be multiple speakers, headphones, or any other device capable of presenting an audio signal to a listener. In some examples, output device 110 can also be capable of outputting visual elements (e.g., a television with speakers).
[0032] 1 is described with reference to implementing dynamic volume control in a vehicle, some or all of the devices included in the exemplary system 100 may be implemented in any environment and in any combination. For example, the system 100 may reside in a recreation room in a home, and the media devices 102, 104 may be game consoles, virtual reality devices, set-top boxes, or any other devices capable of accessing and / or transmitting media. Additionally, in some examples, the media may also include visual elements (e.g., television programs, movies, etc.).
[0033]
[0033] Figure 2 is a block diagram 200 providing further details of an exemplary implementation of the media unit 106 shown in Figure 1. The exemplary media unit 106 is capable of receiving an audio signal, processing the audio signal, and dynamically adjusting the volume of the audio signal within a target volume range. Following the dynamic volume adjustment, the exemplary media unit 106 sends the volume-adjusted audio signal 228 to the audio amplifier 108 to be amplified before being output by the output device 110.
[0034] The exemplary media unit 106 includes an exemplary input audio signal 202 and an exemplary input signal detector 204. The signal detector includes an exemplary compressor gain comparator 206, an exemplary audio volume / power comparator 208, and an exemplary audio sample comparator 210, all of which are used to determine 212 whether a change in audio source has occurred. The exemplary media unit 106 further includes an exemplary input volume detector 214, an exemplary audio signal classifier 216, an exemplary classification database 218, an exemplary volume adjuster 220, an exemplary audio signal identifier 222, an exemplary dynamic range compressor 224, and an exemplary real-time audio monitor 226. The resulting output from the system is an exemplary volume-adjusted audio signal 228.
[0035]
[0035] An exemplary input audio signal 202 is an audio signal that is processed and output for presentation. The input audio signal 202 may be accessed from a radio signal (e.g., an FM signal, an AM signal, a satellite radio signal, etc.), a compact disc, an auxiliary cable (e.g., connected to a media device), a Bluetooth signal, a Wi-Fi signal, or any other media. The input audio signal 202 is accessed by an input signal detector 204, an audio signal classifier 216, and / or a real-time audio monitor 226. The input audio signal 202 is transformed by a volume adjuster 220 and / or a dynamic range compressor 224.
[0036] The exemplary input signal detector 204 detects the input audio signal 202. In some examples, the input signal detector 204 detects whether the input audio signal 202 is associated with a new input audio signal or a new input audio signal source (e.g., an AM signal switches to an FM signal, an auxiliary device signal switches to a CD, etc.). In some examples, the input signal detector 204 detects the input audio signal 202 when the input audio signal 202 begins after the media unit 106 was in an off state (e.g., when the media unit 106 is powered on and the input audio signal 202 begins). In some examples, the input signal detector 204 communicates with the audio signal classifier 216 to begin a classification process if the input audio signal 202 is new (e.g., represents a new type of input audio signal indicating a change in input, represents a signal that began after the media unit had not previously presented any audio signal, etc.). In some examples, the input signal detector 204 determines whether the audio source has changed. For example, the input signal detector 204, via the exemplary compressor gain comparator 206, the exemplary volume / power comparator 208, and the exemplary audio sample comparator 210, can determine if the audio input source has changed, which is used by the exemplary source change determiner to determine 212 if the audio source signal has changed.
[0037] The exemplary compressor gain comparator 206 compares the current gain of the dynamic range compressor 224 to a previous gain of the dynamic range compressor 224. For example, the compressor gain comparator 206 may compare the gain of the dynamic range compressor 224 associated with a current sample block of the input audio signal 202 to an average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with a previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the compressor gain comparator 206 may output a ratio of the current gain of the dynamic range compressor 224 to the average of the previous gain of the dynamic range compressor 224. In other examples, the compressor gain comparator 206 may output any other suitable value (e.g., a difference, etc.) related to the comparison of the current gain of the dynamic range compressor 224 to the average of the previous dynamic gain of the dynamic range compressor 224.
[0038] The exemplary volume / power comparator 208 compares the current power of the input audio signal 202 to a previous power of the input audio signal 202. For example, the power comparator 208 may compare the current power of the input audio signal 202 to an average (e.g., arithmetic mean, median, etc.) power of the input audio signal 202 associated with a previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the power comparator 208 may compare the root-mean-square (RMS) power of a current sample of the input audio signal 202 to the RMS power(s) associated with a previous sample of the input audio signal 202. In some examples, the power comparator 208 may query the peak output of the media unit 106 to determine the RMS power of the audio sample. In some examples, the power comparator 208 may output a ratio of the current RMS power to the average of the previous RMS power(s) after K-weighting has been applied. In other examples, the power comparator 208 may output any other suitable value (e.g., a difference, etc.) related to a comparison between the current RMS power of the input audio signal 202 and the average of the previous RMS power(s) of the input audio signal 202.
[0039] The exemplary audio sample comparator 210 compares the current value of a sample of the input audio signal 202 with a previous value of the input audio signal 202. In some examples, the audio sample comparator 210 determines the value of the audio sample based on the maximum amplitude of the samples in the current block of the input audio signal 202. In some examples, the audio sample comparator 210 determines the value of the audio sample as a normalized value (e.g., between 1 and −1). In other examples, the audio sample comparator 210 may determine the value of the audio sample based on any suitable scale. In some examples, the audio sample comparator 210 determines the absolute value of the determined audio sample value. For example, the audio sample comparator 210 may compare the current maximum audio sample value of the input audio signal 202 with the average (e.g., arithmetic mean, median, etc.) audio sample value of the input audio signal 202 associated with a previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the audio sample comparator 210 may output a ratio of the current maximum audio sample value to the average of the previous audio sample block. In other examples, the audio sample comparator 210 may output any other suitable value (e.g., a difference, etc.) related to a comparison of the current audio sample of the input audio signal 202 with the average of the previous audio sample block of the input audio signal 202.
[0040] The example source change determiner 212 determines whether an audio source of the input audio signal 202 has changed based on the output(s) of the example compressor gain comparator 206, the example power comparator 208, and / or the example audio sample comparator 210. For example, the source change determiner 212 may use regression analysis (e.g., linear regression, binomial regression, least squares, logistic regression, etc.) to determine whether a source change has occurred. In such an example, the source change determiner 212 may further perform the regression analysis based on labeled input data. For example, the labeled input data may include an indication of whether an audio source has changed by making a binary decision of source change or no source change as a result of classification from values corresponding to the power comparison, compressor gain comparison, and / or audio sample comparison. In other examples, the source change determiner 212 may use any other suitable predictive model (e.g., machine learning, neural network, etc.) for determining whether an audio source change has occurred. In some examples, the source change determiner 212 may output a binary value indicating whether a source change occurred within a time frame (e.g., within the previous 3 seconds, etc.). For example, the source change determiner 212 may output a "0" to indicate that a source change has not occurred, or a "1" to indicate that a source change has occurred. In other examples, the source change determiner 212 may output any other suitable indication to indicate that a change in audio source has occurred.
[0041] The exemplary input volume detector 214 determines a volume level associated with the input audio signal 202. In some examples, if the input signal detector 204 indicates that the input audio signal 202 is a new input audio signal, the input volume detector 214 determines an initial input volume level value associated with the input audio signal 202. In some examples, the input volume detector 214 provides a volume level to the dynamic range compressor 224 when the input audio signal is initially received to enable dynamic range compression of the input audio signal 202. For example, the input volume detector 214 can provide the initial volume level of the input audio signal 202 to the dynamic range compressor 224, which can then adjust the dynamic range so that the volume level of the input audio signal 202 falls within a target volume range. The input volume detector 214 in the illustrated example determines the volume level at regular intervals (e.g., every 3 seconds, every 5 seconds, etc.). In some examples, the input volume detector 214 determines an average (e.g., arithmetic mean, median, etc.) volume level for the interval. In some examples, the input volume detector 214 determines the deviation in volume level for the interval.
[0042] The exemplary audio signal classifier 216 identifies a classification of the input audio signal. In some examples, the audio signal classifier 216 analyzes characteristics of the input audio signal 202 to identify a classification group to which the input audio signal 202 belongs. In some examples, the audio signal classifier 216 utilizes a neural network to assist in predicting the dynamic range and to inform the volume adjuster 220 of the amount of volume reduction to apply to the input audio signal 202. For example, a neural network can be utilized to train and output a classification model that can be utilized by and / or incorporated into the audio signal classifier 216. A block diagram illustrating an exemplary audio classification engine capable of providing a trained model for use by the media unit 106 (e.g., by the audio signal classifier 216) is shown in FIG. 3. In some examples, audio characteristics associated with the training data are used by the neural network to identify the classification group and stored in association with the classification group. For example, audio characteristics such as average dynamic range, dynamic range deviation, average volume, average volume deviation, etc. may be identified for the classification groups and stored in the classification database 218 and / or other accessible location (e.g., in a look-up table).
[0043] In some examples, the audio signal classifier 216 and / or the audio classification engine 300 of FIG. 3 accesses loudness profiles and / or other representations of representative audio signals (e.g., representing various instruments, various genres, etc.) and trains a model of the audio signal classifier 216 (e.g., using clustering) to identify classes based on the loudness profiles and / or other representations of the representative audio signals. For example, the loudness profiles and / or other representations may be clustered based on loudness and / or dynamic range. The audio signal classifier 216 may then classify the input audio signal 202 by analyzing the input audio signal 202 to identify loudness, dynamic range, and / or other properties of the input audio signal 202, which may be compared to one or more properties associated with a class.
[0044] The audio signal classifier 216 in the illustrated example identifies one or more classification groups from a plurality of classification groups (e.g., nine classification groups, ten classification groups, etc.) associated with different types of audio signals. For example, a classification group may be associated with a genre of music represented by the input audio signal 202, a period of music represented by the input audio signal 202, different instruments identified in the input audio signal 202, etc. In some examples, a classification group may be associated with spoken content, pop music, rock music, hip hop music, etc. Some example classification groups include speech, music without drums before 1975, music without drums from 1976-1995, music without drums from 1996-present, music with synthetic drums from 1976-1995, music with synthetic drums from 1996-present, music with real drums before 1975, music with real drums from 1976-1995, and / or music with real drums from 1996-present. Thus, the classification groups may correspond to different eras of music / sound production where technical differences in recording and / or playback capabilities correspond to differences in the loudness and / or dynamic range of the music / sound produced. The classification groups may additionally or alternatively be based on observed (e.g., heuristically derived) characteristics of the loudness and / or dynamic range of the audio content.
[0045] The audio signal classifier 216 can utilize any characteristic of the input audio signal 202 to classify the input audio signal 202. For example, the audio signal classifier 216 can use spectral characteristics of the input audio signal 202, constant Q transform (CQT) characteristics of the input audio signal 202, or any other parameters. In some examples, time samples, spectrogram(s), summaries, transforms, and / or descriptions of the audio signal are used as input to the audio signal classifier 216. Such characteristics can be input to a neural network model to identify classification groups of the input audio signal. In some examples, the neural network model can be accessed from a classification database 218.
[0046] The audio signal classifier 216 in the illustrated example can output a single class (e.g., speech, music with drums since 1996, etc.) or can output a probability distribution associated with multiple classes. In some examples, the audio signal classifier 216 identifies the class with the highest probability corresponding to the audio signal and outputs an indication that the audio signal belongs to this class. In other examples, the audio signal classifier 216 outputs a probability associated with the audio signal belonging to each class (e.g., a 60 percent chance that the audio signal belongs to the “speech” class). In some examples, a threshold percentage can be used to identify when a single class is output compared to when a probability distribution is output. For example, if the audio signal classifier 216 identifies an audio signal with a 90 percent chance that it belongs to the speech class, this chance may exceed the threshold percentage, allowing the audio signal classifier 216 to identify the audio signal as belonging to the speech class. In some examples, if the threshold percentage is not met, a probability distribution can be output, or the audio signal classifier 216 can indicate that it is unable to identify a class associated with the audio signal.
[0047] In response to identifying the classification group of the input audio signal 202, the audio signal classifier 216 can select a classification gain value associated with the classification group and communicate the classification gain value to the volume adjuster 220 and / or the dynamic range compressor 224. In some examples, the audio signal classifier 216 accesses the classification gain value from one or more lookup tables associated with the classification group. In some examples, the classification gain value is identified as a combination of values from one or more tables associated with the one or more classification groups. For example, if the audio signal classifier 216 outputs a probability distribution indicating the probability that the audio signal belongs to each classification group, a table associated with each group can be obtained, and gain values or other adjustment values (e.g., EQ values) can be combined and weighted based on the relative probability of each classification group.
[0048] In some examples, the audio signal classifier 216 provides the classification group to the volume adjuster 220 and / or the dynamic range compressor 224, which then access and / or identify the adjustment parameters associated with the classification group. In some examples, the audio signal classifier 216 outputs (1) a classification gain value and / or (2) a time period corresponding to a time at which the volume level of the audio should be reanalyzed.
[0049] The exemplary classification database 218 is a repository of data related to audio signal classification. In some examples, the classification database 218 stores models (e.g., neural network models) used to classify audio signals. In some examples, the classification database 218 accesses and / or retrieves models from an audio classification engine shown in FIG. 3 and described in further detail. In some examples, the classification database 218 may store audio signals, audio fingerprints, and / or any other data utilized by the media unit 106. The classification database 218 stores lookup tables or other storage means, such as for storing audio parameters associated with classification groups. The exemplary classification database 218 may be implemented with volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus dynamic random access memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory). Classification database 218 may additionally or alternatively be implemented by one or more double data rate (DDR) memories, e.g., DDR, DDR2, DDR3, mobile DDR (mDDR), etc. Classification database 218 may additionally or alternatively be implemented by one or more mass storage devices, e.g., hard disk drive(s), compact disk drive(s), digital versatile disk drive(s), etc. Although the illustrated example shows classification database 218 as a single database, classification database 218 may be implemented by any number and / or type(s) of databases. Furthermore, the data stored in classification database 218 may be in any data format, e.g., binary data, comma-separated data, tab-separated data, structured query language (SQL) structures, etc.
[0050] 2 adjusts the volume level of an audio signal. In some examples, the example volume adjuster 220 identifies a single average gain value that converts the volume of the audio signal from a known volume value (e.g., identified by the input volume detector 214) to a desired volume value (e.g., a value near a target volume range). The volume adjuster 220 in the illustrated example communicates with the input volume detector 214 and / or the audio signal classifier 216 to identify the target gain value. The volume adjuster 220 calculates the target gain based on classification gain values corresponding to one or more classification groups identified by the audio signal classifier 216 and the input volume level detected by the input volume detector 214 (e.g., by calculating an average of the classification gain values and the input volume). In some examples, the volume adjuster 220 applies one or more weights to the classification gain values accessed from the audio signal classifier 216 and the input volume accessed from the input volume detector 214.
[0051] In some examples, the volume adjuster 220 resets the gain value applied to the audio signal when a source change is detected (e.g., when the source changes from an FM station to an auxiliary input). In some such examples, the volume adjuster 220 sets the gain value to zero, and the dynamic range compressor 224 performs compression to adjust the volume of the audio signal within the target volume range until the input volume detector 214 and the audio signal classifier 216 provide information about the newly detected audio signal to the volume adjuster 220 to identify a target gain value.
[0052] The volume controller 220 of the illustrated example smoothly transitions between different volume adjustments (e.g., using a smoothing filter, an averaging filter, etc.). In some examples, if the volume controller 220 determines that a large change in the target gain value is needed, the volume controller 220 transitions slowly to the new target gain value. Conversely, the volume controller 220 may transition more quickly between smaller, less perceptible changes in the target gain value. The volume controller 220 of the illustrated example uses a single-pole smoothing filter to transition between target gain values.
[0053] In some examples, the volume adjuster 220 determines whether the updated input volume value from the input volume detector 214 and / or the updated classification output from the audio signal classifier 216 meets a difference threshold relative to the previous input volume value and / or previous classification output. In some such examples, the volume adjuster 220 identifies a new target gain value only if the updated input volume value and / or the updated classification output meets a difference threshold relative to the previous value used to calculate the target gain value.
[0054] The exemplary volume adjuster 220 in the illustrated example applies a target gain value to the audio signal to convert it. In some examples, the volume adjuster 220 performs an initial volume adjustment using fade-in volume adjustment when the input signal detector 204 detects the input audio signal 202 (e.g., minimizes the volume when a new signal is detected and then gradually increases the volume). In some examples, the volume adjuster 220 can set the initial volume value based on a previous volume value for the type of input signal being accessed. For example, if the input audio signal 202 is an FM audio signal, the volume adjuster 220 can identify a previous volume level used for the FM audio signal and set the current initial volume to this value. The volume adjuster 220 can adjust the initial volume of the input audio signal 202 independently or in conjunction with the dynamic range compressor 224 when the input audio signal 202 is first detected.
[0055] 2 identifies the media corresponding to the input audio signal 202. In some examples, the media unit 106 may not include the audio signal identifier 222 and may modify the input audio signal 202 based solely on classification by the audio signal classifier 216. In some examples, the audio signal identifier 222 performs a comparison of a media identifier (e.g., a fingerprint) embedded in the audio signal with a known or reference audio signature to identify the media of the audio signal. In some examples, the example audio signal identifier 222 may find a matching reference media identifier. In such examples, the audio signal identifier 222 may pass identification information specific to the media included in the input audio signal 202 to the volume adjuster 220 and / or the dynamic range compressor 224 to adjust the input audio signal 202. In some examples, the audio signal identifier 222 may interact with an external database (e.g., at a central facility) to find a matching reference signature. In some examples, the audio signal identifier 222 may interact with an internal database (such as, for example, the classification database 218) to find a matching reference signature.
[0056] The exemplary dynamic range compressor 224 in the example shown in FIG. 2 is capable of compressing the input audio signal 202. In some examples, the dynamic range compressor 224 performs audio compression so that the input audio signal 202 has an average volume level that meets a target volume threshold (e.g., associated with a desired volume level). In some examples, the dynamic range compressor 224 is continuously active and performs compression of the input audio signal 202 after any volume adjustments made by the volume adjuster 220 to bring the input audio signal 202 within the target volume threshold (e.g., within ±0.5 dbFS from −21 dbFS). In some examples, the dynamic range compressor 224 functions as the final step in ensuring that the input audio signal 202 is adjusted to fit within the target volume threshold. In some examples, the amount of dynamic range compression performed on the input audio signal 202 is inversely proportional to the output quality of the volume-adjusted audio signal 228 (e.g., the greater the dynamic volume compression, the lower the quality of the volume-adjusted audio signal 228, e.g., the more lossy).
[0057] 2 collects real-time audio monitor 226 real-time loudness measurement data. For example, real-time audio monitor 226 may determine the current audio volume level as an average over a period of time (e.g., 750 milliseconds). In some examples, real-time audio monitor 226 continuously monitors input audio signal 202 for a monitoring period (e.g., 10 seconds, 1 minute, etc.). In such examples, real-time audio monitor 226 may analyze the volume level during the monitoring period to determine whether subsequent adjustment by either volume adjuster 220 or dynamic range compressor 224 is necessary. In some examples, real-time audio monitor 226 continuously monitors input audio signal 202 for the duration of input audio signal 202. In some examples, real-time audio monitor 226 determines whether the average volume level over a period of time (e.g., 750 milliseconds) falls within a target volume range (e.g., within ±0.5 dBFS of −21 dBFS). In response to the volume level not being within the target volume range, the audio signal classifier 216 may reanalyze the characteristics of the input audio signal 202 and attempt to reclassify the input audio signal 202. In some examples, in response to the real-time audio monitor 226 determining that the average volume level over a period of time is not within the target volume range, the volume adjuster 220 and / or the dynamic range compressor 224 further adjust the input audio signal 202.
[0058] The real-time audio monitor 226 in the illustrated example includes and / or accesses a timer to determine whether the period since the previous classification output by the audio signal classifier 216 meets an update time threshold. In some examples, the update time threshold is set by an operator. For example, the real-time audio monitor 226 can be configured with an update time threshold of 3 seconds, meaning that the audio signal classifier 216 will reclassify the audio signal at 3-second intervals (e.g., every 3 seconds, perform the classification process on the past 3 seconds). Additionally or alternatively, the input volume detector 214 in the illustrated example determines the input volume (e.g., average input volume) of the audio signal for the period since the last classification and / or since the last input volume calculation (e.g., 3 seconds in the previous example). In some such examples, after reclassifying the audio signal and / or determining a new input volume, the volume adjuster 220 can determine a new target gain value based on the new classification and / or the new input volume.
[0059]
[0059] Although an exemplary way of implementing the media unit 106 of Figure 2 is illustrated in Figure 4, one or more of the elements, processes, and / or devices illustrated in Figure 2 may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other manner. Additionally, the exemplary source change determiner 212, exemplary input volume detector 214, exemplary audio signal classifier 216, exemplary classification database 218, exemplary volume adjuster 220, exemplary audio signal identifier 222, exemplary dynamic range compressor 224, exemplary real-time audio monitor 226 of Figure 2, and / or more generally, the exemplary input signal detector 204, exemplary compressor gain comparator 206, exemplary volume / power comparator 208, and exemplary audio sample comparator 210 used by the exemplary media unit 106 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. 2, the exemplary source change determiner 212, the exemplary input volume detector 214, the exemplary audio signal classifier 216, the exemplary classification database 218, the exemplary volume adjuster 220, the exemplary audio signal discriminator 222, the exemplary dynamic range compressor 224, the exemplary real-time audio monitor 226, and / or, more generally, the exemplary input signal detector 204, the exemplary compressor gain comparator 206, the exemplary volume / power comparator 208, and the exemplary audio sampler 216 used by the exemplary media unit 106. Any of the comparators 210 may be implemented by one or more analog or digital circuit(s), logic circuit(s), programmable processor(s), programmable controller(s), graphics processing unit(s) (GPU(s)), digital signal processor(s) (DSP(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)).When reading any of the apparatus or system claims of this patent that include a purely software and / or firmware implementation, at least one of the example source change determiner 212, example input volume detector 214, example audio signal classifier 216, example classification database 218, example volume adjuster 220, example audio signal identifier 222, example dynamic range compressor 224, example real-time audio monitor 226 of FIG. 2, and / or, more generally, the example input signal detector 204, example compressor gain comparator 206, example volume / power comparator 208, and example audio sample comparator 210 used by the example media unit 106, is expressly defined herein to include a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disk (DVD), compact disk (CD), Blu-ray disk, etc., that includes software and / or firmware. Further still, the example media unit 106 of Figure 1 may include one or more elements, processes, and / or devices in addition to or instead of those shown in Figure 2, and / or may include two or more of any or all of the illustrated elements, processes, and devices. As used herein, the phrase "communicate," including variations thereof, includes direct communication and / or indirect communication via one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or constant communication, but rather further includes selective communication at regular intervals, scheduled intervals, non-regular intervals, and / or one-time events.
[0060] FIG. 3 is a block diagram illustrating an audio classification engine 300 capable of providing trained models for use by the media unit 106 of FIGS. 1 and 2. Machine learning techniques, whether deep learning networks or other experience / observation-based learning systems, can be used, for example, to optimize results, find objects in images, understand speech and convert speech to text, improve the relevance of search engine results, and so on. While many machine learning systems are seeded with initial features and / or network weights that are then modified through learning and updating of the machine learning network, deep learning networks train themselves to identify features that are “effective” for analysis. The use of multi-layer architectures allows machines employing deep learning techniques to process raw data better than machines using traditional machine learning techniques. The use of various layers of evaluation or abstraction facilitates examining data for groups of highly correlated values or distinctive themes.
[0061] Machine learning techniques, whether neural networks, deep learning networks, and / or other experience / observation-based learning system(s), can be used, for example, to generate optimal results, find objects in images, understand speech and convert it to text, improve the relevance of search engine results, etc. Deep learning is a subset of machine learning that uses a set of algorithms to model high-level abstractions of data using deep graphs with multiple processing layers, including linear and nonlinear transformations. While many machine learning systems are seeded with initial features and / or network weights that are then modified through learning and updating of the machine learning network, deep learning networks train themselves to identify features that are “effective” for analysis. The use of multi-layer architectures allows machines employing deep learning techniques to process raw data better than machines using traditional machine learning techniques. The use of various layers of evaluation or abstraction makes it easier to examine data for groups of highly correlated values or distinctive themes.
[0062]
[0062] For example, deep learning using convolutional neural networks (CNNs) finds and identifies learned observable features in data by segmenting the data using convolutional filters. Each filter or layer in the CNN architecture transforms the input data to increase the selectivity and invariance of the data. This abstraction of the data allows the machine to focus on the features of the data it is trying to classify and ignore irrelevant background information.
[0063] Deep learning works on the condition that many data sets contain high-level features that subsume lower-level features. For example, when examining an image, rather than looking for objects, it is more efficient to look for edges that form motifs that form parts that form the object you are looking for. These feature hierarchies can be found in many different forms of data.
[0064]
[0064] Learned observable features include the objects and quantifiable regularities learned by the machine during supervised learning. A machine provided with a large set of well-classified data is better equipped to distinguish and extract features associated with successful classification of new data.
[0065]
[0065] A deep learning machine utilizing transfer learning can successfully link data features to specific classifications confirmed by human experts. Conversely, the same machine can update classification parameters if informed of an incorrect classification by a human expert. Settings and / or other configuration information can be guided, for example, by learned use of settings and / or other configuration information, which can reduce the number of variations and / or other possibilities for settings and / or other configuration information for a given situation as the system is used more (e.g., repeatedly and / or by multiple users).
[0066]
[0066] An exemplary deep learning neural network can be trained, for example, on a set of expert-classified data. This set of data establishes the initial parameters of the neural network, which constitutes the supervised learning stage. During the supervised learning stage, the neural network can be tested to see if the desired behavior has been achieved.
[0067] Once the desired neural network behavior is achieved (e.g., the machine has been trained to perform according to specified thresholds), the machine can be deployed and used (e.g., by testing the machine on “real” data). During operation, the neural network's classifications can be confirmed or rejected (e.g., by an expert user, an expert system, a reference database, etc.) to continue improving the neural network's behavior. The exemplary neural network then enters a state of transfer learning, as the classification parameters that specify the neural network's behavior are updated based on ongoing interactions. In certain examples, a neural network, such as neural network 302, can provide direct feedback to other processes, such as audio classification scoring engine 304. In certain examples, neural network 302 outputs data, which is buffered (e.g., via the cloud, etc.) and validated before being provided to other processes.
[0068] In the example of FIG. 3 , the neural network 302 receives input from previous results data related to classification training data and outputs an algorithm for predicting a classification group associated with an audio signal. The network 302 can be seeded with some initial correlation and then learn from ongoing experience. In some examples, the neural network 302 continuously receives feedback from at least one classification training data. In the example of FIG. 3 , throughout the operational life of the audio classification engine 300, the neural network 302 is continuously trained via feedback, and the exemplary audio classification scoring engine 304 can be updated based on the neural network 302 and / or based on additional classification training data as desired. The network 302 can learn and evolve based on role, location, situation, etc.
[0069] In some examples, the accuracy of the model generated by the neural network 302 can be determined by an exemplary audio classification scoring engine verifier 306. In such examples, at least one of the audio classification scoring engine 304 and the audio classification scoring engine verifier 306 receives a set of classification training data. Further, in such examples, the audio classification scoring engine 304 receives input related to classification verification data and predicts one or more audio classifications associated with the classification verification data. The predicted results are distributed to the audio classification scoring engine verifier 306. The audio classification scoring engine verifier 306 additionally receives known audio classifications associated with the classification verification data and compares the known audio classifications with the predicted classifications received from the audio classification scoring engine 304. In some examples, this comparison results in an accuracy of the model generated by the neural network 302 (e.g., if 95 comparisons result in a match and 5 results in an error, the model is 95% accurate, etc.). Once the neural network 302 reaches a desired accuracy (e.g., once the network 302 is trained and ready for deployment), the audio classification scoring engine verifier 306 can output the model to the audio signal classifier 216 of FIG. 2 for use in classifying audio other than the classification training data and / or classification validation data.
[0070] Flowcharts depicting exemplary hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the media unit 106 of FIG. 2 are shown in FIGS. 4 and 5. The machine-readable instructions may be an executable program, or portions of an executable program, executed by a computer processor, such as the processor 612 shown in the exemplary processor platform 600 described below in connection with FIG. 6. The program may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 612, although the entire program and / or portions thereof may alternatively be executed by devices other than the processor 612 and / or embodied in firmware or dedicated hardware. Furthermore, although the exemplary program is described with reference to the flowcharts shown in FIGS. 4 and 5, many other ways of implementing the exemplary media unit 106 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the described blocks may be modified, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0071] 4 and 5 may be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on a non-transitory computer- and / or machine-readable medium, such as a hard disk drive, flash memory, read-only memory, compact disc, digital versatile disc, cache, random access memory, and / or any other storage device or storage disk in which information is stored for any period of time (e.g., long-term, permanently, for short moments, for temporary buffering, and / or for caching information). As used herein, the term non-transitory computer-readable medium is defined to include any type of computer-readable storage device and / or storage disk, to exclude propagating signals, and to explicitly exclude transmission media.
[0072]
[0072] "Including" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim uses any form of "include" or "comprise" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within any type of claim recitation, it is to be understood that additional elements, terms, etc. can be present without departing from the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in a claim preamble or the like, the phrase is open-ended, just as the terms "comprising" and "including" are open-ended. The term "and / or," when used in the form, e.g., A, B, and / or C, refers to any combination or subset of A, B, and C, e.g., (1) A only, (2) B only, (3) C only, (4) A and B, (5) A and C, (6) B and C, and (7) A, B, and C. As used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" refers to an implementation that includes any of (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" refers to an implementation that includes any of (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein, when used in the context of describing the implementation or performance of a process, instruction, action, activity, and / or step, the phrase "at least one of A and B" shall refer to an implementation that includes any of: (1) at least one A; (2) at least one B; and (3) at least one A and at least one B.Similarly, herein, when used in the context of describing the implementation or performance of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" shall refer to implementations that include any of: (1) at least one A; (2) at least one B; and (3) at least one A and at least one B.
[0073] 4 and 5 show exemplary machine-readable instructions executable to implement dynamic volume adjustment via audio classification for implementing the media unit 106 of FIGS. 1 and 2. With reference to the foregoing figures and associated description, the exemplary machine-readable instructions 400 begin at block 402. In block 402, the exemplary media unit 106 detects a change in audio signal input. In some examples, the input signal detector 204 detects a change in audio signal input. For example, an audio signal may have started (e.g., the audio signal previously accessed by the media unit 106 is gone and a new one has started) or the audio signal may have changed (e.g., an FM radio signal has changed to an AM radio signal). Execution of block 402 is described in more detail below in connection with FIG. 5.
[0074]
[0074] In block 404, the exemplary media unit 106 compresses the input audio signal 202 to meet the target volume range. In some examples, the dynamic range compressor 224 compresses the input audio signal 202 to meet the target volume range.
[0075] At block 406, the exemplary media unit 106 identifies a classification group for the input audio signal 202. In some examples, the audio signal classifier 216 identifies the classification group for the input audio signal. In some examples, the audio signal classifier 216 identifies the classification group based on a comparison of one or more characteristics of the input audio signal (e.g., a CQT value) to a trained machine learning model. The audio signal classifier 216 can additionally or alternatively determine a probability distribution associated with one or more classification groups.
[0076] At block 408, the exemplary media unit 106 determines the input volume of the input audio signal 202. In some examples, the input volume detector 214 determines the input volume of the input audio signal 202. In some examples, the input volume detector 214 determines an average input volume of the input audio signal 202 over a period of time (e.g., 3 seconds, 5 seconds, etc.). In some examples, the input volume detector 214 determines a deviation in the volume of the input audio signal 202 over a period of time. In some examples, the input volume detector 214 measures one or more instantaneous volume values.
[0077] At block 410, the exemplary media unit 106 identifies classification gain values using a lookup table associated with the classification groups of the input audio signal 202. In some examples, the audio signal classifier 216 identifies classification gain values using a lookup table associated with one or more classification groups identified by the audio signal classifier 216 to be associated with the input audio signal 202. In some examples, the classification gain value is a single value representing the classification group (e.g., based on the average dynamic range observed in the training data for the classification group, based on the average loudness observed in the training data for the classification group, etc.). In some examples, the classification gain value is determined based on a probability distribution output by the audio signal classifier 216 (e.g., one or more gain values are calculated based on the probability that the input audio signal 202 belongs to one or more of the classification groups).
[0078] At block 412, the exemplary media unit 106 weights the input volume and the classification gain values to determine a target gain value. In some examples, the volume adjuster 220 applies a first weight to the input volume and a second weight to the classification gain values, and then determines a target gain value based on the weighted input volume and the weighted classification gain values. In some examples, the input volume indicates the actual state of the audio signal as opposed to an estimate of the classification gain value, so the volume adjuster 220 applies a greater weight to the input than the classification gain value. In some examples, the volume adjuster 220 determines a target gain value as a value between the input volume measurement and a target volume range. In some examples, the volume adjuster 220 calculates an average between the input volume and the volume level obtained by applying the classification gain values, and a target classification gain value is determined as the gain required to bring the volume of the input audio signal 202 to this average volume level.
[0079] At block 414, the exemplary media unit 106 applies the target gain value to the audio signal using a smoothing filter. In some examples, the volume adjuster 220 applies the target gain value to the input audio signal 202 using a smoothing filter. The volume adjuster 220 can utilize different types of filters (e.g., a median filter, a Kalman filter, etc.) to smooth the transition between the first gain value and the updated gain value (e.g., when the classification and / or input volume is updated), or between no gain value and a gain value (e.g., when a new audio signal is detected).
[0080] At block 416, the exemplary media unit 106 adjusts the compression value to meet the target volume range. In some examples, the dynamic range compressor 224 adjusts the compression value to meet the target volume range. For example, if the volume adjuster 220 increases the gain value applied to the input audio signal 202, the dynamic range compressor 224 may decrease the compression value because less dynamic range compression is required to bring the input audio signal 202 within the target volume range. Conversely, if the volume adjuster 220 decreases the gain value applied to the input audio signal 202, the dynamic range compressor 224 may increase the compression value because more dynamic range compression is required to bring the input audio signal 202 within the target volume range.
[0081] At block 418, the exemplary media unit 106 determines whether the time since the last classification meets or exceeds an update time threshold. In some examples, the real-time audio monitor 226 determines whether the time since the last classification was performed meets or exceeds an update time threshold. In some examples, the real-time audio monitor 226 determines whether the time since the last input volume calculation was performed and / or the time since the last volume adjustment performed by the volume adjuster 220 meets or exceeds an update time threshold. In response to the time since the last classification meeting or exceeding the update time threshold, processing proceeds to block 424. Conversely, in response to the time since the last classification neither meets nor exceeds the update time threshold, processing proceeds to block 420.
[0082] At block 420, the exemplary media unit 106 determines whether a change in audio input source has occurred. In some examples, the input signal detector 204 determines whether a change in audio input source has occurred (e.g., the input source has changed from FM radio to an auxiliary input, the input source has changed from CD to AM radio, etc.). In response to a change in audio input source having occurred, processing proceeds to block 422. Conversely, in response to no change in audio input source having occurred, processing proceeds to block 418. Execution of block 420 is described in more detail below in connection with FIG. 5.
[0083] At block 422, the exemplary media unit 106 resets the gain value. In some examples, the volume adjuster 220 resets the gain value. For example, the volume adjuster 220 may set the gain value to zero because a previous target gain value (identified for a previous audio signal from a different input source) may no longer be valid for the new audio signal. Thus, by the time a new target gain value is identified (e.g., following classification and input volume identification), the gain value is reset to one, and the dynamic range compressor 224 compresses the input audio signal 202 to meet the target volume range.
[0084] At block 424, the exemplary media unit 106 determines the input volume over a period of time since the last classification. In some examples, the input volume detector 214 determines the input volume over a period of time since the last classification. For example, if the real-time audio monitor 226 is configured with a 3-second update interval, once the entire update interval has elapsed (e.g., at block 418), the input volume detector 214 determines the input volume for the update interval. In some examples, an average input volume during the update interval is determined.
[0085] At block 426, the exemplary media unit 106 identifies an updated classification group based on the audio signal over a period of time since the last classification. In some examples, the audio signal classifier 216 identifies an updated classification group based on the audio signal over a period of time since the last classification. For example, if the real-time audio monitor 226 is configured with an update interval of 3 seconds, once 3 seconds have passed since the last classification, the audio signal classifier 216 analyzes one or more characteristics of the audio signal to identify an updated classification group. In some examples, the updated classification group is the same as the previously identified classification group.
[0086] At block 428, the exemplary media unit 106 determines whether dynamic volume is enabled. For example, an operator of the media unit 106 can enable or disable dynamic volume (e.g., via a switch, via a setting on the media unit 106, etc.). In response to dynamic volume being enabled, processing proceeds to block 410. Conversely, in response to dynamic volume not being enabled, processing ends.
[0087] 5 is a flowchart illustrating an example process 500 for implementing block 402 and / or block 420 of FIG. 4. The example process 500 begins with block 502. In block 502, the compressor gain comparator 206 compares the current compressor gain to a recent past compressor gain. For example, the compressor gain comparator 206 may compare the gain of the dynamic range compressor 224 associated with the current sample of the input audio signal 202 to an average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with a previous block of samples (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the compressor gain comparator 206 may output a ratio of the current gain of the dynamic range compressor 224 associated with the current sample block of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with the previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.).
[0088] In block 504, the power comparator 208 compares the current volume / power of the input audio signal 202 with the recent past volume / power(s) of the audio signal. For example, the power comparator 208 may compare the current RMS power of the input audio signal 202 with the average (e.g., arithmetic mean, median, etc.) power of the input audio signal 202 associated with a previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the power comparator 208 may query a peak meter output to determine the RMS power. In some examples, the power comparator 208 may output a ratio of the current RMS power to the average of the previous RMS power(s).
[0089] At block 506, the audio sample comparator 210 compares the maximum value of the current audio sample block with the most recent audio sample value(s). For example, the audio sample comparator 210 may compare the current audio sample value of the input audio signal 202 with the average (e.g., arithmetic mean, median, etc.) audio sample value of the input audio signal 202 associated with a previous sample block (e.g., the previous 3 seconds of samples, the previous 5 seconds of samples, the previous 10 seconds of samples, etc.). In some examples, the audio sample comparator 210 may output a ratio of the current audio sample value to the average of the previous sample block.
[0090]
[0090] In block 508, the source change determiner 212 analyzes the audio sample comparison, the compressor gain comparison, and the power comparison to determine whether a source change has occurred. For example, the source change determiner 212 may use regression analysis (e.g., linear regression, binomial regression, least squares, logistic regression, etc.) to determine whether a source change has occurred. In other examples, the source change determiner 212 may use any other suitable means (e.g., neural networks, etc.) to determine whether a source change has occurred.
[0091]
[0091] At block 510, the source change determiner 212 determines whether the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that a source change has occurred. If the source change determiner 212 determines via logistic regression or other classification method that the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that a source change has occurred, the process 500 proceeds to block 512. If the source change determiner 212 determines that the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that a source change has not occurred, the process 500 proceeds to block 514.
[0092]
[0092] In block 512, the source change determiner 212 indicates that a change in source has occurred. For example, the source change determiner 212 can cause the input signal detector 204 to indicate to the media unit 106 that a change in source has occurred.
[0093] At block 514, the source change determiner 212 indicates that a source change has not occurred. For example, the source change determiner 212 may cause the input signal detector 204 to indicate to the media unit 106 that a source change has not occurred. The process 500 then ends.
[0094]
[0094] Figure 6 is a block diagram of an exemplary processor platform 600 configured to execute the instructions of Figure 4 to implement the media unit 106 of Figures 1 and 2. The processor platform 600 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0095] The processor platform 600 of the illustrated example includes a processor 612. The processor 612 of the illustrated example is hardware. For example, the processor 612 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor implements the example source change determiner 212, example input volume detector 214, example audio signal classifier 216, example classification database 218, example volume adjuster 220, example audio signal identifier 222, example dynamic range compressor 224, example real-time audio monitor 226, and / or more generally, the example input signal detector 204, example compressor gain comparator 206, example volume / power comparator 208, and example audio sample comparator 210 of FIG. 2 used by the example media unit 106.
[0096] The processor 612 of the illustrated example includes a local memory 613 (e.g., a cache). The processor 612 of the illustrated example communicates with a main memory, including a volatile memory 614 and a nonvolatile memory 616, via a bus 618. The volatile memory 614 may be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The nonvolatile memory 616 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 614, 616 is controlled by a memory controller.
[0097] The processor platform 600 of the illustrated example also includes an interface circuit 620. The interface circuit 620 can be implemented with any type of interface standard, such as, for example, an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI Express interface.
[0098] In the illustrated example, one or more input devices 622 are connected to interface circuit 620. Input device(s) 622 enable a user to input data and / or commands into processor 612. The input device(s) may be implemented, for example, by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0099]
[0099] One or more output devices 624 are also connected to the interface circuitry 620 of the illustrated example. The output device(s) 624 may be implemented by, for example, a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Accordingly, the interface circuitry 620 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0100] The interface circuitry 620 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate data exchange with external machines (e.g., any type of computing device) via the network 626. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular phone system, etc.
[0101] The processor platform 600 of the illustrated example also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0102]
[0102] The machine-executable instructions 632 of Figure 4 may be stored on the mass storage device 628, the volatile memory 614, the non-volatile memory 616, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0103] From the foregoing, it can be seen that exemplary methods, apparatus, and articles of manufacture are disclosed that adjust the volume of media so that media with different characteristics are played at approximately the same volume, while minimizing the amount of compression required to achieve this volume. While conventional volume equalization implementations rely solely on compression, resulting in noticeable changes to the audio signal, the examples disclosed herein intelligently classify audio signals and enable identification of an average gain value based on a classification associated with the audio signal that, for example, distinguishes between audio signals with a relatively narrow dynamic range that can be significantly modified with gain values and audio signals with a wider dynamic range that may require more compression. The exemplary techniques disclosed herein utilize a combination of input volume measurements and parameters related to the classification of the audio signal to intelligently adjust the volume of an input audio signal in real time. The examples disclosed herein describe techniques for continuously adjusting the volume level when a volume adjustment needs to be corrected after an initial analysis (e.g., due to a change in the classification of the audio signal, a change in the observed input volume, etc.). The exemplary techniques disclosed herein further include techniques for initially adjusting the volume level of an audio signal after a change in the audio signal input. Such techniques are advantageous over conventional implementations because they are nearly imperceptible to the user and allow different media from different or similar sources to be played at substantially the same volume, enabling a seamless media presentation experience.
[0104] In some examples, like the dynamic volume of the present invention, the exemplary audio dynamic range compressor may be active all the time, reducing the signal to a particular range (e.g., -21 dBFS). In other examples, the audio dynamic range compressor may be active for a certain period of time.
[0105] In some examples, an exemplary real-time loudness detector, such as the dynamic volume of the present invention, can be applied to the input to measure the current average level over one or more intervals (e.g., 750 millisecond intervals). In such examples, the current average level can be used as an initial and ongoing guess to guide how much the volume can be reduced.
[0106] In some examples, a neural network-based classifier can also assist in predicting the dynamic range and inform applicable loudness reduction. This can initially be based on current category classifiers (e.g., 9 classifiers, 15 classifiers, etc.) with potential for improvement. In some examples, increasing the number of current category classifiers can make the dynamic range predictor more accurate, using different real-time and neural network approaches. In each example, accuracy associated with the amount of possible loudness reduction can be improved.
[0107]
[0107] In some examples, the goal is to reduce the volume to near a particular level (e.g., -12 dbFS) that the compressor can reach. Once the amount of reduction is identified, a single-pole smoothing filter can be used to reduce the input's current full volume to the specified amount. The compressor will continue to keep the volume at the specified level on average (e.g., -21 dbFS), but the amount the input needs to be reduced can be less because it is reducing the volume to the target.
[0108] In an illustrative example of the operation of the methods, apparatus, and systems disclosed herein, fully normalized loud pop music may be delivered via the input. The compressor may reduce 0.0 dbFS material to -21 dbFS. Substantially in parallel, the input loudness detector may determine that the input is playing at an average of -1 dbFS, and the classifier may determine that music containing synthetic drums and bass from 1996 to the present is being presented. This category may produce a reduction of -15 dbFS, and the loudness detector may produce -20 dbFS. The two values may be averaged, reducing the signal by -17.50 dbFS, and then by another 3.5 decibels to reach the baseline value of -21 dbFS. Because the compressor reduces signals 3.5 decibels above the threshold (e.g., based on the reduction described above), audio quality is improved compared to reducing signals 21 decibels above the threshold, as would be done if only the compressor were utilized.
[0109]
[0109] Exemplary methods, apparatus, systems, and articles of manufacture for dynamic volume adjustment via audio classification are disclosed herein. Further examples and combinations thereof include the following. Example 1 includes an apparatus comprising: an audio signal classifier that uses a neural network to analyze parameters of the audio signal related to a first volume level to identify a classification group associated with the audio signal; an input volume detector that identifies an input volume of the audio signal; a volume adjuster that applies a gain value to the audio signal, the gain value being based on the classification group and the input volume, the gain value modifying the first volume level to a second volume level; and a dynamic range compressor that applies a compression value to the audio signal, the compression value modifying the second volume level to a third volume level that meets a target volume threshold.
[0110]
[0110] Example 2 includes the apparatus of Example 1, further including a source change determiner that determines whether the source of the audio signal has changed.
[0111]
[0111] Example 3 includes the device described in Example 2, in which the source change determiner determines whether the source of the audio signal has changed based on at least one of: (1) a comparison of a current compressor gain associated with the audio signal to a previous compressor gain associated with the audio signal; (2) a comparison of an RMS power associated with the audio signal to a previous RMS power associated with the audio signal; or (3) a comparison of a current audio sample value associated with the audio signal to a previous audio sample value associated with the audio signal.
[0112] Example 4 includes the apparatus of example 2, wherein the volume control further resets a gain value of the audio signal in response to determining that the source of the audio signal has changed.
[0113]
[0113] Example 5 includes the device described in Example 1, wherein the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the period of music represented by the audio signal, or (3) the presence or absence of instruments in the music represented by the audio signal.
[0114]
[0114] Example 6 includes the device described in Example 1, wherein the input volume detector further determines that the fourth volume level over the first period is not within the target volume threshold, the first period occurring after the second period and the third volume level being associated with the second period, and the dynamic range compressor further adjusts the compression value to a fifth volume level, and the adjusted compression value modifies the fourth volume level to a fifth volume level that meets the target volume threshold.
[0115]
[0115] Example 7 includes the device described in Example 1, where the target volume threshold is within 5 dBFS of 21 dBFS in decibels relative to full scale (dBFS).
[0116]
[0116] Example 8 includes a non-transitory computer-readable storage medium containing instructions that, when executed, cause a processor to at least use a neural network to analyze parameters of the audio signal associated with a first volume level to identify a classification group associated with the audio signal; identify an input volume of the audio signal; apply a gain value to the audio signal, where the gain value is based on the classification group and the input volume, modifying the first volume level to a second volume level; and apply a compression value to the audio signal, where the compression value modifies the second volume level to a third volume level that meets a target volume threshold.
[0117]
[0117] Example 9 includes the non-transitory computer-readable storage medium of Example 8, wherein the instructions, when executed, cause the processor to determine whether the source of the audio signal has changed.
[0118]
[0118] Example 10 includes the non-transitory computer-readable storage medium of Example 9, wherein determining whether the source of the audio signal has changed is based on at least one of: (1) a comparison of a current compressor gain associated with the audio signal to a previous compressor gain associated with the audio signal; (2) a comparison of an RMS power associated with the audio signal to a previous RMS power associated with the audio signal; or (3) a comparison of a current audio sample value associated with the audio signal to a previous audio sample value associated with the audio signal.
[0119]
[0119] Example 11 includes the non-transitory computer-readable storage medium of Example 9, wherein the instructions, when executed, cause the processor to reset a gain value of the audio signal in response to determining that the source of the audio signal has changed.
[0120]
[0120] Example 12 includes the non-transitory computer-readable storage medium of Example 11, in which the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the period of music represented by the audio signal, or (3) the presence or absence of musical instruments in the music represented by the audio signal.
[0121]
[0121] Example 13 includes the non-transitory computer-readable storage medium of Example 8, wherein the instructions, when executed, cause the processor to: determine that the fourth volume level over a first period of time is not within a target volume threshold, where the first period of time occurs after a second period of time and the third volume level is associated with the second period of time; and adjust the compression value to a fifth volume level, where the adjusted compression value modifies the fourth volume level to a fifth volume level that meets the target volume threshold.
[0122] Example 14 includes the non-transitory computer-readable storage medium of Example 8, wherein the target volume threshold is within 5 dBFS of 21 dBFS in decibels relative to full scale (dBFS).
[0123]
[0123] Example 15 includes a method including the steps of: using a neural network to analyze parameters of an audio signal related to a first volume level to identify a classification group associated with the audio signal; identifying an input volume of the audio signal; applying a gain value to the audio signal, where the gain value is based on the classification group and the input volume, and the gain value modifies the first volume level to a second volume level; and applying a compression value to the audio signal, where the compression value modifies the second volume level to a third volume level that meets a target volume threshold.
[0124] Example 16 includes the method of example 15, further including determining if the source of the audio signal has changed.
[0125]
[0125] Example 17 includes a method as described in Example 16, in which the step of determining whether the source of the audio signal has changed is based on at least one of: (1) a comparison of a current compressor gain associated with the audio signal to a previous compressor gain associated with the audio signal; (2) a comparison of an RMS power associated with the audio signal to a previous RMS power associated with the audio signal; or (3) a comparison of a current audio sample value associated with the audio signal to a previous audio sample value associated with the audio signal.
[0126] Example 18 includes the method of example 16, further including resetting a gain value of the audio signal in response to determining that the source of the audio signal has changed.
[0127]
[0127] Example 19 includes a method as described in Example 15, in which the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the period of music represented by the audio signal, or (3) the presence or absence of instruments in the music represented by the audio signal.
[0128]
[0128] Example 20 includes a method as described in Example 15, further including a step of determining that a fourth volume level over a first period of time is not within a target volume threshold, where the first period of time occurs after a second period of time and the third volume level is associated with the second period of time, and a step of adjusting a compression value to modify the fourth volume level to a fifth volume level that meets the target volume threshold.
[0129] Although certain exemplary methods, apparatus, and articles of manufacture are disclosed herein, the scope of coverage of this patent is not limited thereto. Rather, this patent includes all methods, apparatus, and articles of manufacture that fairly fall within the scope of the claims of this patent.
Claims
1. A non-transitory computer-readable medium having stored thereon instructions that, when executed, cause one or more processors to: analyzing parameters of the audio signal associated with a first volume level using a neural network to identify a classification group associated with the audio signal; determining an input volume of the audio signal; applying a gain value to the audio signal in response to identifying the classification group and the input volume, the gain value modifying the first volume level to a second volume level; applying a compression value to the audio signal, the compression value modifying the second volume level to a third volume level that meets a target volume threshold; A non-transitory computer-readable medium that causes a set of operations to be performed, the set including:
2. applying a compression value to the audio signal, (i) decreasing the compression value when the gain value increases; and (ii) increasing the compression value if the gain value decreases; The computer-readable medium of claim 1 further comprising:
3. The set of actions may include: determining whether the source of the input audio signal has changed; The non-transitory computer-readable medium of claim 1 , further comprising:
4. determining whether the source of the input audio signal has changed; (1) comparing a current compressor gain associated with the input audio signal with a previous compressor gain associated with the input audio signal; (2) comparing the RMS power associated with the input audio signal with a previous RMS power associated with the input audio signal; and (3) comparing a current audio sample value associated with the input audio signal with a previous audio sample value associated with the input audio signal; The non-transitory computer-readable medium of claim 3 , based on at least one of:
5. The classification group is (1) the genre of music represented by the input audio signal; (2) the duration of the music represented by the input audio signal; and (3) the presence or absence of instruments in the music represented by the input audio signal; 10. The non-transitory computer-readable medium of claim 1, comprising at least one of:
6. the set of operations determining a classification gain value based on the classification group and the input volume. The non-transitory computer-readable medium of claim 1 , further comprising:
7. the set of operations is identifying a target gain value between the input volume and the classification gain value, the target gain value being determined by applying one or more weights to the input volume and the classification gain value. The non-transitory computer-readable medium of claim 6 further comprising:
8. The non-transitory computer-readable medium of claim 7 , wherein the target gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value.
9. 1. A computer-implemented method for adjusting volume, comprising: analyzing parameters of the audio signal associated with a first volume level using a neural network to identify a classification group associated with the audio signal; determining an input volume of the audio signal; In response to identifying the classification group and the input volume, applying a gain value to the audio signal, the gain value modifying the first volume level to a second volume level; applying a compression value to the audio signal, the compression value modifying the second volume level to a third volume level that meets a target volume threshold; A method comprising:
10. applying a compression value to the audio signal, (i) decreasing the compression value when the gain value increases; and (ii) increasing the compression value if the gain value decreases; 10. The method of claim 9, further comprising:
11. The method of claim 9 , further comprising determining whether the source of the input audio signal has changed.
12. determining whether the source of the audio signal has changed; (1) comparing a current compressor gain associated with the input audio signal with a previous compressor gain associated with the input audio signal; (2) comparing the RMS power associated with the input audio signal with a previous RMS power associated with the input audio signal; and (3) comparing a current audio sample value associated with the input audio signal with a previous audio sample value associated with the input audio signal; The method of claim 11 , wherein the method is based on at least one of:
13. The classification group is (1) the genre of music represented by the input audio signal; (2) the duration of the music represented by the input audio signal; and (3) the presence or absence of instruments in the music represented by the input audio signal; The method of claim 9 , comprising at least one of:
14. determining a classification gain value based on the classification group and the input volume; 10. The method of claim 9, further comprising:
15. identifying a target gain value between the input volume and the classification gain value, the target gain value being determined by applying one or more weights to the input volume and the classification gain value; 15. The method of claim 14, further comprising:
16. The method of claim 15 , wherein the target gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value.
17. 1. A computing device comprising: one or more processors; a non-transitory computer-readable medium having instructions stored thereon; the instructions, when executed, cause one or more processors to: analyzing parameters of the audio signal associated with a first volume level using a neural network to identify a classification group associated with the audio signal; determining an input volume of the audio signal; In response to identifying the classification group and the input volume, applying a gain value to the audio signal, the gain value modifying the first volume level to a second volume level; applying a compression value to the audio signal, the compression value modifying the second volume level to a third volume level that meets a target volume threshold; A computing device that causes a set of operations including:
18. applying a compression value to the audio signal, (i) decreasing the compression value when the gain value increases; and (ii) increasing the compression value if the gain value decreases; 20. The computing device of claim 17, further comprising:
19. The set of actions may include: determining whether the source of the input audio signal has changed; 20. The computing device of claim 17, further comprising:
20. determining whether the source of the input audio signal has changed; (1) comparing a current compressor gain associated with the input audio signal with a previous compressor gain associated with the input audio signal; (2) comparing the RMS power associated with the input audio signal with a previous RMS power associated with the input audio signal; and 20. The computing device of claim 19, wherein the step of determining whether the input audio signal is a current audio sample value is based on at least one of: (3) a comparison of a current audio sample value associated with the input audio signal with a previous audio sample value associated with the input audio signal.
Citation Information
Patent Citations
Audio gain control using auditory event detection based on specific loudness
JP2009535897A
Apparatus and method for audio classification and processing
JP2016519784A
Terminal device and method for outputting its audio signal
JP2016522597A
Method and apparatus for audio normalization
US20040264714A1
Reproduction control of an audio signal based on musical genre classification
WO2005106843A1