Computer-readable media, computer-implemented methods, and computing devices for dynamic volume control via audio classification.

JP7899412B2Active Publication Date: 2026-08-03GRACENOTE INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GRACENOTE INC
Filing Date
2025-06-25
Publication Date
2026-08-03

Smart Images

  • Figure 0007899412000001
    Figure 0007899412000001
  • Figure 0007899412000002
    Figure 0007899412000002
  • Figure 0007899412000003
    Figure 0007899412000003
Patent Text Reader

Abstract

To provide a method, apparatus, system, and articles of manufacture that utilize a combination of classification of an audio signal and real-time input audio measurements to determine a targeted gain value that can be applied to the audio signal.SOLUTION: A method according to the present invention includes: a step (406) of analyzing, using a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; a step (408) of determining an input volume of the audio signal; a step (414) of applying a gain value to the audio signal, where the first volume level is modified to a second volume level, by the gain value; and a step (416) of applying a compression value to the audio signal, where the second volume level is modified to a third volume level that satisfies a target volume threshold, by the compression value.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001]

[0001] This patent claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 728,677, filed on September 7, 2018, and U.S. Provisional Patent Application No. 62 / 745,148, filed on October 12, 2018. U.S. Provisional Patent Application No. 62 / 702,734 and U.S. Provisional Patent Application No. 62 / 745,148 are hereby incorporated by reference in their entirety. Field of the Disclosure

[0002]

[0002] This disclosure generally relates to volume adjustment, and more particularly, to methods and apparatuses for dynamic volume adjustment via audio classification. Background

[0003]

[0003] In recent years, a large number of media with various characteristics have been distributed using an increasing number of channels. These media can be received using more traditional channels (e.g., radio) or more recently developed channels, such as by using a streaming device connected to the Internet. As these channels have evolved, systems have also been developed that are capable of processing and outputting audio from multiple sources. For example, some automotive media systems can distribute media from compact discs (CDs), Bluetooth®-connected devices, universal serial bus (USB)-connected devices, Wi-Fi-connected devices, auxiliary inputs, and other sources.

Brief Description of the Drawings

[0004] [Figure 1]

[0004] FIG. 1 is a schematic diagram of an exemplary system constructed in accordance with the teachings of the present disclosure for dynamic volume adjustment via audio classification.

[0005] [Figure 2]

[0005] FIG. 2 is a block diagram showing further details of the media unit of FIG. 1.

[0006] [Figure 3]

[0006] Figure 3 is a block diagram showing an audio classification engine capable of providing trained models for use by the media units of Figures 1 and 2.

[0007] [Figure 4]

[0007] Figure 4 is a flowchart showing exemplary machine-readable instructions that can be used to implement the media unit 106 of Figures 1 and 2 and perform dynamic volume adjustment via audio classification. [Figure 5]

[0007] Figure 5 is a flowchart showing exemplary machine-readable instructions that can be used to implement the media unit 106 of Figures 1 and 2 and perform dynamic volume adjustment via audio classification.

[0008] [Figure 6]

[0008] Figure 6 is a schematic diagram of an exemplary processor platform that can implement the exemplary media unit 106 of Figures 1 and 2 by executing the instructions of Figures 4 and 5.

[0009]

[0009] The drawings are not to scale. Throughout the drawings(s) and the accompanying specification, the same reference numerals are used whenever possible to refer to the same or similar parts. Detailed description

[0010]

[0010] In conventional audio media implementations, audio signals associated with different media may have different volume levels. For example, one CD may be recorded and / or mastered at a significantly different volume than another CD. Similarly, media obtained from a streaming device may have a significantly different volume level than media obtained from a different device, or media obtained from the same device through a different application. As users increasingly listen to media from a variety of different sources, the differences in volume levels between sources and between media from the same source can become very noticeable and potentially frustrate listeners.

[0011]

[0011] Some conventional approaches to volume equalization utilize dynamic range compressors to compress the entire dynamic range of the audio signal to meet a volume threshold. In some conventional implementations, such dynamic range compression continuously monitors and adjusts the volume of the audio signal to meet the volume threshold. Such continuous adjustments have a clearly noticeable effect on the listener's perception of the audio signal, as they significantly alter the original dynamics of the track. In some examples, dynamic range compression significantly degrades the perceived quality of the audio signal (for example, by introducing artifacts into the audio).

[0012]

[0012] The exemplary methods, apparatus, systems, and articles disclosed herein use audio classification to identify the category of an audio signal, and then perform volume adjustment to minimize the amount of dynamic range compression required to bring the audio signal within a target volume range. The exemplary methods, apparatus, systems, and articles disclosed herein utilize a combination of audio signal classification and real-time input audio measurements to identify a target gain value applicable to the audio signal. For example, after identifying the classification group associated with the audio signal, the classification gain value can be obtained (for example, from a lookup table that associates volume gain adjustment values ​​with classification groups). Furthermore, the input volume of the audio signal can be identified. Then, based on the input volume and the recommended classification gain value, a target gain value can be identified. The target gain value is a volume adjustment applied to the input audio signal to bring the volume closer to a target volume range (e.g., within ±1 dBFS from -21 dBFS), resulting in a reduced amount of compression required to bring the gain-adjusted signal within the target volume range when the gain-adjusted signal is provided to the compressor.

[0013]

[0013] The exemplary methods, apparatus, systems, and articles disclosed herein calculate a target gain value based on the classification of an input audio signal and the input volume of the audio signal to reduce the amount of compression required to bring the volume of the audio signal within a target volume range. In some examples, when the input audio signal is first detected, the dynamic range of the audio signal is initially compressed so that the volume of the audio signal is within a target volume range before the input audio signal is classified and the volume of the input audio signal is identified. In some examples, if compression alone is used to adjust the audio signal when the audio signal is first detected, the listener may easily notice that the compression is a reduction in audio level without manual volume adjustment. However, once the initial volume of the audio signal and the classification of the audio signal are identified, a target gain value is calculated to reduce the amount of compression required to bring the volume of the audio signal within a target volume range. In some examples, the classification and identification of the initial volume can be performed quickly enough (e.g., within 5 seconds, within 1 second, etc.) so that the use of initial compression is not noticed by the listener.

[0014]

[0014] Some exemplary methods, apparatus, systems, and articles disclosed herein identify and address changes in the source of an audio signal. In some examples, initial volume adjustment is performed in addition to, or instead of, the use of compression. For example, in response to a change in the audio signal input (e.g., a change from no audio signal to audio signal presentation, a change from one audio signal input source to another, etc.), an initial volume level can be identified (e.g., based on previous volume adjustment settings specific to the source of the audio signal), and an initial volume level adjustment can be performed. In some examples, the initial volume level adjustment is performed using a “fade-in” technique, which gradually increases the audio volume level after the change in the input signal. In some examples, the initial volume level adjustment can be based on stored settings associated with the type of audio input signal (e.g., FM radio, AM radio, CD, auxiliary audio source, etc.).

[0015]

[0015] The exemplary methods, apparatus, systems, and articles disclosed herein classify audio signals into one or more classification groups. When identifying classification groups, the characteristics of the classification groups (e.g., the amount of available headroom, typical dynamic range, etc.) can be used to adjust the volume of the audio signal with minimal loss (e.g., by using minimal dynamic range compression). In some examples, pattern recognition in training data can be used to identify classification groups. For example, audio signals can be grouped based on factors such as the instruments represented in the signal, the year the audio signal was produced, or the genre of music. Once the training data is grouped, characteristics such as the distribution of dynamic range values, the distribution of volume values, or any other arbitrary audio characteristics are stored (e.g., in a lookup table) in association with the classification groups. In some examples, when classifying audio signals, a probability distribution can be determined (as opposed to outputting one specific classification group to which the audio signal belongs, for example). For example, the classification process could output that there is a 50% chance that the audio signal belongs to the group representing drumless music from 1976 to 1995, a 30% chance that the audio signal belongs to the group representing drumless music from 1996 to the present, an 18% chance that the audio signal belongs to the group representing music with synthesized drums from 1976 to 1995, or a 2% chance that it belongs to any other group. In some such examples, selecting gain values ​​associated with the classification groups to perform volume adjustments could involve averaging techniques (e.g., identifying the gain values ​​associated with each group and weighting each value according to the probability that the audio signal belongs to each group).

[0016]

[0016] In some exemplary methods, apparatus, systems, and articles disclosed herein, an audio signal classifier is trained to perform classification of audio signals using a large corpus of volume profiles of a variety of representative audio signals (e.g., representing multiple genres, multiple time periods, etc.). For example, a volume profile contains volume values ​​at multiple times in a song. In some examples, in addition to or instead of volume profiles, other profiles and / or representations of audio signals can be used to train the audio signal classifier. In some examples, clustering is performed on volume profiles to train the audio signal classifier. In some examples, the audio signal classifier is trained to identify clusters of volume profiles based on volume, dynamic range, and / or any other properties of the volume profiles. The audio signal classifier can cluster volume profiles into groups of dynamic range, and then the audio signal classifier can assign incoming audio (e.g., an input audio signal) to one or more of these classification groups.

[0017]

[0017] The exemplary methods, apparatus, systems, and articles disclosed herein allow for the adjustment of the volume level of an audio signal by applying a gain value to the audio signal after identifying a classification group of the audio signal. The gain value can be specific to the classification group. For example, if the classification group is associated with an audio signal having a relatively narrow normalized dynamic range (such as some pop music), a significant volume adjustment can be made to bring the volume level of the audio signal closer to a target volume range (for example, because it is possible to determine the approximate volume deviation of the entire track). Conversely, if the classification group is associated with an audio signal having a relatively wide dynamic range, less volume adjustment can be made to keep the audio signal within the audible range.

[0018]

[0018] Following the application of a gain value based on the classification group associated with the audio signal, compression can be used to bring the volume of the audio signal within a target volume range. Since dynamic range compression can result in a reduction of overall audio quality (e.g., some loss of the audio signal), the exemplary methods, apparatus, systems, and products disclosed herein improve volume adjustment techniques by first applying a gain value specific to the type of audio being presented (e.g., specific to the classification group), and thus reduce the amount of dynamic range compression required to bring the volume level of the audio signal within a target volume range.

[0019]

[0019] In some exemplary methods, apparatus, systems, and articles disclosed herein, when an audio signal is classified in a dynamic volume adjustment situation, the characteristics of the audio signal are estimated from its classification group, and these characteristics are used to determine a target gain value for bringing the volume of the audio signal closer to a target volume threshold with minimal or no compression.

[0020]

[0020] In some exemplary methods, apparatus, systems, and products disclosed herein, input volume measurements are taken into consideration when determining a target gain value. For example, if the input volume is determined to be -15 dBFS and the target volume range is within ±1 dBFS from -21 dBFS (e.g., -20 dBFS to -22 dBFS), the target gain value must be a negative gain value smaller than if the input volume were determined to be -10 dBFS, even if the classification group were constant. In some such examples, input volume measurements are weighted more heavily than classification gain values ​​when determining a target gain value because, ultimately, the actual input volume level of a particular audio signal indicates more than the class-based prediction about how much volume can be adjusted (e.g., real-time measurements may be more accurate than the prediction associated with the class of an audio signal). In some examples, the average of the classification gain value and the input volume is taken to calculate the target gain value. For example, if the input volume is identified as -15 dBFS, and the classification gain value (e.g., one identified based on the average dynamic range of the audio signals in the classification group) indicates that the volume can be adjusted by -6 dBFS, but the target volume range is ±1 dBFS from -21 dBFS, relying solely on the classification gain value results in an extremely small margin of error (e.g., if the dynamic range is wider than expected, the volume is likely to often fall outside the target volume range of -020 dBFS to 22 dBFS). Instead, if the target gain value is calculated as an intermediate value (e.g., the average) between the input volume and the classification gain value, the target gain value will bring the volume of the audio signal closer to the target gain value, while still leaving some margin of error.

[0021]

[0021] In some exemplary methods, apparatus, systems, and products disclosed herein, input volume levels are measured at regular intervals (e.g., every 3 seconds, every 10 seconds, etc.), and classification is performed at regular intervals. A new target gain value can be identified in response to changes in input volume (e.g., changes in the average input volume for that interval, changes in the deviation of the input volume for that interval), and / or changes in the classification group. In some examples, when transitioning between target gain values, a smoothing filter can be used to smoothly transition between the two gain values, thereby avoiding significant fluctuations in volume at each interval. In some examples, larger changes in the target gain value are sloped at a slower rate than relatively smaller changes in the target gain value.

[0022]

[0022] The exemplary methods, apparatus, systems, and articles disclosed herein adjust the volume level of an audio signal to a target volume range. In some examples, the listener can then manually adjust the volume level (e.g., by turning a volume knob, providing a voice command to change the volume level), which is then done by applying a gain value to the volume-adjusted audio signal. Thus, the listener can still choose the volume at which they hear the audio signal, but rather choose the volume from a consistent, standard volume level (e.g., from a target volume range) rather than adjusting for variations between different sources, variations between tracks, etc. Therefore, the techniques disclosed herein enable adjusting input audio to be locked into a consistent volume range. In some exemplary methods, apparatus, systems, and articles disclosed herein, dynamic volume adjustment can be discontinued when manual volume adjustment is performed. For example, if the user manually adjusts the volume level (e.g., by turning a volume knob or providing a voice command to change the volume level), the automatic audio level adjustment (e.g., by classifying the audio, selecting a gain value based on the classification, and monitoring the audio level) can be stopped, allowing the user to have complete control over the audio level.

[0023]

[0023] Some exemplary methods, apparatus, systems, and articles disclosed herein can identify audio signals to further improve volume control. For example, some exemplary techniques disclosed herein utilize audio fingerprinting to identify media in order to obtain metadata associated with audio signals. Audio fingerprinting is a technique used to identify media such as television broadcasts, radio broadcasts, advertisements (television and / or radio), downloadable media, streaming media, and packaged media. Existing audio watermarking techniques identify media by embedding one or more audio codes (e.g., one or more fingerprints), such as media identification information and / or identifiers that can be mapped to media identification information, into audio and / or video components. In some examples, the audio or video components are selected to have signal characteristics sufficient to conceal the watermark. As used herein, the terms “fingerprint,” “code,” “signature,” and “watermark” are used synonymously and are defined to mean any identifying information (e.g., an identifier) ​​that can be inserted into or embedded in the audio or video of media (e.g., a program or advertisement) for the purpose of identifying media or for other purposes such as tuning (e.g., a packet identification header). As used herein, “media” means audio and / or visual (still or moving) content and / or advertisements. To identify the fingerprinted media, the fingerprint(s) are extracted and used to access a table of reference fingerprints, which are mapped to media identification information.

[0024]

[0024] In the examples disclosed herein, volume control can be performed by a component of the vehicle's audio system, or by a component that communicates with the audio system. In some examples, a media unit including a dynamic volume control, or other component capable of dynamic volume control, can be included in the vehicle's head unit. In such examples, the vehicle head unit can receive audio signals from auxiliary inputs, CD inputs, radio signal receiver inputs, external streams from smart devices, Bluetooth inputs, network connections (e.g., internet connection), or any other source. For example, dynamic volume control can be performed on a media system of a home entertainment system, where multiple sources (e.g., DVD player, set-top box, etc.) can transmit audio signals, and the audio signals are dynamically adjusted to attempt to normalize the volume levels between the sources and media. In other examples, dynamic volume control can be performed in any situation or for any media device(s).

[0025]

[0025] In an exemplary procedure for dynamic volume adjustment via audio classification, an audio signal corresponding to normalized loud pop music is accessed. After detecting a change in the audio signal input associated with the audio signal, a dynamic range compressor compresses the audio to a target volume range (e.g., -21 dbFS). In parallel with this compression, an audio signal classifier identifies a classification group corresponding to the audio signal. For example, the classification group can correspond to music containing synthesized drums and bass from the period 1996 to the present. This classification group can be associated with a specific volume adjustment level (e.g., -15 dbFS). In some examples, this volume adjustment level associated with the classification group can be considered in addition to, or instead of, a volume level adjustment value specified based on the current audio volume level. Following the volume adjustment associated with this volume adjustment level, it is possible to reach the target volume range by performing only a small amount of audio compression. For example, if the volume is lowered to a first value (e.g., -X dbFS) by a volume adjustment step and the target volume range is near a second value (e.g., -21 dbFS) that exceeds the first value, a small amount of audio compression can be performed to bring the audio signal to the second value (e.g., near -21 dbFS and within the target volume range). Thus, only dynamic range compression that reduces the signal by a small amount (e.g., 3.5 dbFS) is performed, and the audio quality is significantly better than reducing the signal that requires compression from the original audio input to the target volume range (e.g., compressing the audio signal to -21 dbFS).

[0026]

[0026] FIG. 1 is a schematic diagram of an exemplary system 100 constructed in accordance with the teachings of the present disclosure for dynamic volume adjustment. The exemplary system 100 includes media devices 102, 104 that transmit an audio signal to a media unit 106. The media unit 106 processes the audio signal and transmits the signal to an audio amplifier 108, and subsequently, the audio amplifier 108 outputs the amplified audio signal for presentation via an output device 110.

[0027]

[0027] The exemplary media device 102 in the example shown in FIG. 1 is a portable media player (e.g., an MP3 player). The exemplary media device 102 can store or receive an audio signal corresponding to media and transmit the audio signal to other devices. In the example shown in FIG. 1, the media device 102 transmits an audio signal to the media unit 106 via an auxiliary cable. In some examples, the media device 102 can transmit an audio signal to the media unit 106 via any other interface.

[0028]

[0028] The exemplary media device 104 in the example shown in FIG. 1 is a mobile device (e.g., a mobile phone). The exemplary media device 104 can store or receive an audio signal corresponding to media and transmit the audio signal to other devices. In the example shown in FIG. 1, the media device 104 wirelessly transmits an audio signal to the media unit 106. In some examples, the media device 104 can use Wi-Fi, Bluetooth, and / or any other technology to transmit an audio signal to the media unit 106. In some examples, the media device 104 can communicate with vehicle components or other devices to select media presented by the listener in the vehicle. The media devices 102, 104 can be any device capable of storing and / or accessing an audio signal. In some examples, the media devices 102, 104 can be integrated into a vehicle (e.g., a CD player, a radio, etc.).

[0029]

[0029] The exemplary media unit 106 in the example shown in Figure 1 is capable of receiving and processing audio signals. In the example shown in Figure 1, the exemplary media unit 106 receives media signals from media devices 102 and 104, processes the media signals, and performs dynamic volume adjustment. The exemplary media unit 106 is capable of identifying audio signals based on identifiers embedded in the media (e.g., fingerprints, watermarks, signatures, etc.). The exemplary media unit 106 is further capable of accessing metadata corresponding to the media associated with the audio signals. In some examples, the metadata is stored in the storage device of the media unit 106. In some examples, the metadata is accessed from another location (e.g., from a server over a network). Furthermore, the exemplary media unit 106 is capable of performing dynamic volume adjustment by determining and applying an average gain value based on the metadata to adjust the average volume of the audio signal to meet a volume threshold. The exemplary media unit 106 is further capable of monitoring the audio being output by the output device 110 to determine the average volume level of the audio segment in real time. If an audio signal is not identified as corresponding to a media, and / or metadata containing volume information is not available for the audio signal, the exemplary media unit 106 is capable of dynamic range compression, which provides compression of the audio signal to achieve a desired volume level. In some examples, the exemplary media unit 106 is included as part of another device in the vehicle (e.g., a car radio head unit). In some examples, the exemplary media unit 106 is implemented as software and is included as part of another device available either via a direct connection (e.g., a wired connection) or a network (e.g., one available on the cloud). In some examples, the exemplary media unit 106 can be incorporated into an audio amplifier 108 and an output device 110, and can output the audio signal independently following processing of the audio signal.

[0030]

[0030] The exemplary audio amplifier 108 in the example shown in Figure 1 is a device capable of receiving an audio signal processed by the media unit 106 and performing appropriate amplification of the signal for output by the output device 110. In some examples, the audio amplifier 108 can be integrated with the output device 110. In some examples, the audio amplifier 108 amplifies the audio signal based on the amplified output value from the media unit 106. In some examples, the audio amplifier 108 amplifies the audio signal based on input from a listener (for example, adjustment of a volume selector by a passenger or driver of a vehicle).

[0031]

[0031] The exemplary output device 110 in the example shown in Figure 1 is a speaker. In some examples, the output device 110 can be multiple speakers, headphones, or any other device capable of presenting an audio signal to a listener. In some examples, the output device 110 can also output visual elements (for example, a television with speakers).

[0032]

[0032] The exemplary system 100 shown in Figure 1 is described with reference to the implementation of dynamic volume control in a vehicle, but some or all of the devices included in the exemplary system 100 can be implemented in any combination in any environment. For example, system 100 may be located in a home entertainment room, and media devices 102, 104 may be game consoles, virtual reality devices, set-top boxes, or any other devices capable of accessing and / or transmitting media. Furthermore, in some examples, the media may also include visual elements (e.g., television programs, movies, etc.).

[0033]

[0033] Figure 2 is a block diagram 200 that provides further details of an exemplary implementation of the media unit 106 shown in Figure 1. The exemplary media unit 106 is capable of receiving and processing an audio signal to dynamically adjust the volume of the audio signal within a target volume range. Following the dynamic volume adjustment, the exemplary media unit 106 sends the volume-adjusted audio signal 228 to an audio amplifier 108 to be amplified before being output by the output device 110.

[0034]

[0034] The exemplary media unit 106 includes an exemplary input audio signal 202 and an exemplary input signal detector 204. This signal detector includes an exemplary compressor gain comparator 206, an exemplary audio volume / power comparator 208, and an exemplary audio sample comparator 210, all of which are used to determine whether or not a change has occurred in the audio source 212. The exemplary media unit 106 further includes an exemplary input volume detector 214, an exemplary audio signal classifier 216, an exemplary classification database 218, an exemplary volume control 220, an exemplary audio signal discriminator 222, an exemplary dynamic range compressor 224, and an exemplary real-time audio monitor 226. The resulting output from the system is an exemplary volume-controlled audio signal 228.

[0035]

[0035] An exemplary input audio signal 202 is an audio signal that is processed and output to be presented. The input audio signal 202 can be accessed from a radio signal (e.g., an FM signal, AM signal, satellite radio signal, etc.), a compact disc, an auxiliary cable (e.g., one connected to a media device), a Bluetooth signal, a Wi-Fi signal, or any other media. The input audio signal 202 is accessed by an input signal detector 204, an audio signal classifier 216, and / or a real-time audio monitor 226. The input audio signal 202 is converted by a volume control 220 and / or a dynamic range compressor 224.

[0036]

[0036] An exemplary input signal detector 204 detects the input audio signal 202. In some examples, the input signal detector 204 detects whether the input audio signal 202 is related to a new input audio signal or a new input audio signal source (e.g., an AM signal switching to an FM signal, or an auxiliary device signal switching to a CD). In some examples, the input signal detector 204 detects the input audio signal 202 when the input audio signal 202 begins after the media unit 106 has been off (e.g., when the media unit 106 is powered on and the input audio signal 202 begins). In some examples, if the input audio signal 202 is new (e.g., it represents a new type of input audio signal indicating a change in input, or it represents a signal that began after the media unit had not previously presented any audio signal), the input signal detector 204 communicates with the audio signal classifier 216 to begin the classification process. In some examples, the input signal detector 204 determines whether the audio source has changed. For example, the input signal detector 204 can determine whether the audio input source has changed via an exemplary compressor gain comparator 206, an exemplary volume / power comparator 208, and an exemplary audio sample comparator 210, which an exemplary source change detector can then use to determine whether the audio source signal has changed 212.

[0037]

[0037] The exemplary compressor gain comparator 206 compares the current gain of the dynamic range compressor 224 to the previous gain of the dynamic range compressor 224. For example, the compressor gain comparator 206 can compare the gain of the dynamic range compressor 224 associated with the current sample block of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with previous sample blocks (e.g., previous 3-second samples, previous 5-second samples, previous 10-second samples, etc.). In some examples, the compressor gain comparator 206 can output the ratio of the current gain of the dynamic range compressor 224 to the average of the previous gains of the dynamic range compressor 224. In other examples, the compressor gain comparator 206 can output any other appropriate value (e.g., difference, etc.) related to the comparison of the current gain of the dynamic range compressor 224 to the average of the previous dynamic gains of the dynamic range compressor 224.

[0038]

[0038] An exemplary volume / power comparator 208 compares the current power of the input audio signal 202 to the previous power of the input audio signal 202. For example, the power comparator 208 can compare the current power of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) power of the input audio signal 202 associated with previous sample blocks (e.g., previous 3-second samples, previous 5-second samples, previous 10-second samples, etc.). In some examples, the power comparator 208 can compare the root mean square (RMS) power of the current sample of the input audio signal 202 to the RMS power(s) associated with previous samples of the input audio signal 202. In some examples, the power comparator 208 can query the peak output of the media unit 106 to determine the RMS power of the audio sample. In some examples, the power comparator 208 can output the ratio of the current RMS power to the average of previous RMS power(s) after K-weighting has been applied. In other examples, the power comparator 208 may output any other appropriate value (e.g., difference) related to comparing the current RMS power of the input audio signal 202 with the average of previous RMS powers of the input audio signal 202.

[0039]

[0039] An exemplary audio sample comparator 210 compares the current value of a sample of the input audio signal 202 to a previous value of the input audio signal 202. In some examples, the audio sample comparator 210 identifies the value of an audio sample based on the maximum amplitude of the sample in the current block of the input audio signal 202. In some examples, the audio sample comparator 210 identifies the value of an audio sample as a normalized value (e.g., between 1 and -1). In other examples, the audio sample comparator 210 can identify the value of an audio sample based on any appropriate scale. In some examples, the audio sample comparator 210 finds the absolute value of the identified audio sample value. For example, the audio sample comparator 210 can compare the current maximum audio sample value of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) audio sample value of the input audio signal 202 associated with a previous sample block (e.g., a previous 3-second sample, a previous 5-second sample, a previous 10-second sample, etc.). In some examples, the audio sample comparator 210 can output the ratio of the current maximum audio sample value to the average of previous audio sample blocks. In other examples, the audio sample comparator 210 can output any other appropriate value (e.g., difference) related to the comparison between the current audio sample of the input audio signal 202 and the average of previous audio sample blocks of the input audio signal 202.

[0040]

[0040] The exemplary source change detector 212 determines whether the audio source of the input audio signal 202 has changed based on the outputs of the exemplary compressor gain comparator 206, the exemplary power comparator 208, and / or the exemplary audio sample comparator 210. For example, the source change detector 212 may use regression analysis (e.g., linear regression, binomial regression, least squares method, logistic regression, etc.) to determine whether a source change has occurred. In such an example, the source change detector 212 may further perform regression analysis based on labeled input data. For example, the labeled input data may include an indication of whether the audio source has changed by performing a binary decision of whether there is a source change or not, as a result of classification from values ​​corresponding to the power comparison, compressor gain comparison, and / or audio sample comparison. In other examples, the source change detector 212 may use any other suitable predictive model (e.g., machine learning, neural network, etc.) to determine whether a change in the audio source has occurred. In some examples, the source change detector 212 can output a binary value indicating whether a source change occurred within a time frame (e.g., the previous 3 seconds). For example, the source change detector 212 can output "0" to indicate that no source change has occurred, and "1" to indicate that a source change has occurred. In other examples, the source change detector 212 can output any other appropriate indicator to indicate that a change in the audio source has occurred.

[0041]

[0041] The exemplary input volume detector 214 identifies the volume level associated with the input audio signal 202. In some examples, the input volume detector 214 identifies an initial input volume level value associated with the input audio signal 202 when the input signal detector 204 indicates that the input audio signal 202 is a new input audio signal. In some examples, the input volume detector 214 provides a volume level to the dynamic range compressor 224 when the input audio signal is first received, enabling dynamic range compression of the input audio signal 202. For example, the input volume detector 214 may provide the dynamic range compressor 224 with an initial volume level of the input audio signal 202, which can then adjust the dynamic range so that the volume level of the input audio signal 202 falls within a target volume range. In the illustrated example, the input volume detector 214 identifies the volume level at regular intervals (e.g., 3-second intervals, 5-second intervals, etc.). In some examples, the input volume detector 214 finds the average (e.g., arithmetic mean, median, etc.) volume level over that interval. In some examples, the input volume detector 214 determines the deviation of the volume level over that interval.

[0042]

[0042] An exemplary audio signal classifier 216 identifies the classification of an input audio signal. In some examples, the audio signal classifier 216 analyzes the characteristics of the input audio signal 202 to identify the classification group to which the input audio signal 202 belongs. In some examples, the audio signal classifier 216 uses a neural network to help predict the dynamic range and informs the volume controller 220 of the amount of volume reduction to be applied to the input audio signal 202. For example, a classification model used by and / or incorporated into the audio signal classifier 216 can be trained and output using a neural network. Figure 3 shows a block diagram illustrating an exemplary audio classification engine capable of providing a trained model for use by the media unit 106 (e.g., the audio signal classifier 216). In some examples, audio characteristics related to the training data are used by a neural network to identify classification groups and are stored in association with the classification groups. For example, audio characteristics such as average dynamic range, dynamic range deviation, average volume, and average volume deviation can be identified for classification groups and stored in the classification database 218 and / or other accessible locations (e.g., lookup tables).

[0043]

[0043] In some examples, the audio signal classifier 216 and / or the audio classification engine 300 in Figure 3 access volume profiles and / or other representations of a variety of representative audio signals (e.g., representing different instruments, different genres, etc.) and train the model of the audio signal classifier 216 to identify classes based on the volume profiles and / or other representations of a variety of representative audio signals (e.g., using clustering). For example, volume profiles and / or other representations can be clustered based on volume and / or dynamic range. The audio signal classifier 216 can then classify the input audio signal 202 by analyzing the input audio signal 202 to identify the volume, dynamic range, and / or other properties of the input audio signal 202 that can be compared to one or more properties associated with a class.

[0044]

[0044] The illustrated example audio signal classifier 216 identifies one or more classification groups from a plurality of classification groups (e.g., nine classification groups, ten classification groups, etc.) associated with various types of audio signals. For example, a classification group may be associated with the genre of music represented by the input audio signal 202, the duration of the music represented by the input audio signal 202, different instruments identified in the input audio signal 202, and so on. In some examples, a classification group may be associated with spoken content, pop music, rock music, hip hop music, and so on. Some exemplary classification groups include speech, drumless music from before 1975, drumless music from 1976 to 1995, drumless music from 1996 to the present, music with synthesized drums from 1976 to 1995, music with synthesized drums from 1996 to the present, music with real drums from before 1975, music with real drums from 1976 to 1995, and / or music with real drums from 1996 to the present. Therefore, classification groups can correspond to music / sound production from different eras, where technical differences in recording and / or playback functions correspond to differences in the volume and / or dynamic range of the music / sound produced. Additionally or alternatively, classification groups can be based on observed (e.g., heuristically derived) characteristics of the volume and / or dynamic range of the audio content.

[0045]

[0045] The audio signal classifier 216 can classify the input audio signal 202 using any characteristic of the input audio signal 202. For example, the audio signal classifier 216 can use the spectral characteristics of the input audio signal 202, the constant Q transform (CQT) characteristics of the input audio signal 202, or any other arbitrary parameters. In some examples, time samples, spectrograms, summaries, transformations, and / or descriptions of the audio signal are used as input to the audio signal classifier 216. Such characteristics can be input to a neural network model for identifying classification groups of the input audio signal. In some examples, the neural network model can be accessed from a classification database 218.

[0046]

[0046] The illustrated example audio signal classifier 216 can output a single class (e.g., speech, music including drums since 1996) or a probability distribution associated with multiple classes. In some examples, the audio signal classifier 216 identifies the class with the highest probability corresponding to the audio signal and outputs an indication that the audio signal belongs to this class. In other examples, the audio signal classifier 216 outputs the probabilities associated with the audio signal belonging to each class (e.g., there is a 60 percent chance that the audio signal belongs to the "speech" class). In some examples, a threshold percentage can be used to identify when a single class is output compared to when a probability distribution is output. For example, if the audio signal classifier 216 identifies that there is a 90 percent chance that the audio signal belongs to the speech class, this probability may exceed the threshold percentage, allowing the audio signal classifier 216 to identify the audio signal as belonging to the speech class. In some examples, if the threshold percentage is not met, a probability distribution may be output, or the audio signal classifier 216 may indicate that it cannot identify a class associated with the audio signal.

[0047]

[0047] In response to identifying a classification group of the input audio signal 202, the audio signal classifier 216 can select a classification gain value associated with the classification group and transmit the classification gain value to the volume control 220 and / or dynamic range compressor 224. In some examples, the audio signal classifier 216 accesses the classification gain value from one or more lookup tables associated with the classification group. In some examples, the classification gain value is identified as a combination of values ​​from one or more tables associated with one or more classification groups. For example, if the audio signal classifier 216 outputs a probability distribution showing the probability that an audio signal belongs to each classification group, a table associated with each group can be retrieved and a combination of gain values ​​or other adjustment values ​​(e.g., EQ values) can be weighted based on the relative probability of each classification group.

[0048]

[0048] In some examples, the audio signal classifier 216 provides classification groups to the volume adjusters 220 and / or dynamic range compressors 224, which then access and / or identify the adjustment parameters associated with the classification groups. In some examples, the audio signal classifier 216 outputs (1) classification gain values ​​and / or (2) a period of time at which the audio volume levels should be reanalyzed.

[0049]

[0049] The exemplary classification database 218 is a storage location for data related to audio signal classification. In some examples, the classification database 218 stores models (e.g., neural network models) used to classify audio signals. In some examples, the models are accessed and / or retrieved from an audio classification engine, which is shown in Figure 3 and described in more detail. In some examples, the classification database 218 can store audio signals, audio fingerprints, and / or any other data utilized by the media unit 106. The classification database 218 stores lookup tables or other storage means, for example, one for storing audio parameters associated with classification groups. The exemplary classification database 218 can be implemented by volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), Rambus dynamic random access memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory). The classification database 218 can be implemented by one or more double data rate (DDR) memories, such as DDR, DDR2, DDR3, or mobile DDR (mDDR), either additionally or alternatively. The classification database 218 can also be implemented by one or more mass storage devices, such as one or more hard disk drives, one or more compact disk drives, or one or more digital utility disk drives. In the illustrated example, the classification database 218 is shown as a single database, but the classification database 218 can be implemented by any number and / or types of databases. Furthermore, the data stored in the classification database 218 can be in any data format, such as binary data, comma-separated data, tab-separated data, or structured query language (SQL) structures.

[0050]

[0050] The exemplary volume control 220 in the example shown in Figure 2 adjusts the volume level of an audio signal. In some examples, the exemplary volume control 220 identifies a single average gain value that converts the volume of the audio signal from a known volume value (e.g., one identified by the input volume detector 214) to a desired volume value (e.g., a value near the target volume range). The illustrated example volume control 220 communicates with the input volume detector 214 and / or the audio signal classifier 216 to identify the target gain value. The volume control 220 calculates the target gain based on the classification gain values ​​corresponding to one or more classification groups identified by the audio signal classifier 216 and the input volume level detected by the input volume detector 214 (e.g., by calculating the average of the classification gain values ​​and the input volume). In some examples, the volume control 220 applies one or more weights to the classification gain values ​​accessed from the audio signal classifier 216 and to the input volume accessed from the input volume detector 214.

[0051]

[0051] In some examples, the volume controller 220 resets the gain value applied to the audio signal when a change in source is detected (for example, when the source changes from an FM station to an auxiliary input). In some such examples, the volume controller 220 sets the gain value to zero, and the dynamic range compressor 224 performs compression to adjust the volume of the audio signal to within the target volume range until the input volume detector 214 and the audio signal classifier 216 provide the volume controller 220 with information about the newly detected audio signal to determine the target gain value.

[0052]

[0052] The illustrated example volume control 220 smoothly transitions between different volume settings (for example, by using a smoothing filter, an averaging filter, etc.). In some examples, if the volume control 220 determines that a large change in the target gain value is required, the volume control 220 slowly transitions to the new target gain value. Conversely, the volume control 220 can more quickly transition between smaller, less perceptible changes in the target gain value. The illustrated example volume control 220 uses a unipolar smoothing filter to transition between target gain values.

[0053]

[0053] In some examples, the volume controller 220 determines whether the updated input volume value from the input volume detector 214 and / or the updated classification output from the audio signal classifier 216 meet a difference threshold with respect to the previous input volume value and / or the previous classification output. In some such examples, the volume controller 220 identifies a new target gain value only if the updated input volume value and / or the updated classification output meet a difference threshold with respect to the previous value used to calculate the target gain value.

[0054]

[0054] The illustrative volume control 220 in the illustrated example transforms the audio signal by applying a target gain value to the audio signal. In some examples, the volume control 220 performs initial volume adjustment using fade-in volume adjustment when the input signal detector 204 detects the input audio signal 202 (for example, minimizing the volume when a new signal is detected and then gradually increasing the volume). In some examples, the volume control 220 can set the initial volume value based on previous volume values ​​of the type of input signal being accessed. For example, if the input audio signal 202 is an FM audio signal, the volume control 220 can identify the previous volume level used for the FM audio signal and set the current initial volume to this value. The volume control 220 can adjust the initial volume of the input audio signal 202 independently, or in conjunction with the dynamic range compressor 224, it can adjust when the input audio signal 202 is first detected.

[0055]

[0055] The exemplary audio signal identifier 222 in the example shown in Figure 2 identifies the media corresponding to the input audio signal 202. In some examples, the media unit 106 does not need to include the audio signal identifier 222, and the input audio signal 202 can be modified based solely on classification by the audio signal classifier 216. In some examples, the audio signal identifier 222 identifies the media of the audio signal by performing a comparison between a media identifier (e.g., a fingerprint) embedded in the audio signal and a known or reference audio signature. In some examples, the exemplary audio signal identifier 222 can find a matching reference media identifier. In such examples, the audio signal identifier 222 can adjust the input audio signal 202 by passing identification information specific to the media contained in the input audio signal 202 to the volume adjuster 220 and / or dynamic range compressor 224. In some examples, the audio signal identifier 222 can interact with an external database (e.g., one of a central facility) to find a matching reference signature. In some examples, the audio signal discriminator 222 can interact with an internal database (for example, a classification database 218) to find a matching reference signature.

[0056]

[0056] The exemplary dynamic range compressor 224 in the example shown in Figure 2 is capable of compressing the input audio signal 202. In some examples, the dynamic range compressor 224 performs audio compression so that the input audio signal 202 has an average volume level that satisfies a target volume threshold (e.g., one associated with a desired volume level). In some examples, the dynamic range compressor 224 is continuously active and performs compression of the input audio signal 202 after any volume adjustments made by the volume adjuster 220 to bring the input audio signal 202 within the target volume threshold (e.g., within ±0.5 dBFS from -21 dBFS). In some examples, the dynamic range compressor 224 acts as the final step in ensuring that the input audio signal 202 is adjusted to fall within the target volume threshold. In some examples, the amount of dynamic range compression performed on the input audio signal 202 is inversely proportional to the output quality of the volume-adjusted audio signal 228 (e.g., the greater the dynamic volume compression, the lower the quality of the volume-adjusted audio signal 228, e.g., the greater the loss).

[0057]

[0057] The exemplary real-time audio monitor 226 in the example shown in Figure 2 collects real-time volume measurement data. For example, the real-time audio monitor 226 can determine the current audio volume level as an average over a period of time (e.g., 750 milliseconds). In some examples, the real-time audio monitor 226 continuously monitors the input audio signal 202 for a monitoring period (e.g., 10 seconds, 1 minute, etc.). In such examples, the real-time audio monitor 226 can analyze the volume level during the monitoring period to determine whether subsequent adjustment by either the volume control 220 or the dynamic range compressor 224 is necessary. In some examples, the real-time audio monitor 226 continuously monitors the input audio signal 202 for the duration of the input audio signal 202. In some examples, the real-time audio monitor 226 determines whether the average volume level over a period of time (e.g., 750 milliseconds) falls within a target volume range (e.g., within ±0.5 dBFS from -21 dBFS). In response to the volume level not being within the target volume range, the audio signal classifier 216 may reanalyze the characteristics of the input audio signal 202 and attempt to reclassify it. In some examples, in response to the real-time audio monitor 226 determining that the average volume level over a certain period is not within the target volume range, the volume adjuster 220 and / or dynamic range compressor 224 further adjust the input audio signal 202.

[0058]

[0058] The real-time audio monitor 226 in the illustrated example includes and / or accesses a timer to determine whether the period since the previous classification output by the audio signal classifier 216 meets the update time threshold. In some examples, the update time threshold is set by the operator. For example, the real-time audio monitor 226 may be configured with a 3-second update time threshold, i.e., the audio signal classifier 216 will reclassify the audio signal at 3-second intervals (e.g., every 3 seconds, it will perform the classification process for the past 3 seconds). Additionally or alternatively, the input volume detector 214 in the illustrated example identifies the input volume of the audio signal (e.g., average input volume) for the period since the last classification and / or since the last input volume calculation (e.g., 3 seconds in the previous example). In some such examples, after reclassifying the audio signal and / or identifying the new input volume, the volume control 220 may identify a new target gain value based on the new classification and / or the new input volume.

[0059]

[0059] Figure 4 shows an exemplary method of implementing the media unit 106 of Figure 2, but one or more of the elements, processes, and / or devices shown in Figure 2 can be combined, divided, rearranged, omitted, excluded, and / or implemented in any other way. Furthermore, the exemplary source change detector 212, exemplary input volume detector 214, exemplary audio signal classifier 216, exemplary classification database 218, exemplary volume adjuster 220, exemplary audio signal discriminator 222, exemplary dynamic range compressor 224, exemplary real-time audio monitor 226, and / or, more generally, the exemplary input signal detector 204, exemplary compressor gain comparator 206, exemplary volume / power comparator 208, and exemplary audio sample comparator 210 used by the exemplary media unit 106 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Therefore, for example, the exemplary source change detector 212, exemplary input volume detector 214, exemplary audio signal classifier 216, exemplary classification database 218, exemplary volume adjuster 220, exemplary audio signal discriminator 222, exemplary dynamic range compressor 224, exemplary real-time audio monitor 226, and / or more generally, the exemplary input signal detector 204, exemplary compressor gain comparator 206, exemplary volume / power comparator 208, and exemplary audio sample used by the exemplary media unit 106. Each comparator 210 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field-programmable logic devices (FPLDs).When reading any of the claims of the present invention relating to an apparatus or system that purely involves the implementation of software and / or firmware, the exemplary source change detector 212, exemplary input volume detector 214, exemplary audio signal classifier 216, exemplary classification database 218, exemplary volume control 220, exemplary audio signal discriminator 222, exemplary dynamic range compressor 224, exemplary real-time audio monitor 226, and / or, more generally, at least one of the exemplary input signal detector 204, exemplary compressor gain comparator 206, exemplary volume / power comparator 208, and exemplary audio sample comparator 210 used by the exemplary media unit 106 is expressly defined herein as including a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disc (DVD), compact disc (CD), or Blu-ray disc, including software and / or firmware. Furthermore, the exemplary media unit 106 in Figure 1 may include, in addition to or instead of, those shown in Figure 2, one or more elements, processes, and / or devices, and / or two or more of any or all of the illustrated elements, processes, and devices. As used herein, the phrase “communicate” includes direct communication and / or indirect communication via one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or constant communication, but rather includes selective communication at regular intervals, scheduled intervals, irregular intervals, and / or one-time events.

[0060]

[0060] Figure 3 is a block diagram showing an audio classification engine 300 capable of providing a trained model for use by the media unit 106 in Figures 1 and 2. Machine learning techniques, whether deep learning networks or other experience / observation-based learning systems, can be used to, for example, optimize results, find objects in images, understand speech and convert speech to text, and improve the relevance of search engine results. Many machine learning systems are given initial features and / or network weights as a seed and are modified through learning and updating of the machine learning network, but deep learning networks train themselves to identify features that are "effective" for analysis. By using a multi-layer architecture, machines employing deep learning techniques can process raw data better than machines using conventional machine learning techniques. By using evaluation or abstraction of various layers, it becomes easier to examine data of highly correlated values ​​or groups of characteristic themes.

[0061]

[0061] Machine learning techniques, whether neural networks, deep learning networks, and / or other experience / observation-based learning systems, can be used to, for example, produce optimal results, find objects in images, understand speech and convert speech to text, and improve the relevance of search engine results. Deep learning is a subset of machine learning that uses a set of algorithms to model high-level abstractions of data using a deep graph with multiple processing layers, including linear and nonlinear transformations. While many machine learning systems are given initial features and / or network weights as a seed and modified through learning and updating of the machine learning network, deep learning networks train themselves to identify features that are "effective" for analysis. By using a multi-layer architecture, machines employing deep learning techniques can process raw data better than machines using conventional machine learning techniques. By using evaluation or abstraction of various layers, it becomes easier to examine data of highly correlated values ​​or groups of characteristic themes.

[0062]

[0062] For example, deep learning using a convolutional neural network (CNN) finds and identifies trained observable features in data by segmenting the data using convolutional filters. Each filter or layer in the CNN architecture transforms the input data to increase the selectivity and invariance of the data. This abstraction of the data allows the machine to focus on the features of the data it is trying to classify and ignore irrelevant background information.

[0063]

[0063] Deep learning works under the condition that many datasets contain high-level features that encompass low-level features. For example, when examining an image, it is more efficient to look for edges that form motifs that form parts that form the object being looked for, rather than looking for the object itself. These feature hierarchies can be found in many different forms of data.

[0064]

[0064] Learned observable features include objects and quantifiable regularities learned by the machine during supervised learning. A machine provided with a large set of well-classified data is better equipped to distinguish and extract features in relation to the success of classifying new data.

[0065]

[0065] A deep learning machine utilizing transfer learning can appropriately link data features to specific classifications confirmed by human experts. Conversely, the same machine can update classification parameters if it is informed of an incorrect classification by a human expert. Settings and / or other configuration information can be guided, for example, by learned use of settings and / or other configuration information, and as the system is used more (e.g., repeatedly and / or by multiple users), the number of variability and / or other possibilities of settings and / or other configuration information for a given situation can be reduced.

[0066]

[0066] An exemplary deep learning neural network can be trained, for example, on a set of data classified by an expert. This set of data constructs the initial parameters of the neural network, which constitutes the supervised learning phase. During the supervised learning phase, the neural network can be tested to see whether the desired behavior has been achieved.

[0067]

[0067] Once the desired behavior of the neural network is achieved (for example, the machine has been trained to operate according to a specified threshold), the machine can be deployed and used (for example, to test the machine with "real" data). During operation, the neural network's classifications can be confirmed or rejected (for example, by an expert user, an expert system, a reference database, etc.) to continuously improve the behavior of the neural network. The exemplary neural network then enters a state of transfer learning, as the classification parameters that define the behavior of the neural network are updated based on the ongoing interactions. In a particular example, a neural network such as neural network 302 can provide direct feedback to other processes such as an audio classification scoring engine 304. In a particular example, neural network 302 outputs data, which is buffered (for example, via the cloud, etc.) and validated before being provided to other processes.

[0068]

[0068] In the example in Figure 3, the neural network 302 receives input from previous result data related to classification training data and outputs an algorithm for predicting the classification group associated with an audio signal. The network 302 can be given some initial correlation as a seed and then learn from ongoing experience. In some examples, the neural network 302 continuously receives feedback from at least one classification training data. In the example in Figure 3, throughout the operating lifetime of the audio classification engine 300, the neural network 302 is continuously trained via feedback and the exemplary audio classification scoring engine 304 can be updated based on the neural network 302 and / or additional classification training data as desired. The network 302 can learn and evolve based on roles, locations, contexts, etc.

[0069]

[0069] In some examples, the accuracy of the model generated by the neural network 302 can be determined by an exemplary audio classification scoring engine validater 306. In such examples, at least one of the audio classification scoring engine 304 and the audio classification scoring engine validater 306 receives a set of classification training data. Furthermore, in such examples, the audio classification scoring engine 304 receives input related to classification validation data and predicts one or more audio classifications related to the classification validation data. The predicted results are distributed to the audio classification scoring engine validater 306. The audio classification scoring engine validater 306 receives additional known audio classifications associated with the classification validation data and compares the known audio classifications with the predicted classifications received from the audio classification scoring engine 304. In some examples, this comparison yields the accuracy of the model generated by the neural network 302 (for example, if 95 comparisons are matches and 5 are errors, the model is 95% accurate, etc.). Once the neural network 302 reaches the desired accuracy (for example, once the network 302 has been trained and is ready for deployment), the audio classification scoring engine validater 306 can output the model to the audio signal classifier 216 in Figure 2, so that it can be used to classify audio other than the classification training data and / or classification validation data.

[0070]

[0070] Figures 4 and 5 show flowcharts representing exemplary hardware logic, machine-readable instructions, hardware implementation state machine, and / or any combination thereof for implementing the media unit 106 of Figure 2. The machine-readable instructions may be an executable program or part of an executable program executed by a computer processor, such as the processor 612 shown in the exemplary processor platform 600 described below in relation to Figure 6. The program may be embodied as software stored on a non-temporary computer-readable storage medium such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 612, but the entire program and / or parts thereof may alternatively be executed by a device other than the processor 612 and / or embodied in firmware or dedicated hardware. Furthermore, although the exemplary program is described with reference to the flowcharts shown in Figures 4 and 5, many other methods of implementing the exemplary media unit 106 can be used alternatively. For example, the execution order of blocks can be changed, and / or parts of the described blocks can be changed, excluded, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform the corresponding operations without running software or firmware.

[0071]

[0071] As described above, the exemplary processes in Figures 4 and 5 can be implemented using executable instructions (e.g., computer and / or machine-readable instructions) stored in non-temporary computer and / or machine-readable media, such as hard disk drives, flash memory, read-only memory, compact disks, digital multipurpose disks, caches, random access memory, and / or any other storage device or storage disk where information is stored for any period of time (e.g., for a long period, permanently, for a short moment, for temporary buffering, and / or for caching information). As used herein, the term non-temporary computer-readable media includes any type of computer-readable storage device and / or storage disk, and is explicitly defined to exclude propagating signals and transmission media.

[0072]

[0072] "Including" and "comprising" (and all their forms and tenses) are used herein as open-ended terms. Therefore, whenever a claim uses any form of "include" or "comprise" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within the description of any type of claim, it should be understood that additional elements, terms, etc., may exist without falling outside the scope of the corresponding claim or description. When the phrase "at least" is used herein as a transitional word in the preamble of a claim, etc., that phrase is open-ended, just as the terms "comprising" and "including" are open-ended. The term "and / or" when used in the form, for example, A, B, and / or C, refers to any combination or subset of A, B, and C, for example, (1) A only, (2) B only, (3) C only, (4) A and B, (5) A and C, (6) B and C, and (7) A, B, and C. In this specification, when used in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" refers to an implementation that includes (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, in this specification, when used in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" refers to an implementation that includes (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. In this specification, when used in a context describing the execution or performance of a process, instruction, action, activity, and / or step, the phrase “at least one of A and B” means an implementation that includes (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.Similarly, in the context of describing the implementation or execution of a process, instruction, action, activity, and / or step, the phrase “at least one of A or B” means an implementation that includes (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.

[0073]

[0073] Figures 4 and 5 show exemplary machine-readable instructions that can be executed to implement dynamic volume control via audio classification for the media unit 106 of Figures 1 and 2. Referring to the figures and related descriptions above, the exemplary machine-readable instructions 400 begin in block 402. In block 402, the exemplary media unit 106 detects a change in the audio signal input. In some examples, an input signal detector 204 detects a change in the audio signal input. For example, an audio signal may have started (e.g., there was no audio signal that the media unit 106 had previously accessed and a new one has started), or the audio signal may have been changed (e.g., an FM radio signal has been changed to an AM radio signal). The execution of block 402 will be described in more detail below with reference to Figure 5.

[0074]

[0074] In block 404, an exemplary media unit 106 compresses the input audio signal 202 to satisfy the target volume range. In some examples, a dynamic range compressor 224 compresses the input audio signal 202 to satisfy the target volume range.

[0075]

[0075] In block 406, an exemplary media unit 106 identifies classification groups of the input audio signal 202. In some examples, an audio signal classifier 216 identifies classification groups of the input audio signal. In some examples, the audio signal classifier 216 identifies classification groups based on a comparison of one or more characteristics of the input audio signal (e.g., CQT value) with a trained machine learning model. The audio signal classifier 216 may additionally or alternatively determine probability distributions associated with one or more classification groups.

[0076]

[0076] In block 408, the exemplary media unit 106 determines the input volume of the input audio signal 202. In some examples, an input volume detector 214 determines the input volume of the input audio signal 202. In some examples, the input volume detector 214 determines the average input volume of the input audio signal 202 over a period of time (e.g., 3 seconds, 5 seconds, etc.). In some examples, the input volume detector 214 determines the volume deviation of the input audio signal 202 over a period of time. In some examples, the input volume detector 214 measures one or more instantaneous volume values.

[0077]

[0077] In block 410, the exemplary media unit 106 determines the classification gain value using a lookup table associated with the classification group of the input audio signal 202. In some examples, the audio signal classifier 216 determines the classification gain value using a lookup table associated with one or more classification groups identified by the audio signal classifier 216 to be associated with the input audio signal 202. In some examples, the classification gain value is a single value representing the classification group (e.g., based on the average dynamic range observed in the training data of the classification group, or based on the average volume observed in the training data of the classification group). In some examples, the classification gain value is determined based on a probability distribution output by the audio signal classifier 216 (e.g., one or more gain values ​​are calculated based on the probability that the input audio signal 202 belongs to one or more of the classification groups).

[0078]

[0078] In block 412, the exemplary media unit 106 determines a target gain value by weighting the input volume and the classification gain value. In some examples, the volume controller 220 applies a first weight to the input volume and a second weight to the classification gain value, and then determines a target gain value based on the weighted input volume and the weighted classification gain value. In some examples, since the input volume reflects the actual state of the audio signal as opposed to a prediction of the classification gain value, the volume controller 220 applies a weight greater than the classification gain value to the input. In some examples, the volume controller 220 determines the target gain value as a value between the input volume measurement and the target volume range. In some examples, the volume controller 220 calculates the average of the input volume and the volume level obtained by applying the classification gain value, and the target classification gain value is determined as the gain required to bring the volume of the input audio signal 202 to this average volume level.

[0079]

[0079] In block 414, the exemplary media unit 106 uses a smoothing filter to apply a target gain value to the audio signal. In some examples, the volume control 220 uses a smoothing filter to apply a target gain value to the input audio signal 202. The volume control 220 can utilize different types of filters (e.g., median filter, Kalman filter, etc.) to smooth the transition between the first gain value and the updated gain value (e.g., when the classification and / or input volume has been updated), or between no gain value and a gain value (e.g., when a new audio signal has been detected).

[0080]

[0080] In block 416, the exemplary media unit 106 adjusts the compression value to satisfy the target volume range. In some examples, the dynamic range compressor 224 adjusts the compression value to satisfy the target volume range. For example, if the volume control 220 increases the gain value applied to the input audio signal 202, the dynamic range compressor 224 can decrease the compression value because less dynamic range compression is required to keep the input audio signal 202 within the target volume range. Conversely, if the volume control 220 decreases the gain value applied to the input audio signal 202, the dynamic range compressor 224 can increase the compression value because more dynamic range compression is required to keep the input audio signal 202 within the target volume range.

[0081]

[0081] In block 418, the exemplary media unit 106 determines whether the time since the last classification satisfies or exceeds the update time threshold. In some examples, the real-time audio monitor 226 determines whether the time since the last classification has been performed satisfies or exceeds the update time threshold. In some examples, the real-time audio monitor 226 determines whether the time since the last input volume calculation and / or the time since the last volume adjustment by the volume control 220 has been performed satisfies or exceeds the update time threshold. In response that the time since the last classification satisfies or exceeds the update time threshold, processing moves to block 424. Conversely, in response that the time since the last classification neither satisfies nor exceeds the update time threshold, processing moves to block 420.

[0082]

[0082] In block 420, the exemplary media unit 106 determines whether a change in the audio input source has occurred. In some examples, the input signal detector 204 determines whether a change in the audio input source has occurred (for example, the input source has changed from FM radio to an auxiliary input, or from CD to AM radio, etc.). In response to a change in the audio input source, processing moves to block 422. Conversely, in response to no change in the audio input source, processing moves to block 418. The execution of block 420 will be described in more detail below with reference to Figure 5.

[0083]

[0083] In block 422, the exemplary media unit 106 resets the gain value. In some examples, the volume control 220 resets the gain value. For example, the volume control 220 may set the gain value to zero because a previous target gain value (specified for a previous audio signal from a different input source) may no longer be valid for a new audio signal. Thus, by the time a new target gain value is specified (for example, following classification and input volume specification), the gain value is reset to 1 and the dynamic range compressor 224 compresses the input audio signal 202 to meet the target volume range.

[0084]

[0084] In block 424, the exemplary media unit 106 identifies the input volume over the period since the last classification. In some examples, the input volume detector 214 identifies the input volume over the period since the last classification. For example, if the real-time audio monitor 226 is configured with a 3-second update interval, then (for example, in block 418) once the entire duration of the update interval has elapsed, the input volume detector 214 identifies the input volume for the update interval. In some examples, the average input volume over the update interval is determined.

[0085]

[0085] In block 426, the exemplary media unit 106 identifies updated classification groups based on audio signals over a period of time since the last classification. In some examples, the audio signal classifier 216 identifies updated classification groups based on audio signals over a period of time since the last classification. For example, if the real-time audio monitor 226 is configured with a 3-second update interval, after 3 seconds have elapsed since the last classification, the audio signal classifier 216 analyzes one or more characteristics of the audio signals to identify updated classification groups. In some examples, the updated classification groups are the same as previously identified classification groups.

[0086]

[0086] In block 428, the exemplary media unit 106 determines whether dynamic volume is enabled or not. For example, the operator of the media unit 106 can enable or disable dynamic volume (for example, via a switch, via the settings of the media unit 106, etc.). In response that dynamic volume is enabled, processing moves to block 410. Conversely, in response that dynamic volume is not enabled, processing ends.

[0087]

[0087] Figure 5 is a flowchart illustrating an exemplary process 500 for performing blocks 402 and / or 420 of Figure 4. The exemplary process 500 begins with block 502. In block 502, the compressor gain comparator 206 compares the current compressor gain to a recent past compressor gain. For example, the compressor gain comparator 206 may compare the gain of the dynamic range compressor 224 associated with the current sample of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with a previous sample block (e.g., a previous 3-second sample, a previous 5-second sample, a previous 10-second sample, etc.). In some examples, the compressor gain comparator 206 may output the ratio of the current gain of the dynamic range compressor 224 associated with the current sample block of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) gain of the dynamic range compressor 224 associated with previous sample blocks (e.g., previous 3-second samples, previous 5-second samples, previous 10-second samples, etc.) of the dynamic range compressor 224.

[0088]

[0088] In block 504, the power comparator 208 compares the current volume / power of the input audio signal 202 to the recent past volume / power(s) of the audio signal. For example, the power comparator 208 can compare the current RMS power of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) power of the input audio signal 202 associated with previous sample blocks (e.g., previous 3-second samples, previous 5-second samples, previous 10-second samples, etc.). In some examples, the power comparator 208 can query the peak meter output to determine the RMS power. In some examples, the power comparator 208 can output the ratio of the current RMS power to the average of previous RMS power(s).

[0089]

[0089] In block 506, the audio sample comparator 210 compares the maximum value of the current audio sample block to the most recent audio sample value(s). For example, the audio sample comparator 210 can compare the current audio sample value of the input audio signal 202 to the average (e.g., arithmetic mean, median, etc.) audio sample value of the input audio signal 202 associated with previous sample blocks (e.g., previous 3-second samples, previous 5-second samples, previous 10-second samples, etc.). In some examples, the audio sample comparator 210 can output the ratio of the current audio sample value to the average of the previous sample blocks.

[0090]

[0090] In block 508, the source change detector 212 analyzes the audio sample comparison, compressor gain comparison, and power comparison to determine if a source change has occurred. For example, the source change detector 212 may use regression analysis (e.g., linear regression, binomial regression, least squares method, logistic regression, etc.) to determine if a source change has occurred. In other examples, the source change detector 212 may use any other suitable means (e.g., a neural network) to determine if a source change has occurred.

[0091]

[0091] In block 510, the source change detector 212 determines whether the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that a change in the source has occurred. If the source change detector 212 determines, via logistic regression or other classification method, that the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that a change in the source has occurred, then process 500 proceeds to block 512. If the source change detector 212 determines that the RMS comparison, compressor gain comparison, and / or audio sample comparison indicate that no change in the source has occurred, then process 500 proceeds to block 514.

[0092]

[0092] In block 512, the source change detector 212 indicates that a change in the source has occurred. For example, the source change detector 212 can cause the input signal detector 204 to indicate to the media unit 106 that a change in the source has occurred.

[0093]

[0093] In block 514, the source change detector 212 indicates that no source change has occurred. For example, the source change detector 212 can cause the input signal detector 204 to indicate to the media unit 106 that no source change has occurred. After that, process 500 ends.

[0094]

[0094] Figure 6 is a block diagram of an exemplary processor platform 600 configured to execute the instructions of Figure 4 and implement the media unit 106 of Figures 1 and 2. The processor platform 600 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad®), a personal digital assistant (PDA), an internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.

[0095]

[0095] The illustrated example processor platform 600 includes a processor 612. The illustrated example processor 612 is hardware. For example, the processor 612 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. Hardware processors can be semiconductor-based (e.g., silicon-based) devices. In this example, the processor implements the exemplary source change detector 212, exemplary input volume detector 214, exemplary audio signal classifier 216, exemplary classification database 218, exemplary volume adjuster 220, exemplary audio signal discriminator 222, exemplary dynamic range compressor 224, exemplary real-time audio monitor 226, and / or more generally, exemplary input signal detector 204, exemplary compressor gain comparator 206, exemplary volume / power comparator 208, and exemplary audio sample comparator 210, which are used by the exemplary media unit 106.

[0096]

[0096] The illustrated example processor 612 includes local memory 613 (e.g., cache). The illustrated example processor 612 communicates with main memory, which includes volatile memory 614 and non-volatile memory 616, via bus 618. The volatile memory 614 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The non-volatile memory 616 can be implemented by flash memory and / or any other desired type of memory device. Access to main memory 614, 616 is controlled by a memory controller.

[0097]

[0097] The illustrated example processor platform 600 also includes an interface circuit 620. The interface circuit 620 can be implemented by any type of interface standard, such as an Ethernet® interface, Universal Serial Bus (USB), Bluetooth® interface, Near Field Communication (NFC) interface, and / or PCI Express interface.

[0098]

[0098] In the illustrated example, one or more input devices 622 are connected to the interface circuit 620. The input device(s) 622 allow the user to input data and / or commands to the processor 612. The input device(s) can be implemented, for example, by an audio sensor, microphone, camera (still image or video), keyboard, buttons, mouse, touchscreen, trackpad, trackball, IsoPoint, and / or voice recognition system.

[0099]

[0099] In addition, one or more output devices 624 are connected to the interface circuit 620 of the illustrated example. The output devices 624 can be implemented by, for example, display devices (e.g., light-emitting diodes (LEDs), organic light-emitting diodes (OLEDs), liquid crystal displays (LDCs), cathode ray tube displays (CRTs), in-place switching (IPS) displays, touchscreens, etc.), haptic output devices, printers and / or speakers. Thus, the interface circuit 620 of the illustrated example typically includes a graphics driver card, a graphics driver chip and / or a graphics driver processor.

[0100]

[0100] The illustrated example interface circuit 620 also includes communication devices such as a transmitter, receiver, transceiver, modem, residential gateway, wireless access point, and / or network interface for facilitating data exchange with an external machine (e.g., any type of computing device) via network 626. Communication can be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a mobile phone system, and the like.

[0101]

[0101] The illustrated example processor platform 600 also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, independent disk redundant array (RAID) systems, and digital versatile disk (DVD) drives.

[0102]

[0102] The machine-executable instructions 632 in Figure 4 can be stored in a mass storage device 628, a volatile memory 614, a non-volatile memory 616, and / or a removable non-temporary computer-readable storage medium such as a CD or DVD.

[0103]

[0103] From the above, it will be understood that exemplary methods, apparatus, and products are disclosed for adjusting the volume of media so that media with different characteristics can be played back at approximately the same volume, while minimizing the amount of compression required to achieve this volume. Conventional implementations of volume equalization rely solely on compression, resulting in a clearly discernible change in the audio signal. However, the examples disclosed herein enable the identification of an average gain value based on the classification associated with the audio signal, which intelligently classifies the audio signal and distinguishes, for example, audio signals with a relatively narrow dynamic range that can be significantly modified by the gain value, from audio signals with a wider dynamic range that may require more compression. The exemplary techniques disclosed herein intelligently adjust the volume of an input audio signal in real time by utilizing a combination of input volume measurements and parameters related to the classification of the audio signal. The examples disclosed herein describe techniques for continuously adjusting the volume level when it is necessary to correct the volume adjustment after the initial analysis (for example, due to changes in the classification of the audio signal, changes in the observed input volume, etc.). The exemplary techniques disclosed herein further include techniques for initially adjusting the volume level of the audio signal after a change in the audio signal input. Such technology is virtually imperceptible to the user and has advantages over conventional implementations because it allows for a seamless media presentation experience by playing different media from different or similar sources at substantially the same volume.

[0104]

[0104] In some examples, such as the dynamic volume of the present invention, the exemplary audio dynamic range compressor can be kept active at all times to reduce the signal to a specific range (e.g., -21 dBFS). In other examples, the audio dynamic range compressor can be kept active for a certain period of time.

[0105]

[0105] In some examples, an exemplary real-time volume detector, such as the dynamic volume of the present invention, can be applied to the input to measure the current average level over one or more intervals (e.g., 750 millisecond intervals). In such examples, the current average level can be used as an initial and ongoing estimate to guide how much the volume can be reduced.

[0106]

[0106] In some examples, neural network-based classifiers can also assist in predicting dynamic range and notify of applicable volume reductions. This can initially be based on the current category classifiers that have potential for improvement (e.g., 9 classifiers, 15 classifiers, etc.). In some examples, increasing the number of current category classifiers can make the dynamic range predictor more accurate, using different real-time and neural network approaches. In each example, accuracy can be improved in relation to the amount of volume reduction that can be achieved.

[0107]

[0107] In some examples, the goal is to reduce the volume to near a specific level (e.g., -12 dBFS) that the compressor can reach. Once the amount of reduction is determined, a unipolar smoothing filter can be used to reduce the input from its current full volume to the determined amount. The compressor will continue to maintain the volume at the specific level (e.g., -21 dBFS) on average, but the amount by which the input needs to be reduced can be less, since the amount has been reduced to the target.

[0108]

[0108] In a descriptive example of the operation of the methods, apparatus, and systems disclosed herein, well-normalized loud pop music may be delivered via the input. The compressor can reduce material from 0.0 dBFS to -21 dBFS. Substantially in parallel, the input volume detector determines that the input is flowing at an average of -1 dBFS, and the classifier determines that music containing synthesized drums and bass from 1996 to the present is being presented. This category generates a reduction of -15 dBFS, and the volume detector generates -20 dBFS. The two values ​​are averaged, and the signal can be reduced by -17.50 dBFS, and can be reduced by a further 3.5 decibels to reach the baseline of -21 dBFS. Since the compressor reduces signals that are 3.5 decibels above the threshold (for example, based on the reductions above), the audio quality is improved compared to reducing signals that are 21 decibels above the threshold, which would be done if only the compressor were used.

[0109]

[0109] Exemplary methods, apparatus, systems, and articles for dynamic volume control via audio classification are disclosed herein. Further examples and combinations thereof include: Example 1 includes an audio signal classifier that uses a neural network to analyze parameters of an audio signal related to a first volume level to identify a classification group associated with the audio signal; an input volume detector that identifies the input volume of the audio signal; a volume adjuster that applies a gain value to the audio signal, wherein the gain value is based on the classification group and the input volume, and the gain value modifies the first volume level to a second volume level; and a dynamic range compressor that applies a compression value to the audio signal, wherein the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

[0110]

[0110] Example 2 includes the apparatus described in Example 1, further comprising a source change detector that determines whether the source of the audio signal has changed.

[0111]

[0111] Example 3 includes the apparatus described in Example 2, wherein the source change detector determines whether the source of the audio signal has changed based on at least one of the following: (1) a comparison of the current compressor gain associated with the audio signal with a previous compressor gain associated with the audio signal, (2) a comparison of the RMS power associated with the audio signal with a previous RMS power associated with the audio signal, or (3) a comparison of the current audio sample value associated with the audio signal with a previous audio sample value associated with the audio signal.

[0112]

[0112] Example 4 includes the apparatus described in Example 2, wherein the volume control further resets the gain value of the audio signal in response to a determination that the source of the audio signal has changed.

[0113]

[0113] Example 5 includes the apparatus described in Example 1, wherein the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the duration of the music represented by the audio signal, or (3) the presence or absence of instruments in the music represented by the audio signal.

[0114]

[0114] Example 6 includes the apparatus described in Example 1, wherein the input volume detector further determines that the fourth volume level over a first period is not within the target volume threshold, the first period occurs after the second period, the third volume level is related to the second period, and the dynamic range compressor further adjusts the compression value to a fifth volume level, the adjusted compression value corrects the fourth volume level to a fifth volume level that satisfies the target volume threshold.

[0115]

[0115] Example 7 includes the apparatus described in Example 1, wherein the target sound level threshold is between 21 dBFS and 5 dBFS relative to full scale in decibels (dBFS).

[0116]

[0116] Example 8 is a non-temporary computer-readable storage medium containing instructions, wherein when the instructions are executed, the instructions cause a processor to at least use a neural network to analyze the parameters of an audio signal associated with a first volume level to identify a classification group associated with the audio signal, identify the input volume of the audio signal, apply a gain value to the audio signal such that the gain value modifies the first volume level to a second volume level based on the classification group and the input volume, and apply a compression value to the audio signal such that the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

[0117]

[0117] Example 9 includes the non-temporary computer-readable storage medium described in Example 8, wherein the instruction causes the processor to determine whether the source of the audio signal has changed when the instruction is executed.

[0118]

[0118] Example 10 includes the non-temporary computer-readable storage medium described in Example 9, wherein the determination of whether the source of the audio signal has changed is based on at least one of the following: (1) a comparison of the current compressor gain associated with the audio signal with a previous compressor gain associated with the audio signal, (2) a comparison of the RMS power associated with the audio signal with a previous RMS power associated with the audio signal, or (3) a comparison of the current audio sample value associated with the audio signal with a previous audio sample value associated with the audio signal.

[0119]

[0119] Example 11 includes the non-temporary computer-readable storage medium described in Example 9, wherein the instruction, when executed, causes the processor to reset the gain value of the audio signal in response to a determination that the source of the audio signal has changed.

[0120]

[0120] Example 12 includes the non-temporary computer-readable storage medium described in Example 11, wherein the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the duration of the music represented by the audio signal, or (3) the presence or absence of instruments in the music represented by the audio signal.

[0121]

[0121] Example 13 includes the non-temporary computer-readable storage medium described in Example 8, wherein the instruction, when executed, causes the processor to determine that a fourth volume level over a first period is not within a target volume threshold, that the first period occurs after a second period and the third volume level is related to the second period, and adjust the compression value to a fifth volume level, that the adjusted compression value corrects the fourth volume level to a fifth volume level that satisfies the target volume threshold.

[0122]

[0122] Example 14 includes the non-temporary computer-readable storage medium described in Example 8, wherein the target volume threshold is between 21 dBFS and 5 dBFS relative to full scale in decibels (dBFS).

[0123]

[0123] Example 15 includes a method that uses a neural network to analyze the parameters of an audio signal associated with a first volume level to identify a classification group associated with the audio signal; to identify the input volume of the audio signal; to apply a gain value to the audio signal, wherein the gain value modifies the first volume level to a second volume level based on the classification group and the input volume; and to apply a compression value to the audio signal, wherein the compression value modifies the second volume level to a third volume level that satisfies a target volume threshold.

[0124]

[0124] Example 16 includes the method of Example 15, further comprising the step of determining whether the source of the audio signal has changed.

[0125]

[0125] Example 17 includes the method of Example 16, wherein the step of determining whether the source of the audio signal has changed is based on at least one of (1) comparing the current compressor gain associated with the audio signal with a previous compressor gain associated with the audio signal, (2) comparing the RMS power associated with the audio signal with a previous RMS power associated with the audio signal, or (3) comparing the current audio sample value associated with the audio signal with a previous audio sample value associated with the audio signal.

[0126]

[0126] Example 18 includes the method of Example 16, further comprising the step of resetting the gain value of the audio signal in response to a determination that the source of the audio signal has changed.

[0127]

[0127] Example 19 includes the method of Example 15, wherein the classification group is associated with at least one of (1) the genre of music represented by the audio signal, (2) the duration of the music represented by the audio signal, or (3) the presence or absence of instruments in the music represented by the audio signal.

[0128]

[0128] Example 20 includes the method of Example 15, further comprising the steps of determining that a fourth volume level over a first period is not within a target volume threshold, wherein the first period occurs after a second period and the third volume level is related to the second period, and adjusting the compression value to modify the fourth volume level to a fifth volume level that satisfies the target volume threshold.

[0129]

[0129] While this specification discloses certain exemplary methods, apparatus, and articles, the scope of this patent is not limited to these. Rather, this patent includes all methods, apparatus, and articles that fall appropriately within the claims of this patent.

Claims

1. A non-temporary computer-readable medium in which instructions are stored, wherein, when the instructions are executed, one or more processors, Using a neural network, analyze the parameters of an audio signal to identify the classification group associated with the audio signal. Applying a gain value to the aforementioned audio signal, Based on the aforementioned gain value, the input volume of the audio signal is corrected to a gain-adjusted volume. The aforementioned gain value is between the input volume and the classification gain value. The classification gain value is based on the classification group, The gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value. Applying a compression value to the aforementioned audio signal, wherein the compression value modifies the gain-adjusted volume to a compressed volume that satisfies a target volume threshold. A non-temporary computer-readable medium that causes a set of operations, including [specific actions], to be performed.

2. Applying a compression value to the aforementioned audio signal (i) When the gain value increases, the compression value is reduced, and (ii) If the gain value decreases, increase the compression value. A non-temporary computer-readable medium according to claim 1, further comprising:

3. The set of operations is, To determine whether the source of the aforementioned audio signal has changed, A non-temporary computer-readable medium according to claim 1, further comprising:

4. Determining whether the source of the audio signal has changed is (1) Comparison of the current compressor gain associated with the audio signal with the previous compressor gain associated with the audio signal. (2) Comparison of the RMS power associated with the audio signal with the previous RMS power associated with the audio signal, and (3) Comparison of the current audio sample value associated with the audio signal with the previous audio sample value associated with the audio signal. A non-temporary computer-readable medium according to claim 3, based on at least one of the following.

5. The aforementioned classification group, (1) The genre of music represented by the audio signal, (2) The duration of the music represented by the audio signal, (3) The presence or absence of musical instruments in the music represented by the audio signal, A non-temporary computer-readable medium according to claim 1, comprising at least one of the following.

6. A method performed by a computer for adjusting the volume, The steps include: using a neural network to analyze the parameters of an audio signal and identify the classification group associated with the audio signal; A step of applying a gain value to the audio signal, wherein the input volume of the audio signal is modified to a gain-adjusted volume by the gain value. The aforementioned gain value is between the input volume and the classification gain value. The classification gain value is based on the classification group, The gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value, comprising the steps of: A step of applying a compression value to the audio signal, wherein the compression value modifies the gain-adjusted volume to a compressed volume that satisfies a target volume threshold. Methods that include...

7. The step of applying a compression value to the audio signal is, (i) When the gain value increases, the compression value is reduced, and (ii) If the gain value decreases, increase the compression value. The method according to claim 6, further comprising:

8. The method according to claim 6, further comprising the step of determining whether the source of the audio signal has changed.

9. The step of determining whether the source of the audio signal has changed is, (1) Comparison of the current compressor gain associated with the audio signal with the previous compressor gain associated with the audio signal. (2) Comparison of the RMS power associated with the audio signal with the previous RMS power associated with the audio signal, and (3) Comparison of the current audio sample value associated with the audio signal with the previous audio sample value associated with the audio signal. The method according to claim 8, based on at least one of the following.

10. The aforementioned classification group, (1) The genre of music represented by the audio signal, (2) The duration of the music represented by the audio signal, (3) The presence or absence of musical instruments in the music represented by the audio signal, The method according to claim 6, comprising at least one of the following.

11. A computing device, One or more processors, A non-temporary computer-readable medium on which instructions are stored, It is equipped with such that when the instruction is executed, one or more processors Using a neural network, analyze the parameters of an audio signal to identify the classification group associated with the audio signal. Applying a gain value to the aforementioned audio signal, Based on the aforementioned gain value, the input volume of the audio signal is corrected to a gain-adjusted volume. The aforementioned gain value is between the input volume and the classification gain value. The classification gain value is based on the classification group, The gain value is determined by applying a first weight to the input volume and a second weight to the classification gain value. Applying a compression value to the aforementioned audio signal, wherein the compression value modifies the gain-adjusted volume to a compressed volume that satisfies a target volume threshold. A computing device that performs a set of operations, including [specific actions].

12. Applying a compression value to the aforementioned audio signal (i) When the gain value increases, the compression value is reduced, and (ii) If the gain value decreases, increase the compression value. The computing device according to claim 11, further comprising:

13. The set of operations is, To determine whether the source of the aforementioned audio signal has changed, The computing device according to claim 11, further comprising:

14. Determining whether the source of the audio signal has changed is (1) Comparison of the current compressor gain associated with the audio signal with the previous compressor gain associated with the audio signal. (2) Comparison of the RMS power associated with the audio signal with the previous RMS power associated with the audio signal, and (3) A computing device according to claim 13, based on at least one of the following: a comparison of a current audio sample value associated with the audio signal with a previous audio sample value associated with the audio signal.