Volume leveler controller and control method

The audio processing device automatically classifies and adjusts audio enhancement parameters in real-time, addressing the inconvenience of manual presets and reducing artifacts during content transitions, thereby improving user experience.

JP7862500B2Active Publication Date: 2026-05-19DOLBY LABORATORIES LICENSING CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-10-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing audio processing systems require manual preset selection by users, leading to inconvenient and non-continuous parameter adjustments, which can result in audible artifacts during content transitions.

Method used

An audio processing device that automatically classifies audio content in real-time and adjusts audio enhancement parameters continuously based on the identified content type, using an audio classifier and an adjustment unit to optimize settings for dialogue enhancers, surround virtualizers, and equalizers.

Benefits of technology

Enables seamless and artifact-free audio enhancement by adapting to different content types, enhancing user experience without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007862500000016
    Figure 0007862500000016
  • Figure 0007862500000017
    Figure 0007862500000017
  • Figure 0007862500000018
    Figure 0007862500000018
Patent Text Reader

Abstract

To provide a volume leveler controller and a controlling method.SOLUTION: In one embodiment, a volume leveler controller includes an audio content classifier for identifying a content type of an audio signal in real time; and an adjusting unit for adjusting a volume leveler in a continuous manner on the basis of the identified content type. The adjusting unit may configured to positively correlate dynamic gain of the volume leveler with informative content types of the audio signal, and negatively correlate the dynamic gain of the volume leveler with interfering content types of the audio signal.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - References to Related Applications This application claims priority to Chinese Patent Application No. 201310100422.1, filed on March 26, 2013, and U.S. Provisional Patent Application No. 61 / 811,072, filed on April 11, 2013. Each of these applications is incorporated herein by reference in its entirety.

[0002] Technical Field This application generally relates to audio signal processing. Specifically, embodiments of the present application relate to devices and methods for audio classification and processing, particularly for the control of dialog enhancers, surround virtualizers, volume normalizers, and equalizers.

Background Art

[0003] Some audio improvement devices tend to modify audio signals either in the time domain or the spectral domain in order to improve the overall quality of the audio and correspondingly enhance the user experience. Various audio improvement devices have been developed for various purposes. Some typical examples of audio improvement devices include the following.

[0004] Dialog Enhancer: Dialog is the most important component for understanding the story in movies and radio or television programs. Various methods have been developed to enhance dialog to improve its clarity and comprehensibility, especially for the elderly whose hearing is declining.

[0005] Surround Virtualizer: A surround virtualizer enables surround (multi - channel) sound signals to be rendered through a PC's internal speakers or through headphones. That is, it uses a stereo device (such as speakers and headphones) to virtually generate the effect of surround and provide a movie theater experience for consumers.

[0006] Volume leveler: A volume leveler adjusts the volume of audio content during playback, aiming to ensure that the volume remains nearly consistent throughout the time axis based on a target loudness value.

[0007] Equalizer: An equalizer provides spectral balance consistency, also known as "tone" or "timbre," allowing the user to configure an overall profile (curve or shape) of the frequency response (gain) in individual frequency bands to emphasize certain sounds or remove undesirable ones. Traditional equalizers may offer different equalizer presets for various sounds, such as different musical genres. Once a preset is selected or an equalization profile is set, the same equalization gain is applied to the signal until the equalization profile is manually modified. In contrast, a dynamic equalizer achieves spectral balance consistency by continuously monitoring the spectral balance of the audio, comparing it to the desired tone, and dynamically adjusting the equalization filter to transform the original tone of the audio into the desired tone.

[0008] Generally, audio enhancement devices have their own unique application scenarios / contexts. That is, audio enhancement devices may be suitable not for all possible audio signals, but only for certain sets of content. This is because different content may need to be processed differently. For example, dialogue enhancement methods are typically applied to film content. If applied to music without dialogue, it could incorrectly boost certain frequency subbands, introducing severe timbre changes and perceptual inconsistencies. Similarly, if a noise suppression method were applied to a music signal, strong artifacts would become audible.

[0009] However, for audio processing systems that typically include a collection of audio enhancement devices, the input can inevitably be any possible type of audio signal. For example, an audio processing system integrated into a PC will receive audio content from a variety of sources, including movies, music, VoIP, and games. Therefore, it becomes important to identify or distinguish the content being processed in order to apply a better algorithm or better parameters for each algorithm to the corresponding content.

[0010] To distinguish audio content and apply correspondingly better parameters or audio enhancement algorithms, traditional systems typically pre-design a set of presets, and the user is required to select a preset for the content being played. These presets typically encode a set of the audio enhancement algorithms and / or their best parameters to be applied, such as "movie" presets and "music" presets specifically designed for movie or music playback. [Prior art documents] [Patent Documents]

[0011] [Patent Document 1] International Publication No. 2008 / 106036 (H. Muesch, “Speech Enhancement in Entertainment Audio”) [Patent Document 2] U.S. Patent Application Publication No. 2009 / 0097676A1 (AJ Seefeldt et al., “Calculating and Adjusting the Perceived Loudness and / or the Perceived Spectral Balance of an Audio Signal”) [Patent Document 3] International Publication No. 2007 / 127023 (BG Grockett et al., "Audio Gain Control Using Specific-Loudness-Based Auditory Event Detection")

Patent document 4

Non-licensed literature

[0012]

Non-licensed literature 1

Non-licensed Document 2

Non-licensed Document 4

[0013] However, manual selection is inconvenient for the user. Users typically do not frequently switch between predefined presets, but rather simply stick to one preset for all content. Furthermore, even in some automated solutions, the parameter or algorithm setup in a preset is usually discrete (e.g., turning on or off specific algorithms for specific content), and it is not possible to adjust parameters in a content-based, continuous manner. [Means for solving the problem]

[0014] The first aspect of this invention is to automatically configure and set the audio enhancement device in a continuous manner based on the audio content during playback. In this "automatic" mode, the user can easily enjoy the content without having to select different presets. On the other hand, continuous adjustment is more important in order to avoid audible artifacts at transition points.

[0015] According to one embodiment of the first aspect, the audio processing device includes: an audio classifier for classifying an audio signal into at least one audio type in real time; an audio enhancement device for improving the audience experience; and an adjustment unit for adjusting at least one parameter of the audio enhancement device in a continuous manner based on a confidence value of the at least one audio type.

[0016] The audio enhancement device may be any of the following: a dialogue enhancer, a surround virtualizer, a volume leveler, and an equalizer.

[0017] Correspondingly, the audio processing method includes classifying an audio signal into at least one audio type in real time; and adjusting at least one parameter for audio improvement in a continuous manner based on the confidence value of the at least one audio type.

[0018] According to another embodiment of the first aspect, the volume leveler controller includes an audio content classifier for identifying the content type of an audio signal in real time; and an adjustment unit for adjusting the volume leveler in a continuous manner based on the identified content type. The adjustment unit may be configured to positively correlate the dynamic gain of the volume leveler with the informational content type of the audio signal and negatively correlate the dynamic gain of the volume leveler with the coherent content type of the audio signal.

[0019] An audio processing device having a volume leveling controller as described above is also disclosed.

[0020] Correspondingly, the volume leveler control method includes identifying the content type of the audio signal in real time; and adjusting the volume leveler in a continuous manner based on the identified content type. This adjustment is performed by positively correlating the dynamic gain of the volume leveler with the informational content type of the audio signal, and negatively correlating the dynamic gain of the volume leveler with the coherent content type of the audio signal.

[0021] According to yet another embodiment of the first aspect, the equalizer controller includes an audio classifier for identifying the audio type of an audio signal in real time; and an adjustment unit for adjusting the equalizer in a continuous manner based on a confidence value of the identified audio type.

[0022] An audio processing device having the equalizer controller described above is also disclosed.

[0023] Correspondingly, the equalizer control method includes identifying the audio type of an audio signal in real time; and adjusting the equalizer in a continuous manner based on the confidence value of the identified audio type.

[0024] The present invention also provides a computer-readable medium on which computer program instructions are recorded, wherein, when executed by a processor, the instructions enable the processor to execute the audio processing method, the volume leveler control method, or the equalizer control method described above.

[0025] According to the embodiment of the first aspect, an audio enhancement device, which may be one of a dialogue enhancer, a surround virtualizer, a volume leveler, and an equalizer, may be continuously adjusted according to the type of audio signal and / or the confidence value of said type.

[0026] A second aspect of this invention is the development of a content identification component that identifies multiple audio types. The detection results may be used to control / guide the behavior of various audio enhancement devices in finding better parameters in a continuous manner.

[0027] According to a second aspect of one embodiment, the audio classifier includes: a short-term feature extractor that extracts short-term features from short-term audio segments, each containing a sequence of audio frames; a short-term classifier that classifies sequences of short-term segments within long-term audio segments into various short-term audio types using their respective short-term features; a statistical extractor that calculates statistics of the short-term classifier results as long-term features with respect to sequences of short-term segments within long-term audio segments; and a long-term classifier that classifies the long-term audio segments into long-term audio types using the long-term features.

[0028] An audio processing device having the above-described audio classifier is also disclosed.

[0029] Correspondingly, the audio classification method includes: extracting short-term features from short-term audio segments, each containing a sequence of audio frames; classifying sequences of short-term segments within long-term audio segments into various short-term audio types using their respective short-term features; calculating long-term features as statistics resulting from the classification process for sequences of short-term segments within long-term audio segments; and classifying long-term audio segments into long-term audio types using these long-term features.

[0030] According to another embodiment of the second aspect, the audio classifier includes an audio content classifier that identifies the content type of a short segment of an audio signal; and an audio context classifier that identifies the context type of the short segment based at least in part on the content type identified by the audio content classifier.

[0031] An audio processing device having the above-described audio classifier is also disclosed.

[0032] Correspondingly, the audio classification method includes identifying the content type of a short-term segment of an audio signal; and at least partially identifying the context type of the short-term segment based on the identified content type.

[0033] The present invention also provides a computer-readable medium on which computer program instructions are recorded, wherein, when executed by a processor, the instructions enable the processor to perform the audio classification method described above.

[0034] According to a second aspect embodiment, the audio signal may be classified into various long-term or contextual types, distinct from short-term or content types. The type of the audio signal and / or confidence values ​​of said type may be further used to tune audio enhancement devices such as dialogue enhancers, surround virtualizers, volume levelers, or equalizers. [Brief explanation of the drawing]

[0035] This application is illustrated, not limited to, examples, in the accompanying drawings. In the drawings, similar reference numerals refer to similar elements. [Figure 1] This figure shows an audio processing device based on one embodiment of the present invention. [Figure 2] This figure shows a variation of the embodiment shown in Figure 1. [Figure 3] This figure shows a variation of the embodiment shown in Figure 1. [Figure 4] This diagram shows possible configurations for a classifier that identifies multiple audio types and calculates a confidence score. [Figure 5] This diagram shows possible configurations for a classifier that identifies multiple audio types and calculates a confidence score. [Figure 6]This diagram shows possible configurations for a classifier that identifies multiple audio types and calculates a confidence score. [Figure 7] This figure shows a further embodiment of the audio processing device of the present invention. [Figure 8] This figure shows a further embodiment of the audio processing device of the present invention. [Figure 9] This figure shows a further embodiment of the audio processing device of the present invention. [Figure 10] This figure shows the delays in transitions between various audio types. [Figure 11] This is a flowchart showing an audio processing method based on an embodiment of the present invention. [Figure 12] This is a flowchart showing an audio processing method based on an embodiment of the present invention. [Figure 13] This is a flowchart showing an audio processing method based on an embodiment of the present invention. [Figure 14] This is a flowchart showing an audio processing method based on an embodiment of the present invention. [Figure 15] This figure shows a dialog enhancement controller based on one embodiment of the present invention. [Figure 16] This flowchart shows the use of the audio processing method based on the present invention in the control of a dialogue enhancer. [Figure 17] This flowchart shows the use of the audio processing method based on the present invention in the control of a dialogue enhancer. [Figure 18] This figure shows a surround virtualizer controller based on one embodiment of the present invention. [Figure 19] This flowchart shows the use of the audio processing method according to the present invention in controlling a surround virtualizer. [Figure 20] This figure shows a volume leveler controller based on one embodiment of the present invention. [Figure 21] This figure shows the effect of the volume leveling controller based on the present invention. [Figure 22]This figure shows an equalizer controller based on one embodiment of the present invention. [Figure 23] This figure shows some examples of desired spectral balance presets. [Figure 24] This figure shows an audio classifier based on one embodiment of the present invention. [Figure 25] This figure shows some of the features used by the audio classifier of this invention. [Figure 26] This figure shows some of the features used by the audio classifier of this invention. [Figure 27] This figure shows a further embodiment of the audio classifier based on the present invention. [Figure 28] This figure shows a further embodiment of the audio classifier based on the present invention. [Figure 29] This figure shows a further embodiment of the audio classifier based on the present invention. [Figure 30] This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 31] This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 32] This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 33] This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 34] This figure shows an audio classifier based on another embodiment of the present application. [Figure 35] This figure shows an audio classifier based on yet another embodiment of the present invention. [Figure 36] This figure shows the heuristic rules used in the audio classifier of this application. [Figure 37] This figure shows a further embodiment of the audio classifier based on the present invention. [Figure 38] This figure shows a further embodiment of the audio classifier based on the present invention. [Figure 39]This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 40] This is a flowchart of the audio classification method based on the embodiment of the present invention. [Figure 41] This is a block diagram illustrating an exemplary system implementing an embodiment of the present invention. [Modes for carrying out the invention]

[0036] Embodiments of the present application are described below with reference to the drawings. For clarity, it should be noted that representations and descriptions of components and processes known to those skilled in the art but not necessary for understanding the present application are omitted in the drawings and description.

[0037] As those skilled in the art will understand, aspects of this application may be embodied as systems, devices (e.g., mobile phones, portable media players, personal computers, transsensors, television set-top boxes or digital video recorders or any other media players), methods, or computer program products. Thus, aspects of this application may take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining both software and hardware aspects. These may all be referred to in this paper as “circuits,” “modules,” or “systems.” Furthermore, aspects of this application may take the form of computer program products embodied in one or more computer-readable media in which computer-readable program code is embodied.

[0038] Any combination of one or more computer-readable media may be used. Computer-readable media may be computer-readable signal media or computer-readable storage media. Computer-readable storage media may be, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Further specific examples (not exhaustive) of computer-readable storage media include: electrical connections with one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this paper, computer-readable storage media may be any tangible medium that contains or can store programs for use by or in connection with instruction execution systems, apparatus, or devices.

[0039] Computer-readable signaling media may include propagating data signals in which computer-readable program code is embodied, for example, in the baseband or as part of a carrier wave. Such propagating signals may take any variety of forms, including but not limited to electromagnetic or optical signals or any preferred combination thereof.

[0040] A computer-readable signaling medium may be any computer-readable medium other than a computer-readable storage medium that can communicate, propagate, or carry programs for use by or in connection with an instruction execution system, apparatus, or device.

[0041] Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any preferred combination thereof.

[0042] Computer program code for performing the operations of the aspects of this application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code may run as a standalone software package entirely on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In this last scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection may be made to an external computer (for example, via the Internet using an Internet service provider).

[0043] Aspects of the present invention are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It will be understood that each block of a flowchart and / or block diagram, and combinations of blocks of a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions may be given to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine that produces means for the instructions executed by the processor of the computer or other programmable data processing device to implement one or more functions / processes specified in the flowchart and / or block diagram.

[0044] These computer program instructions are stored on a computer-readable medium that can instruct a computer, other programmable data processing device, or other device to function in a particular way, thereby producing a product that includes instructions for implementing one or more functions / processes identified in the flowchart and / or block diagram.

[0045] Computer program instructions may be loaded into a computer, another programmable data processing device, or other device to perform a series of operational processes on the computer, other programmable device, or other device, thereby creating a computer-implemented process such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / steps identified in one or more blocks of the flowchart and / or block diagram.

[0046] The embodiments of the present invention are described in detail below. For clarity, the description is organized as follows: Part 1: Audio Processing Devices and Methods Section 1.1 Audio Type Section 1.2 Audio-type confidence values ​​and classifier configuration Section 1.3 Smoothing of confidence values ​​for audio type Section 1.4 Parameter Adjustment Section 1.5 Parameter Smoothing Section 1.6 Audio Type Transitions Section 1.7 Combinations of Embodiments and Application Scenarios Section 1.8 Audio Processing Methods Part Two: Dialogue Enhancer Controller and Control Method Section 2.1 Levels of Dialogue Improvement Section 2.2 Thresholds for determining the frequency band to be improved Section 2.3 Adjusting to Background Level Section 2.4 Combinations of Embodiments and Application Scenarios Section 2.5 Control Method for Dialogue Enhancer Part Three: Surround Virtualizer Controller and Control Method Section 3.1 Surround Boost Amount Section 3.2 Starting Frequency Section 3.3 Combinations of Embodiments and Application Scenarios Section 3.4 Surround Virtualizer Control Method Part 4: Volume Leveler Controller and Control Method Section 4.1 Informational and Interferential Content Types Section 4.2 Content Types in Various Contexts Section 4.3 Context Type Section 4.4 Combinations of Embodiments and Application Scenarios Section 4.5 Volume Leveler Control Method Part 5: Equalizer Controller and Control Method Section 5.1 Control based on content type Section 5.2 The certainty of the dominant source in music Section 5.3 Equalizer Presets Section 5.4 Context-based control Section 5.5 Combinations of Embodiments and Application Scenarios Section 5.6 Equalizer Control Method Part Six: Audio Classifiers and Classification Methods Section 6.1 Contextual Classifiers Based on Content Type Classification Section 6.2 Extraction of Long-Term Features Section 6.3 Extraction of Short-Term Features Section 6.4 Combinations of Embodiments and Application Scenarios Section 6.5 Audio Classification Methods Part 7: VoIP Classifiers and Classification Methods Section 7.1 Contextual Classification Based on Short-Term Segments Section 7.2 Classification using VoIP utterances and VoIP noise Section 7.3 Smoothing Fluctuations Section 7.4 Combinations of Embodiments and Application Scenarios Section 7.5 VoIP classification method.

[0047] <Part 1: Audio Processing Devices and Methods> Figure 1 shows an overview framework of a content-adaptive audio processing unit 100 that supports the automatic configuration setting of at least one audio enhancement device with improved parameters based on audio content during playback. It has three main components: an audio classifier 200, an adjustment unit 300, and an audio enhancement device 400.

[0048] The audio classifier 200 classifies audio signals into at least one audio type in real time. It automatically identifies the audio type of content during playback. Any audio classification technique can be applied to identify audio content, including signal processing, machine learning, and pattern recognition. A confidence value representing the probability of the audio content being a predefined set of target audio types is estimated almost simultaneously.

[0049] The audio enhancement device 400 improves the audience experience by processing the audio signal, which will be discussed in detail later.

[0050] The adjustment unit 300 adjusts at least one parameter of the audio enhancement device in a continuous manner based on the confidence value of at least one audio type. It is designed to control the behavior of the audio enhancement device 400. It estimates the most suitable parameter of the corresponding audio enhancement device based on the results obtained from the audio classifier 200.

[0051] Various audio enhancement devices can be applied to this device. Figure 2 shows an exemplary system including four audio enhancement devices, including a Dialog Enhancer (DE) 402, a Surround Virtualizer (SV) 404, a Volume Leveler (VL) 406, and an Equalizer (EQ) 408. Each audio enhancement device can be automatically adjusted in a continuous manner based on the results (audio type and / or confidence value) obtained in the audio classifier 200.

[0052] Of course, the audio processing device does not necessarily have to include all types of audio enhancement devices, and may include only one or more of them. On the other hand, the audio enhancement devices are not limited to those given in this disclosure and may include further types of audio enhancement devices, which are also within the scope of this application. Furthermore, the names of the audio enhancement devices discussed in this disclosure, including the dialogue enhancer (DE) 402, surround virtualizer (SV) 404, volume leveler (VL) 406, and equalizer (EQ) 408, are not limiting and should be interpreted as each covering any other device that performs the same or similar function.

[0053] <Section 1.1 Audio Type> To properly control various types of audio enhancement devices, this invention further provides a new audio-type configuration. However, audio-type configurations in the prior art are also applicable in this invention.

[0054] Specifically, audio types are modeled from different semantic levels, including low-level audio elements representing the basic components of an audio signal and high-level audio genres representing most common audio content in real-world user entertainment applications. The former may be called “content types.” Basic audio content types may include speech, music (including songs), background sounds (or sound effects), and noise.

[0055] The meaning of speech and music is clear. In this application, noise refers to physical noise, not semantic noise. Physical noise in this application may include noise caused by technical reasons, such as noise from an air conditioner or pink noise resulting from a signal transmission path. In contrast, “background noise” in this application is a sound effect that may be an auditory event occurring around the core target of the listener’s attention. For example, in an audio signal from a telephone call, in addition to the speaker’s voice, there may be other unintended sounds such as the voice of some other person unrelated to the call, keyboard sounds, footsteps, etc. These unwanted sounds are referred to as “background noise,” not noise. In other words, “background noise” may be defined as a sound that is not the target (or core target of the listener’s attention), or is even more unwanted, but still has some substantive meaning. On the other hand, “noise” may be defined as any unwanted sound other than target-on and background noise.

[0056] Sometimes, background sounds are not truly "unwanted," but are intentionally generated and carry some useful information. For example, this is the case with background sounds in movies, television programs, or radio broadcasts. Therefore, they are sometimes referred to as "sound effects." In this disclosure, for the sake of brevity, only "background sounds" will be used below, and may be further shortened to "background."

[0057] Furthermore, music can be further classified into music without a dominant source and music with a dominant source. Music is called “music with a dominant source” when there is a source (voice or instrument) that is much stronger than the other sources in a musical piece. Otherwise, it is called “music without a dominant source.” For example, in polyphonic music with singing and various instruments, if it is harmonically balanced or if the energies of some of the most prominent sources are comparable to each other, it is considered music without a dominant source. In contrast, if one source (e.g., voice) is much louder while other sources are much quieter, it is considered to contain a dominant source. Another example is “music with a dominant source” if there is a prominent or conspicuous instrument tone.

[0058] Music can be further classified into various types based on various standards. Music can be classified based on genres such as rock, jazz, rap, and folk, but is not limited to these. Music can also be classified based on instruments, such as vocal music and instrumental music. Instrumental music can include various types of music performed using various instruments, such as piano music and guitar music. Other exemplary standards include any other musical attributes that can be used to group music based on similarities in rhythm, tempo, timbre, and / or attributes. For example, based on timbre, vocal music can be classified into tenor, baritone, bass, soprano, mezzo-soprano, and alto.

[0059] The content type of an audio signal may be classified with respect to short-term audio segments, such as those consisting of multiple frames. Generally, audio frames are several milliseconds long, such as 20ms, while the length of short-term audio segments to be classified by an audio classifier can range from several hundred milliseconds to several seconds, for example, one second.

[0060] To control the audio enhancement device in a content-adaptive manner, the audio signal may be classified in real time. For the content types described above, the content type of the current short-term audio segment represents the content type of the current audio signal. Since the length of the short-term audio segments is not very long, the audio signal may be divided into sequential, non-overlapping short-term audio segments. However, the short-term audio segments may be sampled continuously / semi-continuously along the time axis of the audio signal. That is, the short-term audio segments may be sampled using a window with a predetermined length (the intended length of the short-term audio segment) that moves along the time axis of the audio signal in step sizes of one or more frames.

[0061] High-level audio genres are sometimes referred to as "contextual types" because they represent the long-term type of an audio signal, and may be considered the environment or context of the sound event at that time, and may be classified as content types as described above. According to this application, contextual types may include most common audio applications such as cinematic media, music (including songs), games, and VoIP (Voice over Internet Protocol).

[0062] The meanings of music, games, and VoIP are self-evident. Cinematic media may include films, television programs, radio broadcasts, or any other audio media similar to those mentioned above. The main characteristic of cinematic media is a mixture of possible speech, music, and various kinds of background sounds (sound effects).

[0063] It should be noted that both content-type and context-type music include music (including songs). Hereafter, in this application, the terms "short-term music" and "long-term music" will be used to distinguish between them.

[0064] For some embodiments of the present invention, several other context-type configurations are also proposed.

[0065] For example, audio signals may be classified as high-quality audio (such as cinematic media and music CDs) or low-quality audio (such as VoIP, low-bitrate online streaming audio, and user-generated content). These may be collectively referred to as "audio quality types."

[0066] As another example, audio signals may be classified as VoIP or non-VoIP. These may be considered conversions of the four contextual configurations described above (VoIP, cinematic media, (long-running) music, and games). In relation to the VoIP or non-VoIP context, audio signals may also be classified as VoIP-related audio content types such as VoIP speech, non-VoIP speech, VoIP noise, and non-VoIP noise. The configuration of VoIP audio content types is particularly useful for distinguishing between VoIP and non-VoIP contexts, as the VoIP context is typically the most challenging application scenario for volume levelers (a type of audio enhancement device).

[0067] Generally, the contextual type of an audio signal may be classified with respect to long-term audio segments rather than short-term audio segments. A long-term audio segment consists of multiple frames, more than the number of frames in a short-term audio segment. A long-term audio segment may consist of multiple short-term audio segments. Generally, a long-term audio segment can have a length of several seconds to tens of seconds, for example, 10 seconds.

[0068] Similarly, to control the audio enhancement device in an adaptive manner, the audio signal may be classified into context types in real time. Likewise, the context type of the current long-term audio segment represents the context type of the current audio signal. Because the length of the long-term audio segment is relatively long, the audio signal may be sampled continuously / semi-continuously along the time axis of the audio signal to avoid abrupt changes in its context type, and thus abrupt changes in the operating parameters of the audio enhancement device(s). That is, the long-term audio segment may be sampled using a window with a predetermined length (the intended length of the long-term audio segment) that moves along the time axis of the audio signal in steps of one or more frames or one or more short-term segments.

[0069] In the above, both content types and context types have been described. In embodiments of the present application, the adjustment unit 300 may adjust at least one parameter of the audio enhancement device(s) based on at least one of various content types and / or at least one of various context types. Thus, in one variation of the embodiment shown in Figure 1, as shown in Figure 3, the audio classifier 200 may have an audio content classifier 202 or an audio context classifier 204 or both.

[0070] The above has mentioned various audio types based on various standards (for example, for context types) and various audio types at various hierarchical levels (for example, for content types). However, these standards and hierarchical levels are merely for the convenience of description and are not limiting in any way. In other words, in this application, any two or more of the above-mentioned audio types can be identified simultaneously by the audio classifier 200 and considered simultaneously by the adjustment unit 300. This will be discussed later. In other words, all audio types at various hierarchical levels may be in parallel or at the same level.

[0071] <Section 1.2 Audio-type confidence values ​​and classifier configuration> The audio classifier 200 may output a hard judgment result, or the adjustment unit 300 may consider the result of the audio classifier 200 as a hard judgment result. Even with hard judgments, multiple audio types can be assigned to an audio segment. For example, an audio segment may be a mixed signal of speech and short musical passages and can be labeled by both "speech" and "short musical passages". The resulting labels can be used directly to control one or more audio enhancement devices 400. A simple example is to enable the dialogue enhancer 402 when speech is present and turn it off when speech is absent. However, this hard judgment method may introduce some unnaturalness at the transition point from one audio type to another, without a careful smoothing scheme (described later).

[0072] To allow for more flexible and continuous adjustment of the audio enhancement device parameters, a confidence value can be estimated for each target audio type (soft judgment). The confidence value represents the level of match between the audio content to be identified and the target audio type, on a value from 0 to 1.

[0073] As mentioned earlier, many classification techniques may directly output confidence values. Confidence values ​​can also be calculated from various methods that may be considered part of the classifier. For example, when an audio model is trained using some probabilistic modeling techniques such as Gaussian Mixture Models (GMMs), the posterior probability can be used to represent the confidence value, as follows:

[0074]

number

[0075] On the other hand, when an audio model is trained using discriminative methods such as Support Vector Machines (SVMs) and adaBoost, only a score (a real number) is obtained from the model comparison. In these cases, the following sigmoid function is typically used to map the obtained score (theoretically from -∞ to ∞) to the expected confidence level (from 0 to 1).

[0076]

number

[0077] In some embodiments of the present invention, the adjustment unit 300 may use three or more content types and / or three or more context types. In this case, the audio content classifier 202 must identify three or more content types, and / or the audio context classifier 204 must identify three or more context types. In such a situation, the audio content classifier 202 or the audio context classifier 204 may be a group of classifiers organized in a certain configuration.

[0078] For example, if the adjustment unit 300 requires all four types of contexts—cinematic media, long-running music, games, and VoIP—the audio context classifier 204 can have one of the following various configurations.

[0079] Firstly, the audio context classifier 204 may have six one-to-one binary classifiers (each classifier discriminates one target audio type from another) organized as shown in Figure 4, three one-to-other binary classifiers (each classifier discriminates a target audio type from another audio type) organized as shown in Figure 5, and four one-to-other classifiers organized as shown in Figure 6. Other configurations exist, such as a Decision Directed Acyclic Graph (DDAG) configuration. Note that in Figures 4-6 and the corresponding descriptions below, "film" is used instead of "cinematic media" for brevity.

[0080] Each binary classifier assigns a confidence score H(x) to its output (where x represents an audio segment). After obtaining the outputs of each binary classifier, it is necessary to map them to the final confidence values ​​of the identified context types.

[0081] Generally, audio signals are classified into M context types (where M is a positive integer). A typical one-to-one configuration constructs M(M-1) / 2 classifiers, each trained with data from two classes. Each one-to-one classifier then casts one vote for its preferred class, and the final result is the class with the most votes among the M(M-1) / 2 classifiers. Compared to a typical one-to-one configuration, the hierarchical configuration in Figure 4 also requires constructing M(M-1) / 2 classifiers. However, since segment x is determined to be in / not in the corresponding class at each hierarchical level, and the overall level count is M-1, the test iteration process can be shortened to M-1. The final confidence value for various context types is, for example, the binary classification confidence value H. k (x) may also be used for calculation (k=1,2,...,6 represents various context types).

[0082]

number

[0083]

number

[0084]

number

[0085] It should be noted that in the configurations shown in Figures 4 to 6, the sequences of various binary classifiers are not necessarily as illustrated, and other sequences may be used. Such sequences may be selected by manual assignment or automated learning, depending on the various requirements of different applications.

[0086] The above description is directed to the audio context classifier 204. The situation is similar for the audio content classifier 202.

[0087] Alternatively, the audio content classifier 202 or the audio context classifier 204 may be implemented as a single classifier that simultaneously identifies all content types / context types and provides corresponding confidence values. There are many existing techniques for doing this.

[0088] Using confidence values, the output of the audio classifier 200 can be represented as a vector. Each dimension of the vector represents the confidence value for each target audio type. For example, if the target audio types are sequential (speech, short bursts of music, noise, background), an exemplary output could be (0.9, 0.5, 0.0, 0.0). This indicates that there is 90% certainty that the audio content is speech and 50% certainty that the audio is music. Note that the sum of all dimensions in the output vector does not need to be 1 (for example, the results from Figure 6 are not necessarily normalized). In other words, the audio signal may be a mixture of speech and short bursts of music.

[0089] Later, in Parts 6 and 7, novel implementations of audio context classification and audio content classification will be discussed in detail.

[0090] <Section 1.3 Smoothing of audio-type confidence values> Optionally, after each audio segment has been classified into a predefined audio type, an additional step is to smooth the classification results along the time axis to avoid abrupt jumps from one type to another and to allow for smoother estimation of parameters in the audio enhancement device. For example, if a long excerpt is classified as cinematic media except for one segment classified as VoIP, the abrupt VoIP determination can be corrected to cinematic media through smoothing.

[0091] Therefore, in one variation of the embodiment shown in Figure 7, a type smoothing unit 712 is further provided for each audio type to smooth the confidence value of the audio signal at the current time.

[0092] Common smoothing methods are based on weighted averages, such as calculating a weighted sum of the current actual confidence value and the smoothed confidence value at the last point in time.

[0093] smoothConf(t)=β·smoothConf(t-1)+(1-β)·conf(t) (3) Here, t is the current time (current audio segment), t-1 is the last time (last audio segment), β is the weight, and conf and smoothConf are the confidence values ​​before and after smoothing, respectively.

[0094] From a confidence value perspective, the results from the hard decisions of the classifier can also be expressed using confidence values, which are either 0 or 1. That is, if a target audio type is selected and assigned to an audio segment, the corresponding confidence value is 1; otherwise, the confidence value is 0. Therefore, even if the audio classifier 200 does not provide confidence values ​​and only provides hard decisions regarding audio types, continuous adjustment of the adjustment unit 300 is still possible through the smoothing operation of the type smoothing unit 712.

[0095] A smoothing algorithm can be "asymmetric" by using different smoothing weights for different cases. For example, the weights used to calculate the weighted sum may be adaptively changed based on the confidence value of the audio type of the audio signal. If the confidence value of the current segment is higher, its weight will also be higher.

[0096] From another perspective, the weights for calculating the weighted sum may be adaptively changed based on different pairs of transitions from one audio type to another, especially when the audio enhancement device(s) are adjusted based on multiple content types identified by the audio classifier 200, rather than on the presence or absence of a single content type. For example, for transitions from an audio type that appears more frequently in a given context to another audio type that appears less frequently in that context, the confidence value of the latter may be smoothed so as not to increase too rapidly, as it could be a random occurrence.

[0097] Another factor is the changing (increasing or decreasing) trend, including the rate of change. If we become more concerned with latency as a certain audio type comes into existence (i.e., as its confidence level increases), we can design a smoothing algorithm as follows:

[0098]

number

[0099] From a different perspective, considering a change trend in an audio type is merely a specific example of considering different transition pairs of audio types. For example, increasing the confidence value of type A may be considered a transition from non-A to A, and decreasing the confidence value of type A may be considered a transition from A to non-A.

[0100] <Section 1.4 Parameter Adjustment> The adjustment unit 300 is designed to estimate or adjust the appropriate parameters for one or more audio enhancement devices 400 based on the results obtained from the audio classifier 200. Different adjustment algorithms may be designed for different audio enhancement devices using either content-type or context-type information, or both for joint determination. For example, with context-type information such as cinematic media and long-form music, the presets described above can be automatically selected and applied to the corresponding content. Using the available content-type information, the parameters of each audio enhancement device can be adjusted in a more granular manner, as shown in the following section. Content-type information and context-type information can also be used jointly in the adjustment unit 300 to balance long-term and short-term information. A specific adjustment algorithm for a particular audio enhancement device may be considered a separate adjustment unit, or different adjustment algorithms may be considered as a combined adjustment unit.

[0101] In other words, the adjustment unit 300 may be configured to adjust at least one parameter of the audio enhancement device based on confidence values ​​of at least one content type and / or at least one context type. For a particular audio enhancement device, some audio types are informative and some audio types are coherent. Thus, the parameters of a particular audio enhancement device may be positively or negatively correlated with confidence values ​​of one or more informative audio types or one or more coherent audio types. Here, "positively correlated" means that the parameter increases or decreases linearly or nonlinearly as the confidence value of the audio type increases or decreases. "Negatively correlated" means that the parameter increases or decreases linearly or nonlinearly as the confidence value of the audio type decreases or increases, respectively.

[0102] Here, the decrease and increase in confidence value are directly “transmitted” to the parameter to be adjusted by a positive or negative correlation. In mathematics, such correlations or “transmissions” can be embodied as direct or inverse proportion, plus or minus (addition or subtraction) operations, multiplication or division, or nonlinear functions. All these forms of correlation may be called “transfer functions.” To determine an increase or decrease in confidence value, the current confidence value or its mathematical transformation can also be compared to the last confidence value or a set of historical confidence values ​​or their mathematical transformations. In the context of this application, the term “comparison” means comparison through subtraction or comparison through division. An increase or decrease can be determined by determining whether the difference is greater than 0 or whether the ratio is greater than 1.

[0103] In individual implementations, parameters can be directly related to confidence values ​​or their ratios or differences through appropriate algorithms (such as transfer functions), and it is not necessary for an "external observer" to explicitly know that a particular confidence value and / or a particular parameter has increased or decreased. Several individual examples are given in Parts 2-5 below, which discuss individual audio enhancement devices.

[0104] As described in the previous section, for the same audio segment, the classifier 200 may identify multiple audio types, each with its own confidence value. Since an audio segment may contain multiple components simultaneously, such as music, speech, and background noise, their confidence values ​​may not necessarily sum to 1. In such situations, the parameters of the audio enhancement device need to be balanced among different audio types. For example, the adjustment unit 300 may be configured to consider at least some of the multiple audio types by weighting the confidence value of at least one audio type based on the importance of that at least one audio type. The more important a particular audio type is, the greater the parameter will be affected by it.

[0105] The weights can also reflect the informational and coherent effects of the audio type. For example, coherent audio types may be assigned negative weights. Several specific examples are given in parts two to five below, which describe individual audio enhancement devices.

[0106] It should be noted that in the context of this application, "weight" has a broader meaning than coefficients in a polynomial. In addition to coefficients in a polynomial, "weight" can also take the form of an exponent or power. When it is a coefficient in a polynomial, the weighting coefficient may or may not be normalized. In short, a weight simply represents how much influence the weighted object has on the parameter to be adjusted.

[0107] In some other embodiments, the confidence values ​​of multiple audio types included in the same audio segment may be converted into weights through normalization. The final parameters may then be determined by calculating the sum of parameter preset values ​​that are predefined for each audio type and weighted by the confidence values. That is, the adjustment unit 300 may be configured to consider multiple audio types by weighting the effects of multiple audio types based on their confidence values.

[0108] As an example of weighting, the tuning unit is configured to consider at least one dominant audio type based on its confidence value. Audio types with a confidence value that is too low (below the threshold) may not be considered. This is equivalent to setting the weight of other audio types with confidence values ​​below the threshold to 0. Several specific examples are given in parts two to five below, which describe individual audio enhancement devices.

[0109] Content type and context type can be considered together. In one embodiment, they can be considered to be at the same level, and their confidence values ​​may have their own weights. In another embodiment, as the name suggests, “context type” is the context or environment in which the “context type” is located, and thus the adjustment unit 200 may be configured so that content types in audio signals of different context types are assigned different weights depending on the context type of the audio signal. In general, any audio type can constitute the context of another audio type, and the adjustment unit 200 may be configured to modify the weight of one audio type using the confidence value of another audio type. Several specific examples are given in parts two to five below, which describe specific audio enhancement devices.

[0110] In the context of this application, the term "parameter" has a broader meaning than its literal meaning. In addition to parameters with a single value, a parameter can also mean a set of various parameters, a vector or profile consisting of various parameters, and presets as described above. In particular, the following parameters are discussed in Parts II to V below, but this application is not limited to them: the dialogue enhancement level, the threshold for determining the frequency band to which dialogue enhancement should occur, the background level, the surround boost amount, the starting frequency for the surround virtualizer, the dynamic gain or dynamic gain range of the volume leveler, a parameter indicating the degree to which an audio signal is a new perceptible audio event, the equalization level, the equalization profile, and the spectral balance preset.

[0111] <Section 1.5 Parameter Smoothing> Section 1.3 discussed smoothing the confidence values ​​of the audio type to avoid abrupt changes, and therefore abrupt changes in the parameters of the audio enhancement device. Other measures are also possible. One is to smooth the parameters that are tuned based on the audio type, which will be discussed in this section. The other is to configure the audio classifier and / or tuning unit to slow down changes in the audio classifier's results, which will be discussed in Section 1.6.

[0112] In one embodiment, the parameters can be further smoothed as follows to avoid rapid changes that may introduce audible artifacts at transition points:

[0113]

number

[0114] That is, as shown in Figure 8, the audio processing device may have a parameter smoothing unit 814. This unit smooths the parameter values ​​determined by the adjustment unit 300 at the current time by calculating a weighted sum of the parameter values ​​determined by the adjustment unit at the current time and the smoothed parameter values ​​at the last time, for the parameters of the audio enhancement devices (such as the dialogue enhancer 402, surround virtualizer 404, volume leveler 406, and equalizer 408) that are adjusted by the adjustment unit 300.

[0115] The time constant τ can be a fixed value based on the specific requirements of the application and / or the implementation of the audio enhancement device 400. The time constant τ may be adaptively changed based on the audio type, in particular based on various transition types from one audio type to another, such as from music to speech, or from speech to music.

[0116] Let's take an equalizer as an example (further details may be discussed in Part Five). Equalization works well for musical content but not so well for spoken content. Therefore, to smooth the level of equalization, the time constant can be relatively small when the audio signal transitions from music to speech, thereby allowing a smaller level of equalization to be applied more quickly to the spoken content. On the other hand, the time constant for the transition from speech to music can be relatively large to avoid audible artifacts at the transition point.

[0117] To estimate the transition type (e.g., speech to music or music to speech), the content classification results can be used directly. That is, if audio content is classified as music or speech, obtaining the transition type becomes straightforward. Instead of directly comparing hard decisions of audio type, we can also rely on the estimated unsmoothed equalization level so that transitions can be estimated in a more continuous manner. The general idea is that if the unsmoothed equalization level increases, it indicates a transition from speech to music (or more musical), and if not, it is closer to a transition from music to speech (or more speech). By distinguishing between different transition types, time constants can be set accordingly. One example is as follows:

[0118]

number

[0119] <Section 1.6 Audio Type Transitions> Referring to Figures 9 and 10, another method for avoiding abrupt changes in the audio type, and thus abrupt changes in the parameters of the audio enhancement device, is described.

[0120] As shown in Figure 9, the audio processing unit 100 may further have a timer 916 for measuring the duration for which the audio classifier 200 continuously outputs the same new audio type. The adjustment unit 300 may be configured to continue using the current audio type until the duration of the new audio type reaches a threshold.

[0121] In other words, an observation (or maintenance) phase is introduced, as shown in Figure 10. Before the adjustment unit 300 actually uses the new audio type, the audio type change is further monitored over a continuous amount of time in the observation phase (corresponding to a threshold of duration length) to confirm whether the audio type has actually changed.

[0122] As shown in Figure 10, arrow (1) indicates that the current state is type A and the result of the audio classifier 200 remains unchanged.

[0123] If the current state is type A and the result of the audio classifier 200 is type B, the timer 916 starts timing, or, as shown in Figure 10, the process enters the observation phase (arrow (2)), and the initial value of the hangover count cnt is set. This indicates the length of the observation period (equal to the threshold).

[0124] Next, if the audio classifier 200 continuously outputs type B, and cnt continuously decreases (arrow (3)), until cnt eventually becomes equal to 0 (i.e., the duration of the new type B reaches a threshold), then the adjustment unit 300 may use the new audio type B (arrow (4)). In other words, only at this point can the audio type be considered to have truly changed to type B.

[0125] Otherwise, if the output of the audio classifier 200 returns to the original type A before cnt becomes 0 (before the duration reaches the threshold), the observation phase is terminated and the adjustment unit 300 continues to use the original type A (arrow (5)).

[0126] The change from type B to type A may be the same as the process described above.

[0127] In the process described above, the threshold (or remaining count) may be set based on the requirements of the application. This may be a predefined fixed value, or it may be set adaptively. In some variations, the threshold may differ for different transition pairs from one audio type to another. For example, when changing from type A to type B, the threshold may be the first value, and when changing from type B to type A, the threshold may be the second value.

[0128] In another variation, the hangover count (threshold) may be negatively correlated with the confidence value of the new audio type. The general idea is that if the confidence value indicates confusion between the two types (for example, when the confidence value is only about 0.5), the observation duration needs to be long. Otherwise, the duration can be relatively short. Following this guideline, an exemplary hangover count can be set by the following formula:

[0129] HangCnt=C·|0.5-Conf|+D Here, HangCnt is the remaining duration or threshold, and C and D are two parameters that can be set based on the requirements of the application, typically C being a negative value and D being a positive value.

[0130] Note that while timer 916 (and thus the transition process described above) is part of the audio processing unit, it has been described as being external to the audio classifier 200. In some other embodiments, it may be considered part of the audio classifier 200, as described in Section 7.3.

[0131] <Section 1.7 Combinations of Embodiments and Application Scenarios> All embodiments and their variations discussed above may be implemented in any combination thereof, and any components that are mentioned in different parts / embodiments but have the same or similar function may be implemented as the same or separate components.

[0132] Specifically, when describing embodiments and their variations above, components with the same reference numerals as those already described in previous embodiments or variations are omitted, and only different components are described. In fact, these different components can be combined with components of other embodiments or variations, or they can constitute separate solutions on their own. For example, any two or more of the solutions described with reference to Figures 1 to 10 may be combined with each other. In the most complete solution, the audio processing unit may have both an audio content classifier 202 and an audio context classifier 204, as well as a smoothing unit 712, a parameter smoothing unit 814, and a timer 916.

[0133] As previously mentioned, the audio enhancement device 400 may include a dialogue enhancer 402, a surround virtualizer 404, a volume leveler 406, and an equalizer 408. The audio processing unit 100 may include any one or more of these, and the adjustment unit 300 may be adapted accordingly. When involving multiple audio enhancement devices 400, the adjustment unit 300 may be considered to include multiple subunits 300A to 300D (Figures 15, 18, 20, and 22) specific to each audio enhancement device 400, or it may still be considered as a single combined adjustment unit. When specific to a particular audio enhancement device, the adjustment unit 300 may be considered together with the audio classifier 200 and other possible components as the controller for that particular audio enhancement device. This will be discussed in detail in Parts Two to Five below.

[0134] Furthermore, the audio enhancement device 400 is not limited to the examples described above and may include any other audio enhancement device.

[0135] Furthermore, any solutions already discussed or any combination thereof may be further combined with any embodiments described or implied in other parts of this disclosure. In particular, the embodiments of audio classifiers discussed in Parts VI and VII may be used in audio processing devices.

[0136] <Section 1.8 Audio Processing Methods> It is clear that in the process of describing the audio processing apparatus in the above embodiments, several processes or methods are also disclosed. An overview of these methods is given below, but some of the details already discussed above will not be repeated. However, although these methods are disclosed in the process of describing the audio processing apparatus, these methods do not necessarily employ the components described, nor are they necessarily performed by such components. For example, embodiments of the audio processing apparatus may be implemented partially or entirely using hardware and / or firmware, while the audio processing methods discussed below may employ the hardware and / or firmware of the audio processing apparatus, but may also be implemented entirely by a computer executable program.

[0137] These methods are described below with reference to Figures 11-14. Note that when these methods are implemented in real time, various operations are repeated in response to the streaming attributes of the audio signal, and the different operations are not necessarily for the same audio segment.

[0138] In the embodiment shown in Figure 11, an audio processing method is provided. First, the audio signal to be processed is classified in real time into at least one audio type (operation 1102). Based on the confidence value of the at least one audio type, at least one parameter for audio improvement can be continuously adjusted (operation 1104). Audio improvement may be dialogue enhancement (operation 1106), surround virtualization (operation 1108), volume leveling (operation 1110), and / or equalization (operation 1112). Correspondingly, the at least one parameter may include at least one parameter for at least one of the dialogue enhancement process, surround virtualization process, volume leveling process, and equalization process.

[0139] Here, "in real time" and "continuously" mean that the audio type, and therefore the parameters, change in real time along with the specific content of the audio signal. "Continuously" also means that the adjustment is a continuous adjustment based on confidence values, rather than an abrupt or discrete adjustment.

[0140] The audio type may include a content type and / or a context type. Correspondingly, the adjustment operation 1104 may be configured to adjust the at least one parameter based on a confidence value of at least one content type and at least one context type. The content type may further include at least one of the content types of short-term music, speech, background sounds, and noise. The context type may further include at least one of the context types of long-term music, cinematic media, games, and VoIP.

[0141] Several other context-type schemes are also proposed. For example, a VoIP relationship context type that includes VoIP and non-VoIP, and an audio quality type that includes high-quality audio or low-quality audio.

[0142] Short-term music may be further classified into subtypes according to various standards. Depending on the presence of a dominant source, it may include music without a dominant source and music with a dominant source. Furthermore, short-term music may include at least one genre-based cluster or at least one instrument-based cluster or at least one music cluster classified based on the rhythm, tempo, timbre and / or any other musical attribute of the music.

[0143] When both content types and context types are identified, the importance of a content type may be determined by the context type in which it is located. That is, content types in audio signals of different context types are assigned different weights depending on the context type of the audio signal. More generally, one audio type may influence another audio type, or may be a prerequisite for another audio type. Therefore, the operation of adjustment 1104 may be configured to modify the weight of one audio type using the confidence value of another audio type.

[0144] When an audio signal is classified into multiple audio types simultaneously (i.e., with respect to the same audio segment), the operation of the tuner 1104 may take into account some or all of the identified audio types for tuning parameters (one or more) to improve that audio segment. For example, the operation of the tuner 1104 may be configured to weight the confidence value of at least one audio type based on the importance of that at least one audio type. Alternatively, the operation of the tuner 1104 may be configured to take into account at least some of the audio types by weighting them based on their confidence values. In a special case, the operation of the tuner 1104 may be configured to take into account at least one dominant audio type based on its confidence value.

[0145] To avoid abrupt changes in the results, a smoothing method may be introduced.

[0146] The adjusted parameter values ​​may be smoothed (action 1214 in Figure 12). For example, the parameter value determined by action 1104 at the current time may be replaced by a weighted sum of the parameter value determined by action 1104 at the current time and the smoothed parameter value at the last time. In this way, the parameter values ​​are smoothed over time through sequentially iterative smoothing actions.

[0147] The weights used to calculate the weighted sum may be adaptively changed based on the audio type of the audio signal or on different transition pairs from one audio type to another. Alternatively, the weights used to calculate the weighted sum may be adaptively changed based on an increasing or decreasing trend in the parameter values ​​determined by the tuning operation.

[0148] Another smoothing method is shown in Figure 13. This method may further include smoothing the confidence value of the audio signal at the current time by calculating a weighted sum of the actual confidence value at the current time and the smoothed confidence value at the last time for each audio type (operation 1303). Similar to parameter smoothing operation 1214, the weights for calculating the weighted sum may be adaptively changed based on the confidence values ​​of the audio types of the audio signal, or based on different transition pairs from one audio type to another.

[0149] Another smoothing method is a buffer mechanism that delays the transition from one audio type to another, even if the output of the audio classification operation 1102 changes. In other words, the operation of the adjustment 1104 does not immediately use the new audio type, but waits for the output of the audio classification operation 1102 to stabilize.

[0150] Specifically, this method may include measuring the duration for which the classification operation continuously outputs the same new audio type (operation 1403 in Figure 14). Here, the operation of adjustment 1104 is configured to continue using the current audio type (operation "N" in operation 14035 and operation 11041) until the duration of the new audio type reaches a certain threshold ("Y" in operation 14035 and operation 11042). Specifically, when the audio type output from audio classification operation 1102 changes with respect to the current audio type used in audio parameter adjustment operation 1104 ("Y" in operation 1403), timing begins (operation 14032). If audio classification operation 1102 continues to output new audio types, that is, if the judgment in operation 14031 remains "Y", timing continues (operation 14032). Ultimately, when the duration of the new audio type reaches a threshold ("Y" in operation 14035), adjustment operation 1104 uses the new audio type (operation 11042), and the timer is reset in preparation for the next audio type switch (operation 14034). Until the threshold is reached ("N" in operation 14035), adjustment operation 1104 continues to use the current audio type (operation 11041).

[0151] Here, timing may be implemented using a timer mechanism (count-up or count-down). After timing has started, but before the threshold is reached, if the output of audio classification operation 1102 returns to the current audio type used in adjustment operation 1104, then it should be considered that there has been no change with respect to the current audio type used by adjustment operation 1104 ("N" in operation 14031). However, the current classification result (corresponding to the current audio segment to be classified in the audio signal) has changed with respect to the previous output of audio classification operation 1102 (corresponding to the previous audio segment to be classified in the audio signal) ("Y" in operation 14033), and therefore timing is reset until the next change ("Y" in operation 14031) starts timing (operation 14034). Of course, if the classification result of audio classification operation 1102 does not change with respect to the current audio type used by audio parameter adjustment operation 1104 ("N" in operation 14031), and also does not change with respect to the previous classification ("N" in operation 14033), then this indicates that the audio classification is in a stable state and the current audio type will continue to be used.

[0152] The thresholds used here may differ for different transition pairs from one audio type to another, because, when the state is not very stable, it may generally be preferable for the audio enhancement device to be in its default state rather than other states. On the other hand, if the confidence level of the new audio type is relatively high, it is safer to transition to the new audio type. Therefore, the threshold may be negatively correlated with the confidence level of the new audio type. The higher the confidence level, the lower the threshold, meaning the audio type can transition to the new audio type more quickly.

[0153] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is feasible. On the other hand, every aspect of embodiments and variations of audio processing methods may be a separate solution. In particular, audio classification methods, such as those discussed in Parts VI and VII, may be used in all audio processing methods.

[0154] <Part Two: Dialogue Enhancer Controller and Control Method> One example of an audio enhancement device is a dialogue enhancer (DE). This device, particularly for elderly individuals with declining hearing, intermittently monitors audio during playback, detects the presence of dialogue, and aims to enhance dialogue clarity and intelligibility (making it easier to hear and understand). In addition to detecting the presence of dialogue, if dialogue is present and therefore improved accordingly (using dynamic spectral rebalancing), the device also detects the frequencies most important for intelligibility. An exemplary dialogue enhancement method is presented in Patent Document 1, which is incorporated herein by reference.

[0155] A common manual configuration setting for dialogue enhancers is typically enabled for cinematic media content but disabled for musical content. This is because dialogue enhancers can sometimes trigger too frequently in response to musical signals.

[0156] Using available audio type information, the level of dialogue enhancement and other parameters can be adjusted based on the confidence value of the identified audio type. As individual examples of the audio processing devices and methods discussed above, the dialogue enhancer may use all the embodiments discussed in Part I and any combination thereof. In particular, when controlling the dialogue enhancer, the audio classifier 200 and adjustment unit 300 in the audio processing device 100 as shown in Figures 1 to 10 may constitute a dialogue enhancer controller 1500 as shown in Figure 15. In this embodiment, the adjustment unit is specific to the dialogue enhancer and may therefore be referred to as 300A. As discussed in the preceding section, the audio classifier 200 may include at least one of the audio content classifier 202 and the audio context classifier 204, and the dialogue enhancer controller 1500 may further include at least one of the type smoothing unit 712, the parameter smoothing unit 814, and the timer 916.

[0157] Therefore, this section will not repeat what has already been described in the previous section, but will simply provide some examples specific to this section.

[0158] Regarding the dialogue enhancer, the adjustable parameters include, but are not limited to, the level of dialogue enhancement, the background level, and thresholds for determining the frequency band to be enhanced. See Patent Document 1. The whole is incorporated herein by reference.

[0159] <Section 2.1: Levels of Dialogue Improvement> When it comes to the level of dialogue enhancement, the adjustment unit 300A may be configured to positively correlate the dialogue enhancement level of the dialogue enhancer with the confidence value of the utterance. Additionally or alternatively, the level may be negatively correlated with the confidence values ​​of other content types. Thus, the level of dialogue enhancement can be set to be proportional (linearly or nonlinearly) to the confidence level of the utterance. Therefore, dialogue enhancement is not very effective for non-speech signals such as music and background sounds (sound effects).

[0160] For context-based configurations, the adjustment unit 300A may be configured to positively correlate the dialogue enhancement level of the dialogue enhancer with the confidence value of cinematic media and / or VoIP, and negatively correlate the dialogue enhancement level of the dialogue enhancer with the confidence value of long-term music and / or games. For example, the dialogue enhancement level may be set to be proportional (linearly or non-linearly) to the confidence value of cinematic media. When the cinematic media confidence value is 0 (for example, in music content), the dialogue enhancement level is also 0, which is equivalent to disabling the dialogue enhancer.

[0161] As mentioned in the previous section, content type and context type may be considered together.

[0162] <Section 2.2 Thresholds for Determining the Frequency Band to be Improved> During the operation of the dialogue enhancer, there is a threshold (typically an energy or loudness threshold) for each frequency band to determine whether it needs to be enhanced. That is, frequency bands above each energy / loudness threshold are enhanced. To adjust these thresholds, the adjustment unit 300A may be configured to positively correlate the thresholds with the confidence level of short-term music and / or noise and / or background sounds and / or negatively correlate the thresholds with the confidence level of speech. For example, if the confidence level of speech is high, the thresholds can be lowered assuming more reliable speech detection, allowing for the enhancement of more frequency bands. On the other hand, if the confidence level of music is high, the thresholds can be raised so that fewer frequency bands are enhanced (thus reducing artifacts).

[0163] <Section 2.3 Adjusting to Background Level> Another component in the dialogue enhancer is the minimum tracking unit 4022, as shown in Figure 15. This is used to estimate the background level in the audio signal (for SNR estimation and frequency band threshold estimation as described in Section 2.2). This can also be adjusted based on the confidence level of the audio content type. For example, if the speech confidence level is high, the minimum tracking unit can be more confident in setting the background level to the current minimum. If the music confidence level is high, the background level can be set slightly higher than its current minimum, or in other words, it can be set to a weighted average of the current minimum and the energy of the current frame, with the current minimum being given a larger weight. If the noise and background confidence level is high, the background level can be set much higher than the current minimum, or in other words, it can be set to a weighted average of the current minimum and the energy of the current frame, with the current minimum being given a smaller weight.

[0164] Thus, the adjustment unit 300A may be configured to assign adjustments to background levels estimated by the minimum tracking unit. Here, the adjustment unit is further configured to positively correlate the adjustments with confidence values ​​of short-term music and / or noise and / or background sounds and / or negatively correlate the adjustments with confidence values ​​of speech. In one variation, the adjustment unit 300A may be configured to more positively correlate the adjustments with confidence values ​​of noise and / or background sounds than with short-term music.

[0165] <Section 2.4 Combinations of Embodiments and Application Scenarios> As with Part 1, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components mentioned in different parts / embodiments but having the same or similar function may be implemented as the same or separate components.

[0166] For example, any two or more of the solutions described in Sections 2.1 through 2.3 may be combined with each other. These combinations may then be further combined with any embodiments described or implied in Part I and other parts described later. In particular, many formulas are actually applicable to each type of audio improvement device or method, but they are not necessarily described or discussed in each part of this disclosure. In such cases, cross-referencing may be made between parts of this disclosure to apply a particular formula discussed in one part to another part, with only the relevant parameters, coefficients, powers (exponents) and weights being appropriately adjusted according to the specific requirements of a particular application.

[0167] <Section 2.5 Control Method for Dialogue Enhancer> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the dialogue enhancer controller in the above embodiments. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0168] First, the embodiments of the audio processing method discussed in Part 1 may be used for the dialogue enhancer. The parameters (one or more) of the dialogue enhancer are one of the targets to be adjusted by the audio processing method. From this perspective, the audio processing method is also a dialogue enhancer control method.

[0169] This section discusses only aspects specific to the control of the dialogue enhancer. For general aspects of the control method, please refer to Part 1.

[0170] According to one embodiment, the audio processing method may further include dialogue enhancement processing, the operation of adjustment 1104 including positively correlating the level of dialogue enhancement with the confidence level of cinematic media and / or VoIP and / or negatively correlating the level of dialogue enhancement with the confidence level of long-term music and / or game. That is, the dialogue enhancement is primarily directed towards audio signals in the context of cinematic media or VoIP.

[0171] More specifically, the operation of adjustment 1104 may include positively correlating the level of dialogue enhancement of the dialogue enhancer with the confidence value of the utterance.

[0172] The present invention may adjust the frequency bands to be improved in the dialogue improvement process. As shown in Figure 16, a threshold (typically energy or loudness) for determining whether each frequency band should be improved may be adjusted according to the present invention based on one or more identified audio type confidence values ​​(operation 1602). Then, within the dialogue improver, based on the adjusted thresholds, frequency bands above each threshold are selected (operation 1604) and improved (operation 1606).

[0173] Specifically, the operation of adjustment 1104 may include positively correlating those thresholds with confidence values ​​for short-term music and / or noise and / or background sounds and / or negatively correlating those thresholds with confidence values ​​for speech.

[0174] Audio processing methods (particularly dialogue enhancement processing) generally further include estimating background levels in the audio signal. This is generally implemented by a minimum tracking unit 4022 realized in the dialogue enhancer 402 and used in SNR estimation or frequency band threshold estimation. The present invention may be used to adjust background levels. In such a situation, after the background levels have been estimated (operation 1702), the background levels are first adjusted based on confidence values ​​(single or multiple) of audio types (operation 1704), and then used in SNR estimation and / or frequency band threshold estimation (operation 1706). In particular, the operation of adjustment 1104 may be configured to assign adjustments to the estimated background levels, where the operation of adjustment 1104 may be configured to positively correlate the adjustments with confidence values ​​of short-term music and / or noise and / or background sounds and / or negatively correlate the adjustments with confidence values ​​of speech.

[0175] More specifically, the operation of adjustment 1104 may be configured to correlate the adjustment more positively with the confidence value of noise and / or background than with short-term music.

[0176] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is feasible. On the other hand, every aspect of embodiments and variations of audio processing methods may be a separate solution. Furthermore, any two or more of the solutions described in this section may be combined with each other, and these combinations may be further combined with any embodiments described or implied in Part I or other parts described later.

[0177] <Part Three: Surround Virtualizer Controller and Control Method> A surround sound virtualizer allows surround sound signals (such as multi-channel 5.1 and 7.1) to be rendered through a PC's internal speakers or headphones. In other words, it virtually generates surround effects using stereo devices such as built-in laptop speakers or headphones, providing consumers with a cinematic experience. Surround sound virtualizers typically utilize head-related transfer functions (HRTFs) to simulate the arrival of sounds to the ear from various speaker positions associated with a multi-channel audio signal.

[0178] Current surround sound virtualizers work well on headphones, but they function differently with built-in speakers and different types of content. Generally, cinematic media content makes surround sound virtualizers work well with speakers, while music does not, because it can sound too thin.

[0179] Since the same parameters in a surround virtualizer cannot simultaneously produce a good sound image for both cinematic media and musical content, the parameters need to be more precisely adjusted based on the content. The functionality can be performed in conjunction with this invention using available audio type information, particularly music confidence values ​​and speech confidence values, as well as any other content type information and contextual information.

[0180] As in Part II, the surround virtualizer 404 may use any of the embodiments discussed in Part I and any combination thereof as individual examples of the audio processing devices and methods discussed in Part I. In particular, when controlling the surround virtualizer, the audio classifier 200 and adjustment unit 300 in the audio processing device 100 as shown in Figures 1 to 10 may constitute a surround virtualizer controller 1800 as shown in Figure 18. In this embodiment, the adjustment unit may be referred to as 300B as it is specific to the surround virtualizer. As in Part II, the audio classifier 200 may include at least one of the audio content classifier 202 and the audio context classifier 204, and the surround virtualizer controller 1800 may further include at least one of the type smoothing unit 712, the parameter smoothing unit 814 and the timer 916.

[0181] Therefore, this section will not repeat what has already been described in Part One, but will simply provide some examples specific to this section.

[0182] For the surround virtualizer, the adjustable parameters include, but are not limited to, the surround boost amount and the start frequency of the surround virtualizer 404.

[0183] <Section 3.1 Surround Boost Amount> When relating to the surround boost amount, the adjustment unit 300B may be configured to positively correlate the surround boost amount of the surround virtualizer 404 with the confidence value of noise and / or background and / or speech, and / or negatively correlate the surround boost amount with the confidence value of short-term music.

[0184] In particular, to modify the surround virtualizer 404 so that music (content type) sounds acceptable, an exemplary implementation of the adjustment unit 300B can adjust the amount of surround boost based on short-term music confidence. For example, SB ∝ (1 - Conf music ) (5) Here, SB is the surround boost amount, and Conf music is the reliability value of short-term music.

[0185] It reduces the surround boost for the music, preventing the music from sounding dull.

[0186] Similarly, the speech reliability value can also be used. For example, SB ∝ (1 - Conf music ) * Conf speech α (6) Here, Conf speech is the reliability value of the speech, α is a weighting coefficient in exponential form, and it may be in the range of 1 to 2. This formula indicates that the surround boost amount will only be high for pure speech (high speech reliability and low music reliability).

[0187] Alternatively, only the speech reliability value can be considered SB ∝ Conf speech (7) Various modifications can be designed in the same way. In particular, for noise or background sound, formulas similar to formulas (5) to (7) may be constructed. Furthermore, the effects of those four content types may be considered together in any combination. In such a situation, noise and background are ambient sounds, and it is safer to have a large boost amount. Speech can have a medium boost amount, assuming that the speaker usually sits in front of the screen. Therefore, the adjustment unit 300B may be configured to make the surround boost amount more positively correlated with the reliability values of noise and / or background than with the content type speech.

[0188] Assuming that the expected boost amount (which is weight equalization) for each content type is defined in advance, another alternative can be applied.

[0189]

Equation

[0190] From another perspective, α with a content type subscript in formula (8) is the expected / predefined boost amount of that content type, and the quotient obtained by dividing the confidence value of the corresponding content type by the sum of the confidence values of all identified content types may be regarded as the normalized weight of the predefined / expected boost amount of the corresponding content type. That is, the adjustment unit 300B may be configured to consider at least some of the plurality of content types by weighting the predefined boost amounts of the plurality of content types based on their confidence values.

[0191] For the context type, the adjustment unit 300B may be configured to positively correlate the surround boost amount of the surround virtualizer 404 with the confidence value of movie media and / or games, and negatively correlate the surround boost amount with the confidence value of long-term music and / or VoIP. Then, formulas similar to (5) to (8) can be constructed.

[0192] As a special case, the surround virtualizer 404 may be enabled for purely cinematic media and / or games, and disabled for music and / or VoIP. On the other hand, the boost amount of the surround virtualizer 404 may be set differently for cinematic media and games. Cinematic media may use a higher boost amount, and games a lower one. Therefore, the adjustment unit 300B may be configured to correlate the surround boost amount more positively with the confidence value of cinematic media than with games.

[0193] Similar to the content type, the amount of boost for the audio signal can also be set to a weighted average of the confidence values ​​of the context type.

[0194]

number

[0195] Similar to content types, α with a context type subscript in formula (9) is the expected / predefined boost amount for that context type, and the quotient obtained by dividing the confidence value of the corresponding context type by the sum of the confidence values ​​of all identified context types may be considered as the normalized weight of the predefined / expected boost amount for the corresponding context type. That is, the adjustment unit 300B may be configured to consider at least some of the multiple context types by weighting the predefined boost amounts of multiple context types based on those confidence values.

[0196] <Section 3.2 Starting Frequency> Other parameters, such as the starting frequency, can also be modified in the surround virtualizer. Generally, higher frequency components in an audio signal are more suitable for spatial rendering. For example, in music, it sounds strange if the bass is rendered to have more surround effect. Therefore, for a given audio signal, the surround virtualizer needs to determine a frequency threshold above which components are spatially rendered and components below are preserved. The frequency threshold is the starting frequency.

[0197] According to one embodiment of the present invention, the starting frequency for the surround virtualizer can be increased for music content, thereby allowing more bass to be retained in the music signal. The adjustment unit 300B can then be configured to positively correlate the starting frequency of the surround virtualizer with the confidence value of the short-term music.

[0198] <Section 3.3 Combinations of Embodiments and Application Scenarios> As with Part 1, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components mentioned in different parts / embodiments but having the same or similar function may be implemented as the same or separate components.

[0199] For example, any two or more of the solutions described in Sections 3.1 and 3.2 may be combined with each other. Any of these combinations may be further combined with any embodiments described or implied in Part I, Part II and other parts described later.

[0200] <Section 3.4 Surround Virtualizer Control Method> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the surround virtualizer controller in the embodiments described above. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0201] First, the embodiments of the audio processing method discussed in Part 1 may be used for surround virtualizers. The parameters (one or more) of the surround virtualizer are one of the targets to be adjusted by the audio processing method. From this perspective, the audio processing method is also a surround virtualizer control method.

[0202] This section discusses only aspects specific to the control of surround virtualizers. For general aspects of control methods, please refer to Part 1.

[0203] According to one embodiment, the audio processing method may further include surround virtualization processing, and the adjusting operation 1104 may be configured to positively correlate the surround boost amount of the surround virtualization processing with the confidence value of noise and / or background and / or speech and / or negatively correlate the surround boost amount with the confidence value of short-term music.

[0204] In particular, the adjustment operation 1104 may be configured to correlate the surround boost amount more positively with the confidence values ​​of noise and / or background and / or speech than with content-type speech.

[0205] Alternatively or additionally, the surround boost amount may be adjusted based on a context-type(single or multiple) confidence value(s). In particular, the adjustment operation 1104 may be configured to positively correlate the surround boost amount for surround virtualization processing with a confidence value for cinematic media and / or games and / or negatively correlate the surround boost amount with a confidence value for long-term music and / or VoIP.

[0206] In particular, the adjustment operation 1104 may be configured to more positively correlate the surround boost amount with movie media than with games.

[0207] Another parameter to be adjusted is the starting frequency for the surround virtualization process. As shown in FIG. 19, the starting frequency is first adjusted (operation 1902) based on one or more reliability values of the audio type(s), and then the surround virtualizer processes the audio components above the starting frequency (operation 1904). In particular, the adjustment operation 1104 may be configured to positively correlate the starting frequency of the surround virtualization process with the reliability value of short-term music.

[0208] Similar to the embodiments of the audio processing apparatus, embodiments of the audio processing method and any combination of its variations are realistic. On the other hand, all aspects of the embodiments of the audio processing method and its variations may be separate solutions. Furthermore, any two or more of the solutions described in this section may be combined with each other, and these combinations may also be combined with any embodiments described or implied in other parts of the present disclosure.

[0209] <Part 4: Volume Leveling Controller and Control Method> The volume of different audio sources or different pieces of the same audio source sometimes varies greatly. This is cumbersome because the user has to frequently adjust the volume. The Volume Leveler (VL) aims to adjust the volume of the audio content during playback so that it is mostly consistent over the time axis based on a target loudness value. Exemplary volume levelers are described in Patent Document 2, Patent Document 3, and Patent Document 4. These three documents are hereby incorporated by reference in their entirety.

[0210] A volume leveler continuously measures the loudness of an audio signal in some way and then modifies the signal by an amount of gain. Gain is a scaling factor for modifying the loudness of an audio signal and is typically a function of the measured loudness, the desired target loudness, and several other factors. Several factors need to be considered to estimate the appropriate gain, along with the underlying criteria for approaching the target loudness while maintaining dynamic range. This typically includes several sub-elements such as automatic gain control (AGC), auditory event detection, and dynamic range control (DRC).

[0211] Control signals are commonly applied in volume levelers to control the "gain" of an audio signal. For example, a control signal can be an indicator of changes in the magnitude of an audio signal, derived by pure signal analysis. A control signal can also be an audio event indicator, indicating whether a new audio event is appearing, through psychoacoustic analysis such as auditory scene analysis or specific-loudness-based auditory event detection. Such control signals are applied in volume levelers for gain control, for example, by ensuring that the gain remains nearly constant within an auditory event to reduce possible audible artifacts caused by abrupt changes in gain in the audio signal, and by constraining much of the gain change to the vicinity of the event boundary.

[0212] However, conventional methods for deriving control signals fail to distinguish between informative auditory events and non-informative (intrusive) auditory events. Here, informative auditory events represent audio events containing meaningful information, such as dialogue and music, which users may pay more attention to. Non-informative signals, on the other hand, do not contain meaningful information for the user, such as noise in VoIP. As a result, non-informative signals can also be boosted to near the target loudness by applying a large gain. This is undesirable in some applications. For example, in VoIP calls, noise signals appearing during pauses in conversation are often boosted to a large volume after being processed by a volume leveler. This is undesirable by the user.

[0213] To address this problem at least partially, the present application proposes controlling a volume leveler based on the embodiments discussed in Part I.

[0214] As with Parts Two and Three, as an individual example of the audio processing apparatus and method discussed in Part One, the volume leveler 406 may use all the embodiments discussed in Part One and any combination of those embodiments disclosed therein. In particular, when controlling the volume leveler 406, the audio classifier 200 and adjustment unit 300 in the audio processing apparatus 100 as shown in Figures 1 to 10 may constitute a volume leveler 406 controller 2000 as shown in Figure 20. In this embodiment, the adjustment unit is specific to the volume leveler 406 and may therefore be referred to as 300C.

[0215] That is, based on the disclosure in Part 1, the volume leveler controller 2000 may have an audio classifier 200 that continuously identifies the audio type (such as content type and / or context type) of an audio signal, and an adjustment unit 300C that continuously adjusts the volume leveler based on the confidence value of the identified audio type. Similarly, the audio classifier 200 may include at least one of an audio content classifier 202 and an audio context classifier 204, and the volume leveler controller 2000 may further include at least one of a type smoothing unit 712, a parameter smoothing unit 814, and a timer 916.

[0216] Therefore, this section will not repeat what has already been described in Part One, but will simply provide some examples specific to this section.

[0217] Based on the classification results, various parameters of the volume leveler 406 can be adaptively adjusted. For example, parameters directly related to the dynamic gain or the range of the dynamic gain can be adjusted by reducing the gain for non-informational signals. The dynamic gain can also be indirectly controlled by adjusting a parameter that indicates the degree to which a signal is a new perceptible audio event (the gain may change slowly within an audio event but abruptly at the boundary between two audio events). Several embodiments of parameter adjustment or volume leveler control mechanisms are presented here.

[0218] <Section 4.1 Informational and Interferential Content Types> As mentioned above, in relation to the control of the volume leveler, audio content types can be classified into informational content types and coherent content types. The adjustment unit 300C may be configured to positively correlate the dynamic gain of the volume leveler with the informational content type of the audio signal and negatively correlate the dynamic gain of the volume leveler with the coherent content type of the audio signal.

[0219] For example, suppose the noise is coercive (non-informational) and becomes annoying when boosted to a high volume. Parameters that directly control dynamic gain or indicate new audio events are: GainControl∝1-Conf noise (10) As shown, the noise confidence value (Conf noise It can be set to be proportional to the decreasing function of ).

[0220] Here, for simplicity, we use the symbol GainControl to represent all parameters (or their effects) related to gain control in a volume leveler. This is because different implementations of volume levelers may use different names for parameters with different fundamental meanings. Using a single term, GainControl, allows for a concise expression without loss of generality. Essentially, adjusting these parameters is equivalent to applying linear or nonlinear weights to the original gain. As an example, GainControl can be used directly to scale the gain such that a smaller GainControl results in a smaller gain. In another specific example, the gain is indirectly controlled by scaling an event control signal using GainControl, as described in Patent Document 3, which is incorporated hereby by reference in its entirety. In this case, when GainControl is small, the gain control of the volume leveler is modified to prevent the gain from changing significantly over time. When GainControl is large, the control is modified to allow the leveler's gain to change more freely.

[0221] Using the gain control described in formula (10) (directly scaling the original gain or event control signal), the dynamic gain of an audio signal is correlated (linearly or nonlinearly) to its noise confidence value. If the signal is noisy with high confidence, the sampling gain is correlated to the factor (1-Conf noiseTherefore, the volume decreases. In this way, we avoid boosting the noise signal to an unpleasantly high volume.

[0222] As a variation of formula (10), if background noise is not a concern in applications such as VoIP, background noise can be treated similarly and applied with a small gain. The control function is the noise confidence value (Conf noise ) and background confidence (Conf bkg Both of these can be taken into consideration. For example, GainControl∝(1-Conf noise )·(1-Conf bkg ) (11) In the above formula, since both noise and background sound are undesirable, GainControl is equally influenced by the confidence values ​​of the noise and the background. This can be considered to mean that noise and background sound have the same weight. Depending on the situation, they may have different weights. For example, different coefficients or different exponents (α and γ) may be given to the confidence values ​​of noise and background sound (or their difference from 1). That is, formula (11) GainControl∝(1-Conf noise ) α ·(1-Conf bkg ) γ (12) or GainControl∝(1-Conf noise α )·(1-Conf bkg γ ) (13) It can also be rewritten as follows.

[0223] Alternatively, the adjustment unit 300C may be configured to consider at least one dominant content type based on a confidence value. For example, GainControl∝1-max(Conf noise Conf bkg ) (14) Both formula (11) (and its variations) and formula (14) exhibit small gains for noise and background signals, and the original behavior of the volume leveler is maintained only when both the noise and background confidence values ​​are small (as in speech and music signals) and GainControl is close to 1.

[0224] The above example considers the dominant interfering content type. Depending on the situation, the adjustment unit 300C may be configured to consider the dominant informative content type based on confidence. More generally, the adjustment unit 300C may be configured to consider at least one dominant content type based on confidence, regardless of whether the identified audio type is / includes an informative and / or interfering audio type.

[0225] As another exemplary variation of formula (10), if the speech signal is the most informative content and requires less modification to the default behavior of the volume leveler, the control function can be expressed as a noise confidence value (Conf noise ) and speech confidence value (Conf speech ) both GainControl∝1-Conf noise ·(1-Conf speech ) (15) This can be considered as follows. Using this number, a small GainControl is obtained only for signals with high noise reliability and low speech reliability (e.g., pure noise), and when speech reliability is high, GainControl approaches 1 (thus maintaining the original behavior of the volume leveler). More generally, a certain content type (Conf noise The weight of (etc.) is at least one other content type (Conf speech It can be considered that it can be corrected by (etc.). In formula (15) above, the confidence of the utterance can be considered to change the weight coefficient of the confidence of the noise (a different kind of weight compared to the weights in formulas (12) and (13)). In other words, in formula (10) Conf noiseThe coefficient of can be considered as 1, while in formula (15), several other audio types (such as but not limited to speech) affect the importance of the confidence value of the noise. Therefore, Conf noise It can be said that the weights are modified by the confidence value of the utterance. In the context of this disclosure, the term “weights” is interpreted to include this; that is, they indicate the importance of a value, but are not necessarily standardized. See Section 1.4.

[0226] From another perspective, similar to formulas (12) and (13), exponential weights can be applied to the confidence values ​​in the above functions to indicate the priority (or importance) of different audio signals. For example, formula (15) can be modified as follows:

[0227] GainControl∝1-Conf noise α ·(1-Conf speech ) γ (16) Here, α and γ are two weights. These can be set smaller if a larger response is expected to be required to correct the leveler parameters.

[0228] Formulas (10) to (16) can be freely combined to form a variety of control functions that may be suitable for different applications. Other audio content-type confidence values, such as music confidence values, can be easily incorporated into control functions in a similar manner.

[0229] If GainControl is used to adjust a parameter indicating the degree to which a signal is a new, perceptible audio event, and to indirectly control the dynamic gain (where the gain changes slowly within an audio event but can change abruptly at the boundary between two audio events), then another transfer function between the content type confidence value and the final dynamic gain may be considered.

[0230] <Section 4.2 Content Types in Various Contexts> The control functions in formulas (10) to (16) above take into account confidence values ​​for audio content types such as noise, background sounds, short bursts of music, and speech, but do not consider the audio context from which the sound comes, such as in cinematic media and VoIP. The same audio content type, for example, background sounds, may need to be handled differently in different audio contexts. Background sounds include a variety of sounds such as car engines, explosions, and applause. These may not be meaningful in VoIP but may be important in cinematic media. This indicates that the audio context of interest needs to be identified, and different control functions should be designed for different audio contexts.

[0231] Therefore, the adjustment unit 300C may be configured to consider the content type of the audio signal as either informative or intrusive based on the context type of the audio signal. For example, by considering noise confidence and background confidence values ​​and distinguishing between VoIP and non-VoIP contexts, the audio context-dependent control function could be as follows:

[0232] if the audio context is VoIP GainControl∝1-max(Conf noise Conf bkg ) else GainControl∝1-Conf noise (17) In other words, in a VoIP context, noise and background sounds are considered coercive content types, while in a non-VoIP context, background sounds are considered informative content types.

[0233] As another example, considering confidence values ​​for speech, noise, and background, an audio context-dependent control function that distinguishes between VoIP and non-VoIP contexts could be as follows:

[0234] if the audio context is VoIP GainControl∝1-max(Conf noise Conf bkg ) else GainControl∝1-Conf noise ·(1-Conf speech ) (18) Here, utterances are emphasized as informational content.

[0235] If music is also considered important information in a non-VoIP context, then the latter half of formula (18) GainControl∝1-Conf noise ·(1-max(Conf speech Conf music ) (19) It can be expanded in this way.

[0236] In fact, each of the control functions in (10) to (16), or variations thereof, can be applied in different / corresponding audio contexts. Thus, it is possible to generate numerous combinations that form audio context-dependent control functions.

[0237] In addition to the VoIP and non-VoIP contexts distinguished and utilized in Officials (17) and (18), other audio contexts such as cinematic media, long-running music and games, or low-quality and high-quality audio may be utilized in a similar manner.

[0238] <Section 4.3 Contextual Type> Contextual types can also be used directly to control volume levelers to prevent annoying noises from being boosted too much. For example, a VoIP confidence value can be used to manipulate the volume leveler to be less sensitive when its confidence value is high.

[0239] Specifically, VoIP confidence value Conf VOIPUsing this, the level of the volume leveler is (1-Conf VOIP It can be set to be proportional to ). In other words, the volume leveler is mostly deactivated for VoIP content (when the VoIP confidence value is high). This is consistent with the traditional manual setup (preset) that disables the volume leveler for VoIP contexts.

[0240] Alternatively, different dynamic gain ranges can be set for various contexts of the audio signal. Generally, the VL (volume leveler) quantity can be seen as another (non-linear) weight on the gain, further adjusting the amount of gain applied to the audio signal. In one embodiment, the setup could be as follows:

[0241] [Table 1] Furthermore, assume that the expected VL amount is predefined for each context type. For example, the VL amount may be set to 1 for cinematic media, 0 for VoIP, 0.6 for music, and 0.3 for games, but the present invention is not limited to this. According to this example, if the dynamic gain range for cinematic media is 100%, then the dynamic gain range for VoIP will be 60%, and so on. If the classification of the audio classifier 200 is based on hard judgment, the dynamic gain range may be set directly as in the example above. If the classification of the audio classifier 200 is based on soft judgment, the range may be adjusted based on the confidence value of the context type.

[0242] Similarly, the audio classifier 200 may identify multiple context types from the audio signal, and the adjustment unit 300C may be configured to adjust the range of dynamic gain by weighting the confidence values ​​of the multiple content types based on their importance.

[0243] In general, for context types as well, functions similar to (10) to (16) can be used here, replacing the content type with the context type, in order to adaptively set the appropriate VL amount. In fact, Table 1 reflects the importance of different context types.

[0244] From another perspective, confidence values ​​may be used to derive the normalized weights discussed in Section 1.4. If certain quantities are predefined for each context type in Table 1, a formula similar to formula (9) can also be applied. Note that a similar solution may be applied to multiple content types and any other audio types.

[0245] <Section 4.4 Combinations of Embodiments and Application Scenarios> As with Part I, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components that are mentioned in different parts / embodiments but have the same or similar function may be implemented as the same or separate components. For example, any two or more of the solutions described in Sections 4.1 to 4.3 may be combined with each other. Any of these combinations may be further combined with any embodiments described or implied in Parts I to III and other parts described later.

[0246] Figure 21 illustrates the effect of the volume leveler controller proposed in this application by comparing the original short-term segment (Figure 21(A)), the short-term segment processed by a conventional volume leveler without parameter modification (Figure 21(B)), and the short-term segment processed by the volume leveler presented in this application (Figure 21(C)). As can be seen, the conventional volume leveler shown in Figure 21(B) also boosts the volume of the noise (the latter half of the audio signal), which is annoying. In contrast, the new volume leveler shown in Figure 21(C) boosts the volume of the effective portion of the audio signal without apparent boosting the volume of the noise, providing a better experience for the listener.

[0247] <Section 4.5 Volume Leveler Control Method> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the volume leveler controller in the embodiments described above. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0248] First, the embodiments of the audio processing method discussed in Part 1 may be applied to a volume leveler. The parameters (one or more) of the volume leveler are one of the targets to be adjusted by the audio processing method. From this perspective, the audio processing method is also a volume leveler control method.

[0249] This section discusses only aspects specific to the control of volume levelers. For general aspects of control methods, please refer to Part 1.

[0250] According to the present invention, a volume leveler control method is provided, which involves identifying the content type of an audio signal in real time, positively correlating the dynamic gain of the volume leveler with the informational content type of the audio signal, and negatively correlating the dynamic gain of the volume leveler with the coherent content type of the audio signal, thereby continuously adjusting the volume leveler based on the identified content type.

[0251] Content types may include speech, short bursts of music, noise, and background sounds. Generally, noise is considered an intrusive content type.

[0252] When adjusting the dynamic gain of a volume leveler, it may be adjusted directly based on the content type confidence value, or it may be adjusted via the transfer function of the content type confidence value.

[0253] As already mentioned, an audio signal can be classified into multiple audio types simultaneously. When dealing with multiple content types, the adjustment operation 1104 may be configured to consider at least some of the multiple audio content types by weighting the confidence values ​​of the multiple content types based on their importance, or by weighting the effects of the multiple content types based on their confidence values. In particular, the adjustment operation 1104 may be configured to consider at least one dominant content type based on its confidence value. When the audio signal includes both one or more coherent content types and one or more informative content types, the adjustment operation may be configured to consider at least one dominant coherent content type and / or at least one dominant informative content type based on its confidence value.

[0254] Different audio types may influence each other. Therefore, the adjustment operation 1104 may be configured to modify the weight of one content type using the confidence value of at least one other content type.

[0255] As mentioned in Part 1, the audio type confidence values ​​of audio signals may be smoothed. For details on the smoothing process, please refer to Part 1.

[0256] The method may further include identifying the context type of the audio signal, where the adjustment operation 1104 may be configured to adjust the range of dynamic gain based on a confidence value of the context type.

[0257] The role of a content type is limited by the context type in which it is located. Therefore, when both content type information and context type information are available for an audio signal simultaneously (i.e., for the same audio segment), the content type of the audio signal may be determined to be informative or coherent based on the context type of the audio signal. Furthermore, content types in audio signals of different context types may be assigned different weights depending on the context type of the audio signal. From another perspective, different weights (greater or smaller, positive or negative values) can be used to reflect the informative or coherent nature of the content type.

[0258] The context types of an audio signal may include VoIP, cinematic media, long-running music, and games. In contextual VoIP audio signals, background sound may be considered an intrusive content type. On the other hand, in contextual non-VoIP audio signals, background and / or speech and / or music are considered an informative content type. Other context types may include high-quality audio or low-quality audio.

[0259] Similar to multiple content types, when an audio signal is classified into multiple context type information simultaneously (i.e., for the same audio segment), the adjustment operation 1104 may be configured to consider at least some of the multiple audio content types by weighting the confidence values ​​of the multiple context types based on their importance, or by weighting the effects of the multiple context types based on their confidence values. In particular, the adjustment operation may be configured to consider at least one dominant content type based on its confidence value.

[0260] Finally, the methods discussed in this section may be implemented using the audio classification methods discussed in Parts 6 and 7. A detailed description is omitted here.

[0261] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is practical. On the other hand, any aspect of embodiments and variations of audio processing methods may be separate solutions. Furthermore, any two or more solutions described in this section may be combined with each other, and these combinations may further be combined with any embodiments described or implied in other parts of this disclosure.

[0262] <Part 5: Equalizer Controllers and Control Methods> Equalization is typically applied to musical signals to adjust or modify their spectral balance, known as "tone" or "timbre." Traditional equalizers allow users to configure the overall profile (curve or shape) of the frequency response (gain) in individual frequency bands to emphasize certain instruments or remove unwanted sounds. Common music players, such as Windows Media Player, provide graphic equalizers for adjusting the gain in each frequency band to achieve the best listening experience for various genres of music, and also offer sets of equalizer presets for various music genres such as rock, rap, jazz, and folk. Once a preset is selected and a profile is set, the same equalization gain is applied to the signal until the profile is manually modified.

[0263] In contrast, a dynamic equalizer provides a means of automatically adjusting the equalization gain in each frequency band to maintain overall consistency of the spectral balance with respect to the desired timbre or tone. This consistency is achieved by continuously monitoring the spectral balance of the audio, comparing it to a desired preset spectral balance, and dynamically adjusting the applied equalization gain to convert the original spectral balance of the audio to the desired spectral balance. The desired spectral balance is selected manually or preset before processing.

[0264] Both types of equalizers share the following drawbacks: The best equalization profile, desired spectral balance, or related parameters must be selected manually and cannot be automatically corrected based on the audio content during playback. Determining the type of audio content is crucial for providing good overall quality for various audio signals. For example, different musical pieces, such as those of different genres, require different equalization profiles.

[0265] In an equalizer system that can accept any type of audio signal (not just music), the equalizer parameters need to be adjusted based on the content type. For example, the equalizer is typically enabled for music signals but disabled for speech signals because it can alter the timbre of the speech too much, making the signal sound unnatural.

[0266] To address this problem at least partially, the present invention proposes controlling the equalizer based on the embodiments discussed in Part I.

[0267] Similar to Parts II to IV, as an individual example of the audio processing apparatus and methods discussed in Part I, the equalizer 408 may use all the embodiments discussed in Part I and any combination of those embodiments disclosed therein. In particular, when controlling the equalizer 408, the audio classifier 200 and adjustment unit 300 in the audio processing apparatus 100 as shown in Figures 1 to 10 may constitute an equalizer 408 controller 2000 as shown in Figure 22. In this embodiment, the adjustment unit is specific to the equalizer 408 and may therefore be referred to as 300D.

[0268] That is, based on the disclosure in Part 1, the equalizer controller 2200 may have an audio classifier 200 that continuously identifies the audio type of an audio signal, and an adjustment unit 300D that continuously adjusts the equalizer based on the confidence value of the identified audio type. Similarly, the audio classifier 200 may include at least one of an audio content classifier 202 and an audio context classifier 204, and the volume equalizer controller 2200 may further include at least one of a type smoothing unit 712, a parameter smoothing unit 814 and a timer 916.

[0269] Therefore, this section will not repeat what has already been described in Part One, but will simply provide some examples specific to this section.

[0270] <Section 5.1 Control based on content type> In general, for common audio content types such as music, speech, background noise, and noise, the equalizer should be configured differently for different content types. Similar to traditional setups, the equalizer can be automatically enabled for music signals but disabled for speech. Alternatively, a more continuous approach can be taken, setting a higher equalization level for music signals and a lower equalization level for speech signals. In this way, the equalizer's equalization level can be automatically configured for different audio content.

[0271] In particular, with regard to music, it has been observed that the equalizer does not function very well for musical pieces with a dominant source. This is because, when inappropriate equalization is applied, the timbre of the dominant source can change significantly and sound unnatural. Considering this, it is better to set the playing level for musical pieces with a dominant source. On the other hand, for musical pieces without a dominant source, the equalization level can be kept high. Using this information, the equalizer can automatically set the equalization level for various musical content.

[0272] Music can also be grouped based on various attributes, including genre, instruments and rhythm, tempo, and timbre. Just as different equalizer presets are used for different musical genres, these musical groups / clusters may also have their own optimal equalization profiles or equalizer curves (in the case of traditional equalizers) or optimal desired spectral balances (in the case of dynamic equalizers).

[0273] As mentioned above, equalizers are generally enabled for music content but disabled for speech. This is because equalizers can sometimes make dialogue sound less audible due to timbre changes. One way to achieve this automatically is to associate the equalizer level with the content, particularly with the music confidence and / or speech confidence values ​​obtained from the Audio Content Classification Module. Here, the equalization level can be described as the weight of the equalizer gain applied. The higher the level, the stronger the equalization applied. For example, if the equalization level is 1, a full equalization profile is applied. If the equalization level is 0, all gains correspond to 0 dB, and therefore deequalization is applied. The equalization level may be represented by various parameters in various implementations of the equalizer algorithm. An example of these parameters is the equalizer weight, as implemented in Patent Document 2, which is incorporated hereby by reference in its entirety.

[0274] Various control schemes can be designed to adjust the equalization level. For example, in audio content information, speech confidence or music confidence can be used to set the equalization level. L eq ∝Conf music (20) or L eq ∝1-Conf speech (twenty one) It can be used as follows: Here, L eq This is the level of equalization, Conf music and Conf speechThis represents the confidence level of the music and speech.

[0275] In other words, the adjustment unit 300D may be configured to positively correlate the equalization level with the confidence value of short-term music or negatively correlate the equalization level with the confidence value of speech.

[0276] Speech confidence and musical confidence can be further integrated and used to set the equalization level. The general idea is that a high equalization level is only achieved when musical confidence is high and speech confidence is low; otherwise, the equalization level is low. For example, L eq =Conf music (1-Conf speech α ) (twenty two) Here, the speech confidence value is raised to the power of α to accommodate the frequently occurring non-zero speech confidence values ​​in musical signals. Using the formula above, the equalization is fully applicable (at a level equal to 1) to pure musical signals without speech components. As mentioned in Part 1, α can also be considered a weighting coefficient based on the importance of the content type and can typically be set to 1 or 2.

[0277] If greater weight is placed on the confidence level of the utterance, the adjustment unit 300D may be configured to disable the equalizer 408 when the confidence level for content-type utterances is greater than a certain threshold.

[0278] The above description uses music and speech content types as examples. Alternatively or additionally, confidence values ​​for background sound and / or noise may also be considered. In particular, the adjustment unit 300D may be configured to positively correlate the equalization level with the background confidence value and / or negatively correlate the equalization level with the noise confidence value.

[0279] As another example, confidence values ​​may be used to derive the normalized weights discussed in Section 1.4. If the expected level of equalization is predefined for each content type (for example, 1 for music, 0 for speech, and 0.5 for noise and background), then a formula similar to formula (8) can be strictly applied.

[0280] The equalization level may be further smoothed to avoid abrupt changes that could introduce audible artifacts at transition points. This can be done using the parameter smoothing unit 814 described in Section 1.5.

[0281] <Section 5.2 The certainty of the dominant source in music> To avoid applying a high level of equalization to music with a dominant source, the level of equalization is further defined by a confidence value (Conf) indicating whether a musical piece contains a dominant source. dom It may be correlated with, for example, L eq =1-Conf dom (twenty three).

[0282] In this way, the level of equalization is low for musical pieces with a dominant source and high for musical pieces without a dominant source.

[0283] Here, confidence values ​​for music with a dominant source are described, but confidence values ​​for music without a dominant source can also be used. That is, the adjustment unit 300D may be configured to positively correlate the equalization level with the confidence value of short-term music without a dominant source and / or negatively correlate the equalization level with the confidence value of short-term music with a dominant source.

[0284] As stated in Section 1.1, music and speech, and music with or without a dominant source, are content types at different hierarchical levels, but can be considered in parallel. By considering the confidence values ​​of the dominant source and the confidence values ​​of speech and music together, the level of equalization can be set by combining at least one of formulas (20) to (21) with (23). One example is combining all three of these formulas. L eq =Conf music (1-Conf speech )(1-Conf dom ) (twenty four) That is the idea.

[0285] For generality, different weights based on the importance of content types can be applied to yet different confidence values, as in formula (22).

[0286] As another example, Conf dom Assuming that this is calculated only when the audio signal is music, the step function can be designed as follows:

[0287]

number

[0288] The same smoothing method discussed in Section 1.5 can also be applied, and the time constant α can be further determined based on the transition type, such as the transition from music with a dominant source to music without a dominant source, or the transition from music without a dominant source to music with a dominant source. For this purpose, a formula similar to formula (4') can also be applied.

[0289] <Section 5.3 Equalizer Presets> In addition to adaptively adjusting the equalization level based on the confidence level of the audio content type, the system can also automatically select an appropriate equalization profile or desired spectral balance preset for various audio content, depending on its genre, instruments, or other characteristics. Music of the same genre, containing the same instruments, or having the same musical characteristics can share the same equalization profile or desired spectral balance preset.

[0290] For generality, the term “music cluster” is used to represent musical groups that share the same genre, the same instruments, or similar musical attributes. This can be considered another hierarchical level of audio content types as described in Section 1.1. Appropriate equalization profiles, equalization levels, and / or desired spectral balance presets may be associated with each music cluster. The equalization profile is a gain curve applied to the music signal and can be any of the equalizer presets used for different music genres (e.g., classical, rock, jazz, and folk), while the desired spectral balance preset represents the desired timbre for each cluster. Figure 23 shows some examples of desired spectral balance presets implemented in Dolby Home Theater technology. Each describes the desired spectral shape across the audible frequency range. This shape is continuously compared to the spectral shape of the incoming audio, and the equalization gain is calculated from the comparison of the spectral shape of the incoming audio with its transformation to the spectral shape of the preset.

[0291] For new musical pieces, the nearest cluster can be determined (hard determination) or a confidence value for each musical cluster can be calculated (soft determination). Based on this information, an appropriate equalization profile or a desired spectral balance preset can be determined for a given musical piece. The simplest method is, P eq =P c* (26) The goal is to assign the corresponding profile of the best-matched cluster. Here, P eq This is the estimated equalization profile or desired spectral balance preset, and c * This is an index of the best-matched music clusters (dominant audio types), obtained by picking the cluster with the highest confidence value.

[0292] Furthermore, there may be two or more music clusters with a confidence value greater than 0. That is, a piece of music will have attributes that are more or less similar to those clusters. For example, a piece of music may have multiple instruments or attributes of multiple genres. This leads to the idea of ​​another method for estimating a proper equalization profile by considering all clusters rather than just the nearest cluster. For example, a weighted sum can be used.

[0293]

number

[0294] In some applications, it may not be desirable to involve all the clusters shown in formula (27). Only a subset of the clusters—those most relevant to the current musical piece—need to be considered, and formula (27) can be slightly modified as follows:

[0295]

number

[0296] The above description uses music clusters as an example. In fact, these solutions are applicable to any audio type at any hierarchical level discussed in Section 1.1. Therefore, the adjustment unit 300D may generally be configured to assign equalization levels and / or equalization profiles and / or spectral balance presets to each audio type.

[0297] <Section 5.4 Context-Based Control> The previous sections have focused on various content types. In the further embodiments discussed in this section, context types may be considered alternatively or additionally.

[0298] Generally, the equalizer is effective for music but ineffective for cinematic media because it may make dialog in cinematic media less audible due to obvious timbre changes. This indicates that the equalization level can be related to the long-term confidence value of music and / or the confidence value of cinematic media: L eq ∝Conf MUSIC (29) Or L eq ∝1-Conf MOVIE (30) Here, L eq is the equalization level, Conf MUSIC and Conf MOVIE represent the long-term confidence values of music and cinematic media.

[0299] That is, the adjustment unit 300D may be configured to positively correlate the equalization level with the long-term confidence value of music or negatively correlate the equalization level with the confidence value of cinematic media.

[0300] That is, for cinematic media signals, the cinematic media confidence value is high (or the music confidence level is low), so the equalization level is low. On the other hand, for music signals, the cinematic media confidence value is low (or the music confidence level is high), so the equalization level is high.

[0301] The solutions shown in formulas (29) and (30) may be modified in the same way as formulas (22) to (25), and / or may be combined with any of the solutions shown in formulas (22) to (25).

[0302] [[ID=(37)]] Additionally or alternatively, the adjustment unit 300D may be configured to negatively correlate the equalization level with the confidence value of the game.

[0303] In another embodiment, confidence values ​​may be used to derive the normalized weights discussed in Section 1.4. If the expected equalization levels / profiles are predefined for each context type (the equalization profiles are shown in Table 2 below), then a formula similar to formula (9) can also be applied.

[0304] [Table 2] Here, in some profiles, all gains can be set to 0 as a way to disable the equalizer for certain context types, such as cinematic media and games.

[0305] <Section 5.5 Combinations of Embodiments and Application Scenarios> As with Part 1, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components mentioned in different parts / embodiments but having the same or similar function may be implemented as the same or separate components.

[0306] For example, any two or more of the solutions described in Sections 5.1 through 5.4 may be combined with each other. Any of these combinations may be further combined with any embodiments described or implied in Parts I through IV and other parts described later.

[0307] <Section 5.6 Equalizer Control Method> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the equalizer controller in the embodiments described above. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0308] First, the embodiments of the audio processing method discussed in Part 1 may be applied to the equalizer. The parameters (one or more) of the equalizer are one of the targets to be adjusted by the audio processing method. From this perspective, the audio processing method is also an equalizer control method.

[0309] This section discusses only aspects specific to the control of equalizers. For general aspects of control methods, please refer to Part 1.

[0310] According to various embodiments, the equalizer control method may include identifying the audio type of an audio signal in real time and adjusting the equalizer in a continuous manner based on the confidence value of the identified audio type.

[0311] As with other parts of the present invention, when multiple audio types with corresponding confidence values ​​are involved, the adjustment operation 1104 may be configured to consider at least some of the multiple audio types by weighting the confidence values ​​of the multiple audio types based on their importance, or by weighting the effects of the multiple audio types based on their confidence values. In particular, the adjustment operation 1104 may be configured to consider at least one dominant content type based on its confidence value.

[0312] As mentioned in Part 1, the adjusted parameter values ​​may be smoothed. See Sections 1.5 and 1.8; further details are omitted here.

[0313] The audio type can be content-type or context-type. When involving the content type, the adjustment operation 1104 may be configured to positively correlate the equalization level with the confidence value of short-term music and / or negatively correlate the equalization level with the confidence value of speech. Additionally or alternatively, the adjustment operation may be configured to positively correlate the equalization level with the confidence value of background and / or negatively correlate the equalization level with the confidence value of noise.

[0314] When dealing with contextual types, the adjustment operation 1104 may be configured to positively correlate the equalization level with the confidence level of long-term music and / or negatively correlate the equalization level with the confidence level of cinematic media and / or games.

[0315] For short-term music content types, adjustment operation 1104 may be configured to positively correlate the equalization level with the confidence value of short-term music without a dominant source and / or negatively correlate the equalization level with the confidence value of short-term music with a dominant source. This can only be done when the confidence value for short-term music is greater than a certain threshold.

[0316] In addition to adjusting the equalization level, other aspects of the equalizer may be adjusted based on the confidence value(s) of the audio type(s) of the audio signal. For example, adjustment operation 1104 may be configured to assign equalization levels and / or equalization profiles and / or spectral balance presets to each audio type.

[0317] For specific examples of audio types, please refer to Part 1.

[0318] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is practical. On the other hand, any aspect of embodiments and variations of audio processing methods may be separate solutions. Furthermore, any two or more solutions described in this section may be combined with each other, and these combinations may further be combined with any embodiments described or implied in other parts of this disclosure.

[0319] <Part Six: Audio Classifiers and Classification Methods> As described in Sections 1.1 and 1.2, the audio types discussed in this application, including various hierarchical levels of content and context types, can be classified or identified using some existing classification scheme, including machine learning-based methods. In this and subsequent parts, the application proposes some novel aspects of classifiers and methods for classifying the context types mentioned in previous parts.

[0320] <Section 6.1 Contextual Classifier Based on Content Type Classification> As described in previous sections, the audio classifier 200 is used to identify the content type of an audio signal and / or the context type of an audio signal. Therefore, the audio classifier 200 may have an audio content classifier 202 and / or an audio context classifier 204. When employing existing techniques for implementing the audio content classifier 202 and / or the audio context classifier 204, the two classifiers may be independent of each other but may share some features and therefore some methods for extracting those features.

[0321] In this section and the following seventh section, the audio context classifier 204 may utilize the results of the audio content classifier 202, in accordance with the novel aspects proposed herein. That is, the audio classifier 200 includes an audio content classifier 202 that identifies the content type of an audio signal; and an audio context classifier 204 that identifies the context type of an audio signal based on the results of the audio content classifier 202. Thus, the classification results of the audio content classifier 202 may be used by both the audio context classifier 204 and the adjustment unit 300 (or adjustment units 300A to 300D) discussed in the preceding sections. However, although not shown in the drawings, the audio classifier 200 may include two audio content classifiers 202, which are used by the adjustment unit 300 and the audio context classifier 204, respectively.

[0322] Furthermore, as discussed in Section 1.2, particularly when classifying multiple audio types, the audio content classifier 202 or the audio context classifier 204 may form a group of classifiers that cooperate with each other. However, it is also possible to implement them as a single classifier.

[0323] As discussed in Section 1.1, content types are generally audio types relating to short-term audio segments with a length of several to tens of frames (e.g., 1 second), while context types are generally audio types relating to long-term audio segments with a length of several to tens of seconds (e.g., 10 seconds). Therefore, "short-term" and "long-term" are used as needed, corresponding to "content type" and "context type," respectively. However, as will be discussed in Part VII, while context types indicate attributes of audio signals on a relatively long time scale, they can also be identified based on features extracted from short-term audio segments.

[0324] Now, referring to Figure 24, let us look at the structures of the audio content classifier 202 and the audio context classifier 204.

[0325] As shown in Figure 24, the audio content classifier 202 may have a short-term feature extractor 2022 that extracts short-term features from short-term audio segments, each containing a sequence of audio frames; and a short-term classifier 2024 that classifies sequences of short-term segments within long-term audio segments into short-term audio types using their respective short-term features. Both the short-term feature extractor 2022 and the short-term classifier 2024 may be implemented using existing techniques, but some modifications to the short-term feature extractor 2022 are also proposed in Section 6.3 below.

[0326] The short-term classifier 2024 may be configured to classify each segment of a sequence of short-term segments into at least one of the following short-term audio types (content types): speech, short-term music, background noise, and noise, as described in Section 1.1. Each content type may be further classified into lower-level content types, as, but not limited to, those discussed in Section 1.1.

[0327] As is known in the art, the confidence value of the classified audio type may be obtained by a short-term classifier 2024. In this application, when referring to the operation of any classifier, it is understood that confidence values ​​may be obtained simultaneously, if necessary, whether or not they are explicitly recorded. An example of audio type classification can be found in Non-Patent Document 1, which is incorporated herein by reference in its entirety.

[0328] On the other hand, as shown in Figure 24, the audio context classifier 204 may have a statistics extractor 2024 for calculating the statistics of the short-term classifier result with respect to the sequence of short-term segments within a long-term audio segment as long-term features; and a long-term classifier 2044 for classifying the long-term audio segment into a long-term audio type using the long-term features. Similarly, both the statistics extractor 2042 and the long-term classifier 2044 may be implemented using existing techniques, although some modifications to the statistics extractor 2042 are proposed in the following section 6.2.

[0329] The long-term classifier 2044 may be configured to classify long-term audio segments into at least one of the following long-term audio types (context types): cinematic media, long-term music, games, and VoIP, as described in Section 1.1. Alternatively or additionally, the long-term classifier 2044 may be configured to classify long-term audio segments into VoIP or non-VoIP as described in Section 1.1. Alternatively or additionally, the long-term classifier 2044 may be configured to classify long-term audio segments into high-quality audio or low-quality audio as described in Section 1.1. In practice, various target audio types can be selected and trained based on application / system requirements.

[0330] For the meaning and selection of short-term and long-term segments (and the framework discussed in Section 6.3), please refer to Section 1.1.

[0331] <Section 6.2 Extraction of Long-Term Characteristics> As shown in Figure 24, in one embodiment, only the statistic extractor 2042 is used to extract long-term features from the results of the short-term classifier 2024. As long-term features, at least one of the following may be calculated by the statistic extractor 2042: the mean and variance of the confidence values ​​of short-term audio types in the short-term segments within the long-term segment to be classified; the mean and variance weighted by the importance of the short-term segments; the frequency of occurrence of each short-term audio type; and the frequency of transitions between various short-term audio types within the long-term segment to be classified.

[0332] Figure 25 shows the average confidence values ​​for speech and short-term music in each short-term segment (length 1 s). For comparison, segments are drawn from three different audio contexts: cinematic media (Figure 25(A)), long-term music (Figure 25(B)), and VoIP (Figure 25(C)). For the cinematic media context, high confidence values ​​are obtained for either speech or music, and frequent alternation between these two audio types can be observed. In contrast, the long-term music segments yield consistently high confidence values ​​for short-term music and relatively consistently low confidence values ​​for speech. On the other hand, the VoIP segments yield consistently low confidence values ​​for short-term music, but yield fluctuating confidence values ​​for speech due to pauses during VoIP conversations.

[0333] The variance of confidence values ​​for each audio type is also an important feature for classifying various audio contexts. Figure 26 gives a histogram of the variance of confidence values ​​for speech, short-term music, background, and noise in cinematic media, long-term music, and VoIP audio contexts (the horizontal axis is the variance of confidence values ​​in the dataset, and the vertical axis is the number of occurrences in each bin of the variance value s in the dataset, which can be normalized to show the normal probability of each bin of the variance value). For cinematic media, the variance of confidence values ​​for speech, short-term music, and background are all relatively high and widely distributed. This indicates that the confidence values ​​for these audio types change significantly. For long-term music, the variance of confidence values ​​for speech, short-term music, background, and noise are all relatively low and narrowly distributed. This indicates that the confidence values ​​for these audio types remain stable. The confidence value for speech remains consistently low, and the confidence value for music remains consistently high. For VoIP, the variance of confidence value for short-term music is low and narrowly distributed, while the variance of confidence value for speech is relatively widely distributed. This is due to the frequent pauses during VoIP conversations.

[0334] The weights used when calculating the weighted mean and variance are determined based on the importance of each short-term segment. The importance of a short-term segment may be measured by its energy or loudness. Energy and loudness can be estimated using many existing techniques.

[0335] The frequency of occurrence of each short-term audio type in the long-term segment to be classified is calculated by normalizing the count of each audio type classified within that long-term segment by the length of the long-term segment.

[0336] The frequency of transitions between various short-term audio types within a long-term segment to be classified is the count of audio type changes between adjacent short-term segments within the long-term segment to be classified, normalized by the length of the long-term segment.

[0337] When discussing the mean and variance of confidence values ​​with reference to Figure 25, the frequency of occurrence of each short-term audio type and the frequency of transitions between these various short-term audio types are also relevant. These characteristics are also important for audio context classification. For example, long-term music mostly contains short-term music audio types and therefore has a high frequency of occurrence of short-term music. On the other hand, VoIP mostly contains speech and pauses and therefore has a high frequency of occurrence of speech or noise. As another example, cinematic media transitions between different short-term audio types more frequently than long-term music or VoIP and therefore generally has a higher frequency of transitions between short-term music, speech and background. VoIP usually transitions between speech and noise more frequently than others and therefore has a higher frequency of transitions between speech and noise.

[0338] Generally, it is assumed that long-term segments are of the same length within the same application / system. If so, the occurrence count for each short-term audio type and the transition count between different short-term audio types within a long-term segment may be used directly without normalization. If the length of the long-term segment is variable, the occurrence and transition frequencies described above should be used. The claims of this application should be construed to cover both situations.

[0339] Additionally or alternatively, the audio classifier 200 (or audio context classifier 204) may further include a long-term feature extractor 2046 (Figure 27) for extracting further long-term features from the long-term audio segment based on the short-term features of the sequence of short-term segments within the long-term audio segment. In other words, the long-term feature extractor 2046 does not use the classification results of the short-term classifier 2024, but directly uses the short-term features extracted by the short-term feature extractor 2022 to derive several long-term features to be used by the long-term classifier 2044. The long-term feature extractor 2046 and the statistics extractor 2042 may be used independently or jointly. In other words, the audio classifier 200 may include either or both of the long-term feature extractor 2046 or the statistics extractor 2042.

[0340] Any feature can be extracted by the long-term feature extractor 2046. In this application, it is proposed to calculate at least one of the following statistics of the short-term features from the short-term feature extractor 2022 as the long-term features: mean, variance, weighted mean, weighted variance, high mean, low mean, and the ratio between the high mean and the low mean (contrast).

[0341] The mean and variance of short-term features extracted from short-term segments within the long-term segments to be classified.

[0342] The weighted mean and variance of short-term features extracted from short-term segments within the long-term segments to be classified. Each short-term feature is weighted based on its importance, measured using the energy or loudness described above for each short-term segment.

[0343] High average: The average of selected short-term features extracted from short-term segments within the long-term segment to be classified. Short-term features are selected when they meet at least one of the following conditions: greater than a certain threshold; or within a predetermined percentage of short-term features, e.g., within the top 10% of short-term features, and not lower than all other short-term features.

[0344] Low average: The average of selected short-term features extracted from the short-term segments within the long-term segments to be classified. Short-term features are selected if they meet at least one of the following conditions: less than a certain threshold; or not higher than all other short-term features; and within a predetermined percentage of short-term features, e.g., within the bottom 10% of short-term features.

[0345] Contrast: The ratio between the high average and the low average. It represents the dynamics of short-term features within a long-term segment.

[0346] Short-Term Feature Extractor 2022 may be implemented using existing techniques, and any features can be extracted by such implementation. Nevertheless, several modifications for Short-Term Feature Extractor 2022 are proposed in Section 6.3.

[0347] <Section 6.3 Extraction of Short-Term Characteristics> As shown in Figures 24 and 27, the short-term feature extractor 2022 may be configured to directly extract at least one of the following features as short-term features from each short-term audio segment: rhythmic characteristics, interruption / mute characteristics, and short-term audio quality features.

[0348] Rhythmic characteristics may include rhythmic intensity, rhythmic regularity, rhythmic clarity (see Non-Patent Document 2, incorporated hereby by reference), and 2D subband modulation (see Non-Patent Document 3, incorporated hereby by reference).

[0349] Interruption / mute characteristics may include speech interruption, sharp decay, mute length, unnatural silence, average unnatural silence, and total energy of unnatural silence.

[0350] Short-term audio quality features are audio quality features relating to short-term segments, and are similar to the audio quality features extracted from audio frames, which will be discussed below.

[0351] Alternatively or additionally, as shown in Figure 28, the audio classifier 200 may have a frame-level feature extractor 2012 that extracts frame-level features from each frame of a sequence of audio frames included in the short-term segment. The short-term feature extractor 2022 may be configured to compute short-term features based on the frame-level features extracted from the sequence of audio frames.

[0352] As a preprocessing step, the input audio signal may be downmixed to a mono audio signal. This preprocessing step is unnecessary if the audio signal is already a mono signal. The signal is then divided into frames of a predetermined length (typically 10 to 25 milliseconds). Correspondingly, frame-level features are extracted from each frame.

[0353] The frame-level feature extractor 2012 may be configured to extract at least one of the following features: features characterizing the attributes of various short-term audio types, cutoff frequencies, static signal-to-noise ratio (SNR) characteristics, segment signal-to-noise ratio (SNR) characteristics, basic speech descriptors, and vocal tract characteristics.

[0354] Features characterizing the attributes of various short-term audio types (in particular speech, short-term music, background noise, and noise) may include at least one of the following features: frame energy, subband spectral distribution, spectral flux, Mel-frequency cepstral coefficient (MFCC), bass, residual information, chroma features, and zero-crossing rate.

[0355] For details on MFCC, see Non-Patent Document 1, which is incorporated in its entirety by reference. For details on chromatic features, see Non-Patent Document 4, which is incorporated in its entirety by reference.

[0356] The cutoff frequency represents the highest frequency of an audio signal above which the content's energy is close to zero. It is designed to detect content with limited bandwidth, and in this application, it is useful for audio context classification. The cutoff frequency is typically caused by encoding, as most encoders discard high frequencies at low or medium bitrates. For example, the MP3 codec has a cutoff frequency of 16kHz at 128kbps. Another example is many common VoIP codecs, which have cutoff frequencies of 8kHz or 16kHz.

[0357] In addition to the cutoff frequency, signal degradation during the audio encoding process is considered as a further characteristic for distinguishing between various audio contexts, such as VoIP versus non-VoIP contexts, and high-quality versus low-quality audio contexts. To capture richer characteristics, features representing audio quality, such as those for objective speech quality evaluation (see Non-Patent Document 5, incorporated in its entirety by reference here), may be further extracted at multiple levels. Examples of audio quality features include: a) Static SNR characteristics including estimated background noise level and spectral clarity. b) Segment SNR characteristics including spectral level deviation, spectral level range, relative noise floor, etc. c) Basic speech descriptors including pitch average, speech section level variation, and speech level. d) Vocal tract characteristics, including robotization and pitch cross power.

[0358] To derive short-term features from frame-level features, the short-term feature extractor 2022 may be configured to calculate statistics of frame-level features as short-term features.

[0359] Examples of frame-level feature statistics include mean and standard deviation. These capture rhythmic attributes for distinguishing between various audio types such as short-term music, speech, background, and noise. For example, speech typically alternates between voiced and unvoiced consonants at a syllable rate, while music lacks such alternation; therefore, the variability of frame-level features in speech is typically greater than in music.

[0360] Another example of a statistic is the weighted average of frame-level features. For example, for a cutoff frequency, the weighted average of the cutoff frequencies derived from all audio frames within a short-term segment, weighted by the energy or loudness of each frame, would be the cutoff frequency for the short-term segment.

[0361] Alternatively or additionally, as shown in Figure 29, the audio classifier 200 may include a frame-level feature extractor 2012 that extracts frame-level features from audio frames, and a frame-level classifier 2014 that uses the respective frame-level features to classify each frame in the sequence of audio frames into a frame-level audio type. Here, a short-term feature extractor 2022 may be configured to calculate short-term features based on the results of the frame-level classifier 2014 for the audio frames of the sequence.

[0362] In other words, in addition to the audio content classifier 202 and the audio context classifier 204, the audio classifier 200 may also have a frame classifier 201. In such a configuration, the audio content classifier 202 classifies short-term segments based on the frame-level classification results of the frame classifier 201, and the audio context classifier 204 classifies long-term segments based on the short-term classification results of the audio content classifier 202.

[0363] The frame-level classifier 2014 may be configured to classify each frame in a sequence of audio frames into some class, which may be referred to as a "frame-level audio type". In some embodiments, frame-level audio types may have a configuration similar to that of the content types discussed above, and may have the same meaning as the content types. The only difference is that frame-level audio types and content types are classified at different levels of the audio signal, namely frame level and short-term segment level. For example, frame-level classifier 2014 may be configured to classify each frame of a sequence of audio frames into at least one of the following frame-level audio types: speech, music, background noise, and noise. On the other hand, frame-level audio types may have a configuration that is partially or entirely different from that of the content types, and is more suitable for frame-level classification and for use as short-term features for short-term classification. For example, frame-level classifier 2014 may be configured to classify each frame of a sequence of audio frames into at least one of the following frame-level audio types: voiced, silent, and pause.

[0364] A similar method may be employed to derive short-term features from the results of frame-level classification, by referring to the description in Section 6.2.

[0365] Alternatively, the short-term classifier 2024 may use short-term features based on the results of the frame-level classifier 2014 and short-term features directly based on frame-level features obtained by the frame-level feature extractor 2012. Thus, the short-term feature extractor 2022 may be configured to compute short-term features based on both frame-level features extracted from the sequence of audio frames and the results of the frame-level classifier on the sequence of audio frames.

[0366] In other words, the frame-level feature extractor 2012 may be configured to compute both short-term features that include at least one of the following features, as described in relation to the statistics discussed in Section 6.2 and in connection with Figure 28: features that characterize the attributes of various short-term audio types, cutoff frequencies, static signal-to-noise ratio characteristics, segment signal-to-noise ratio characteristics, basic speech descriptors, and vocal tract characteristics.

[0367] To work in real time, in all embodiments, the short-term feature extractor 2022 may be configured to act on short-term audio segments formed using a moving window that slides within the time dimension of the long-term audio segment by a predetermined step length. For details on the moving window for short-term audio segments and the moving window for long-term audio segments, see Section 1.1.

[0368] <Section 6.4 Combinations of Embodiments and Application Scenarios> As with Part 1, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components mentioned in different parts / embodiments but having the same or similar function may be implemented as the same or separate components.

[0369] For example, any two or more of the solutions described in sections 6.1 to 6.3 may be combined with each other. Any of these combinations may be further combined with any embodiments described or implied in Parts I to Five and other parts described later. In particular, the type smoothing unit 712 discussed in Part I may be used in this part as a component of the audio classifier 200 to smooth the results of the frame classifier 2014 or the audio content classifier 202 or the audio context classifier 204. Furthermore, a timer 916 may also play a role as a component of the audio classifier 200 to avoid abrupt changes in the output of the audio classifier 200.

[0370] <Section 6.5 Audio Classification Methods> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the audio classifier in the embodiments described above. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0371] In one embodiment shown in Figure 30, an audio classification method is provided. To identify the long-term audio type (i.e., content type) of a long-term audio segment consisting of a sequence of short-term audio segments (overlapping or not overlapping), the short-term audio segments are first classified into short-term audio types, i.e., content types (operation 3004), and long-term features are obtained by calculating statistics of the results of the classification operation on the sequence of short-term segments within the long-term audio segment (operation 3006). Then, long-term classification (operation 3008) may be performed using the long-term features. Short-term audio segments may include a sequence of audio frames. Of course, in order to identify the short-term audio type of the short-term segments, short-term features need to be extracted from those segments.

[0372] Short-term audio content (content type) may include, but is not limited to, speech, short-term music, background sounds, and noise.

[0373] The long-term characteristics include the mean and variance of the confidence values ​​of those short-term segments, the mean and variance weighted by the importance of those short-term segments, the frequency of occurrence of each short-term audio type, and the frequency of transitions between different short-term audio types.

[0374] In some cases, as shown in Figure 31, further long-term features may be obtained directly based on the short-term features of the sequence of short-term segments within the long-term audio segment (operation 3107). Such long-term features may include, but are not limited to, the following statistics of the short-term features: mean, variance, weighted mean, weighted variance, high mean, low mean, and the ratio between the high mean and the low mean.

[0375] There are various methods for extracting short-term features. One is to directly extract short-term features from the short-term audio segments to be classified. Such features include, but are not limited to, rhythmic characteristics, interruption / mute characteristics, and short-term audio quality characteristics.

[0376] The second method involves extracting frame-level features from the audio frames contained in each short-term segment (operation 3201 in Figure 32), and then calculating short-term features based on the frame-level features, for example, calculating statistics of the frame-level features as short-term features. Frame-level features may include, but are not limited to, features characterizing the attributes of various short-term audio types, cutoff frequencies, static signal-to-noise ratio characteristics, segment signal-to-noise ratio characteristics, basic speech descriptors, and vocal tract characteristics. Features characterizing the attributes of various short-term audio types may further include frame energy, subband spectral distribution, spectral flux, Mel-frequency cepstral coefficients (MFCCs), bass, residual information, chroma features, and zero-crossing rate.

[0377] A third method involves extracting short-term features in a similar manner to extracting long-term features: after extracting frame-level features from audio frames within the short-term segment to be classified (operation 3201), each audio frame is classified into a frame-level audio type using each frame-level feature (operation 32011 in Figure 33); short-term features can be extracted by calculating short-term features based on frame-level audio types (optionally including confidence values) (operation 3002). Frame-level audio types may have similar attributes and configurations to short-term audio types (content types) and may also include speech, music, background sounds, and noise.

[0378] The second and third methods described above may be combined, as shown by the dashed arrows in Figure 33.

[0379] As discussed in Part 1, both short-term and long-term audio segments may be sampled using a moving window. That is, the operation to extract short-term features (operation 3002) may be performed on a short-term audio segment formed using a moving window that slides in the time dimension of the long-term audio segment by a predetermined step length, and the operation to extract long-term features (operation 3107) and the operation to calculate short-term audio type statistics (operation 3006) may be performed on a long-term audio segment formed using a moving window that slides in the time dimension of the audio signal by a predetermined step length.

[0380] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is feasible. On the other hand, any aspect of embodiments and variations of audio processing methods may be separate solutions. Furthermore, any two or more solutions described in this section may be combined with each other, and these combinations may further be combined with any embodiments described or implied in other parts of this disclosure. In particular, as already discussed in Section 6.4, audio type smoothing schemes and transition schemes may be part of the audio classification methods discussed herein.

[0381] <Part 7: VoIP Classifiers and Classification Methods> Part 6 proposes a novel audio classifier for classifying audio signals into audio context types, at least partially based on the results of content type classification. In the embodiments discussed in Part 6, long-term features are extracted from long-term segments lasting several seconds to tens of seconds. Therefore, audio context classification can result in high latency. It is desirable that audio contexts can be classified in real time or near real time, for example, at the short-term segment level.

[0382] <Section 7.1 Contextual Classification Based on Short-Term Segments> Therefore, as shown in Figure 34, an audio classifier 200A is provided, which includes an audio content classifier 202A for identifying the content type of a short-term segment of an audio signal, and an audio context classifier 204A for identifying the context type of a short-term segment, at least in part, based on the content type identified by the audio content classifier.

[0383] Here, the audio content classifier 202A may employ the techniques already described in Part 6, or it may employ the various techniques discussed in Section 7.2 below. Similarly, the audio context classifier 204A may employ the techniques already described in Part 6, but with the difference that the context classifier 204A may use the results of the audio content classifier 202A directly, rather than using the statistical results from the audio content classifier 202A. This is because both the audio context classifier 204 and the audio content classifier 202A classify the same short-term segments. Furthermore, as in Part 6, in addition to the results from the audio content classifier 202A, the audio context classifier 204 may use other features directly extracted from the short-term segments. That is, the audio context classifier 204A may be configured to classify short-term segments based on a machine learning model, using the confidence value of the content type of the short-term segments and other features extracted from the short-term segments as features. Part 6 can be referenced for features extracted from short-term segments.

[0384] The audio content classifier 200A may simultaneously label short-term segments as more audio types than VoIP speech / noise and / or non-VoIP speech / noise (VoIP speech / noise and / or non-VoIP speech / noise are discussed in Section 7.2 below), and each of the multiple audio types may have its own confidence value as discussed in Section 1.2. This allows for richer information to be supplemented, resulting in better classification accuracy. For example, combined confidence values ​​for speech and short-term music can reveal to what extent the audio content is likely to be a mixture of speech and background music, thereby discriminating it from pure VoIP content.

[0385] <Section 7.2 Classification using VoIP utterances and VoIP noise> This aspect of the present invention is particularly useful in VoIP / non-VoIP classification systems that require the classification of current short-term segments due to short decision latency.

[0386] For this purpose, as shown in Figure 34, the audio classifier 200A is specifically designed for VoIP / non-VoIP classification. To classify VoIP / non-VoIP, a VoIP speech classifier 2026 and / or a VoIP noise classifier are developed to generate intermediate results for the final robust VoIP / non-VoIP classification by the audio context classifier 204.

[0387] VoIP short-term segments will likely contain alternating VoIP utterances and VoIP noise. While high accuracy can be achieved in classifying short-term segments of speech as either VoIP or non-VoIP utterances, it is observed that this is not the case when classifying short-term segments of noise as either VoIP noise or non-VoIP noise. Thus, it can be concluded that discriminative accuracy is blurred by directly classifying short-term segments into VoIP (containing VoIP utterances and VoIP noise, but without individual identification of VoIP utterances and VoIP noise) and non-VoIP, without considering the difference between speech and noise, and thus leaving the features of these two content types (speech and noise) mixed together.

[0388] For a classifier, it makes sense to achieve higher accuracy in VoIP speech / non-VoIP speech classification than in VoIP noise / non-VoIP noise classification. This is because speech contains more information than noise, and features such as cutoff frequency are more effective in classifying speech. According to the weight ranking obtained from the Adaboost training / process, the top-weighted short-term features for VoIP / non-VoIP classification are: standard deviation of log energy, cutoff frequency, standard deviation of rhythmic intensity, and standard deviation of spectral flux. The standard deviation of log energy, standard deviation of rhythmic intensity, and standard deviation of spectral flux are generally higher for VoIP speech than for non-VoIP speech. One possible reason is that many short-term speech segments in non-VoIP contexts such as cinematic media or games are typically mixed with other sounds such as background music or sound effects, and the values ​​of the above features are lower for such background music or sound effects. On the other hand, the cutoff characteristic is generally lower for VoIP utterances than for non-VoIP utterances. This indicates the low cutoff frequency introduced by many common VoIP codecs.

[0389] Therefore, in one embodiment, the audio content classifier 202A may have a VoIP utterance classifier 2026 that classifies short-term segments into content-type VoIP utterances or content-type non-VoIP utterances, and the audio context classifier 204 may be configured to classify short-term segments into context-type VoIP or context-type non-VoIP based on confidence values ​​of VoIP and non-VoIP utterances.

[0390] In another embodiment, the audio content classifier 202A may further include a VoIP noise classifier 2028 that classifies short-term segments into content-type VoIP noise or content-type non-VoIP noise, and the audio context classifier 204 may be configured to classify short-term segments into context-type VoIP or context-type non-VoIP based on confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise.

[0391] The content types of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise may be identified using existing techniques discussed in Part VI, Sections 1.2 and 7.1.

[0392] Alternatively, the audio content classifier 202A may have a hierarchical structure as shown in Figure 35. That is, it first classifies short segments into speech or noise / background using the results from the speech / noise classifier 2025.

[0393] Based on an embodiment that simply uses the VoIP utterance classifier 2026, if a short-term segment is determined to be an utterance by the utterance / noise classifier 2025 (which in such a situation is simply an utterance classifier), the VoIP utterance classifier 2026 proceeds to classify whether it is a VoIP utterance or a non-VoIP utterance and computes a binary classification result. Otherwise, the confidence value of the VoIP utterance may be low, or the decision about the VoIP utterance may be considered uncertain.

[0394] Based on an embodiment that simply uses the VoIP noise classifier 2028, if a short-term segment is determined to be noise by the speech / noise classifier 2025 (in such a situation this is simply a noise (background) classifier), the VoIP noise classifier 2028 proceeds to classify it as either VoIP noise or non-VoIP noise and compute a binary classification result. Otherwise, the confidence value for VoIP noise may be low, or the decision about VoIP noise may be considered uncertain.

[0395] Here, since speech is generally an informational content type and noise / background is an intrusive content type, even if a short-term segment is not noise, it cannot be definitively determined in the preceding embodiment that the short-term segment is not contextual VoIP. If a short-term segment is not speech, then in embodiments that simply use the VoIP speech classifier 2026, it is probably not contextual VoIP. Therefore, embodiments that use the primary VoIP speech classifier 2026 can generally be implemented independently. On the other hand, other embodiments that use the unit VoIP noise classifier 2028 can be used, for example, as a supplementary embodiment that works in cooperation with embodiments that use the VoIP speech classifier 2026.

[0396] In other words, both the VoIP speech classifier 2026 and the VoIP noise classifier 2028 may be used. If the short-term segment is determined to be a speech / noise classifier 2025, the VoIP speech classifier 2026 proceeds to classify it as either a VoIP or non-VoIP speech and calculates a binary classification result. If the short-term segment is determined to be noise by the speech / noise classifier 2025, the VoIP noise classifier 2028 proceeds to classify it as either VoIP noise or non-VoIP noise and calculates a binary classification result. Otherwise, the short-term segment may be considered as potentially classifiable as non-VoIP.

[0397] The implementation of the speech / noise classifier 2025, the VoIP speech classifier 2026, and the VoIP noise classifier 2028 may employ any existing technique, including the audio content classifier 202 discussed in Parts I through VI.

[0398] If the audio content classifier 202A implemented in accordance with the above ultimately fails to classify the short-term segment as speech, noise, or background, or as VoIP speech, non-VoIP speech, VoIP noise, or non-VoIP noise, i.e., if all relevant confidence values ​​are low, the audio content classifier 202A (and audio context classifier 204) may classify the short-term segment as non-VoIP.

[0399] To classify short-term segments into VoIP or non-VoIP context types based on the results of VoIP speech classifier 2026 and VoIP noise classifier 2028, the audio context classifier 204 may employ the machine learning-based techniques discussed in Section 7.1, and as a modification, more features may be used, including short-term features directly extracted from short-term segments and / or the results of other audio content classifiers (one or more) directed towards content types other than VoIP-related content types, as already discussed in Section 7.1.

[0400] In addition to the machine learning-based techniques described above, an alternative approach to VoIP / non-VoIP classification can be heuristic rules that leverage domain knowledge and utilize the classification results in relation to VoIP utterances and VoIP noise.

[0401] If the current short-term segment at time t is determined to be either a VoIP or non-VoIP utterance, the classification result is taken directly as the VoIP / non-VoIP classification result. This is because the VoIP / non-VoIP utterance classification is robust, as discussed earlier. That is, if the short-term segment is determined to be a VoIP utterance, it is a contextual VoIP. If the short-term segment is determined to be a non-VoIP utterance, it is a contextual non-VoIP.

[0402] When VoIP speech classifier 2026 makes a binary determination between VoIP and non-VoIP utterances based on utterances classified by the aforementioned speech / noise classifier 2025, the confidence values ​​for VoIP and non-VoIP utterances can be complementary. That is, their sum is 1 (where 0 represents 100% negative and 1 represents 100% positive). The confidence thresholds for distinguishing between VoIP and non-VoIP utterances can actually point to the same point. If VoIP speech classifier 2026 is not a binary classifier, the confidence values ​​for VoIP and non-VoIP utterances are not complementary, and the confidence thresholds for distinguishing between VoIP and non-VoIP utterances may not necessarily point to the same point.

[0403] However, if the confidence level of VoIP or non-VoIP utterances is close to a threshold and fluctuates around that threshold, the VoIP / non-VoIP classification result may switch too frequently. To avoid such fluctuations, a buffering mechanism may be provided. Both thresholds for VoIP and non-VoIP utterances may be set higher, so that switching from one content type to the other becomes less easy. For simplicity of description, confidence levels for non-VoIP utterances may be converted to confidence levels for VoIP utterances. That is, a higher confidence level indicates that the short-term segment is considered closer to a VoIP utterance, and a lower confidence level indicates that the short-term segment is considered closer to a non-VoIP utterance. For such a non-binary classifier, a high confidence level for non-VoIP utterances does not necessarily mean a low confidence level for VoIP utterances, but such a simplification reflects the essence of the solution well, and relevant claims described using the language of binary classifiers are interpreted as covering equivalent solutions for non-binary classifiers.

[0404] The buffer scheme is shown in Figure 36. There is a buffer region between two thresholds Th1 and Th2 (Th1 ≥ Th2). When the confidence value v(t) of a VoIP utterance decreases in this region, the context classification does not change. This is indicated by the left and right arrows in Figure 36. A short-term segment is classified as VoIP only when the confidence value v(t) is greater than the larger threshold Th1 (as indicated by the lower arrow in Figure 36), and a short-term segment is classified as non-VoIP only when the confidence value is not greater than the smaller threshold Th2 (as indicated by the upper arrow in Figure 36).

[0405] The situation is similar if VoIP noise classifier 2028 is used instead. To make the solution more robust, VoIP speech classifier 2026 and VoIP noise classifier 2028 may be used together. Audio context classifier 204A may then be configured to: classify a short-term segment as contextual VoIP if the confidence value of the VoIP speech is greater than a first threshold or the confidence value of the VoIP noise is greater than a third threshold; classify a short-term segment as contextual non-VoIP if the confidence value of the VoIP speech is less than a second threshold that is not greater than the first threshold or the confidence value of the VoIP noise is less than a fourth threshold that is not greater than the third threshold; otherwise, classify the short-term segment as contextual for the last short-term segment.

[0406] Here, the first threshold may be equal to the second threshold, and the third threshold may be equal to the fourth threshold. This is especially true for, but not limited to, binary VoIP speech classifiers and binary VoIP noise classifiers. However, since VoIP noise classification results are generally not very robust, it would be preferable for the third and fourth thresholds not to be equal to each other. Both should be far from 0.5 (0 indicates a high confidence that it is not VoIP noise, and 1 indicates a high confidence that it is VoIP noise).

[0407] <Section 7.3 Smoothing Fluctuations> To avoid rapid fluctuations, another solution is to smooth the confidence values ​​determined by the audio content classifier. Therefore, as shown in Figure 37, a type smoothing unit 203A may be included in the audio classifier 200A. For each of the confidence values ​​of the content types of the four VoIP relationships discussed earlier, the smoothing method discussed in Section 1.3 may be employed.

[0408] Alternatively, as in Section 7.2, VoIP and non-VoIP utterances may be considered as pairs with complementary confidence values. VoIP noise and non-VoIP noise may also be considered as pairs with complementary confidence values. In such situations, only one of each pair needs to be smoothed, and the smoothing schemes discussed in Section 1.3 may be employed.

[0409] Taking the confidence value of a VoIP utterance as an example, formula (3) can be rewritten as follows:

[0410] v(t)=β·v(t-1)+(1-β)·voipSpeechConf(t) (3") Here, v(t) is the smoothed VoIP utterance confidence score at time t, v(t-1) is the smoothed VoIP utterance confidence score at the last time point, voipSpeechConf is the VoIP utterance confidence score at the current time t before smoothing, and α is the weighting coefficient.

[0411] In one variation, if we have the speech / noise classifier 2025 as described above, and the confidence level of the speech for a short-term segment is low, then that short-term segment cannot be robustly classified as a VoIP utterance, and we can directly set voipSeechConf(t)=v(t-1) without actually making the VoIP speech classifier 2026 function.

[0412] Alternatively, in the above situation, you can indicate an uncertain case by setting voipSpeechConf(t)=0.5 (or another value not higher than 0.5, such as 0.4 to 0.5) (where confidence=1 indicates high confidence that it is VoIP, and confidence=0 indicates high confidence that it is not VoIP).

[0413] Therefore, according to this modification, as shown in Figure 37, the audio content classifier 200A may further have an utterance / noise classifier 2025 for identifying the content type of short-term segment utterances, and the type smoothing unit 203A may be configured to set the confidence value of the VoIP utterance for the current short-term segment before smoothing as a predetermined confidence value (such as 0.5 or other values ​​such as 0.4-0.5), or as the smoothed confidence value of the last short-term segment whose confidence value for content type utterances classified by the utterance / noise classifier is lower than a fifth threshold. In such a situation, the VoIP utterance classifier 2026 may or may not be in operation. Alternatively, the setting of the confidence value may be performed by the VoIP utterance classifier 2026. This is equivalent to the solution in which this operation is performed by the type smoothing unit 203A, and the claim should be interpreted as covering both situations. Furthermore, while the phrase "the confidence level for content-type utterances classified by the speech / noise classifier is lower than the fifth threshold" is used here, the scope of protection is not limited to that, and is equivalent to situations where short-term segments are classified as content-type other than speech.

[0414] The situation regarding the confidence level of VoIP noise is similar, so a detailed explanation will be omitted here.

[0415] To avoid abrupt fluctuations, another solution is to smooth the confidence values ​​determined by the audio context classifier 204A, and the smoothing methods discussed in Section 1.3 may be used.

[0416] To avoid abrupt fluctuations, another solution is to delay the transition of context types between VoIP and non-VoIP, using the same method described in Section 1.6. As described in Section 1.6, Timer 916 may be external to the audio classifier or internal to it as part of the audio classifier. Thus, as shown in Figure 38, the audio classifier 200A may also have Timer 916. The audio classifier is configured to continue outputting the current context type until the duration of the new context type reaches a sixth threshold (the context type is an instance of the audio type). A detailed explanation can be omitted here by referring to Section 1.6.

[0417] Additionally or alternatively, another way to delay the transition between VoIP and non-VoIP is that the first and / or second thresholds for VoIP / non-VoIP classification may differ depending on the context type of the last short-term segment. That is, when the context type of the new short-term segment differs from that of the last short-term segment, the first and / or second thresholds become larger, while they become smaller when the context type of the new short-term segment is the same as that of the last short-term segment. In this way, the context type tends to be maintained at its current state, and thus rapid fluctuations in the context type can be suppressed to some extent.

[0418] <Section 7.4 Combinations of Embodiments and Application Scenarios> As with Part 1, all embodiments and their variations discussed above may be implemented in any combination thereof, and any components mentioned in different parts / embodiments but having the same or similar function may be implemented as the same or separate components.

[0419] For example, any two or more of the solutions described in Sections 7.1 to 7.3 may be combined with each other. Any of these combinations may be further combined with any embodiment described or implied in Parts I to VI. In particular, the embodiments discussed in this Part and any combination thereof may be combined with the embodiments of the audio processing apparatus / method or volume leveler / control method discussed in Part IV.

[0420] <Section 7.5 VoIP classification method> As with Part 1, it is clear that several processes or methods are disclosed in the process of describing the audio classifier in the embodiments described above. An overview of these methods is given below, but some of the details already discussed above will not be repeated.

[0421] In one embodiment shown in Figure 39, the audio classification method includes identifying the content type of a short-term segment of an audio signal (operation 4004), and then identifying the context type of the short-term segment based on at least the identified content type (operation 4008).

[0422] To dynamically and rapidly identify the context type of an audio signal, the audio classification method in this section is particularly useful when distinguishing between context-based VoIP and non-VoIP. In such situations, short segments may first be classified as either content-based VoIP utterances or content-based non-VoIP utterances, and the above operation for identifying the context type is configured to classify short segments as either context-based VoIP or context-based non-VoIP based on the confidence values ​​of VoIP and non-VoIP utterances.

[0423] Alternatively, short-term segments may first be classified as content-based VoIP noise or content-based non-VoIP noise, and the above operation for identifying context types may be configured to classify short-term segments as context-based VoIP or context-based non-VoIP based on the confidence values ​​of VoIP noise and non-VoIP noise.

[0424] Utterances and noise may be considered together. In such situations, the above operation for identifying context types may be configured to classify short-term segments as context-type VoIP or context-type non-VoIP based on confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise.

[0425] To identify the context type of the short-term segment, a machine learning model may be used that incorporates both the confidence value of the content type of the short-term segment and other features extracted from the short-term segment as features.

[0426] The above operation for identifying context types may be implemented based on heuristic rules. When only VoIP and non-VoIP utterances are involved, the heuristic rules are as follows: if the confidence value of the VoIP utterance is greater than the first threshold, classify the short-term segment as context-type VoIP; if the confidence value of the VoIP utterance is less than the second threshold, classify the short-term segment as context-type non-VoIP, where the second threshold is less than the first threshold; otherwise, classify the short-term segment as the context type for the last short-term segment.

[0427] The same applies to heuristic rules for situations involving only VoIP noise and non-VoIP noise.

[0428] When both speech and noise are involved, the heuristic rule is as follows: classify the short-term segment as contextual VoIP if the confidence value of the VoIP speech is greater than the first threshold or the confidence value of the VoIP noise is greater than the third threshold; classify the short-term segment as contextual non-VoIP if the confidence value of the VoIP speech is less than the second threshold but not greater than the first threshold or the confidence value of the VoIP noise is less than the fourth threshold but not greater than the third threshold; otherwise, classify the short-term segment as contextual for the last short-term segment.

[0429] The smoothing methods discussed in Sections 1.3 and 1.8 may be used here, and a detailed explanation is omitted. As a modification to the smoothing method described in Section 1.3, the method may further include a step of identifying content-type utterances from short-term segments (operation 40040 in Figure 40) before smoothing operation 4106. Here, if the confidence value for content-type utterances is lower than the fifth threshold ("No" in operation 40041), the confidence value of the VoIP utterance for the current short-term segment before smoothing is set to a predetermined confidence value or the smoothed confidence value of the last short-term segment (operation 40044 in Figure 40).

[0430] Instead, if the content-type utterance identification operation reliably determines that a short-term segment is an utterance (Yes in operation 40041), the short-term segment is further classified as either a VoIP utterance or a non-VoIP utterance before the smoothing operation 4106 (operation 40042).

[0431] In fact, even without using a smoothing scheme, this method may first identify content-type utterances and / or noise, and when a short segment is classified as utterance or noise, further classification is implemented to classify the short segment as either VoIP utterance or non-VoIP utterance, or VoIP noise or non-VoIP noise. Then, the process of identifying the context type is performed.

[0432] As described in sections 1.6 and 1.8, the transition schemes discussed therein may be incorporated as part of the audio classification method described here, and details are omitted. In short, the method may further include measuring the duration for which the context type identification operation continuously outputs the same context type. The audio classification method is configured to continue outputting the current context type until the duration of a new context type reaches a sixth threshold.

[0433] Similarly, different sixth thresholds may be set for different transition pairs from one context type to another. Furthermore, the sixth threshold may be negatively correlated with the confidence value of the new context type.

[0434] As a modification to the transition scheme in an audio classification method specifically for VoIP / non-VoIP classification, one or more of the first to fourth thresholds for the current short-term segment may be set to differ depending on the context type of the last short-term segment.

[0435] Similar to embodiments of audio processing devices, any combination of embodiments and variations of audio processing methods is practical. On the other hand, any aspect of embodiments and variations of audio processing methods may be separate solutions. Furthermore, any two or more solutions described in this section may be combined with each other, and these combinations may further be combined with any embodiments described or implied in other parts of this disclosure. In particular, the audio classification methods described herein may be used in the aforementioned audio processing methods, especially the volume leveler control methods.

[0436] As discussed at the beginning of the “Modes for Carrying Out the Invention” section of this application, embodiments of this application can be embodied in hardware, software, or both. Figure 41 is a block diagram illustrating an exemplary system that implements aspects of this application.

[0437] In Figure 41, the central processing unit (CPU) 4201 executes various processes according to programs stored in read-only memory (ROM) 4202 or programs loaded from memory section 4208 into random access memory (RAM) 4203. RAM 4203 also stores data required by the CPU 4201 when executing these various processes, etc., as needed.

[0438] The CPU 4201, ROM 4202, and RAM 4203 are connected to each other via bus 4204. The input / output interface 4205 is also connected to bus 4204.

[0439] The following components are connected to the input / output interface 4205: an input unit 4206 including a keyboard, mouse, etc.; an output unit 4207 including a display such as a cathode ray tube (CRT), liquid crystal display (LCD), and loudspeaker, etc.; a storage unit 4208 including a hard disk, etc.; and a communication unit 4209 including a network interface card such as a LAN card or modem. The communication unit 4209 performs communication processes over a network such as the Internet.

[0440] Drive 4210 is also connected to the input / output interface 4205 as needed. Removable media 4211, such as magnetic disks, optical disks, magneto-optical disks, and semiconductor memory, are mounted on drive 4210 as needed. Computer programs to be read from there are then installed in storage unit 4208 as needed.

[0441] If the above components are implemented by software, the program making up the software is installed from a network such as the Internet or from a storage medium such as the removable media 4211.

[0442] It should be noted that the terms used herein are merely for the purpose of describing specific embodiments and are not intended to limit the application. In its usage herein, the singular form is intended to include the plural unless the context explicitly indicates otherwise. Furthermore, as used herein, the terms “includes” and / or “have” indicate the presence of the described features, integers, actions, stages, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, actions, stages, elements, components, and / or groups thereof.

[0443] The corresponding structures, materials, processes, and equivalents of elements that add function to any means or actions in a claim are intended to include any structures, materials, or processes for performing the function in combination with other claim-defined elements explicitly stated in the claims. The descriptions in this application are presented for illustrative and explanatory purposes, but are not intended to be exhaustive or to limit the applications to the forms disclosed. Many modifications and variations will be obvious to those skilled in the art without deviating from the scope and spirit of this application. The embodiments have been selected and described to best illustrate the principles and practical applications of this application and to enable those skilled in the art to understand this application in terms of various embodiments with various modifications suitable for the specific applications considered.

[0444] Several aspects are described below. [Aspect 1] An audio content classifier for identifying the content type of an audio signal in real time; A volume leveler controller having an adjustment unit that adjusts the volume leveler in a continuous manner based on identified content types, The adjustment unit is configured to positively correlate the dynamic gain of the volume leveler with the content type of the informationality of the audio signal, and negatively correlate the dynamic gain of the volume leveler with the content type of the coherence of the audio signal. Volume leveling controller. [Aspect 2] The volume leveling controller according to embodiment 1, wherein the content type of the audio signal includes one of speech, short-term music, noise, and background sound. [Aspect 3] A volume leveler controller according to embodiment 1, wherein the noise is considered to be of the coercive content type. [Aspect 4] The volume leveler controller according to embodiment 1, wherein the adjustment unit is configured to adjust the dynamic gain of the volume leveler based on the confidence value of the content type. [Aspect 5] The volume leveler controller according to embodiment 4, wherein the adjustment unit is configured to adjust the dynamic gain via the content-type confidence value transfer function. [Aspect 6] Volume leveler controller according to embodiment 1, wherein the audio content classifier is configured to classify the audio signal into a plurality of content types having corresponding confidence values, and the adjustment unit is configured to take into account at least some of the plurality of audio types by weighting the confidence values ​​of the plurality of content types based on the importance of the plurality of content types. [Aspect 7] The volume leveler controller according to embodiment 1, wherein the audio content classifier is configured to classify the audio signal into a plurality of content types having corresponding confidence values, and the adjustment unit is configured to modify the weight of one content type using the confidence value of at least one other content type. [Aspect 8] Volume leveler controller according to embodiment 1, wherein the audio content classifier is configured to classify the audio signal into a plurality of content types having corresponding confidence values, and the adjustment unit is configured to take into account at least some of the plurality of audio types by weighting the effects of the plurality of content types based on the confidence values. [Aspect 9] The volume leveler controller according to embodiment 8, wherein the adjustment unit is configured to consider at least one dominant content type based on the confidence value. [Aspect 10] The volume leveler controller according to embodiment 1, wherein the audio content classifier is configured to classify the audio signal into a plurality of coherent content types and / or informative content types having corresponding confidence values, and the adjustment unit is configured to consider at least one dominant coherent content type and / or at least one dominant informative content type based on the confidence values. [Aspect 11] A volume leveling controller according to any one of embodiments 1 to 10, further comprising a type smoothing unit for smoothing the current confidence value of the audio signal based on past confidence values ​​of the audio signal for each content type. [Aspect 12] The volume leveler controller according to embodiment 11, wherein the smoothing unit is configured to determine the current smoothed confidence value of the audio signal by calculating a weighted sum of the current actual confidence value and the smoothed confidence value at the last point in time. [Aspect 13] A volume leveler controller according to any one of embodiments 1 to 10, further comprising an audio context classifier that identifies the context type of the audio signal, wherein the adjustment unit is configured to adjust the range of the dynamic gain based on a confidence value of the context type. [Aspect 14] A volume leveler controller according to any one of embodiments 1 to 10, further comprising an audio context classifier that identifies the context type of the audio signal, wherein the adjustment unit is configured to consider the content type of the audio signal as informative or intrusive based on the context type of the audio signal. [Aspect 15] A volume leveler controller according to embodiment 14, wherein the context type of the audio signal includes one of VoIP, cinematic media, long-running music, and games. [Aspect 16] A volume leveler controller according to embodiment 14, wherein in context-type VoIP audio signals, background sounds are considered coercive content, while in context-type non-VoIP audio signals, background sounds and / or speech and / or music are considered informative content. [Aspect 17] The volume leveler controller according to embodiment 14, wherein the context type of the audio signal includes high-quality audio or low-quality audio. [Aspect 18] A volume leveler controller according to embodiment 14, wherein the content type in audio signals of different context types is assigned different weights depending on the context type of the audio signal. [Aspect 19] Volume leveler controller according to embodiment 14, wherein the audio context classifier is configured to classify the audio signal into a plurality of context types having corresponding confidence values, and the adjustment unit is configured to take into account at least some of the plurality of context types by weighting the confidence values ​​of the plurality of context types based on the importance of the plurality of context types. [Aspect 20] Volume leveler controller according to embodiment 14, wherein the audio context classifier is configured to classify the audio signal into a plurality of context types having corresponding confidence values, and the adjustment unit is configured to take into account at least some of the plurality of context types by weighting the effects of the plurality of context types based on the confidence values. [Aspect 21] The audio content classifier is configured to identify the content type based on a short-term segment of the audio signal, The volume leveler controller according to embodiment 14, wherein the audio context classifier is configured to identify the context type based at least in part on a short-term segment of the audio signal based on the content type identified by the audio content classifier. [Aspect 22] The aforementioned audio content classifier includes a VoIP utterance classifier that classifies short-term segments into content-type VoIP utterances or content-type non-VoIP utterances. The volume leveler controller according to embodiment 21, wherein the audio context classifier is configured to classify the short-term segments as contextual VoIP or contextual non-VoIP based on confidence values ​​of VoIP and non-VoIP utterances. [Aspect 23] The aforementioned audio content classifier further, It has a VoIP noise classifier that classifies short-term segments into VoIP noise content types and non-VoIP noise content types. The audio context classifier is configured to classify the short-term segment as contextual VoIP or contextual non-VoIP based on the confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise. A volume leveling controller according to embodiment 22. [Aspect 24] The aforementioned audio context classifier: A volume leveler controller according to embodiment 22, configured to classify a short-term segment as contextual VoIP if the confidence value of a VoIP utterance is greater than a first threshold; to classify a short-term segment as contextual non-VoIP if the confidence value of a VoIP utterance is less than a second threshold that is less than the first threshold; and otherwise, to classify a short-term segment as contextual for the last short-term segment. [Aspect 25] The aforementioned audio context classifier: A volume leveler controller according to embodiment 23, configured to classify a short-term segment as contextual VoIP if the confidence value of the VoIP utterance is greater than a first threshold or the confidence value of the VoIP noise is greater than a third threshold; to classify a short-term segment as contextual non-VoIP if the confidence value of the VoIP utterance is less than a second threshold that is not greater than the first threshold or the confidence value of the VoIP noise is less than a fourth threshold that is not greater than the third threshold; and otherwise to classify the short-term segment as contextual for the last short-term segment. [Aspect 26] A volume leveling controller according to any one of embodiments 21 to 25, further comprising a type smoothing unit for smoothing the current confidence value of the content type based on past confidence values ​​of the content type. [Aspect 27] The volume leveler controller according to embodiment 26, wherein the smoothing unit is configured to determine the smoothed confidence value of the current short-term segment by calculating a weighted sum of the confidence value of the current short-term segment and the smoothed confidence value of the last short-term segment. [Aspect 28] Volume leveler controller according to embodiment 26, wherein the audio content classifier further comprises an utterance / noise classifier that identifies the content type of the utterances in the short-term segment, and the type smoothing unit is configured to set the confidence value of the VoIP utterance for the current short-term segment before smoothing as a predetermined confidence value, or as the smoothed confidence value of the last short-term segment in which the confidence value for the content type utterance classified by the utterance / noise classifier is lower than a fifth threshold. [Aspect 29] Volume leveler controller according to embodiment 22 or 23, wherein the audio context classifier is configured to classify the short-term segments based on a machine learning model, using, as features, a confidence value of the content type of the short-term segments and other features extracted from the short-term segments. [Aspect 30] A volume leveler controller according to any one of embodiments 14 to 29, further comprising a timer for measuring the duration for which the audio context classifier continuously outputs the same context type, and the adjustment unit being configured to continue using the current context type until the duration of a new context type reaches a sixth threshold. [Aspect 31] A volume leveler controller according to embodiment 30, wherein different sixth thresholds are set for different transition pairs from one context type to another context type. [Aspect 32] The volume leveler controller according to embodiment 30, wherein the sixth threshold is negatively correlated with the confidence value of the new context type. [Aspect 33] A volume leveler controller according to embodiment 24 or 25, wherein the first and / or second thresholds vary depending on the context type of the last short-term segment. [Aspect 34] An audio processing apparatus having a volume leveling controller according to any one of embodiments 1 to 33. [Aspect 35] An audio content classifier that identifies the content type of short-term segments of an audio signal; The system includes an audio context classifier that identifies the context type of the short-term segment based at least partially on the content type identified by the audio content classifier, Audio classifier. [Aspect 36] The audio content classifier includes a VoIP utterance classifier that classifies the short-term segment into content-type VoIP utterances or content-type non-VoIP utterances. The audio context classifier is configured to classify the short-term segment as either contextual VoIP or contextual non-VoIP based on the confidence values ​​of VoIP and non-VoIP utterances. The audio classifier according to embodiment 35. [Aspect 37] The aforementioned audio content classifier further, The system includes a VoIP noise classifier that classifies the aforementioned short-term segments into content-type VoIP noise and content-type non-VoIP noise. The audio context classifier is configured to classify the short-term segment as contextual VoIP or contextual non-VoIP based on the confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise. The audio classifier according to embodiment 36. [Aspect 38] The aforementioned audio context classifier: An audio classifier according to embodiment 37, configured to classify a short-term segment as contextual VoIP if the confidence value of a VoIP utterance is greater than a first threshold; to classify a short-term segment as contextual non-VoIP if the confidence value of a VoIP utterance is less than a second threshold that is less than the first threshold; and otherwise, to classify a short-term segment as contextual for the last short-term segment. [Aspect 39] The aforementioned audio context classifier: The system is configured to classify a short-term segment as contextual VoIP if the confidence value of the VoIP utterance is greater than a first threshold or the confidence value of the VoIP noise is greater than a third threshold; to classify a short-term segment as contextual non-VoIP if the confidence value of the VoIP utterance is less than a second threshold that is less than the first threshold or the confidence value of the VoIP noise is less than a fourth threshold that is less than the third threshold; and otherwise, to classify the short-term segment as contextual for the last short-term segment. The audio classifier according to embodiment 37. [Aspect 40] An audio classifier according to any one of embodiments 35 to 39, further comprising a type smoothing unit for smoothing the current confidence value of the content type based on past confidence values ​​of the content type. [Aspect 41] The audio classifier according to embodiment 40, wherein the smoothing unit is configured to determine the smoothed confidence value of the current short-term segment by calculating a weighted sum of the confidence value of the current short-term segment and the smoothed confidence value of the last short-term segment. [Aspect 42] The audio classifier according to embodiment 41, wherein the audio content classifier further comprises an utterance / noise classifier that identifies content-type utterances from the short-term segments, and the type smoothing unit is configured to set the confidence value of the VoIP utterance for the current short-term segment before smoothing as a predetermined confidence value, or as the smoothed confidence value of the last short-term segment in which the confidence value for the content-type utterance classified by the utterance / noise classifier is lower than a fifth threshold. [Aspect 43] The audio classifier according to embodiment 36 or 37, wherein the audio context classifier is configured to classify the short-term segments based on a machine learning model, using, as features, a confidence value of the content type of the short-term segments and other features extracted from the short-term segments. [Aspect 44] The audio classifier according to embodiment 38 or 39, further comprising a timer for measuring the duration for which the audio context classifier continuously outputs the same context type, wherein the audio classifier is configured to continue outputting the current context type until the duration of a new context type reaches a sixth threshold. [Aspect 45] An audio classifier according to embodiment 44, wherein different sixth thresholds are set for different transition pairs from one context type to another. [Aspect 46] The audio classifier according to embodiment 44, wherein the sixth threshold is negatively correlated with the confidence value of the new context type. [Aspect 47] The audio classifier according to embodiment 38 or 39, wherein the first and / or second thresholds vary depending on the context type of the last short-term segment. [Aspect 48] An audio processing apparatus having an audio classifier according to any one of embodiments 35 to 47. [Aspect 49] The stage of identifying the content type of the audio signal in real time; The process includes adjusting the volume leveler in a continuous manner based on the identified content type, by positively correlating the dynamic gain of the volume leveler with the content type of the informationality of the audio signal and negatively correlating the dynamic gain of the volume leveler with the content type of the coherence of the audio signal. Volume leveling device control method. [Aspect 50] The volume leveler control method according to embodiment 49, wherein the content type of the audio signal includes one of speech, short-term music, noise, and background sound. [Aspect 51] The volume leveler control method according to embodiment 49, wherein the noise is considered to be of the coherent content type. [Aspect 52] The volume leveler control method according to embodiment 49, wherein the adjustment operation is configured to adjust the dynamic gain of the volume leveler based on the confidence value of the content type. [Aspect 53] The volume leveler control method according to embodiment 52, wherein the adjustment operation is configured to adjust the dynamic gain via the transfer function of the content type confidence value. [Aspect 54] The volume leveler control method according to embodiment 49, wherein the audio signal is classified into a plurality of content types having corresponding confidence values, and the adjustment operation is configured to take into account at least some of the plurality of audio types by weighting the confidence values ​​of the plurality of content types based on the importance of the plurality of content types. [Aspect 55] The volume leveler control method according to embodiment 49, wherein the audio signal is classified into a plurality of content types having corresponding confidence values, and the adjustment operation is configured to modify the weight of one content type using the confidence value of at least one other content type. [Aspect 56] The volume leveler control method according to embodiment 49, wherein the audio signal is classified into a plurality of content types having corresponding confidence values, and the adjustment operation is configured to take into account at least some of the plurality of content types by weighting the effects of the plurality of content types based on the confidence values. [Aspect 57] The volume leveler control method according to embodiment 56, wherein the adjustment operation is configured to consider at least one dominant content type based on the confidence value. [Aspect 58] The volume leveler control method according to embodiment 56, wherein the audio signal is classified into a plurality of coherence content types and / or a plurality of informative content types having corresponding confidence values, and the adjusting operation is configured to consider at least one dominant coherence content type and / or at least one dominant informative content type based on the confidence values. [Aspect 59] A volume leveler control method according to any one of embodiments 49 to 58, further comprising the step of smoothing the current confidence value of the audio signal based on past confidence values ​​of the audio signal for each content type. [Aspect 60] The volume leveler control method according to embodiment 59, wherein the type smoothing operation is configured to determine the current smoothed confidence value of the audio signal by calculating a weighted sum of the current actual confidence value and the smoothed confidence value at the last point in time. [Aspect 61] A volume leveler control method according to any one of embodiments 49 to 58, further comprising the step of identifying the context type of the audio signal, wherein the adjusting operation is configured to adjust the range of the dynamic gain based on a confidence value of the context type. [Aspect 62] A volume leveler control method according to any one of embodiments 49 to 58, further comprising the step of identifying the context type of the audio signal, wherein the adjusting operation is configured to consider the content type of the audio signal as informative or interfering based on the context type of the audio signal. [Pattern 63] The volume leveler control method according to embodiment 62, wherein the context type of the audio signal includes one of VoIP, cinematic media, long-running music, and games. [Aspect 64] The volume leveler control method according to embodiment 62, wherein in context-type VoIP audio signals, background sounds are considered to be coercive content, while in context-type non-VoIP audio signals, background sounds and / or speech and / or music are considered to be informative content. [Aspect 65] The volume leveler control method according to embodiment 62, wherein the context type of the audio signal includes high-quality audio or low-quality audio. [Pattern 66] The volume leveler control method according to embodiment 62, wherein the content type in audio signals of different context types is assigned different weights depending on the context type of the audio signal. [Pattern 67] The volume leveler control method according to embodiment 62, wherein the audio signal is classified into a plurality of context types having corresponding confidence values, and the adjustment operation is configured to take into account at least some of the plurality of context types by weighting the confidence values ​​of the plurality of context types based on the importance of the plurality of context types. [Pattern 68] The volume leveler control method according to embodiment 62, wherein the audio signal is classified into a plurality of context types having corresponding confidence values, and the adjustment operation is configured to take into account at least some of the plurality of context types by weighting the effects of the plurality of context types based on the confidence values. [Pattern 69] The operation for identifying the content type is configured to identify the content type based on a short-term segment of the audio signal, The operation for identifying the context type is configured to identify the context type based at least partially on a short-term segment of the audio signal based on the identified content type. A volume leveling device control method according to embodiment 62. [Aspect 70] The process of identifying content type includes classifying short-term segments as either content-type VoIP utterances or content-type non-VoIP utterances. The volume leveler control method according to embodiment 69, wherein the operation for identifying the context type is configured to classify the short-term segment as context-type VoIP or context-type non-VoIP based on the confidence values ​​of VoIP and non-VoIP utterances. [Aspect 71] The process of identifying content types is further enhanced by: This includes classifying short-term segments into content-based VoIP noise and content-based non-VoIP noise, The context type identification operation is configured to classify the short-term segment as either context-type VoIP or context-type non-VoIP based on the confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise. A volume leveling device control method according to embodiment 70. [Aspect 72] The behavior of identifying the context type is: If the confidence level of the VoIP utterance is greater than the first threshold, the short-term segment is classified as contextual VoIP; If the confidence value of a VoIP utterance is not greater than the second threshold but not greater than the first threshold, the short-term segment is classified as contextual non-VoIP; Otherwise, the system is configured to classify the short-term segment as a contextual type for the last short-term segment. A volume leveling device control method according to embodiment 70. [Aspect 73] The behavior of identifying context: If the confidence level of the VoIP utterance is greater than the first threshold or the confidence level of the VoIP noise is greater than the third threshold, the short-term segment is classified as contextual VoIP; If the confidence value of the VoIP utterance is not greater than the second threshold but not greater than the first threshold, or if the confidence value of the VoIP noise is not greater than the fourth threshold but not greater than the third threshold, the short-term segment is classified as contextual non-VoIP; The volume leveler control method according to embodiment 71, wherein in other cases, the short-term segment is configured to be classified as a context type for the last short-term segment. [Aspect 74] A volume leveler control method according to any one of embodiments 69 to 73, further comprising the step of smoothing the confidence value of the content type at the present time based on the past confidence value of the content type. [Aspect 75] The volume leveler control method according to embodiment 74, wherein the type smoothing operation is configured to determine the smoothed confidence value of the current short-term segment by calculating a weighted sum of the confidence value of the current short-term segment and the smoothed confidence value of the last short-term segment. [Aspect 76] The volume leveler control method according to embodiment 75, further comprising the step of identifying the content type of the utterance of the short-term segment, wherein the confidence value of the VoIP utterance for the current short-term segment before smoothing is set as a predetermined confidence value, or as the smoothed confidence value of the last short-term segment in which the confidence value for the content type utterance is lower than a fifth threshold. [Aspect 77] A volume leveler control method according to embodiment 70 or 71, characterized in that the short-term segments are classified based on a machine learning model using the content type confidence value of the short-term segments and other features extracted from the short-term segments. [Aspect 78] A volume leveler control method according to any one of embodiments 62 to 77, further comprising measuring the duration for which the operation of identifying a context type continuously outputs the same context type, wherein the adjusting operation is configured to continue using the current context type until the duration of a new context type reaches a sixth threshold. [Aspect 79] A volume leveler control method according to embodiment 78, wherein different sixth thresholds are set for different transition pairs from one context type to another context type. [Aspect 80] The volume leveler control method according to embodiment 78, wherein the sixth threshold is negatively correlated with the new context-type confidence value. [Aspect 81] The volume leveler control method according to embodiment 72 or 73, wherein the first and / or second thresholds differ depending on the context type of the last short-term segment. [Aspect 82] The step of identifying the content type of a short segment of an audio signal; This includes the step of identifying the context type of the short-term segment based at least partially on the identified content type, Audio classification methods. [Aspect 83] The operation of classifying content types includes classifying the short-term segment into a content-type VoIP utterance or a content-type non-VoIP utterance, The context type identification operation is configured to classify the short-term segment as either context-type VoIP or context-type non-VoIP based on the confidence values ​​of VoIP and non-VoIP utterances. The audio classification method described in embodiment 82. [Aspect 84] The process of classifying content types is further enhanced by: This includes classifying the aforementioned short-term segment as either content-based VoIP noise or content-based non-VoIP noise. The context type identification operation is configured to classify the short-term segment as either context-type VoIP or context-type non-VoIP based on the confidence values ​​of VoIP utterances, non-VoIP utterances, VoIP noise, and non-VoIP noise. The audio classification method described in aspect 83. [Aspect 85] The behavior of identifying the context type is: If the confidence level of the VoIP utterance is greater than the first threshold, the short-term segment is classified as contextual VoIP; If the confidence value of a VoIP utterance is not greater than the second threshold but not greater than the first threshold, the short-term segment is classified as contextual non-VoIP; Otherwise, the system is configured to classify the short-term segment as a contextual type for the last short-term segment. The audio classification method described in aspect 83. [Aspect 86] The behavior of identifying the context type is: If the confidence level of the VoIP utterance is greater than the first threshold or the confidence level of the VoIP noise is greater than the third threshold, the short-term segment is classified as contextual VoIP; If the confidence value of the VoIP utterance is not greater than the second threshold but not greater than the first threshold, or if the confidence value of the VoIP noise is not greater than the fourth threshold but not greater than the third threshold, the short-term segment is classified as contextual non-VoIP; Otherwise, the system is configured to classify the aforementioned short-term segment as a contextual type for the last short-term segment. Audio classification method as described in aspect 84. [Aspect 87] An audio classification method according to any one of embodiments 82 to 86, further comprising the step of smoothing the confidence value of the content type at the present time based on the past confidence value of the content type. [Pattern 88] The audio classification method according to embodiment 87, wherein the type smoothing operation is configured to determine the smoothed confidence value of the current short-term segment by calculating a weighted sum of the confidence value of the current short-term segment and the smoothed confidence value of the last short-term segment. [Aspect 89] The audio classification method according to embodiment 88, further comprising the step of identifying content-type utterances from the short-term segments, wherein the confidence value of the VoIP utterance for the current short-term segment before smoothing is set as a predetermined confidence value, or as the smoothed confidence value of the last short-term segment in which the confidence value for the content-type utterance is lower than a fifth threshold. [Aspect 90] The audio classification method according to embodiment 83 or 84, wherein the operation for identifying context types is configured to classify the short-term segments based on a machine learning model, characterized by using the confidence value of the content type of the short-term segments and other features extracted from the short-term segments. [Aspect 91] The audio classification method according to embodiment 85 or 86, further comprising the step of measuring the duration for which an operation to identify a context type continuously outputs the same context type, wherein the audio classification method is configured to continue outputting the current context type until the duration of a new context type reaches a sixth threshold. [Aspect 92] The audio classification method according to embodiment 91, wherein different sixth thresholds are set for different transition pairs from one context type to another context type. [Aspect 93] The audio classification method according to embodiment 91, wherein the sixth threshold is negatively correlated with the confidence value of the new context type. [Aspect 94] The audio classification method according to embodiment 85 or 86, wherein the first and / or second thresholds vary depending on the context type of the last short-term segment. [Pattern 95] A computer-readable medium on which computer program instructions are recorded that, when executed by a processor, enable the processor to execute a volume leveler control method, wherein the volume leveler control method is The stage of identifying the content type of the audio signal in real time; The process includes adjusting the volume leveler in a continuous manner based on the identified content type, by positively correlating the dynamic gain of the volume leveler with the content type of the informationality of the audio signal and negatively correlating the dynamic gain of the volume leveler with the content type of the coherence of the audio signal. Computer-readable media. [Aspect 96] A computer-readable medium on which computer program instructions are recorded that, when executed by a processor, enable the processor to perform an audio classification method, wherein the audio classification method is The step of identifying the content type of a short segment of an audio signal; This includes the step of identifying the context type of the short-term segment based at least partially on the identified content type, Computer-readable media.

Claims

1. A loudness normalization method based on a target loudness value, wherein the method is: A step of determining dynamic gain parameters applied to audio frames of an audio signal based on short-term or long-term characteristics of the audio signal, wherein the determination includes determining at least one dynamic gain parameter for a first audio frame based on the short-term characteristics of the audio signal and the target loudness value, and determining at least one dynamic gain parameter for a second audio frame based on the long-term characteristics of the audio signal and the target loudness value; A step of scaling at least a portion of the first audio frame by the at least one dynamic gain parameter for the first audio frame to modify the loudness of the first audio frame so that it matches the target loudness value during playback; The process includes scaling at least a portion of the second audio frame by the at least one dynamic gain parameter for the second audio frame to modify the loudness of the second audio frame to match the target loudness value during playback, The aforementioned long-term characteristics are determined in a different manner than the aforementioned short-term characteristics. Loudness normalization method.

2. The loudness normalization method according to claim 1, wherein the dynamic gain parameter is identified and applied in real time.

3. The loudness normalization method according to claim 1, wherein a dialogue enhancement is applied that has the effect of making the dialogue more prominent within a specific context.

4. The loudness normalization method according to claim 1, wherein loudness equalization having an effect on one or more playback levels with respect to tonal balance is applied.

5. The loudness normalization method according to claim 1, wherein parameter smoothing is applied to the dynamic gain parameter.

6. An audio processing device configured to normalize loudness based on a target loudness value, wherein: At least one processor; It has at least one memory that stores computer programs; The at least one memory having the computer program, together with the at least one processor, provides to the audio processing device at least: A step of determining dynamic gain parameters applied to audio frames of an audio signal based on short-term or long-term characteristics of the audio signal, wherein the determination includes determining at least one dynamic gain parameter for a first audio frame based on the short-term characteristics of the audio signal and the target loudness value, and determining at least one dynamic gain parameter for a second audio frame based on the long-term characteristics of the audio signal and the target loudness value; A step of scaling at least a portion of the first audio frame by the at least one dynamic gain parameter for the first audio frame to modify the loudness of the first audio frame so that it matches the target loudness value during playback; The system is configured to perform the steps of scaling the second audio frame by at least one dynamic gain parameter for the second audio frame to adjust the loudness of the second audio frame to match the target loudness value during playback, The aforementioned long-term characteristics are determined in a different manner than the aforementioned short-term characteristics. Device.

7. The apparatus according to claim 6, wherein the dynamic gain parameter is identified and applied in real time.

8. The apparatus according to claim 6, wherein a dialogue enhancement is applied that has the effect of making the dialogue more prominent within a specific context.

9. The apparatus according to claim 6, wherein loudness equalization having an effect on one or more playback levels is applied with respect to the tonal balance.

10. The apparatus according to claim 6, wherein parameter smoothing is applied to the dynamic gain parameter.

11. A storage medium having a software program adapted for execution on a processor for performing a step of any one of claims 1 to 5 when executed on a computing device.

12. A computer program for causing a computer to perform the method described in any one of claims 1 to 5.