Method and device for speech / music classification and core encoder selection in a sound codec
A two-stage classification method optimizes stereo audio encoding by categorizing sound signals and selecting core encoders, addressing inefficiencies in existing codecs to reduce bit rates and improve sound quality in immersive audio systems.
Patent Information
- Application Number
- JP2022562835
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-16
- Filing Date
- 2021-04-08
- Publication Date
- 2026-02-12
- Estimated Expiration
- 2041-04-08
AI Technical Summary
Existing audio codecs face challenges in efficiently encoding stereo sound signals, leading to increased bit rates and reduced sound quality due to the lack of exploiting redundancy between left and right channels, particularly in immersive audio scenarios.
A two-stage speech/music classification method and device that classifies input sound signals into specific categories and selects an appropriate core encoder based on high-level features, utilizing a Gaussian Mixture Model (GMM) for initial classification and additional features for refined encoder selection, optimizing encoding for stereo signals in immersive audio systems.
Enhances encoding efficiency by reducing bit rates while maintaining high sound quality, supporting immersive audio experiences by seamlessly adapting to different audio content types.
Smart Images

Figure 0007813238000130 
Figure 0007813238000131 
Figure 0007813238000132
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to sound coding, and more particularly but not exclusively to speech / music classification and core encoder selection, for example in multi-channel sound codecs capable of producing good sound quality in complex audio scenes at low bitrates and low delay. [Background technology]
[0002] In this disclosure and the accompanying claims: The term "sound" can relate to voice, audio and any other sound. - The term "stereo" is an abbreviation for "stereophonic." - The term "mono" is an abbreviation for "monophonic."
[0003] Historically, conversational telephony has been implemented with handsets that have only one transducer for outputting sound to only one of the user's ears. In the last decade, users have begun to use their portable handsets along with headphones to receive sound through both of their ears, primarily for listening to music, but occasionally for listening to voices as well. However, when a portable handset is used to send and receive conversational voice, the content is still mono, but when headphones are used, it is presented to both of the user's ears.
[0004] Using the latest 3GPP® speech coding standard EVS (Enhanced Voice Services), as described in Reference [1], the entire contents of which are incorporated herein by reference, the quality of coded sound, e.g., voice and / or audio, sent and received through portable handsets has been significantly improved. The next logical step is to send stereo information that will allow the receiver to get as close as possible to the real-life audio scene captured at the other end of the communication link.
[0005] For example, transmission of stereo information is commonly used in audio codecs such as those described in reference [2], the entire contents of which are incorporated herein by reference.
[0006] Mono signals are the standard for conversational speech codecs. Because both the left and right channels of a stereo sound signal are coded using a mono codec, the bit rate often doubles when a stereo sound signal is transmitted. While this works well in most scenarios, it presents the drawback of doubling the bit rate and failing to exploit any potential redundancy between the two channels (the left and right channels of a stereo sound signal). Furthermore, to keep the overall bit rate at a reasonable level, very low bit rates for each of the left and right channels are used, thus affecting the overall sound quality. To reduce the bit rate, efficient stereo coding techniques have been developed and used. As non-limiting examples, two stereo coding techniques that can be used efficiently at low bit rates are described in the following paragraphs.
[0007] The first stereo coding technique is called parametric stereo. Parametric stereo encodes two inputs (left and right channels) as a mono signal using a common mono codec with the addition of some amount of stereo side information (corresponding to stereo parameters) that represents the stereo image. The two inputs are downmixed to a mono signal, and then the stereo parameters are calculated. This is usually performed in the frequency domain (FD), for example, in the discrete Fourier transform (DFT) domain. The stereo parameters are related to so-called binaural cues or inter-channel cues. Binaural cues (see, for example, Reference [3], the entire contents of which are incorporated herein by reference) comprise the interaural level difference (ILD), the interaural time difference (ITD), and the interaural correlation (IC). Depending on the sound signal characteristics, the stereo scene configuration, etc., some or all of the binaural cues are coded and transmitted to the decoder. Information about which binaural cues are coded and transmitted is usually sent as signaling information that is part of the stereo side information. Also, a given binaural cue can be quantized using various coding techniques, which results in a variable number of bits being used. Then, in addition to the quantized binaural cues, the stereo side information may include a quantized residual signal resulting from downmixing, usually at a moderate to high bit rate. The residual signal may be coded using an entropy coding technique, for example, an arithmetic encoder.
[0008] Another stereo coding technique operates in the time domain. This stereo coding technique mixes two inputs (left and right channels) into so-called primary and secondary channels. For example, according to the method described in Reference [4], the entire content of which is incorporated herein by reference, time-domain mixing can be based on a mixing ratio, which determines the respective contributions of the two inputs (left and right channels) to generating the primary and secondary channels. The mixing ratio is derived from several metrics, such as the normalized correlation of the two inputs (left and right channels) with a mono signal or the long-term correlation difference between the two inputs (left and right channels). The primary channel can be coded by a common mono codec, while the secondary channel can be coded by a codec with a lower bit rate. Coding of the secondary channel may exploit the coherence between the primary and secondary channels and may reuse some parameters from the primary channel.
[0009] Moreover, in recent years, audio generation, recording, presentation, coding, transmission, and playback have been moving toward enhanced, interactive, and immersive experiences for listeners. An immersive experience can be described, for example, as a state of being deeply engaged or involved in a sound scene with sounds coming from all directions. In immersive audio (also called 3D (three-dimensional) audio), a sound image is reproduced around the listener in all three dimensions, taking into account a wide range of sound characteristics such as timbre, directionality, reverberation, transparency, and auditory spaciousness. Immersive audio is generated for a specific sound playback or reproduction system, such as a loudspeaker-based system, an integrated reproduction system (sound bar), or headphones. The interactivity of the sound reproduction system may then include, for example, the ability to adjust sound levels, change the sound location, or select different languages for playback.
[0010] There are three basic approaches to achieving an immersive experience:
[0011] The first approach to achieving an immersive experience is a channel-based audio approach, which uses spaced microphones to capture sound from different directions, one microphone corresponding to one audio channel in a particular loudspeaker layout. Each recorded channel is then fed to a loudspeaker in a given location. Examples of channel-based audio approaches are, for example, stereo, 5.1 surround, 5.1+4, etc.
[0012] A second approach to achieving an immersive experience is the scene-based audio approach, which represents a desired sound field over a localized space as a function of time by a composition of dimensional components. The sound signals representing scene-based audio are independent of the location of the audio sources, but the sound field is transformed in a renderer into a chosen layout of loudspeakers. An example of scene-based audio is Ambisonics.
[0013] A third approach to achieving an immersive experience is the object-based audio approach, which represents an auditory scene as a set of individual audio elements (e.g., singers, drums, guitars, etc.) accompanied by information such as their position so that they can be rendered by a sound playback system in their intended location. This gives the object-based audio approach great flexibility and interactivity, as each object is kept separate and can be manipulated individually.
[0014] Each of the above-mentioned audio techniques for achieving an immersive experience presents pros and cons. Therefore, rather than using only one audio technique, several audio techniques are typically combined in a complex audio system to create an immersive auditory scene. One example would be an audio system that combines scene-based or channel-based audio with object-based audio, e.g., Ambisonics, with several separate audio objects. Summary of the Invention [Means for solving the problem]
[0015] According to a first aspect, the present disclosure provides a two-stage speech / music classification device for classifying an input sound signal and selecting a core encoder for encoding the sound signal, comprising: a first stage for classifying an input sound signal into one of several final classes; and a second stage for extracting high-level features of the input sound signal and for selecting a core encoder for encoding the input sound signal depending on the extracted high-level features and the final class selected in the first stage.
[0016] According to a second aspect, there is provided a two-stage speech / music classification method for classifying an input sound signal and for selecting a core encoder for encoding the sound signal, comprising: in a first stage, classifying the input sound signal into one of several final classes; and in a second stage, extracting high-level features of the input sound signal and selecting a core encoder for encoding the input sound signal depending on the extracted high-level features and the final class selected in the first stage.
[0017] The above and other objects, advantages, and features of the sound codec including two-stage voice / music classification device and two-stage voice / music classification method will become more apparent upon reading the following non-limiting description of exemplary embodiments thereof, given by way of example only with reference to the accompanying drawings. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a schematic block diagram of a sound processing and communication system illustrating a possible context for implementation of a sound codec including a two-stage speech / music classification device and a two-stage speech / music classification method. [Figure 2] 1 is a schematic block diagram showing simultaneously the first stage of a two-stage audio / music classification device and the first stage of a corresponding two-stage audio / music classification method; [Figure 3] 1 is a schematic block diagram showing simultaneously the second stage of a two-stage audio / music classification device and the second stage of a corresponding two-stage audio / music classification method; [Figure 4] 1 is a schematic block diagram illustrating simultaneously the operation of a state machine of the first stage of a two-stage voice / music classification device and a signal segment of the first stage of a two-stage voice / music classification method. [Figure 5] 1 is a graph showing a non-limiting example of onset / attack detection based on relative frame energy. [Figure 6] 1 is a histogram of selected features in the training database. [Figure 7] 10 is a graph illustrating detection of outlier features based on histogram values. [Figure 8] 1 is a graph showing Box-Cox transformation curves for various values of the power conversion index λ. [Figure 9] 1 is a graph illustrating, by way of non-limiting example, the behavior of rising and falling edge detection used to calculate the forgetting factor of an adaptive IIR filter; [Figure 10]10 is a graph showing the distribution of smoothed difference scores wdlp(n) for the training database, as well as thresholds for transitioning between the SPEECH / NOISE, UNCLEAR, and MUSIC classes. [Figure 11] 10 is a graph showing the ordering of samples in the ENTRY state during the calculation of the weighted average of the difference scores. [Figure 12] FIG. 1 is a class transition diagram showing the complete set of rules for transitioning between the classes SPEECH / NOISE, UNCLEAR, and MUSIC. [Figure 13] FIG. 2 is a schematic diagram illustrating segment attack detection performed on several short segments in a current frame of an input sound signal. [Figure 14] 4 is a schematic diagram illustrating the mechanism of core encoder initial selection used by the core encoder initial selector of the second stage of the two-stage speech / music classification device of FIG. 3. [Figure 15] 1 is a simplified block diagram of an exemplary configuration of hardware components for implementing a sound codec including a two-stage speech / music classification device and method; DETAILED DESCRIPTION OF THE INVENTION
[0019] Recently, 3GPP (Third Generation Partnership Project) has started work on developing a 3D (three-dimensional) sound codec for immersive services called IVAS (Immersive Voice and Audio Services), based on the EVS codec (see reference [5], the entire contents of which are incorporated herein by reference).
[0020] This disclosure describes a speech / music classification technique and a core encoder selection technique in an IVAS coding framework. Both techniques are part of a two-stage speech / music classification method whose result is core encoder selection.
[0021] The speech / music classification method and speech / music classification device are based on those in EVS (see References [6] and [1], Section 5.1.13.6, the entire contents of which are incorporated herein by reference), but with some improvements and developments. Also, the two-stage speech / music classification method and two-stage speech / music classification device are described in this disclosure by way of example only, with reference to the IVAS coding framework, which is referred to throughout this disclosure as the IVAS codec (or IVAS sound codec). However, it is within the scope of this disclosure to incorporate such two-stage speech / music classification method and two-stage speech / music classification device into any other sound codec.
[0022] FIG. 1 is a schematic block diagram of a stereo sound processing and communication system 100 illustrating a possible context for the implementation of a sound codec (IVAS codec) including a two-stage voice / music classification device and method.
[0023] The stereo sound processing and communication system 100 of Figure 1 supports the transmission of stereo sound signals across a communication link 101. The communication link 101 may comprise, for example, a wire link or an optical fiber link. Alternatively, the communication link 101 may comprise, at least in part, a radio frequency link. Radio frequency links often support multiple simultaneous communications requiring shared bandwidth resources, such as may be found with cellular telephony. Although not shown, the communication link 101 may be replaced by a storage device in a single-device implementation of the system 100 that records and stores the coded stereo sound signals for later playback.
[0024] Still referring to Figure 1, for example, a pair of microphones 102 and 122 generate left channel 103 and right channel 123 of an original analog stereo sound signal. As indicated in the above description, the sound signal may comprise, among other things, but not limited to, speech and / or audio.
[0025] The left channel 103 and right channel 123 of the original analog stereo sound signal are provided to an analog-to-digital (A / D) converter 104 for converting them into the left channel 105 and right channel 125 of the original digital stereo sound signal. The left channel 105 and right channel 125 of the original digital stereo sound signal may also be recorded and provided from a storage device (not shown).
[0026] The stereo sound encoder 106 codes the left channel 105 and the right channel 125 of the original digital stereo sound signal, thereby generating a set of coding parameters that are multiplexed under the form of a bitstream 107 that is delivered to the optional error correction encoder 108. When present, the optional error correction encoder 108 adds redundancy to the binary representation of the coding parameters in the bitstream 107 before transmitting the resulting bitstream 111 over the communication link 101.
[0027] At the receiver side, an optional error correction decoder 109 utilizes the aforementioned redundant information in the received bitstream 111 to detect and correct errors that may have occurred during transmission over the communication link 101 and generates a bitstream 112 using the received coding parameters. A stereo sound decoder 110 converts the received coding parameters in the bitstream 112 to create a combined left channel 113 and a right channel 133 of a digital stereo sound signal. The left channel 113 and right channel 133 of the digital stereo sound signal reconstructed in the stereo sound decoder 110 are converted in a digital-to-analog (D / A) converter 115 to a combined left channel 114 and a right channel 134 of an analog stereo sound signal.
[0028] The combined left channel 114 and right channel 134 of the analog stereo sound signal are played back on a pair of loudspeaker units or binaural headphones 116 and 136, respectively. Alternatively, the left channel 113 and right channel 133 of the digital stereo sound signal from the stereo sound decoder 110 may also be fed to and recorded in a storage device (not shown).
[0029] For example, the stereo sound encoder 106 of FIG. 1 may be implemented by an IVAS codec encoder that includes the two-stage speech / music classification device of FIGS.
[0030] 1. Two-stage speech / music classification As indicated in the above description, this disclosure describes a speech / music classification technique and a core encoder selection technique in an IVAS coding framework. Both techniques are part of a two-stage speech / music classification method (and corresponding device) whose result is the selection of a core encoder for coding the primary (dominant) channel (in the case of time-domain (TD) stereo coding) or the downmixed mono channel (in the case of frequency-domain (FD) stereo coding). The basis for the development of this technology is speech / music classification in the EVS codec (Reference [1]). This disclosure describes modifications and improvements implemented in this disclosure that are part of the baseline IVAS codec framework.
[0031] The first stage of the speech / music classification method and device in the IVAS codec is based on the Gaussian Mixture Model (GMM). An earlier model brought over from the EVS codec has been extended, improved, and optimized for processing stereo signals.
[0032] In summary, - The GMM model takes a feature vector as input and provides probabilistic estimates for three classes including speech, music, and background noise. - The parameters of the GMM model are trained on a large set of manually labeled vectors of sound signal features. The GMM model provides a probabilistic estimate for each of the three classes within every frame, e.g., a 20 ms frame. Sound signal processing frames, including subframes, are well known to those skilled in the art, and more information on such frames can be found, for example, in reference [1]. - An outlier detection logic ensures proper processing of frames in which one or more features of the sound signal do not satisfy the conditions of normal distribution. - Individual probabilities are converted into a single unbounded score using logistic regression. - The two-stage voice / music classification device has its own state machine that is used to classify the incoming signal into one of four states. - Adaptive smoothing is applied to the output scores depending on the current state of the two-stage speech / music classification method and the two-stage speech / music classification device. - Fast response of the two-stage speech / music classification method and device in rapidly changing content is realized using onset / attack detection logic based on relative frame energy. The smoothed scores are used to perform a selection between three categories of signal type: pure speech, pure music, speech with music.
[0033] FIG. 2 is a schematic block diagram showing simultaneously a first stage 200 of a two-stage audio / music classification device and a first stage 250 of a corresponding two-stage audio / music classification method.
[0034] Referring to FIG. 2, the first stage of the two-stage speech / music classification device includes: - a state machine 201 for signal division; - an onset / attack detector 202 based on relative frame energies, - feature extractor 203, - a histogram-based outlier detector 204; - Short-term feature vector filter 205, - Nonlinear feature vector transformer 206 (Box-Cox), - Principal Component Analyzer (PCA) 207, - Gaussian Mixture Model (GMM) calculator 208, - an adaptive smoother 209, and - State-dependent category classifier 210 Equipped with.
[0035] The core encoder selection technique in the IVAS codec (the second stage of the two-stage speech / music classification device and method) is built on top of the first stage of the two-stage speech / music classification device and method, and delivers the final output for performing core encoder selection from ACELP (Algebraic Code Excited Linear Prediction), TCX (Transform Coded Excitation), and GSC (Generic Audio Signal Coder) as described in Reference [7], the entire contents of which are incorporated herein by reference. Other suitable core encoders may also be implemented within the scope of this disclosure.
[0036] In summary, The selected core encoder is then applied to encode the primary (dominant) channel (in the case of TD stereo coding) or the downmixed mono channel (in the case of FD stereo coding). The core encoder selection uses additional high level features, typically computed over a longer window than that used in the first stage of the two stage speech / music classification device method. - The core encoder selection uses its own attack / onset detection logic, which is optimized for seamless switching. The output of this attack / onset detector is different from the output of the first stage attack / onset detector. - A core encoder is initially selected based on the output of the first-stage state-dependent category classifier 210. Such a selection is then refined by examining additional high-level features and the output of this second-stage onset / attack detector.
[0037] FIG. 3 is a schematic block diagram illustrating simultaneously the second stage 300 of a two-stage audio / music classification device and the second stage 350 of a corresponding two-stage audio / music classification method.
[0038] Referring to FIG. 3, the second stage of the two-stage voice / music classification device: - an additional high-level feature extractor 301; an initial selector 302 for the core encoder; and - Core Encoder Initial Selection Improver 303 Equipped with.
[0039] 2. First Stage of Two-Stage Audio / Music Classification Device and Two-Stage Audio / Music Classification Method First, it should be mentioned that the GMM model is trained using the Expectation-Maximization (EM) algorithm on a large, manually labeled database of training samples. The database includes mono items used in the EVS codec and several additional stereo items. The total size of the mono training database is approximately 650 MB. The original mono files are converted to their corresponding dual-mono variants before being used as input to the IVAS codec. The total size of the additional stereo training database is approximately 700 MB. The additional stereo database includes real recordings of speech signals from simulated conversations, music samples downloaded from open sources on the Internet, and several artificially created items. The artificially created stereo items are obtained by convolving mono speech samples with pairs of real binaural room impulse responses (BRIRs). These impulse responses correspond to several typical room configurations, such as a small office, a seminar room, or an auditorium. Labels for the training items are generated semi-automatically using Voice Activity Detection (VAD) information extracted from the IVAS codec. This is not optimal, but manual frame-by-frame labeling is not possible given the size of the database.
[0040] 2.1 State Machine for Signal Section 2, the first stage 250 of the two-stage voice / music classification method comprises a signal segmentation operation 251. To perform this operation, the first stage 200 of the two-stage voice / music classification device comprises a state machine 201.
[0041] The state machine concept in the first stage is brought over from the EVS codec. No significant modifications have been made to the IVAS codec. The purpose of the state machine 201 is to classify the incoming sound signal into one of four states: INACTIVE, ENTRY, ACTIVE, and UNSTABLE.
[0042] FIG. 4 is a schematic block diagram showing simultaneously the state machine 201 of the first stage 200 of the two-stage voice / music classification device and the signal segmentation operation 251 of the first stage 250 of the two-stage voice / music classification method.
[0043] The schematic diagram of FIG. 4 also shows the transition conditions used by the state machine 201 to transition the input sound signal from one of its states to another, these transition conditions relating to characteristics of the input sound signal.
[0044] The INACTIVE state 401, which indicates background noise, is selected as the initial state.
[0045] When the VAD flag 403 (see reference [1]) changes from "0" to "1", the state machine 201 switches from the INACTIVE state 401 to the ENTRY state 402. Any VAD detector or SAD (Sound Activity Detection) detector may be utilized to generate the VAD flag used by the first stage of the two-stage speech / music classification method and device. After an extended period of silence, the ENTRY state 402 marks the first onset or attack in the input sound signal.
[0046] For example, after eight frames 405 in ENTRY state 402, state machine 201 enters ACTIVE state 404, which marks the beginning of a stable sound signal with sufficient energy (for a given level of energy). If the signal's energy 409 suddenly becomes low while state machine 201 is in ENTRY state 402, state machine 201 changes from the ENTRY state to UNSTABLE state 407, which corresponds to an input sound signal with a level of energy approaching background noise. Also, if VAD flag 403 changes from "1" to "0" while state machine 201 is in ENTRY state 402, state machine 201 returns to INACTIVE state 401. This ensures continuity of classification during short interruptions.
[0047] If the energy 406 of the stable signal (ACTIVE state 404) suddenly drops closer to the level of the background noise or the VAD flag 403 changes from "1" to "0", the state machine 201 switches from the ACTIVE state 404 to the UNSTABLE state 407.
[0048] For example, after a period of 12 frames 410 in UNSTABLE state 407, state machine 201 returns to INACTIVE state 401. If the energy 408 of the unstable signal suddenly goes high or the VAD flag 403 changes from '0' to '1' while state machine 201 is in UNSTABLE state 407, state machine 210 returns to ACTIVE state 404. This ensures continuity of classification during short interruptions.
[0049] In the following description, the current state of the state machine 201 is f SM The constants assigned to each state can be defined as follows: INACTIVE f SM =-8 UNSTABLE f SM ∈<-7,-1> ENTRY f SM ∈<0,7> ACTIVE f SM =+8
[0050] In the INACTIVE and ACTIVE states, f SM corresponds to a single constant, but in the UNSTABLE and ENTRY states, f SM takes multiple values depending on the progress of the state machine 201. Therefore, in the UNSTABLE and ENTRY states, f SM can be used as a short-term counter.
[0051] 2.2 Onset / Attack Detector 2, the first stage 250 of the two-stage speech / music classification method comprises an operation of onset / attack detection based on relative frame energy 252. To perform this operation, the first stage 200 of the two-stage speech / music classification device comprises an onset / attack detector 202.
[0052] The onset / attack detector 202 and the corresponding onset / attack detection operation 252 are adapted to the purpose and function of the IVAS codec's speech / music classification. Its purpose includes, but is not limited to, the localization of both the onsets (attacks) of speech utterances and the onsets of music clips. These events are typically associated with abrupt changes in the characteristics of the input sound signal. Successful detection of signal onsets and attacks after periods of signal inactivity allows for the reduction of the influence of past information in the process of score smoothing (described below). The onset / attack detection logic plays a role in the state machine 201 (FIG. 2) similar to the ENTRY state 402 in FIG. 4. The difference between these two concepts relates to their input parameters. The state machine 201 primarily uses the VAD flag 403 (FIG. 4) from the HE-SAD (High Efficiency Sound Activity Detection) technique (see Reference [1]), while the onset / attack detector 252 uses relative frame energy differences.
[0053] Relative frame energy E rmay be calculated as the difference between the frame energy in dB and the long-term average energy. The frame energy in dB is calculated using the following relationship:
number
number
number
number
[0054] The parameter used by the onset / attack detector 252 is the cumulative sum, updated in every frame, of the difference between the relative energy of the input sound signal in the current frame and the relative energy of the input sound signal in the previous frame. This parameter is initialized to 0 and the relative energy E r (n) is the relative energy E in the previous frame r The onset / attack detector 252 is updated only when the onset / attack time is higher than (n-1). v run (n)=v run (n-1)+(E r (n)-E r (n-1)) Using the cumulative sum v run (n), where n is the index of the current frame. The onset / attack detector 252 updates the cumulative sum v run (n) is used to calculate the onset / attack frame counter v cntThe counter of the onset / attack detector 252 is initialized to 0 and incremented by 1 every frame in the ENTRY state 402, except that v run > 5. If not, it is reset to 0.
[0055] The output of the attack / onset detector 202 is a binary flag f att and the binary flag f att For example, 0 to indicate onset / attack detection. <v run <3, then it is set to 1. Otherwise, this binary flag is set to 0 to indicate non-detection of onset / attack. This can be expressed as:
number
[0056] The graph of FIG. 5 demonstrates the operation of the onset / attack detector 202 as a non-limiting example.
[0057] 2.3 Feature Extractor 2, the first stage 250 of the two-stage speech / music classification method comprises the operation of extraction of features of the input sound signal 253. To perform this operation, the first stage 200 of the two-stage speech / music classification device comprises a feature extractor 203.
[0058] In the training stage of the GMM model, training samples are resampled to 16 kHz, normalized to -26 dBov (dBov is the dB level relative to the overload point of the system), and concatenated. The resampled and concatenated training samples are then fed to the encoder of the IVAS codec for collecting features using a feature extractor 203. For the purpose of feature extraction, the IVAS codec may be run in FD stereo coding mode, TD stereo coding mode, or any other stereo coding mode, and at any bit rate. As a non-limiting example, the feature extractor 203 runs in TD stereo coding mode at 16.4 kbps. The feature extractor 203 extracts the following features used in the GMM model for speech / music / noise classification:
[0059] [Table 1]
[0060] All of the above features, except for the MFCC features, are already present in the EVS codec (see reference [1]).
[0061] The feature extractor 203 extracts the open-loop pitch T OL and vocalization measures
number
[0062] MFCC features are N corresponding to Mel-frequency cepstral coefficients. mel is a vector of values, and the Mel-frequency cepstral coefficients are the result of the real logarithmic cosine transform of the short-term energy spectrum, expressed on the Mel-frequency scale (see reference [8], the entire contents of which are incorporated herein by reference).
[0063] The last two features diff and P sta The calculation of is, for example,
number
number
[0064] Power spectrum difference P diff teeth,
number
[0065] Spectral stationarity feature P stais expressed by the following relation:
number
[0066] 2.4 Outlier detector based on individual feature histograms 2, the first stage 250 of the two-stage speech / music classification method comprises an operation 254 of detecting outlier features based on the individual feature histograms. To perform operation 254, the first stage 200 of the two-stage speech / music classification device comprises an outlier detector 204.
[0067] The GMM model is trained on a vector of features collected from an IVAS codec on a large training database. The accuracy of the GMM model is affected to a significant extent by the statistical distribution of the individual features. Best results are achieved when the features are normally distributed, e.g., X~N(μ,σ), where N represents a statistical distribution with mean μ and variance σ. Figure 6 shows histograms of selected features on the large training database. As can be seen, the histograms of some features in Figure 6 do not indicate that they are drawn from a normal distribution.
[0068] GMM models can represent features with non-normal distributions to some extent. If the value of one or more features differs significantly from its mean value, the feature vector is determined to be an outlier. Outliers usually lead to incorrect probability estimates. Rather than discarding the feature vector, it is possible to replace the outlier features, for example, with feature values from the previous frame, average feature values over several previous frames, or global mean values over a significant number of previous frames.
[0069] The detector 204 detects outliers in the first stage 200 of the two-stage speech / music classification device based on the analysis of individual feature histograms computed on the training database (see, e.g., FIG. 6 showing feature histograms and FIG. 7 showing a graph illustrating outlier detection based on histogram values). For each feature, e.g., the following relationship:
number
number
number
[0070] f xs (x|0,σ 2 ) to the threshold thr H By substituting and rearranging the variables, the following relations are obtained:
number
[0071] thr H =1e -4For , the following is obtained: x≒2.83σ
[0072] Therefore, 1e -4 Applying the threshold value f xs (0|0,σ 2 This results in truncating the probability density function to a range of ±2.83σ around the mean, provided that the data is scaled so that σ = 1. The probability of a feature value being outside the truncated range is given by, for example, the following relation:
number
[0073] If the variance of the feature values is σ=1, the percentage of detected outliers will be approximately 0.47%. The above calculation is only an approximation because the true distribution of the feature values is not normal. This is evident from the non-stationary feature n in Figure 6. sta where the tail to the right of the mean is "heavier" than the tail to the left of the mean. If the sample variance σ is used as the criterion for outlier detection and the interval is set to, say, ±3σ, then many of the "good" values to the right of the mean will be classified as outliers.
[0074] Lower limit H low and upper bound H high is calculated for each feature used by the first stage 250 / 200 of the two-stage speech / music classification method and device and stored in the memory of the IVAS codec. When the IVAS codec encoder is executed, the outlier detector 204 calculates the value X of each feature j in the current frame n. j (n) is the boundary H of the feature low and H highand marking a feature j having a value outside the corresponding range defined between the lower and upper bounds as an outlier feature.
number
number
[0075] If the number of outlier features is, for example, equal to or greater than 2, the outlier detector 204 sets a binary flag f out is set to 1. This can be expressed as follows:
number
[0076] Flag F out is used to signal that a vector of features is an outlier. The flag f out If is equal to 1, the outlier feature X j (n) is replaced with the value from the previous frame, for example: f odv If (j)=1, then for j=1,..,F, X j (n)=X j (n-1)
[0077] 2.5 Short-term feature vector filter 2, the first stage 250 of the two-stage speech / music classification method comprises a short-term feature vector filtering operation 255. To perform operation 255, the first stage 200 of the two-stage speech / music classification device comprises a short-term feature vector filter 205 for smoothing the short-term vector of extracted features.
[0078] The speech / music classification accuracy is improved using feature vector smoothing. This is achieved by using the following short-term infinite impulse response (IIR) filter as the short-term feature vector filter 205:
number
number
[0079] To avoid smearing effects on strong attacks or outliers at the beginning of the ACTIVE signal segment, where the informative potential of feature vectors in previous frames is limited, feature vector smoothing (operation 255 of filtering short-term feature vectors) is performed by att =1 or f out = 1. Smoothing is also not performed in the ACTIVE state 404 (stable signal) of Figure 4 in frames classified as onset / transient by the IVAS transient classification algorithm (see reference [1]). When the short-term feature vector filtering operation 255 is not performed, the feature values X of the unfiltered vector j (n) is simply copied over and used. This can be expressed by the following relationship:
number
[0080] In the following description,
number
number
[0081] 2.6 Nonlinear Feature Vector Transformation (Box-Cox) 2, the first stage 250 of the two-stage speech / music classification method comprises a non-linear feature vector transformation operation 256. To perform operation 256, the first stage 200 of the two-stage speech / music classification device comprises a non-linear feature vector transformer 206.
[0082] As shown by the histogram in Figure 6, the features used in speech / music classification are not normally distributed, and as a result, the best accuracy of the GMM cannot be achieved. As a non-limiting example, the nonlinear feature vector converter 206 can use a Box-Cox transformation, such as that described in Reference [9], the entire contents of which are incorporated herein by reference, to convert the non-normal features into features with a normal shape. The Box-Cox transformation X of feature X box is as follows, i.e.,
number
number
[0083] During the training process, the nonlinear feature vector converter 206 considers and tests all values of the exponent λ to select the optimal value of the exponent λ based on a normality test. The normality test is based on the method of D'Agostino and Pearson, as described in Reference
[10] , the entire contents of which are incorporated herein by reference, and combines the skew and kurtosis of the probability distribution function. The normality test uses the following skew and kurtosis measures r sk (SK measure) r sk =s 2 +k 2 where s is the z-score returned by the skew test and k is the z-score returned by the kurtosis test. For more information about skew and kurtosis tests, see reference
[11] , the entire contents of which are incorporated herein by reference.
[0084] The normality test also returns a two-sided chi-squared probability for the null hypothesis, i.e., that the feature values are drawn from a normal distribution. The optimal value of the exponent λ minimizes the SK measure. This is due to the following relationship:
number
[0085] In the encoder, the nonlinear feature vector transformer 206 satisfies the following condition related to the SK measure:
number
number
[0086] In the following description, X box,j Instead of (n), the feature value X j The original symbol for (n) is used, i.e., For the selected features, X j (n)←X box,j (n) It is assumed that this is the case.
[0087] 2.7 Principal component analyzer 2, the first stage 250 of the two-stage speech / music classification method comprises a principal component analysis (PCA) operation 257 to reduce sound signal feature dimensionality and increase sound signal class distinctiveness. To perform operation 257, the first stage 200 of the two-stage speech / music classification device comprises a principal component analyzer 207 of the principal components.
[0088] After the short-term feature vector filtering operation 255 and the nonlinear feature vector transformation operation 256, the principal component analyzer 207 standardizes the feature vectors by removing the mean of the features and scaling them to unit variance. To that end, the following relationship is used:
number
number
[0089] Feature X j The average μ j and deviation s j is as follows, i.e.,
number
number
[0090] In the following description,
number
number
[0091] The principal component analyzer 207 then processes the feature vectors using PCA, where the dimensionality is increased from, for example, F=15 to F=15. PCA = 12. PCA is an orthogonal transformation that transforms a set of possibly correlated features into a set of linearly uncorrelated variables called principal components (see reference
[12] , the entire contents of which are incorporated herein by reference). In a speech / music classification method, analyzer 207 may, for example, use the following relationship:
number
number
number
number
number
[0092] In the following description,
number
number
number
[0093] 2.8 Gaussian Mixture Models (GMM) 2, the first stage 250 of the two-stage speech / music classification method comprises a Gaussian Mixture Model (GMM) calculation operation 258. To perform operation 258, the first stage 200 of the two-stage speech / music classification device comprises a GMM calculator 208. As can be seen, the GMM calculator 208 estimates a decision bias parameter by maximizing harmonic balance accuracy on the training database. The decision bias is a parameter that is added to the GMM to improve the accuracy of the "MUSIC" class decision due to insufficient training data.
[0094] A multivariate GMM is parameterized by a mixture of component weights, component means, and covariance matrices. The speech / music classification method uses three GMMs, each trained on its own training database: a "speech" GMM, a "music" GMM, and a "noise" GMM. In a GMM with K components, each component has its own mean.
number
number
number
number
number
number
[0095] To reduce the complexity of the probability calculations, the above relation may be simplified by taking the logarithm of the inner term inside the summation term Σ as follows:
number
[0096] The output of the above simplified equation is called the "score." The score is an unbounded variable proportional to the log-likelihood. The larger the score, the greater the probability that a given feature vector was generated by the GMM. Scores are calculated by the GMM calculator 208 for each of the three GMMs. Score on the "Audio" GMM
number
number
number
number
number
number
number
[0097]
number
number
[0098] After EM training, i.e., when the parameters of the GMM are known, the difference score
number
number
number
[0099] The accuracy of this binary predictor is measured using four statistical measures:
number
number
number
number
[0100] The statistic defined above is the true positive rate, usually called recall.
number
number
number
[0101] Decision bias b s The values of y are the labels / predictors pred (n), where b s is selected from the interval (-2, 2) in successive steps. The spacing of candidate values for the decision bias is approximately logarithmic, with values of higher concentration around 0.
[0102] Decision bias b s The difference score calculated using the found values of
number
number
[0103] 2.9 Adaptive Smoother 2, the first stage 250 of the two-stage speech / music classification method comprises an adaptive smoothing operation 259. To perform operation 259, the first stage 200 of the two-stage speech / music classification device comprises an adaptive smoother 209.
[0104] The adaptive smoother 209 calculates the difference score dlp(X, b s ) is an adaptive IIR filter for smoothing the signal. The adaptive smoothing or filtering operation 259 includes the following operations: wdlp(n)=wght(n)·wdlp(n-1)+(1-wght(n))·dlp(n) where wdlp(n) is the resulting smoothed difference score, wght(n) is the so-called forgetting factor of the adaptive IIR filter, and n represents the frame index.
[0105] The forgetting factor is the product of three individual parameters as shown in the following relationships: wght(n)=wrelE(n) · wdrop(n) · wrise(n)
[0106] The parameter wrelE(n) is the relative energy E of the current frame. r It is linearly proportional to (n) and may be calculated using the following relationship:
number
[0107] The parameter wdrop(n) is proportional to the derivative of the difference score dlp(n). First, the short-term average dlp ST (n) is calculated, for example, using the following relationship: dlp ST (n)=0.8·dlp ST (n-1)+0.2·dlp(n)
[0108] The parameter wdrop(n) is set to 0 and is only modified in frames where the following two conditions are met: dlp(n)<0 dlp(n) <dlp ST (n)
[0109] Therefore, the adaptive smoother 209 updates the parameter wdrop(n) only when the difference score dlp(n) has a decreasing trend and when the difference score dlp(n) indicates that the current frame belongs to the SPEECH class. ST If (n)>0, the parameter wdrop(n) is wdrop(n)=-dlp(n) is set to
[0110] If not, the adaptive smoother 209 steadily increases the parameter wdrop(n), for example, using the following relationship: wdrop(n)=wdrop(n-1)+(dlp ST (n-1)-dlp(n))
[0111] If the two conditions defined above are not true, the parameter wdrop(n) is reset to 0. Thus, the parameter wdrop(n) reacts to a sudden drop in the difference score dlp(n) below the 0 level, which indicates a potential speech onset. The final value of the parameter wdrop(n) is linearly mapped to the interval, for example (0.7, 1.0), as shown in the following relation:
number
[0112] Note that for simplicity of notation, the value of wdrop(n) is "overwritten" in the above equation.
[0113] The adaptive smoother 209 calculates the parameter wrise(n) as well as the parameter wdrop(n) using the difference where the parameter wdrop(n) responds to sudden rises in the difference score dlp(n), which indicates potential music onsets. The parameter wrise(n) is set to 0 but is modified in frames that satisfy the following condition: f SM (n)=8(ACTIVE) dlp ST (n)>0 dlp ST (n)>dlp ST (n-1)
[0114] Therefore, when the difference score dlp(n) has an increasing trend and indicates that the current frame n belongs to the MUSIC class, the adaptive smoother 209 updates the parameter wrise(n) only in the ACTIVE state 404 of the input sound signal (see Figure 4).
[0115] In the first frame, when the above three specified conditions are met, and the short-term average dlp ST If (n-1)<0, the third parameter wrise(n) is wrise(n)=-dlp ST (n) is set to
[0116] If not, the adaptive smoother 209 may, for example, use the following relationship: wrise(n)=wrise(n-1)+(dlp ST (n)-dlp ST (n-1)) According to this, the parameter wrise(n) is steadily increased.
[0117] If the above three conditions are not true, the parameter wrise(n) is reset to 0. Thus, the third parameter wrise(n) reacts to sudden rises in the difference score dlp(n) above the 0 level, which indicate potential musical onsets. The final value of the parameter wrise(n) is linearly mapped to the interval, for example (0.95, 1.0), as follows:
number
[0118] Note that for simplicity of notation, the value of the parameter wrise(n) is "overwritten" in the above equation.
[0119] 9 is a graph showing, by way of non-limiting example, the behavior of the parameters wdrop(n) and wrise(n) for a short segment of a speech signal with background music. The peak of the parameter wdrop(n) is typically located near the speech onset, while the peak of the parameter wrise(n) is generally located where the speech slowly backs off and the background music begins to dominate the signal content.
[0120] The forgetting factor wght(n) of the adaptive IIR filter of adaptive smoother 209 is decreased in response to strong SPEECH or MUSIC signal content. To that end, adaptive smoother 209 reduces the long-term average of the difference score dlp(n), calculated, for example, using the following relationship:
number
number
number
number
[0121] In the ENTRY state 402 (Figure 4) of the input sound signal, the long-term average is
number
number
number
[0122] formula r m2v (n) corresponds to the long-term standard deviation of the difference score. The forgetting factor wght(n) of the adaptive IIR filter of the adaptive smoother 259 is calculated by, for example, using the following relation: m2v It is made smaller in frames where (n)>15. wght(n)←0.9 wght(n)
[0123] The final value of the forgetting factor wght(n) of the adaptive IIR filter of the adaptive smoother 209 is limited to the range (0.01, 1.0), for example. tot For frames where (n) is less than 10 dB, the forgetting factor wght(n) is set to, for example, 0.92, which ensures proper smoothing of the difference score dlp(n) during silence.
[0124] The filtered and smoothed difference score wdlp(n) is a parameter for category determination in the speech / music classification method, as explained below.
[0125] 2.10 State-Dependent Category Classifiers 2, the first stage 250 of the two-stage speech / music classification method comprises an operation 260 of state-dependent categorization of an input sound signal as a function of the difference score distribution and a direction-dependent threshold. To perform operation 260, the first stage 200 of the two-stage speech / music classification device comprises a state-dependent category classifier 210.
[0126] Operation 260 is the final operation of the first stage 250 of the two-stage speech / music classification method and comprises categorizing the input sound signal into three final classes: SPEECH / NOISE (0) UNCLEAR (1) MUSIC (2)
[0127] In the above, the numbers in parentheses are numeric constants associated with the three final classes. The above set of classes differs slightly from the classes described so far in terms of difference scores. The first difference is that the SPEECH and NOISE classes are combined. This is to facilitate the core encoder selection mechanism (described in the following description) in which an ACELP encoder core is typically selected to code both speech signals and background noise. A new class, the UNCLEAR final class, has been added to the set. Frames that fall into this category are typically found in speech segments with high levels of additive background music. The smoothed difference scores wdlp(n) of frames in the UNCLEAR class are mostly close to 0. Figure 10 is a graph showing the distribution of smoothed difference scores wdlp(n) for the training database and their relationship to the final classes SPEECH / NOISE, UNCLEAR, and MUSIC.
[0128] The final class selected by the state-dependent category classifier 210 is denoted by d. SMC (n) shall be as indicated.
[0129] When the input sound signal is in the ENTRY state 402 (see FIG. 4) in the current frame, the state-dependent category classifier 210 determines the final class d based on the weighted average of the difference scores dlp(n) calculated in the frames preceding the current frame that belong to the ENTRY state 402. SMC (n) is selected. The weighted average is calculated using the following equation:
number
[0130] Absolute Frame Energy E tot If, for example, dlp(n) is lower than 10 dB in the current frame, the state-dependent category classifier 210 determines the final class d SMC Set (n) to SPEECH / NOISE. This is to avoid misclassification during silence.
[0131] Weighted average of difference scores in the ENTRY state (wdlp) ENTRY If (n) is less than, for example, 2.0, the state-dependent category classifier 210 selects the final class d SMC Set (n) to SPEECH / NOISE.
[0132] Weighted average of difference scores in the ENTRY state (wdlp) ENTRY If (n) is greater than, for example, 2.0, the state-dependent category classifier 210 determines the final class d based on the unsmoothed difference score dlp(n) in the current frame. SMC Set dlp(n). If dlp(n) is greater than, say, 2.0, then the final class is MUSIC. Otherwise, the final class is UNCLEAR.
[0133] In other states of the input sound signal (see FIG. 4), the state-dependent category classifier 210 selects a final class in the current frame based on the smoothed difference score wdlp(n) and the final class selected in the previous frame. The final class in the current frame is first initialized to the class from the previous frame, i.e., d SMC (n)=d SMC (n-1) is.
[0134] If the smoothed difference score wdlp(n) crosses a threshold (see Table 3) for a class different from the class selected in the previous frame, the decision may be changed by the state-dependent category classifier 210. These transitions between classes are shown in FIG. 10. For example, if the final class d selected in the previous frame SMC If (n) is SPEECH / NOISE, but the smoothed difference score wdlp(n) at the current frame is greater than, for example, 1.0, then the final class d at the current frame is SMC(n) is changed to UNCLEAR. The graph in Figure 10 shows a histogram of the smoothed difference score wdlp(n) for the SPEECH / NOISE and MUSIC final classes, calculated on the training database excluding INACTIVE frames. As can be seen from the graph in Figure 10, there are two sets of thresholds, one for the SPEECH / NOISE->UNCLEAR->MUSIC transition and the other for the opposite direction, i.e., MUSIC->UNCLEAR->SPEECH / NOISE transition. The final class d SMC There is no switching of (n). The value of the decision threshold indicates that the state-dependent category classifier 210 prefers the SPEECH / NOISE final class. Examples of transitions between classes and associated thresholds are summarized in Table 3 below. [Table 3]
[0135] As noted hereinabove, transitions between classes are driven not only by the value of the smoothed difference score wdlp(n), but also by the final class selected in the previous frame. The complete set of rules for transitions between classes is shown in the class transition diagram of Figure 12.
[0136] The arrows in Figure 12 indicate the direction in which the class may be changed if the condition inside the corresponding diamond is satisfied. In case of multiple conditions inside a diamond, a logical AND is assumed between them, i.e., all must be satisfied for the transition to occur. If an arrow is qualified by the notation "≥ X frames", it means that the class may only be changed after at least X frames. This adds a short hysteresis to some transitions.
[0137] In Figure 12, the symbol f spindicates the short pitch flag, a by-product of the stable high pitch analysis module of the IVAS codec (see reference [1]). The short pitch flag indicates a high value of the voicing measure.
number
number
number
[0138] Short Pitch Flag f sp is as follows, i.e.,
number
number
number
number
number
number
number
number
[0139] In Figure 12, the parameter c VAD is the counter of ACTIVE frames. Counter c VAD is initialized to 0 and reset to 0 in all frames where the VAD flag is 0. VAD increases by 1 only in frames where the VAD flag is 1, until a threshold of, for example, 50 is reached or the VAD flag returns to 0.
[0140] Parameter v run (n) is defined in Section 2.2 (Onset / Attack Detection) of this disclosure.
[0141] 3. Core Encoder Selection FIG. 3 is a schematic block diagram illustrating simultaneously the second stage 300 of a two-stage audio / music classification device and the second stage 350 of a corresponding two-stage audio / music classification method.
[0142] The final class d selected by the state-dependent category classifier 210 in the second stage 350 / 300 of the two-stage speech / music classification method and device SMC(n) is "mapped" to one of the three core encoder technologies of the IVAS codec: ACELP (Algebraic Code Excited Linear Prediction), GSC (Generalized Audio Signal Coding), or TCX (Transform Coded Excitation). This is called three-way classification. This does not guarantee that the selected technology will be used as the core encoder, as there are other factors that influence the decision, such as bit rate or bandwidth limitations. However, for common types of input sound signals, the initial selection of the core encoder technology will be used.
[0143] The class d selected by the state-dependent category classifier 210 in the first stage SMC Besides (n), the core encoder selection mechanism takes into account some additional high-level features.
[0144] 3.1 Additional High-Level Feature Extractors 3, the second stage 350 of the two-stage speech / music classification method comprises an operation of extraction of additional high-level features of the input sound signal 351. To perform operation 351, the second stage 300 of the two-stage speech / music classification device comprises an additional high-level feature extractor 301.
[0145] In the first stage 200 / 250 of the two-stage speech / music classification device and method, most features are calculated over short segments (frames) of the input sound signal, typically not exceeding 80 ms. This allows for rapid response to events, such as speech onsets or offsets in the presence of background music. However, it also leads to a relatively high rate of misclassification. While misclassification can be mitigated to some extent using adaptive smoothing, as described in Section 2.9 above, for some types of signals, this is not efficient enough. Therefore, as part of the second stage 300 / 350 of the two-stage speech / music classification device and method, class d is used to select the most appropriate core encoder technology for some types of signals.SMC (n) may be modified. To detect such types of signals, the detector computes additional high-level features and / or flags, usually for longer segments of the input signal.
[0146] 3.1.1 Long-term signal stability Long-term signal stability is a feature of the input sound signal that can be used for successful discrimination between vocals from opera and music. In the context of core encoder selection, signal stability is understood as segmental long-term stationarity with high autocorrelation. An additional high-level feature extractor 301 extracts a "voicing" measure
number
number
number
number
number
[0147] To obtain greater robustness, the vocalization parameters at the current frame n are
number
number
[0148] Smoothed vocalization parameters cor LT (n) is sufficiently large and the variance of the vocalization parameters is var If (n) is small enough, the input signal is considered "stable" for the purposes of core encoder selection. This is because the value cor LT (n) and cor var It is measured by comparing (n) to a predetermined threshold and setting a binary flag, for example, using the following rules:
number
[0149] Binary flag f STAB (n) is an indicator of long term signal stability and is used in the core encoder selection described later in this disclosure.
[0150] 3.1.2 Segmentation Attack Detection The extractor 301 extracts segment attack features from some, for example, 32 short segments of the current frame n, as shown in FIG.
[0151] For each segment, an additional high-level feature extractor 301 extracts, for example,
number
number
[0152] The additional high-level feature extractor 301 extracts the energy E of the input signal s(n) from the beginning (segment 0) to 3 / 4 (segment 24) of the current frame n. ata The attack of the current frame n (segment k=k) is calculated based on the average of (k) (denominator of the equation below). ata ) to the end (segment 31) of the input sound signal s(n), E ata By comparing the average of (k) (the numerator in the equation below), the strength of the attack is ata Estimate the strength str ata This estimation is done, for example, using the following relation:
number
[0153] str ata If the value of is, for example, greater than 8, the attack is considered strong enough and the segment k ata is used as an indicator to signal the location of the attack inside the current frame n. Otherwise, indicator k ata is set to 0, indicating that no attack was identified. Attacks are only detected in the GENERIC frame type signaled by the IVAS frame type selection logic (see [1]). To reduce false attack detections, the segment k where an attack was identified is selected, for example, using the following relationship: ata Energy E ata (k ata ) is the first
number
number
[0154] Comparison value str for segment k=2,..,213_4 (k) is, for example, 2(k≠k ata ), then k ata is set to 0, indicating that no attack was identified. In other words, the energy of the segment containing the attack is
number
[0155] The mechanism described above is mainly for the last
number
[0156] For unvoiced frames, which are classified as UNVOICED_CLAS, UNVOICED_TRANSITION, or ONSET by the IVAS FEC classification module (see reference [1]), the additional high-level feature extractor 301 may, for example, use the relation
number
[0157] In the above equation, a negative index in the denominator indicates the segmental energy E ata (k) value. The strength str calculated using the above formula ataIf, for example,,is larger than 16, the attack is strong enough,and,k, ata is used to signal the location of the attack inside the current frame; otherwise, k ata is set to 0, indicating that no attack was identified. If the last frame was classified as UNVOICED_CLAS by the IVAS FEC classification module, the threshold is set to, for example, 12 instead of 16.
[0158] For unvoiced frames, classified as UNVOICED_CLAS, UNVOICED_TRANSITION, or ONSET by the IVAS FEC classification module (see reference [1]), there is another condition that must be met in order for the detected attack to be considered strong enough:
number
number
number
number
[0159] If an attack has already been detected in the previous frame, k ata is reset to 0 at the current frame n to prevent attack smearing effects.
[0160] For other frame types (excluding UNVOICED and GENERIC as explained above), the additional high level feature extractor 301 may, for example, use the following ratio:
number
[0161] Therefore, for segment attack detection, the final output of the additional high-level feature detector 301 is the index k=k of the segment containing the attack. ata or k ata = 0. If the index is positive, the attack is detected; otherwise, the attack is not identified.
[0162] 3.1.3 Signal Tonality Estimation The tonality of the input sound signal in the two-stage speech / music classification device and the second stage of the two-stage speech / music classification method is represented as a tonality binary flag that reflects both the spectral stability and the harmonicity in the lower frequency range of the input signal up to 4 kHz. The additional high-level feature extractor 301 extracts the correlation map S, which is a by-product of the sound stability analysis in the IVAS encoder. map From (n,k), we calculate this tonality binary flag (see reference [1]).
[0163] The correlation map is a measure of both signal stability and harmonicity. The correlation map is calculated from the first, say, 80 bins E of the residual energy spectrum in the logarithmic domain. dB,res (k) (k=0,..,79) (see reference [1]). The correlation map is calculated at segments of the residual energy spectrum where peaks exist. These segments are determined by the parameter i min (p), where p=1,...,N min is the segment index, and N min is the total number of segments.
[0164] The set of indices belonging to a particular segment x is PK(p)={i|i≧i min (p) and i min (p+1) and i<80} Then the correlation map may be calculated as follows:
number
[0165] Correlation Map M cor (PK(p)) is, for example, the following two relations:
number
number
[0166] S mass thr mass If greater than, the tonality flag f ton is set to 1. Otherwise, it is set to 0.
[0167] 3.1.4 Spectral peak-to-average ratio Another high-level feature used in the core encoder selection mechanism is the spectral peak-to-mean ratio. This feature is a measure of the spectral sharpness of the input sound signal s(n). The extractor 301 extracts the power spectrum S of the input signal s(n) in the logarithmic domain, for example in the range 0 to 4 kHz. LT (n,k) (k=0,...,79), where the power spectrum S LT (n,k) is, for example, the following relation
number
number
[0168] 3.2 Core Encoder Initial Selector 3, the second stage 350 of the two-stage speech / music classification method comprises an initial core encoder selection operation 352. To perform operation 352, the second stage 300 of the two-stage speech / music classification device comprises a core encoder initial selector 302.
[0169] The initial selection of the core encoder by selector 302 is based on (a) the relative frame energy E r , (b) the final class d selected in the first stage of the two-stage speech / music classification device and two-stage speech / music classification method. SMC (n), and (c) additional high-level features r p2a (n), S mass , and thr mass The selection mechanism used by the core encoder initial selector 302 is shown in the schematic diagram of FIG.
[0170] "0" represents ACELP technology, "1" represents GSC technology, and "2" represents TCX technology, d core Let ∈{0,1,2} denote the core encoder technique selected by the mechanism in Figure 14. Thus, the initial selection of the core encoder technique is determined by the final class d from the first stage of the two-stage speech / music classification device and method. SMC (n) Strictly follow the allocation. An exception concerns signals with strong tones where TCX technology should be selected, as it yields better quality.
[0171] 3.3 Core Encoder Selection Improver 3, the second stage 350 of the two-stage speech / music classification method comprises an operation of refining the initial selection of a core encoder 353. To perform operation 353, the second stage 300 of the two-stage speech / music classification device comprises a core encoder selection refiner 303.
[0172] d core = 1, i.e. when the GSC core encoder is initially selected for core coding, the core encoder selection refiner 303 may change the core encoder technique. This situation may occur, for example, for musical items classified as MUSIC, which have low energy below 400 Hz. The affected segments of the input signal have an energy ratio of
number
[0173] The sum in the numerator extends over the first eight frequency bins of the energy spectrum, which corresponds to the frequency range 0 to 400 Hz. The Core Encoder Selection Refinement 303 selects the energy ratios rat in frames previously classified as MUSIC with a reasonably high degree of certainty. LF The core encoder technology calculates and analyzes the following conditions:
number
[0174] For signals with very short and stable pitch periods, GSC is not the optimal core coder technique. sp= 1, the core encoder selection improver 303 changes the core encoder technique from GSC to ACELP or TCX as follows:
number
[0175] Highly correlated signals with only small energy variations are another type of signal for which the GSC core encoder technique is not suitable. For these signals, the core encoder selection improver 303 switches the core encoder technique from GSC to TCX. As a non-limiting example, this change in core encoder may be performed under the following conditions:
number
number
[0176] Finally, in a non-limiting example, the core encoder selection refiner 303 may change the initial core encoder selection from GSC to ACELP in a frame in which an attack is detected, if the following condition is met:
number
[0177] The above condition ensures that this change of the core encoder from GSC to ACELP occurs only in segments with rising energy. If the above condition is met and at the same time, the transition frame counter TC cnt If is set to 1 (reference document [1]), the core encoder selection improver 303 changes the core encoder to ACELP, i.e.,
number
[0178] As explained in Section 3.1.2 above, if an attack is detected by the segmentation attack detection procedure of the additional high-level feature detection operation 351, the index (position) k of this attack ata is further investigated. If the location of the detected attack is within the last subframe of frame n, the core encoder selection improver 303 changes the core encoder technique to ACELP, for example, when the following condition is met:
number
[0179] This means that the attacks are coded using the TRANSITION mode of the ACELP core encoder.
[0180] If the detected attack location is not within the last subframe but is at least beyond the first quarter of the first subframe, the core encoder selection is not changed and the attack is encoded using the GSC core encoder. As in the previous case, a new attack "flag" f ata may be set as follows: f no_GSC =1 and TC cnt ≠ 1 and k ata If >4, f ata =kata +1
[0181] Parameter k ata is intended to reflect the location of the detected attack, and therefore the attack flag f ata is somewhat redundant, but it is used in this disclosure for consistency with other documents and with the IVAS codec source code.
[0182] Finally, the core encoder selection refiner 303 changes the frame type from GENERIC to TRANSITION in speech frames for which the ACELP core coder technique was selected during the initial selection. This situation applies if the local VAD flag is set to 1 and an attack has been detected in it by the segmentation attack detection procedure of the additional high-level feature detection operation 351 described in section 3.1.2, i.e., k ata >0, occurs only in the active frame.
[0183] The attack flags are now similar to those in the previous situation: f ata =k ata +1 is.
[0184] 4. Exemplary Configurations of Hardware Components FIG. 15 is a simplified block diagram of an exemplary configuration of hardware components forming the above-described IVAS codec including a two-stage voice / music classification device.
[0185] The IVAS codec including two-stage audio / music classification device may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The IVAS codec including two-stage audio / music classification device (identified as 1500 in FIG. 15) comprises an input 1502, an output 1504, a processor 1506, and a memory 1508.
[0186] The input unit 1502 is configured to receive an input sound signal s(n), e.g., in the case of an IVAS codec encoder, the left and right channels of an input stereo sound signal in digital or analog form. The output unit 1504 is configured to provide an encoded multiplexed bitstream in the case of an IVAS codec encoder. The input unit 1502 and the output unit 1504 may be implemented in a common module, e.g., a serial input / output device.
[0187] The processor 1506 is operatively connected to the input 1502, to the output 1504, and to the memory 1508. The processor 1506 is implemented as one or more processors for executing code instructions supporting the functionality of the various elements and operations of the IVAS codec described above, including the two-stage speech / music classification device and two-stage speech / music classification method, as shown in the accompanying drawings and / or described in this disclosure.
[0188] The memory 1508 may comprise non-transitory memory for storing code instructions executable by the processor 1506, in particular processor-readable memory that stores non-transitory instructions that, when executed, cause the processor to implement the elements and operations of the IVAS codec, including the two-stage voice / music classification device and two-stage voice / music classification method. The memory 1508 may also comprise random access memory or buffers for storing intermediate processed data from the various functions performed by the processor 1506.
[0189] Those skilled in the art will appreciate that the description of the IVAS codec, including the two-stage voice / music classification device and two-stage voice / music classification method, is illustrative only and is not intended to be limiting in any way. Other embodiments will readily suggest themselves to such skilled artisans having the benefit of this disclosure. Furthermore, the disclosed IVAS codec, including the two-stage voice / music classification device and two-stage voice / music classification method, may be customized to provide useful solutions to existing needs and problems of encoding and decoding sound, e.g., stereo sound.
[0190] For clarity, not all of the routine features of an implementation of an IVAS codec including a two-stage audio / music classification device and method have been shown and described. Of course, it will be appreciated that in developing any such actual implementation of an IVAS codec including a two-stage audio / music classification device and method, numerous implementation-specific decisions may need to be made to achieve the developer's particular goals, such as meeting application-related, system-related, network-related, and business-related constraints, and that these particular goals will vary from implementation to implementation and developer to developer. Moreover, it will be appreciated that the development effort may be complex and time-consuming, but will nevertheless be a routine undertaking of engineering for one of ordinary skill in the art of sound processing having the benefit of this disclosure.
[0191] In accordance with this disclosure, the elements, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that devices of a less general-purpose nature, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. When a method comprising a series of operations and sub-operations is performed by a processor, computer, or machine, and the operations and sub-operations can be stored as a series of non-transitory code instructions readable by a processor, computer, or machine, they may be stored on a tangible and / or non-transitory medium.
[0192] The elements and processing operations of the IVAS codec, including the two-stage audio / music classification device and two-stage audio / music classification method as described herein, may comprise software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0193] In the IVAS codec, including the two-stage speech / music classification device and method, various processing operations and sub-operations may be performed in various orders, and some of the processing operations and sub-operations may be optional.
[0194] While the present disclosure has been described above with respect to its non-limiting exemplary embodiments, these embodiments may be freely modified within the scope of the appended claims without departing from the spirit and essence of the present disclosure.
[0195] References This disclosure cites the following references, the entire contents of which are incorporated herein by reference: [1] 3GPP TS 26.445, v.12.0.0, "Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description", Sep 2014. [2] M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robillard, J. Lecompte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefevbre, P. Gournay, et al., "The ISO / MPEG Unified Speech and Audio Coding Standard - Consistent High Quality for All Content Types and at All Bit Rates", J. Audio Eng. Soc., vol. 61, no. 12, pp. 956-977, Dec. 2013. [3] F. Baumgarte, C. Faller, "Binaural cue coding - Part I: Psychoacoustic fundamentals and design principles," IEEE Trans. Speech Audio Processing, vol. 11, pp. 509-519, Nov. 2003. [4] Tommy Vaillancourt, "Method and system using a long-term correlation difference between left and right channels for time domain down mixing a stereo sound signal into primary and secondary channels," PCT Application WO2017 / 049397A1. [5] 3GPP SA4 contribution S4-170749 "New WID on EVS Codec Extension for Immersive Voice and Audio Services", SA4 meeting #94, June 26-30, 2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip [6] V. Malenovsky, T. Vaillancourt, W. Zhe, K. Choo and V. Atti, "Two-stage speech / music classifier with decision smoothing and sharpening in the EVS codec," 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, 2015, pp. 5718-5722. [7] T. Vaillancourt and M. Jelinek, "Coding generic audio signals at low bitrates and low delay", U.S. Patent No. 9,015,038 B2. [8] K.S. Rao and A.K. Vuppala, Speech Processing in Mobile Environments, Appendix A: MFCC features, Springer International Publishing, 2014 [9] Box, G. E. P. and Cox, D. R. (1964). An analysis of transformations, Journal of the Royal Statistical Society, Series B, 26, 211-252.
[10] D'Agostino, R. and Pearson, ES (1973), "Tests for departure from normality", Biometrika, 60, 613-622.
[11] D'Agostino, AJ Belanger and RB D'Agostino Jr., "A suggestion for using powerful and informative tests of normality", American Statistician 44, pp. 316-321, 1990.
[12] I. Jolliffe, Principal component analysis. New York: Springer Verlag, 2002. [Explanation of symbols]
[0196] 100 Stereo Sound Processing and Communications System 101 Communication Links 102 microphones 103 Left Channel 104 Analog-to-Digital (A / D) Converter 105 left channel 106 Stereo Sound Encoder 107 Bitstream 108 Error Correction Encoder 109 Error Correction Decoder 110 Stereo Sound Decoder 111 Bitstream 112 bitstream 113 Left Channel 114 left channel 115 Digital-to-Analog (D / A) Converter 116 loudspeaker unit or binaural headphones 122 microphones 123 Right Channel 125 Right Channel 133 Right Channel 134 Right Channel 136 loudspeaker unit or binaural headphones 200 First Stage 201 State Machine 202 Onset / Attack Detector 203 Feature Extractor 204 Outlier Detector 205 Short-Term Feature Vector Filter 206 Nonlinear Feature Vector Transformer 207 Principal component analyzer 208 Gaussian Mixture Model (GMM) Calculator 209 Adaptive Smoother 210 State-Dependent Category Classifier 300 Second Stage 301 Additional High-Level Feature Extractors 302 Core Encoder Initial Selector 303 Core Encoder Selection Improver 401 INACTIVE status 402 ENTRY state 404 ACTIVE status 407 Unstable 1500 IVAS codec 1502 Input section 1504 Output section 1506 processor 1508 memory
Claims
1. 1. A two-stage speech / music classification device for classifying an input sound signal and for selecting a core encoder for encoding said input sound signal, comprising: a first stage for classifying the input sound signal into one of several final classes, the first stage comprising an extractor of features of the input sound signal for classifying the input sound signal; a second stage for extracting additional features of the input sound signal and for selecting the core encoder for encoding the input sound signal depending on the extracted additional features and the final class selected in the first stage; Equipped with the first stage comprising: (a) detecting outlier features among the features extracted in the first stage based on a histogram of the extracted features; (b) determining a vector of the features extracted in the first stage as outliers based on a number of detected outlier features; (c) replacing the outlier features in the vector with feature values obtained from at least one previous frame, rather than discarding the outlier vector; Two-stage speech / music classification device with outlier detector for speech recognition.
2. 2. The two-stage speech / music classification device of claim 1, wherein the first stage comprises a detector of onset / attack in the input sound signal based on relative frame energy.
3. 3. The two-stage speech / music classification device of claim 2, wherein the onset / attack detector updates a cumulative sum of differences between the relative energy of the input sound signal in the current frame and the relative energy of the input sound signal in the previous frame when the relative energy of the input sound signal in the current frame is greater than the relative energy of the input sound signal in the previous frame.
4. 4. The two-stage speech / music classification device of claim 3, wherein the onset / attack detector outputs a binary flag that is set to a first value if the cumulative sum lies within a given range to indicate detection of an onset / attack, and otherwise is set to a second value to indicate non-detection of an onset / attack.
5. 5. The two-stage speech / music classification device of claim 1, wherein the outlier detector calculates a lower and upper bound for each feature extracted in the first stage, compares the values of the features with the lower and upper bounds, and marks as outlier features those features whose values are outside a range defined between the lower and upper bounds.
6. 6. The two-stage speech / music classification device of claim 5, wherein the outlier detector calculates the lower and upper bounds using the histogram of a normalized version of the feature, an index of a frequency bin containing a maximum value in the histogram for the feature, and a threshold value.
7. 7. The two-stage speech / music classification device of claim 1, wherein the first stage comprises a non-linear feature vector converter for converting non-regular features extracted in the first stage from the input sound signal into features having a regular shape.
8. the nonlinear feature vector transformer uses a Box-Cox transformation to transform non-regular features into features having regular shapes; the Box-Cox transformation performed by the nonlinear feature vector converter uses a power transformation with an exponent, different values of the exponent define different Box-Cox transformation curves, and the nonlinear feature vector converter selects the value of the exponent for the Box-Cox transformation based on a normality test; 8. The two-stage speech / music classification device of claim 7, wherein the Box-Cox transformation performed by the nonlinear feature vector transformer uses a bias to ensure that all input values of the extracted features are positive.
9. 9. The two-stage speech / music classification device of claim 8, wherein the normality test produces skew and kurtosis measures, and the nonlinear feature vector transformer applies the Box-Cox transformation only to features that satisfy conditions related to the skew and kurtosis measures.
10. 10. A two-stage speech / music classification device according to any one of claims 1 to 9, wherein the first stage comprises a component analyzer for reducing sound signal feature dimensionality and increasing sound signal class distinctiveness, the component analyzer performing an orthogonal transformation for transforming a set of possibly correlated features extracted from the input sound signal into a set of linearly uncorrelated variables forming components.
11. 11. The two-stage speech / music classification device of claim 1, wherein the first stage comprises a Gaussian Mixture Model (GMM) calculator for determining a first score proportional to the probability that a given vector of features extracted from the input sound signal was produced by an audio Gaussian Mixture Model (GMM) and a second score proportional to the probability that the given vector of features was produced by a music GMM, and wherein the GMM calculator combines the first score and the second score by calculating a difference between the first and second scores to produce a difference score.
12. 12. The two-stage speech / music classification device of claim 11, wherein a negative difference score indicates that the input sound signal is speech and a positive difference score indicates that the input sound signal is music.
13. 13. The two-stage speech / music classification device of claim 11 or 12, wherein the GMM calculator uses a decision bias in the calculation of the difference between the first score and the second score.
14. 14. The two-stage speech / music classification device of claim 13, wherein the GMM calculator predicts a label in an active frame of a training database indicating that the input sound signal is a speech, music, or noise signal, and the GMM calculator uses the label to find the decision bias.
15. 15. The two-stage speech / music classification device of claim 1, wherein the number of final classes comprises a first final class related to speech, a second final class related to music, and a third final class related to speech with background music.
16. 15. The two-stage speech / music classification device of claim 11, wherein the first stage comprises a state-dependent category classifier of the input sound signal into one of three final classes including SPEECH / NOISE, MUSIC, and UNCLEAR, the final class SPEECH / NOISE relating to speech and background noise, the final class MUSIC relating to music, and the final class UNCLEAR relating to speech with background music.
17. 17. The two-stage speech / music classification device of claim 16, wherein when, in a current frame, the input sound signal is in an ENTRY state as determined by a state machine to mark a first onset in the input sound signal after a period of silence, the state-dependent category classifier selects one of the three final classes SPEECH / NOISE, MUSIC, and UNCLEAR based on a weighted average of the difference scores calculated in frames in the ENTRY state preceding the current frame.
18. 18. The two-stage speech / music classification device of claim 17, wherein in states of the input sound signal other than ENTRY as determined by the state machine, the state-dependent category classifier selects the final class SPEECH / NOISE, MUSIC, or UNCLEAR based on a smoothed version of the difference score and the final class SPEECH / NOISE, MUSIC, or UNCLEAR selected in a previous frame.
19. 19. The two-stage speech / music classification device of claim 16, wherein the state-dependent category classifier first initializes the final class in a current frame to the final class set in a previous frame: SPEECH / NOISE, MUSIC, or UNCLEAR.
20. 20. The two-stage speech / music classification device of claim 18, wherein the state-dependent category classifier initially initializes the final class in the current frame to the final class set in the previous frame, SPEECH / NOISE, MUSIC, or UNCLEAR, and within the current frame, the state-dependent category classifier transitions from the final class set in the previous frame, SPEECH / NOISE, MUSIC, or UNCLEAR, to another of the final classes in response to the smoothed difference score crossing a given threshold.
21. 21. The two-stage speech / music classification device of claim 20, wherein the state-dependent category classifier transitions from the final class SPEECH / NOISE set in the previous frame to the final class UNCLEAR when a counter of ACTIVE frames in which voice activity is detected is less than a first threshold, a cumulative sum of difference frame energies equals zero, and the smoothed difference score is greater than a second threshold.
22. 22. The two-stage speech / music classification device of claim 16, wherein the state-dependent category classifier transitions from the final class SPEECH / NOISE set in a previous frame to the final class UNCLEAR if a short pitch flag generated from an open-loop pitch analysis of the input sound signal is equal to a given value and the difference score of a smoothed version is greater than a given threshold.
23. 23. The two-stage speech / music classification device of claim 1, wherein the second stage comprises an extractor of the additional features of the input sound signal in a current frame, the additional features comprising tonality of the input sound signal.
24. The second stage comprises an extractor of the additional features of the input sound signal in the current frame, the additional features comprising: (a) the tonality of the input sound signal; (b) long-term stability of the input sound signal, wherein the extractor of the additional features generates a flag indicating the long-term stability of the input sound signal; (c) a segmental attack in the input sound signal, wherein the extractor of the additional features generates an indicator of (a) the location of the segmental attack within a current frame of the input sound signal, or (b) the absence of a segmental attack; and (d) a spectral peak-to-mean ratio calculated from the power spectrum of the input sound signal, forming a measure of the spectral sharpness of the input sound signal; having a feature selected from the group consisting of:
23. A two-stage speech / music classification device according to any one of claims 1 to 22.
25. 25. The two-stage speech / music classification device of claim 1, wherein the second stage comprises a core encoder initial selector for making an initial selection of the core encoder using (a) relative frame energies, (b) the final class into which the input sound signal is classified by the first stage, and (c) the additional features extracted in the second stage.
26. 15. The two-stage speech / music classification device of claim 11, wherein the second stage comprises a core encoder initial selector for making an initial selection of the core encoder depending on the additional features extracted in the second stage and the final class selected in the first stage, and an initial core encoder selection refiner if a GSC core encoder is initially selected by the core encoder initial selector.
27. 1. A two-stage speech / music classification method for classifying an input sound signal and for selecting a core encoder for encoding said input sound signal, comprising: in a first stage, classifying the input sound signal into one of several final classes, the classifying step including extracting features of the input sound signal; in a second stage, extracting additional features of the input sound signal, and selecting the core encoder for encoding the input sound signal depending on the extracted additional features and the final class selected in the first stage; Equipped with The two-stage speech / music classification method comprises, in the first stage for classifying the input sound signal, Detecting outlier features among the features extracted in the first stage based on a histogram of the extracted features; determining a vector of the features extracted in the first stage as outliers based on a number of detected outlier features; replacing the outlier features in the vector with feature values obtained from at least one previous frame, rather than discarding the outlier vector; A two-stage speech / music classification method comprising:
28. 28. The two-stage speech / music classification method of claim 27, comprising in the first stage detecting onsets / attacks in the input sound signal based on relative frame energies.
29. In the first stage, (a) Open-loop pitch characteristics, (b) Vocalization measure features, (c) features related to line spectrum frequencies from LP analysis; (d) residual energy related features from said LP analysis; (e) Short-term correlation map features, (f) non-stationarity features, (g) Mel-frequency cepstral coefficient features, (h) power spectrum difference features, and (i) Spectral stationarity feature extracting features of the input sound signal selected from the group consisting of:
29. A two-stage speech / music classification method according to claim 27 or 28.
30. 30. The two-stage speech / music classification method of any one of claims 27 to 29, wherein detecting outlier features comprises the steps of: calculating a lower and upper bound for each feature extracted in the first stage; comparing the values of the features with the lower and upper bounds; and marking as outlier features those features whose values lie outside a range defined between the lower and upper bounds.
31. 31. A two-stage speech / music classification method according to any one of claims 27 to 30, comprising in the first stage a non-linear transformation of non-regular features extracted in the first stage from the input sound signal into features having a regular shape.
32. the nonlinear transformation comprising using a Box-Cox transformation to transform non-regular features into features having a regular shape; the Box-Cox transformation comprises using an exponential power transformation; different values of the exponent define different Box-Cox transformation curves; selecting the index value for the Box-Cox transformation based on a normality test; 32. The two-stage speech / music classification method of claim 31, wherein the Box-Cox transformation comprises using a bias to ensure that all input values of the extracted features are positive.
33. 33. The two-stage speech / music classification method of claim 32, wherein the normality tests produce skew and kurtosis measures, and the Box-Cox transformation is applied only to features that satisfy conditions related to the skew and kurtosis measures.
34. 34. A two-stage speech / music classification method according to any one of claims 27 to 33, comprising in the first stage a step of analyzing principal components in order to reduce sound signal feature dimensionality and increase sound signal class discrimination, the step of analyzing principal components comprising an orthogonal transformation to transform a set of possibly correlated features extracted from the input sound signal into a set of linearly uncorrelated variables forming the principal components.
35. 34. The two-stage speech / music classification method of claim 27, comprising in the first stage a Gaussian Mixture Model (GMM) computation to determine a first score proportional to the probability that a given vector of features extracted from the input sound signal was produced by an audio GMM and a second score proportional to the probability that the given vector of features was produced by a music GMM, the GMM computation comprising combining the first score and the second score by calculating a difference between the first and second scores to produce a difference score.
36. 36. The two-stage speech / music classification method of claim 35, wherein a negative difference score indicates that the input sound signal is speech and a positive difference score indicates that the input sound signal is music.
37. 37. A two-stage speech / music classification method according to claim 35 or 36, wherein said GMM calculation comprises using a decision bias in said calculation of said difference between said first score and said second score.
38. 38. The two-stage speech / music classification method of claim 37, wherein the GMM computation comprises predicting a label in an active frame of a training database indicating that the input sound signal is a speech, music, or noise signal, and the GMM computation uses the label to find the decision bias.
39. 39. The two-stage speech / music classification method of claim 27, wherein the number of final classes comprises a first final class related to speech, a second final class related to music, and a third final class related to speech with background music.
40. 39. The two-stage speech / music classification method of claim 35, wherein the first stage comprises a state-dependent categorization of the input sound signal into one of three final classes including SPEECH / NOISE, MUSIC, and UNCLEAR, the final class SPEECH / NOISE relating to speech and background noise, the final class MUSIC relating to music, and the final class UNCLEAR relating to speech with background music.
41. 41. The two-stage speech / music classification method of claim 40, wherein when, in a current frame, the input sound signal is in an ENTRY state as determined by a state machine to mark a first onset in the input sound signal after a period of silence, the state-dependent categorization comprises selecting one of the three final classes SPEECH / NOISE, MUSIC, and UNCLEAR based on a weighted average of the difference scores calculated in frames in the ENTRY state preceding the current frame.
42. 42. The two-stage speech / music classification method of claim 41, wherein in states of the input sound signal other than ENTRY as determined by the state machine, the state-dependent categorization comprises selecting the final class SPEECH / NOISE, MUSIC, or UNCLEAR based on a smoothed version of the difference score and the final class SPEECH / NOISE, MUSIC, or UNCLEAR selected in a previous frame.
43. 43. The two-stage speech / music classification method of claim 40, wherein the state-dependent categorization comprises a step of first initializing the final class in a current frame to the final class set in a previous frame: SPEECH / NOISE, MUSIC, or UNCLEAR.
44. 44. The two-stage speech / music classification method of claim 27, further comprising the step of extracting, in the second stage, the additional features of the input sound signal in a current frame, the additional features comprising tonality of the input sound signal.
45. In the second stage, extracting the additional features of the input sound signal in the current frame, the additional features comprising: (a) the tonality of the input sound signal; (b) long-term stability of the input sound signal, wherein extracting the additional features comprises generating a flag indicative of long-term stability of the input sound signal; (c) a segmental attack in the input sound signal, wherein extracting the additional feature comprises generating an indicator of (a) a position of the segmental attack within a current frame of the input sound signal, or (b) an absence of a segmental attack; and (d) a spectral peak-to-mean ratio forming a measure of the spectral sharpness of the input sound signal, wherein the step of extracting additional features comprises calculating the spectral peak-to-mean ratio from the power spectrum of the input sound signal. having a feature selected from the group consisting of:
44. A two-stage speech / music classification method according to any one of claims 27 to 43.
46. (A) extracting the tonality of the input sound signal comprises expressing the tonality by a tonality flag reflecting both spectral stability and harmonicity within a given frequency range of the input sound signal; (B) the step of extracting the tonality flag comprises: (i) calculating the tonality flag using a correlation map forming a measure of signal stability and harmonicity within a number of first frequency bins within the given frequency range of the residual energy spectrum of the input sound signal and within a segment of the residual energy spectrum in which a peak is present; (ii) applying smoothing of the correlation map and calculating a weighted sum of the correlation map across the frequency bins within the given frequency range of the input sound signal in the current frame to result in a single number; (iii) setting the tonality flag by comparing the single number to an adaptive threshold; Equipped with (C) said two-stage speech / music classification method, In the second stage, the following conditions are met: (a) if the relative frame energy is greater than a first value, the spectral peak-to-mean ratio is greater than a second value, and the unity number is greater than the adaptive threshold, a TCX core encoder is initially selected; (b) if condition (a) does not exist and the final class into which the input sound signal is classified by the first stage is SPEECH / NOISE, an ACELP core encoder is first selected; (c) if conditions (a) and (b) do not exist and the final class into which the input sound signal is classified by the first stage is UNCLEAR, a GSC core encoder is first selected; (d) If conditions (a), (b), and (c) do not exist, the TCX core encoder is selected first. an initial selection of the core encoder; 46. The two-stage speech / music classification method of claim 45.
47. 39. The two-stage speech / music classification method of claim 35, further comprising the step of: in the second stage, initial selection of the core encoder according to the extracted additional features and the final class selected in the first stage; and refining the initial core encoder selection if the initial core encoder selection initially selects a GSC core encoder.
48. 48. The two-stage speech / music classification method of claim 47, wherein refining the initial core encoder selection comprises changing the initial selection of a GSC core encoder to a selection of an ACELP core encoder if (a) a ratio of energy in some first frequency bins of a signal segment to the total energy of the signal segment is less than a first value, and (b) the short-term average of the difference scores is greater than a second value.
49. 48. The two-stage speech / music classification method of claim 47, wherein refining the initial core encoder selection comprises changing the initial selection of a GSC core encoder to a selection of an ACELP core encoder (a) if the difference score of a smoothed version is less than a given value, or otherwise (b) to a selection of a TCX core encoder if the smoothed difference score is greater than or equal to the given value.
50. 48. The two-stage speech / music classification method of claim 47, wherein refining the initial core encoder selection comprises changing the initial selection of a GSC core encoder to a selection of a TCX core encoder in response to (a) long-term stability of the input sound signal, and (b) an open-loop pitch greater than a given value.
51. 48. The two-stage speech / music classification method of claim 47, wherein if a segmental attack is detected in the input sound signal, refining the initial core encoder selection comprises changing the initial selection of a GSC core encoder to a selection of an ACELP core encoder, provided that an indicator that a core encoder selection change is enabled has a first value and a transition frame counter has a second value.
52. 48. The two-stage speech / music classification method of claim 47, wherein refining the initial core encoder selection comprises changing the initial selection of a GSC core encoder to a selection of an ACELP core encoder if the segmental attack is detected in the input sound signal, provided that an indicator that a change in core encoder selection is enabled has a first value, a transition frame counter does not have a second value, and an indicator identifying a segment corresponding to a position of the segmental attack in the current frame is greater than a third value.
Citation Information
Patent Citations
Method and classifier for classifying different segments of a signal
JP2011527445A
Signal classification method and device, and audio encoding method and device using the same
JP2017511905A
Apparatus and method for classification and segmentation of audio content, based on the audio signal
US20100004926A1
Audio signal classifier
US20110016077A1
Multi-mode audio recognition and auxiliary data encoding and decoding
US20140108020A1