Method and device for uncorrelated stereo content classification, crosstalk detection, and stereo mode selection in sound codecs - Patents.com
By classifying uncorrelated stereo content and detecting crosstalk in stereo signals, the method and device adaptively switch between encoding modes, addressing the inefficiencies of existing codecs and ensuring effective stereo encoding at low bit rates with maintained sound quality.
Patent Information
- Application Number
- JP2023515652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-09
- Filing Date
- 2021-09-08
- Publication Date
- 2026-01-28
- Estimated Expiration
- 2041-09-08
AI Technical Summary
Existing audio codecs struggle with efficiently encoding stereo signals at low bit rates, particularly in complex acoustic situations with uncorrelated channels or crosstalk, leading to increased bit rates and compromised sound quality.
A method and device for classifying uncorrelated stereo content and detecting crosstalk in stereo sound signals, using features extracted from the left and right channels to switch between LRTD and DFT stereo modes based on logistic regression models, ensuring efficient encoding and maintaining sound quality.
The solution enables efficient stereo encoding by dynamically selecting the appropriate stereo mode, reducing bit rates while maintaining sound quality and handling complex acoustic scenarios effectively.
Smart Images

Figure 0007808095000103 
Figure 0007808095000104 
Figure 0007808095000105
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to audio coding, and in particular, but not exclusively, to classification of uncorrelated stereo content, crosstalk detection, and stereo mode selection, such as in multi-channel audio codecs that can produce good sound quality in complex acoustic situations at low bit rates and low delays.
[0002] In this disclosure and the accompanying claims: The term "sound" can relate to voice, acoustics, and any other sound. - The term "stereo" is an abbreviation for "stereophonic." - The term "monaural" is an abbreviation for "monophonic." [Background technology]
[0003] Historically, conversational telephones have been implemented with handsets having only one transducer to output sound to only one of the user's ears. Over the last decade, users have begun to use their mobile handsets in combination with headphones to receive sound into their two ears, primarily for listening to music and occasionally for listening to speech. Nevertheless, when a mobile handset is used to send and receive conversational audio, the content is still mono, but is presented to the user's two ears when headphones are used.
[0004] The latest 3GPP speech coding standard, Enhanced Voice Service (EVS), as described in Reference [1], the entire contents of which are incorporated herein by reference, has significantly improved coded sounds, such as voice and / or audio, transmitted and received through mobile handsets. The next natural step is to transmit stereo information in a way that the receiver matches as closely as possible to the real-life acoustic situation that will be experienced at the other end of the communication link.
[0005] For example, transmission of stereo information is commonly used in audio codecs such as those described in reference [2], the entire contents of which are incorporated herein by reference.
[0006] For speech codecs, mono signals are the norm. When stereo signals are transmitted, both the left and right channels of the stereo signal are coded using a mono codec, often doubling the bit rate. While this works well in most scenarios, it presents the drawback of doubling the bit rate and failing to take advantage of the potential redundancy between the two channels (between the left and right channels of the stereo signal). Furthermore, to keep the overall bit rate at a reasonable level, very low bit rates are used for each of the left and right channels, which impacts the overall sound quality. To reduce the bit rate, efficient stereo coding techniques have been developed and used. As non-limiting examples, two stereo coding techniques that can be used efficiently at low bit rates are discussed in the following paragraphs.
[0007] The first stereo coding technique is called parametric stereo. Parametric stereo encodes two inputs (left and right channels) into a mono signal using a common mono codec, adding a certain amount of stereo side information (corresponding to stereo parameters) that represents the stereo image. The left and right channels of the two inputs are downmixed to a mono signal, and then the stereo parameters are calculated. This is usually performed in the frequency domain (FD), for example, in the discrete Fourier transform (DFT) domain. The stereo parameters are related to so-called binaural or interchannel cues. Binaural cues (see, for example, Reference [3], the entire contents of which are incorporated herein by reference) include interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC). Depending on the sound signal characteristics, stereo situation configuration, etc., some or all of the binaural cues are coded and transmitted to the decoder. Information about which binaural cues are coded and transmitted is typically sent as signal information that is part of the stereo side information. Also, given binaural cues may be quantized using different coding techniques, resulting in a variable number of bits being used. Therefore, in addition to the quantized binaural cues, the stereo side information may include a quantized residual signal resulting from the downmix, typically at medium to high bit rates. The residual signal may be coded using an entropy coding technique, such as an arithmetic coder. In the remainder of this disclosure, parametric stereo will be referred to as "DFT stereo" because parametric stereo coding techniques typically operate in the frequency domain, and this disclosure will describe non-limiting embodiments using DFT.
[0008] Another stereo coding technique operates in the time domain. This stereo coding technique mixes two inputs (left and right channels) into a so-called main channel and a secondary channel. For example, according to the method described in Reference [4], the entire contents of which are incorporated herein by reference, the time-domain mixing can be based on a mixing ratio, which determines the respective contributions of the two inputs (left and right channels) in generating the main channel and the secondary channel. The mixing ratio is derived from several criteria, such as the normalized correlation of the two inputs (left and right channels) with the mono signal or the difference in long-term correlation between the two inputs (left and right channels). While the main channel can be coded by a common mono codec, the secondary channel can be coded by a lower bitrate codec. The coding of the secondary channel may exploit the coherence between the main and secondary channels and may reuse some parameters of the main channel. For certain sounds where the left and right channels exhibit little correlation, it is better to code the left and right channels of a stereo input signal in the time domain either separately or with minimal inter-channel parameterization. Such an approach in the encoder is a special case of time-domain TD stereo and is referred to throughout this disclosure as "LRTD stereo."
[0009] Moreover, in recent years, sound generation, recording, representation, coding, transmission, and reproduction have progressed toward improved interactive and immersive experiences for listeners. An immersive experience can be described, for example, as being deeply engaged or involved in a sound situation while sounds come from all directions. In immersive sound (also called 3D (three-dimensional) sound), sound images are reproduced in all three dimensions around the listener, taking into account a wide range of sound characteristics, such as timbre, directionality, reverberation, transparency, and accuracy of the (auditory) space. Immersive sound is generated for specific sound reproduction or playback systems, such as speaker-based systems, integrated playback systems (sound bars), or headphones. Thus, interactivity in a sound reproduction system may include, for example, the ability to adjust sound levels, change the location of the sound, or select different languages for playback.
[0010] There are three basic approaches to achieving an immersive experience.
[0011] The first approach to achieving an immersive experience is a channel-based audio approach, which uses multiple spaced microphones to capture sound from different directions, one microphone corresponding to one audio channel in a particular speaker arrangement. Each recorded channel is then fed to a speaker in a given location. Examples of channel-based audio approaches include stereo, 5.1 surround, 5.1+4, etc.
[0012] A second approach to achieving an immersive experience is situational audio, which uses a combination of dimensional components to represent a desired sound field for a local space as a function of time. The sound signals representing situational audio are independent of the location of the sound sources, but the sound field is transformed into a selected arrangement of loudspeakers in a renderer. An example of situational audio is Ambisonics.
[0013] A third approach to achieving an immersive experience is the object-based audio approach, which represents the audio situation as a set of individual audio elements (e.g., singers, drums, guitars, etc.) along with information such as their location so that they are provided by a sound reproduction system in their intended location. This gives the object-based audio approach great flexibility and interactivity, as each object remains discrete and can be manipulated individually.
[0014] Each of the above acoustic approaches to achieving an immersive experience presents advantages and disadvantages. Therefore, in complex acoustic systems, instead of using only one acoustic approach, several acoustic approaches are typically combined to create an immersive acoustic situation. As an example, there may be an acoustic system that combines situation-based or channel-based audio with object-based audio, such as Ambisonics with several discrete audio objects.
[0015] Recently, the 3rd Generation Partnership Project (3GPP)® has started work to develop a 3D (three-dimensional) sound codec for immersive services called IVAS (Immersive Voice and Audio Services), based on the EVS codec (see reference [5], the entire contents of which are incorporated herein by reference).
[0016] DFT stereo mode is efficient for coding single-talk speech. With two or more talkers, parametric stereo techniques have difficulty fully representing the spatial characteristics of the situation. This problem is particularly evident when two talkers are speaking simultaneously (a crosstalk scenario) and when the signals in the left and right channels of the stereo input signal are weakly correlated or completely uncorrelated. In this situation, it is better to code the left and right channels of the stereo input signal separately in the time domain using LRTD stereo mode, or with minimal inter-channel parameterization. As the situation captured in the stereo input signal evolves, it is desirable to switch between DFT and LRTD stereo modes based on a classification of the stereo situation. Summary of the Invention [Means for solving the problem]
[0017] According to a first aspect, the present disclosure relates to a method for classifying uncorrelated stereo content in a stereo sound signal comprising a left channel and a right channel, responsive to features extracted from the stereo sound signal comprising a left channel and a right channel, the method comprising the steps of: calculating a score representative of the uncorrelated stereo content in the stereo sound signal responsive to the extracted features; and switching between a first class indicating one of uncorrelated and correlated stereo content in the stereo sound signal and a second class indicating another of uncorrelated and correlated stereo content in the stereo sound signal responsive to the score.
[0018] According to a second aspect, the present disclosure relates to a device for classifying uncorrelated stereo content in a stereo sound signal comprising a left channel and a right channel, the device being responsive to features extracted from the stereo sound signal comprising a left channel and a right channel, the device comprising: a device for calculating a score representative of uncorrelated stereo content in the stereo sound signal, responsive to the extracted features; and a class switching mechanism responsive to the score for switching between a first class indicating one of uncorrelated and correlated stereo content in the stereo sound signal and a second class indicating the other of uncorrelated and correlated stereo content.
[0019] The present disclosure also relates to a method for detecting crosstalk in a stereo sound signal comprising a left channel and a right channel in response to features extracted from the stereo sound signal comprising the left channel and the right channel, the method comprising the steps of: calculating a score representative of crosstalk in the stereo sound signal in response to the extracted features; calculating an auxiliary parameter for use in detecting crosstalk in the stereo sound signal; and switching between a first class indicating the presence of crosstalk in the stereo sound signal and a second class indicating the absence of crosstalk in the stereo sound signal in response to the crosstalk score and the auxiliary parameter.
[0020] According to a further aspect, the present disclosure provides an apparatus for detecting crosstalk in a stereo sound signal comprising a left channel and a right channel, responsive to features extracted from the stereo sound signal comprising the left channel and the right channel, the apparatus comprising: an apparatus for calculating a score representative of crosstalk in the stereo sound signal, responsive to the extracted features; an apparatus for calculating an auxiliary parameter for use in detecting crosstalk in the stereo sound signal; and a class switching mechanism responsive to the crosstalk score and the auxiliary parameter for switching between a first class indicating the presence of crosstalk in the stereo sound signal and a second class indicating the absence of crosstalk in the stereo sound signal.
[0021] The present disclosure also relates to a method for selecting one of a first stereo mode and a second stereo mode for coding a stereo sound signal including a left channel and a right channel, the method including the steps of: generating a first output indicating a presence or absence of uncorrelated stereo content in the stereo sound signal; generating a second output indicating a presence or absence of crosstalk in the stereo sound signal; calculating auxiliary parameters for use in selecting a stereo mode for coding the stereo sound signal; and selecting the stereo mode for coding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters.
[0022] According to still a further aspect, the present disclosure provides a device for selecting one of a first stereo mode and a second stereo mode for coding a stereo sound signal including a left channel and a right channel, the device comprising: a classifier for generating a first output indicating the presence or absence of uncorrelated stereo content in the stereo sound signal; a detector for generating a second output indicating the presence or absence of crosstalk in the stereo sound signal; an analysis processor for calculating auxiliary parameters for use in selecting the stereo mode for coding the stereo sound signal; and a stereo mode selector for selecting the stereo mode for coding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters.
[0023] The foregoing and other objects, advantages, and features of the decorrelated stereo content classification apparatus, decorrelated stereo content classification method, crosstalk detection apparatus, crosstalk detection method, stereo mode selection device, and stereo mode selection method will become more apparent from a reading of the following non-limiting description of illustrative embodiments, given by way of example only with reference to the accompanying drawings. [Brief explanation of the drawings]
[0024] [Figure 1] 1 is a schematic block diagram illustrating simultaneously a device for coding a stereo sound signal and a corresponding method for coding a stereo sound signal; [Figure 2] FIG. 1 is a schematic diagram showing a planar view of a crosstalk situation with two opposing talkers captured by a pair of hypercardioid microphones. [Figure 3] 1 is a graph showing the location of peaks in the GCC-PHAT function. [Figure 4] FIG. 1 is a plan view from above of the stereo situation set up for the actual recording. [Figure 5] 10 is a graph showing a normalization function applied to the output of a LogReg model in classifying uncorrelated stereo content in LRTD stereo mode. [Figure 6] 2 is a state machine diagram illustrating a mechanism for switching between stereo content classes in a classifier for uncorrelated stereo content forming part of the device of FIG. 1 for coding a stereo sound signal; FIG. [Figure 7] FIG. 1 is a schematic floor plan of a large conference room with an AB microphone setup, where the AB microphones consist of a pair of cardioid or omnidirectional microphones spaced apart in such a way that they cover the space without creating phase issues with each other, conditions being simulated for crosstalk detection. [Figure 8] FIG. 1 illustrates automatic labeling of crosstalk examples using VAD (Voice Activity Detection). [Figure 9] 10 is a graph showing a function for scaling the raw output of the LogReg model in crosstalk detection in LRTD stereo mode. [Figure 10] 2 is a graph showing a mechanism for detecting rising edges in a part forming a crosstalk detection unit of the device of FIG. 1 for encoding a stereo sound signal in LRTD stereo mode; [Figure 11] FIG. 10 is a logic diagram illustrating the mechanism for switching between states of the output of the crosstalk detection device in LRTD stereo mode. [Figure 12] FIG. 10 is a logic diagram illustrating the mechanism for switching between states of the output of the crosstalk detector in DFT stereo mode. [Figure 13] FIG. 1 is a schematic block diagram illustrating a mechanism for selecting between LRTD and DFT stereo modes. [Figure 14] FIG. 1 is a simplified block diagram of an example configuration of hardware components implementing a method and device for encoding a stereo sound signal. DETAILED DESCRIPTION OF THE INVENTION
[0025] This disclosure describes classification of uncorrelated stereo content (hereinafter "UNCLR classification") and crosstalk detection (hereinafter "XTALK detection") in an input stereo sound signal. This disclosure also describes stereo mode selection, e.g., automatic LRTD / DFT stereo mode selection.
[0026] FIG. 1 is a schematic block diagram illustrating simultaneously a device 100 for encoding a stereo sound signal 190 and a corresponding method 150 for encoding the stereo sound signal 190 .
[0027] Specifically, FIG. 1 shows how UNCLR classification, XTALK detection, and stereo mode selection are incorporated into a method 150 and device 100 for encoding a stereo sound signal.
[0028] UNCLR classification and XTALK detection form two independent technologies. However, they are based on the same statistical model and share some features and parameters. Furthermore, both UNCLR classification and XTALK detection are designed and trained separately for LRTD stereo mode and DFT stereo mode. In this disclosure, LRTD stereo mode is provided as a non-limiting example of a time-domain stereo mode, and DFT stereo mode is provided as a non-limiting example of a frequency-domain stereo mode. Implementing other time-domain and frequency-domain stereo modes is within the scope of this disclosure.
[0029] The UNCLR classification analyzes features extracted from the left and right channels of the stereo sound signal 190 and detects weak or zero correlation between the left and right channels. On the other hand, the XTALK detection detects the presence of two speakers speaking simultaneously in a stereo situation. For example, both the UNCLR classification and the XTALK detection provide binary outputs. These binary outputs are combined together in the stereo mode selection logic. As a non-limiting general rule, the stereo mode selection selects the LRTD stereo mode when the UNCLR classification and the XTALK detection indicate the presence of two speakers standing on either side of the capture device (e.g., a microphone). This situation typically results in weak correlation between the left and right channels of the stereo sound signal 190. The selection of the LRTD stereo mode or the DFT stereo mode is performed on a frame-by-frame basis (as is well known in the art, the stereo sound signal 190 is sampled at a given sampling rate and processed by a group of these samples, called a "frame," which is divided into several "subframes"). The stereo mode selection logic is also designed to avoid frequent switching between LRTD and DFT stereo modes and stereo mode switching during perceptually important signal sections.
[0030] Non-limiting exemplary embodiments of UNCLR classification, XTALK detection, and stereo mode selection are described in this disclosure by way of example only with reference to the IVAS coding framework, referred to as the IVAS codec (or IVAS sound codec). However, incorporating such classification, detection, and selection with any other sound codec is within the scope of this disclosure.
[0031] 1. Feature Extraction UNCLR classification is based on a logistic regression (LogReg) model, such as that described in Reference [9], the entire contents of which are incorporated herein by reference. The LogReg model is trained separately for LRTD stereo mode and DFT stereo mode. Training is performed using a large database of features extracted from a stereo sound signal coding device 100 (stereo codec). Similarly, XTALK detection is based on a LogReg model trained separately for LRTD stereo mode and DFT stereo mode. The features used in XTALK detection are different from those used in UNCLR classification. However, certain features are shared by both techniques.
[0032] The features used in UNCLR classification and XTALK detection are: - Inter-channel correlation analysis, - TD preprocessing, and - DFT stereo parameterization Extracted from
[0033] The method 150 for coding a stereo sound signal includes the above-mentioned feature extraction operation (not shown). To perform the feature extraction operation, the device 100 for coding a stereo sound signal comprises a feature extraction unit (not shown).
[0034] 2. Inter-channel correlation analysis The feature extraction operation (not shown) includes an inter-channel correlation analysis operation 151 for LRTD stereo mode and an inter-channel correlation analysis operation 152 for DFT stereo mode. To perform operations 151 and 152, the feature extraction device (not shown) includes an inter-channel correlation analysis device 101 and an inter-channel correlation analysis device 102, respectively. Operations 151 and 152 and analysis devices 101 and 102 are similar and will be described simultaneously.
[0035] The analyzer 101 / 102 receives as input the left and right channels of the current stereo sound signal frame. The left and right channels are first downsampled to 8 kHz. For example, the downsampled left and right channels can be represented as follows: X L (n),X R (n), n=0, .., N-1 (1) where n is the sample index in the current frame and N=160 is the length of the current frame (160 samples long). The downsampled left and right channels are used to calculate the inter-channel correlation function. First, the absolute energy of the left and right channels is calculated, for example, using the following relationship:
[0036]
number
[0037] The analysis device 101 / 102 calculates the numerator of the inter-channel correlation function from the dot product between the left and right channels over time lags <-40, 40>. For negative time lags, the dot product between the left and right channels is calculated using, for example, the following relationship:
[0038]
number
[0039] For positive time lags, the dot product is given, for example, by the following relationship:
[0040]
number
[0041] The analysis device 101 / 102 then calculates the inter-channel correlation function, for example, using the following relationship:
[0042]
number
[0043] where the superscript [-1] indicates a reference to the previous frame. A passive mono signal is calculated by averaging the left and right channels.
[0044]
number
[0045] The side signal is calculated as the difference between the left and right channels using, as a non-limiting example, the following relationship:
[0046]
number
[0047] Finally, it is also useful to define the per-sample product of the left and right channels as follows: X P (n)=X L (n)·X R (n), n=0, .., N-1 (8)
[0048] The analysis device 101 / 102 includes an infinite impulse response (IIR) filter (not shown) to smooth the inter-channel correlation function, for example using the following relationship:
[0049]
number
[0050] where the superscript [n] indicates the current frame, the superscript [n-1] indicates the previous frame, and α ICA is the smoothing coefficient.
[0051] Smoothing coefficient α ICA is adaptively set in the inter-channel correlation analysis (ICA) module (reference [1]) of the stereo sound signal coding device 100 (stereo codec). The inter-channel correlation function is then weighted at its location in the region of the predicted peak. The mechanisms for peak finding and local window generation are implemented within the ICA module and are not described in this document; see reference [1] for additional information about the ICA module. The inter-channel correlation function after ICA weighting is denoted as R, where k∈<-40, 40>. W This will be denoted as (k).
[0052] The location of the maximum of the inter-channel correlation function is an important indicator of the direction from which the dominant sound is coming to the capture position and is used as a feature by UNCLR classification and XTALK detection in LRTD stereo mode. The analysis unit 101 / 102 calculates the maximum of the inter-channel correlation function, which is also used as a feature by XTALK detection in LRTD stereo mode, for example, using the following relationship:
[0053]
number
[0054] This maximum position uses the following relationship, as a non-limiting example:
[0055]
number
[0056] Maximum R of inter-channel correlation function max If is negative, it is set to 0. The maximum value R in the current frame max The difference between the current frame and the previous frame is calculated, for example, as follows:
[0057]
number
[0058] Here, the superscript [-1] indicates a reference to the previous frame.
[0059] The position of the maximum of the inter-channel correlation function determines which channels will be the "reference" channel (REF) and the "target" channel (TAR) in the ICA module. max If ≧0, the left channel (L) becomes the reference channel (REF) and the right channel (R) becomes the target channel (TAR). max <0, the right channel (R) becomes the reference channel (REF) and the left channel (L) becomes the target channel (TAR). The target channel (TAR) is then shifted to offset its delay relative to the reference channel (REF). The number of samples used to shift the target channel (TAR) is, for example, |k max However, the position k between successive frames can be set directly to | max To eliminate artifacts arising from absolute changes in σ, the number of samples used to shift the target channel (TAR) can be smoothed with an appropriate filter within the ICA module.
[0060] The number of samples used to shift the target channel (TAR) is k. shift where k shift >0. The reference channel signal is X ref (n) and the target channel signal is denoted as X tar(n). The instantaneous target gain reflects the ratio of energy between the reference channel (REF) and the shifted target channel (TAR). The instantaneous target gain may be calculated, for example, using the following relationship:
[0061]
number
[0062] where N is the frame length. The instantaneous target gain is used as a feature by UNCLR classification in LRTD stereo mode.
[0063] 2.1 Inter-channel characteristics The analysis unit 101 / 102 derives a first set of features used in UNCLR classification and XTALK detection directly from the inter-channel analysis. The value of the inter-channel correlation function at zero time lag R(0) is used as a feature by UNCLR classification and XTALK detection in LRTD stereo mode. By calculating the logarithm of the absolute value of C(0), another feature used by UNCLR classification and XTALK detection in LRTD stereo mode is obtained as follows:
[0064]
number
[0065] The ratio of the energy of the side signal to the energy of the mono signal is also used as a feature by UNCLR classification and XTALK detection in LRTD stereo mode. This ratio is calculated, for example, using the following relationship:
[0066]
number
[0067] The energy ratio of relation (15) is smoothed over time, for example as follows:
[0068]
number
[0069] where c hang is the counter of VAD hangover frames calculated as part of the VAD (Voice Activity Detection) module (see, for example, reference [1]) of the stereo sound signal coding device 100 (stereo codec). The smoothed ratio of relation (16) is used as a feature by XTALK detection in LRTD stereo mode.
[0070] The analysis device 101 / 102 derives the following dot products from the left channel and the mono signal and between the right channel and the mono signal: First, the dot product between the left channel and the mono signal is expressed as, for example:
[0071]
number
[0072] Then, the dot product between the right channel and the mono signal can be expressed as, for example:
[0073]
number
[0074] Both dot products are positive with a lower bound of 0. A criterion based on the maximum and minimum difference of these two dot products is used as a feature by UNCLR classification and XTALK detection in LRTD stereo mode. This can be calculated using the following relation: d mmLR =max[C LM , C RM ]-min[C LM , C RM ] (19)
[0075] A similar criterion used as an independent feature by UNCLR classification and XTALK detection in LRTD stereo mode is based directly on the absolute difference between two dot products, calculated, for example, using the following relationship, in both the linear and logarithmic domains: Δ LRM =C LM -C RM d LRM =log 10 |C LM- C RM | (20)
[0076] The final feature used by UNCLR classification and XTALK detection in LRTD stereo mode is calculated as part of the inter-channel correlation analysis operation 151 / 152 and reflects the evolution of the inter-channel correlation function, which is calculated as follows:
[0077]
number
[0078] Here, the superscript [-2] indicates a reference to the frame two frames before the current frame.
[0079] 3. Time Domain (TD) Preprocessing In LRTD stereo mode, there is no mono downmix, and both the left and right channels of the input stereo sound signal 190 are analyzed with respective time-domain preprocessing operations to extract features, i.e., operation 153 for time-domain preprocessing the left channel of the stereo sound signal 190 and operation 154 for time-domain preprocessing the right channel. To perform operations 153, 154, a feature extraction device (not shown) comprises respective time-domain preprocessing devices 103 and 104, as shown in Fig. 1. Operations 153 and 154 and the corresponding preprocessing devices 103 and 104 are similar and will be described simultaneously.
[0080] The time domain pre-processing operation 153 / 154 performs several sub-operations to generate specific parameters that are used as extracted features for performing UNCLR classification and XTALK detection. Such sub-operations include: - Spectral analysis, - Linear predictive analysis, - open-loop pitch estimation, - Voice Activity Detection (VAD), - background noise estimation, and - Frame Error Concealment (FEC) classification is possible.
[0081] The time-domain preprocessor 103 / 104 performs linear prediction analysis using the Levinson-Durbin algorithm. The output of the Levinson-Durbin algorithm is a set of linear prediction coefficients (LPCs). The Levinson-Durbin algorithm is an iterative method, and the total number of iterations in the Levinson-Durbin algorithm can be denoted as M. At each iteration, i=1, .., M, the residual error energy
[0082]
number
[0083] is calculated.
[0084] In this disclosure, as a non-limiting example implementation, it is assumed that the Levinson-Durbin algorithm is performed with M=16 iterations. The difference in residual error energy between the left and right channels of the input stereo sound signal 190 is used as a feature for XTALK detection in LRTD stereo mode. The difference in residual error energy may be calculated as follows:
[0085]
number
[0086] Here, the subscripts L and R are added to denote the left and right channels, respectively, of the input stereo sound signal 190. In this non-limiting embodiment, the features (difference d LPC13 ) is calculated using the residual energy from the 14th iteration instead of the last iteration, since this iteration has been experimentally found to have the greatest characteristic potential for UNCLR classification. Further information on the Levinson-Durbin algorithm, and details on residual error energy calculations, can be found, for example, in reference [1].
[0087] The LPC coefficients estimated by the Levinson-Durbin algorithm are converted to line spectral frequencies LSF(i), i=0, .., M-1. The sum of the LSF values can serve as an estimate of the gravity point of the envelope of the input stereo sound signal 190. The difference between the sum of the LSF values in the left channel and the sum of the LSF values in the right channel contains information about the similarity of the two channels. For that reason, this difference is used as a feature in XTALK detection in LRTD stereo mode. The difference between the sum of the LSF values in the left channel and the sum of the LSF values in the right channel can be calculated using the following relationship:
[0088]
number
[0089] Further information on the LPC to LSF conversion mentioned above can be found, for example, in reference [1].
[0090] The time domain pre-processing unit 103 / 104 performs open-loop pitch estimation and uses an autocorrelation function from which the left channel (L) / right channel (R) open-loop pitch difference is calculated. The left channel (L) / right channel (R) open-loop pitch difference can be calculated using the following relationship:
[0091]
number
[0092] where T [k] is the open-loop pitch estimate for the kth partition of the current frame. In this disclosure, as a non-limiting illustrative example, it is assumed that open-loop pitch analysis is performed on three adjacent half-frames (partitions), indexed k=1, 2, 3, where two partitions are located in the current frame and one partition is located in the second half of the previous frame. In addition to using different numbers of partitions, different partition lengths and overlaps are possible. Additional information on open-loop pitch estimation can be found, for example, in Reference [1].
[0093] The difference in maximum autocorrelation value (determined by the above autocorrelation function) (voice) between the left and right channels of the input stereo sound signal 190 is also used as a feature by XTALK detection in LRTD stereo mode. The difference between the maximum autocorrelation value of the left channel and the maximum autocorrelation value of the right channel can be calculated using the following relationship:
[0094]
number
[0095] where ν [k] denotes the maximum autocorrelation value of the left (L) and right (R) channels in the kth half frame.
[0096] Background noise estimation is part of the voice activity detection (VAD) detection algorithm (see reference [1]). Specifically, background noise estimation uses an active / inactive signal detector (not shown) that relies on a set of features, some of which are used by UNCLR classification and XTALK detection. For example, the active / inactive signal detector (not shown) estimates the left channel (L) and right channel (R) nonstationarity parameters f staas a measure of spectral stability. The difference in non-stationarity between the left and right channels of the input stereo sound signal 190 is used as a feature by XTALK detection in LRTD stereo mode. The difference in non-stationarity between the left (L) and right (R) channels can be calculated using the following relationship: d sta =|f sta,L -f sta,R | (26)
[0097] An active / inactive signal detector (not shown) measures the correlation map parameter C map The correlation map is a measure of the timbral stability of the input stereo sound signal 190 and is used by UNCLR classification and XTALK detection. The difference between the correlation map of the left (L) channel and the correlation map of the right (R) channel is used as a feature by XTALK detection in LRTD stereo mode and is calculated, for example, using the following relationship: d cmap =|C map,L -C map,R | (27)
[0098] Finally, an active / inactive signal detector (not shown) performs regular measurements of the spectral diversity and noise characteristics in each frame. These two parameters are also used as features by UNCLR classification and XTALK detection in LRTD stereo mode. Specifically, (a) the difference in spectral diversity between the left channel (L) and the right channel (R) can be calculated as follows: d sdiv =|log(S div,L )-log(S div,R )| (28) where S div represents a measure of spectral diversity in the current frame, and (b) the difference in noise characteristics between the left channel (L) and the right channel (R) can be calculated as: d nchar =|log(n char,L )-log(nchar,R )| (29) where n char represents a measure of the noise characteristics in the current frame. For details on the calculation of the correlation map, non-stationarity, spectral diversity, and noise characteristics parameters, reference can be made to [1].
[0099] The ACELP (Algebraic Code-Excited Linear Prediction) core encoder, which is part of the stereo sound signal coding device 100, is equipped with specific settings for coding unvoiced sounds, as described in Reference [1]. The use of these settings is conditioned by several factors, including a measure of sudden energy increases in short segments inside the current frame. The settings for unvoiced sound coding in the ACELP core encoder are only applied when there are no sudden energy increases inside the current frame. By comparing the measurement of sudden energy increases in the left channel with the measurement of sudden energy increases in the right channel, it is possible to locate the start of a crosstalk segment. The sudden energy increases are measured using the EVS codec, as described in Reference [1], for the 3GPP EVS codec. d The difference in the sudden energy increase between the left channel (L) and the right channel (R) can be calculated using the following relationship: d dE =|log(E d,L )-log(E d,R )| (30) Here, the subscripts L and R have been added to indicate the left and right channels, respectively, of the input stereo sound signal 190 .
[0100] The time-domain preprocessing unit 103 / 104 and the preprocessing operation 153 / 154 use an FEC classification module that includes a state machine for FEC techniques. The FEC class for each frame is selected from predetermined classes based on a function of merit. The difference between the FEC classes selected for the left channel (L) and the right channel (R) in the current frame is used as a feature by XTALK detection in LRTD stereo mode. However, for the purposes of such classification and detection, the FEC classes can be restricted as follows:
[0101]
number
[0102] where t class is the selected FEC class for the current frame. Therefore, the FEC classes are limited to only voiced and unvoiced sounds. The difference between the class in the left channel (L) and the class in the right channel (R) can be calculated as follows: d class =|t class,L -t class,R | (32)
[0103] For additional details about FEC classification, reference may be made to [1].
[0104] The time domain preprocessing unit 103 / 104 and the preprocessing operation 153 / 154 implement a speech / music classification and a corresponding speech / music classifier. This speech / music classification makes a binary decision at each frame according to the power spectrum divergence and the power spectrum stability. The difference in power spectrum divergence between the left channel (L) and the right channel (R) is calculated, for example, using the following relationship: d Pdiff =|P diff,L -P diff,R | (33) where P diffwhere σ represents the power spectral divergence in the left (L) and right (R) channels in the current frame, and the difference in power spectral stability between the left (L) and right (R) channels is calculated using, for example, the following relationship: d Psta =|P sta,L -P sta,R | (34) where P sta represents the power spectrum stability in the left channel (L) and right channel (R) in the current frame.
[0105] Reference [1] provides details on power spectral divergence and power spectral stability calculated during speech / music classification.
[0106] 4. DFT Stereo Parameters The method 150 for coding a stereo sound signal 190 includes an operation 155 of calculating a Fast Fourier Transform (FFT) of the left channel (L) and the right channel (R). To perform operation 155, the device 100 for coding a stereo sound signal 190 comprises an FFT transform calculation unit 105.
[0107] The feature extraction operation (not shown) includes a DFT stereo parameter calculation operation 156. To perform operation 156, the feature extraction device (not shown) comprises a DFT stereo parameter calculation device 106.
[0108] In DFT stereo mode, the transform computation unit 105 transforms the left channel (L) and right channel (R) of the input stereo sound signal 190 into the frequency domain using an FFT transform.
[0109] The complex spectrum of the left channel (L) is shown as follows:
[0110]
number
[0111] Then, the complex spectrum of the right channel (R) is given as follows:
[0112]
number
[0113] where k=0, .., N FFT -1 is the frequency bin index and N FFT is the length of the FFT transform. For example, when the sampling rate of the input stereo sound signal is 32 kHz, the DFT stereo parameter calculation unit 106 calculates the complex spectrum for a 40 ms window, and N FFT = 1280 samples. The complex cross-channel spectrum can then be calculated using, as a non-limiting example, the following relationship:
[0114]
number
[0115] The asterisk superscript indicates the complex conjugate. The complex cross-channel spectrum can be decomposed into real and imaginary parts using the following relationship:
[0116]
number
[0117] Using decomposition into real and imaginary parts, it is possible to express the absolute magnitude of the complex cross-channel spectrum as:
[0118]
number
[0119] The DFT stereo parameter calculation unit 106 obtains the total absolute magnitude of the complex inter-channel spectrum by summing the absolute magnitudes of the complex inter-channel spectrum for frequency bins using the following relationship:
[0120]
number
[0121] The energy spectrum of the left channel (L) and the energy spectrum of the right channel (R) can be expressed as follows:
[0122]
number
[0123] The total energy of the left channel (L) and right channel (R) can be obtained by summing the left channel (L) energy spectrum and the right channel (R) energy spectrum for a frequency bin using the following relationship:
[0124]
number
[0125] UNCLR classification and XTALK detection in DFT stereo mode use the overall absolute magnitude of the complex inter-channel spectrum as one of their features, but not in the direct form as defined above, but in the logarithmic domain in an energy normalized form, for example, as expressed using the following relationship:
[0126]
number
[0127] The DFT stereo parameter calculation unit 106 may calculate the mono downmix energy using, for example, the following relationship:
[0128]
number
[0129] The inter-channel level difference (ILD) is a feature used by UNCLR classification and XTALK detection in DFT stereo mode because it contains information about the angle from which the dominant sound comes. For the purposes of UNCLR classification and XTALK detection, the inter-channel level difference (ILD) can be expressed in the form of a gain coefficient. The DFT stereo parameter calculation unit 106 calculates the inter-channel level difference (ILD) gain using, for example, the following relationship:
[0130]
number
[0131] The inter-channel phase difference (IPD) contains information that allows a listener to infer the direction of an incoming sound signal. The DFT stereo parameter calculation unit 106 calculates the inter-channel phase difference (IPD), for example, using the following relationship:
[0132]
number
[0133] Here,
[0134]
number
[0135] The derivative of the inter-channel phase difference (IPD) with respect to the previous frame is calculated using, for example, the following relationship:
[0136]
number
[0137] The superscript n is used to denote the current frame, and the superscript n-1 is used to denote the previous frame. Finally, the calculation unit 106 calculates the IPD gain by dividing the phase-aligned (IPD=0) downmix energy (the numerator of relation (47)) by the mono downmix energy E M It is possible to calculate the ratio between the energy of
[0138]
number
[0139] IPD gain g IPD_lin is restricted to the interval <0, 1>. If the value exceeds an upper threshold of 1.0, the value of the IPD gain from the previous frame is substituted for it. UNCLR classification and XTALK detection in DFT stereo mode use the IPD gain in the logarithmic domain as a feature. The calculation unit 106 determines the IPD gain in the logarithmic domain using, for example, the following relationship: g IPD =log(1-g IPD_lin ) (48)
[0140] The inter-channel phase difference (IPD) can also be expressed in the form of an angle, which is used as a feature by UNCLR classification and XTALK detection in DFT stereo mode, and is calculated, for example, as shown below:
[0141]
number
[0142] The side channel can be calculated as the difference between the left channel (L) and the right channel (R). The mono downmix energy E MThe energy difference (E L -E R ) can be used to express the gain of the side channel.
[0143]
number
[0144] Gain g side The larger the gain g of the left channel, the greater the difference between the energy of the left channel (L) and the energy of the right channel (R). side is restricted to the interval <0.01, 0.99>. Values outside this range are clamped.
[0145] The phase difference between the left channel (L) and the right channel (R) of the input stereo sound signal 190 can also be analyzed from the predicted gain, which is calculated using, for example, the following relationship: g pred_lin =(1-g side )E L +(1+g side )E R -2|X LR | (51) Here, the prediction gain g pred_lin The value of g is restricted to the interval <0, ∞>, i.e., it is restricted to positive values. pred_lin The above formula for the cross-channel spectrum (X LR ) Energy and Mono Downmix Energy E M =E L +E R +2|X LR The computing device 106 captures the difference between |g| and |g| for use as a feature by UNCLR classification and XTALK detection in DFT stereo mode, for example using relation (52). pred_lin to the logarithmic domain. g pred =log(g pred_lin +1) (52)
[0146] The calculation unit 106 also uses the channel energy per bin of the relationship (39) to calculate the average energy of the inter-channel coherence (ICC), which forms a cue for determining differences between the left channel (L) and the right channel (R) that are not captured by the inter-channel time difference (ITD) and inter-channel phase difference (IPD) described hereinafter. First, the calculation unit 106 calculates the total energy of the inter-channel spectrum, for example, using the following relationship: E X =Re(X LR ) 2 +IM(X LR ) 2 (53)
[0147] To express the average energy of the inter-channel coherence (ICC), it is useful to calculate the following parameter:
[0148]
number
[0149] The average energy of the inter-channel coherence (ICC) is then used as a feature by UNCLR classification and XTALK detection in DFT stereo mode and can be expressed as:
[0150]
number
[0151] If the inner term is less than 1.0, the average energy E coh The value of is set to 0. Another possible interpretation of the inter-channel coherence (ICC) is the side-to-mono energy ratio, calculated as follows:
[0152]
number
[0153] Finally, the computing device 106 calculates the ratio r between the maximum and minimum inter-channel amplitude products used for UNCLR classification and XTALK detection. pp This feature, which is used as a feature by UNCLR classification and XTALK detection in DFT stereo mode, is calculated, for example, using the following relationship:
[0154]
number
[0155] Here, the inter-channel amplitude product is defined as follows:
[0156]
number
[0157] A parameter used in stereo signal reconstruction is the inter-channel time difference (ITD). In DFT stereo mode, the DFT stereo parameter calculation unit 106 estimates the inter-channel time difference (ITD) from the generalized cross-channel correlation function with phase difference (GCC-PHAT). The inter-channel time difference (ITD) corresponds to the time delay of arrival (TDOA) estimation. The GCC-PHAT function is a robust method for estimating the inter-channel time difference (ITD) in reverberant signals. The GCC-PHAT is calculated, for example, using the following relationship:
[0158]
number
[0159] Here, IFFT stands for inverse fast Fourier transform.
[0160] The inter-channel time difference (ITD) is then estimated from the GCC-PHAT function, for example using the following relationship:
[0161]
number
[0162] where d is the time lag in samples corresponding to a time delay ranging from -5ms to +5ms. ITD The maximum value of the GCC-PHAT function corresponding to is used as a feature by UNCLR classification and XTALK detection in DFT stereo mode and can be retrieved using the following relation:
[0163]
number
[0164] In a single-talk scenario, there is typically a single dominant peak in the GCC-PHAT function corresponding to the inter-channel time difference (ITD). However, in a crosstalk situation where two talkers are positioned on either side of the capture microphone, there are typically two dominant peaks positioned far apart from each other. Figure 2 illustrates such a situation. Specifically, by way of a non-limiting illustrative example, Figure 2 is a plan view of a crosstalk situation in which two opposing talkers S1 and S2 are captured by a pair of hypercardioid microphones M1 and M2, and Figure 3 is a graph showing the locations of the two dominant peaks in the GCC-PHAT function.
[0165] First peak G ITD The amplitude of is calculated using the relation (61) and its position d ITD is calculated using the relation (60). The amplitude of the second peak can be located by searching for a second maximum of the GCC-PHAT function in the opposite direction to the first peak. More specifically, the direction s in which to search for the second peak is ITD is the position of the first peak d ITD is determined by the sign of s ITD =sgn(d ITD ) (62) where sgn(.) is the sign function.
[0166] The DFT stereo parameter calculation unit 106 then calculates the direction s ITD The second maximum of the GCC-PHAT function at (second highest peak) can be extracted.
[0167]
number
[0168] As a non-limiting example, the threshold thr xt = 8 is the start of the second peak of the GCC-PHAT function (d ITD = 0). As far as crosstalk (XTALK) detection is concerned, this means that any potential secondary talkers in the situation must be at least a certain minimum distance away from both the first "dominant" talker and the midpoint (d = 0).
[0169] The position of the second highest peak of the GCC-PHAT function is calculated using the relationship (63) by replacing the max(.) function with the arg max(.) function. The position of the second highest peak of the GCC-PHAT function is d ITD2 is shown as:
[0170] The relationship between the amplitude of the first peak and the amplitude of the second highest peak of the GCC-PHAT function is used as a feature by XTALK detection in DFT stereo mode and can be evaluated using the following ratio:
[0171]
number
[0172] Ratio r GITD12has a high discriminatory ability, but by using it as a feature, XTALK detection eliminates occasional false alarms arising from the limited time resolution applied during the frequency transformation in DFT stereo mode. This can be done, for example, by using the following relationship to determine the rate r in the current frame: GITD12 This can be done by multiplying the value of by the same percentage value from the previous frame. r GITD12 ←r GITD12 (n)·r GITD12 (n-1) (65) The index n is added to indicate the current frame, and the index n-1 is added to indicate the previous frame. For simplicity, the parameter name r GITD12 is reused to identify the output parameter.
[0173] The amplitude of the second highest peak alone constitutes an indicator of the strength of the secondary speaker in the situation. GITD12 Similarly, the value G ITD2 The occasional random "spikes" in are reduced using, for example, the following relation (66) to obtain another feature used by XTALK detection in DFT stereo mode: m ITD2 =G ITD2 (n)·G ITD2 (n-1) (66)
[0174] Another feature used in XTALK detection in DFT stereo mode is the position d of the second highest peak in the current frame relative to the previous frame, calculated for example using the following relationship: ITD2 (n) is the difference. Δ ITD2 =|d ITD2 (n)-d ITD2 (n-1)| (67)
[0175] 5. Downmix and Inverse Fast Fourier Transform (IFFT) In DFT stereo mode, the method 150 for coding a stereo audio signal includes an operation 157 of downmixing a left channel (L) and a right channel (R) of the stereo audio signal 190 and an operation 158 of computing an IFFT transform of the downmixed signal. To perform operations 157 and 158, the device 100 for coding a stereo audio signal 190 comprises a downmix unit 107 and an IFFT transform computation unit 108.
[0176] The downmixer 107 downmixes the left channel (L) and right channel (R) of the stereo sound signal into a mono channel (M) and a side channel (S), for example as described in reference [6], the entire contents of which are incorporated herein by reference.
[0177] IFFT transform computation unit 108 then computes the IFFT transform of the downmixed mono channel (M) from downmix unit 107 to generate a time-domain mono channel (M) that is processed in TD preprocessor 109. The IFFT transform used in computation unit 108 is the inverse of the FFT transform used in computation unit 105.
[0178] 6. TD Preprocessing in DFT Stereo Mode In DFT stereo mode, a feature extraction operation (not shown) includes a TD preprocessing operation 159 to extract features used in UNCLR classification and XTALK detection. To perform operation 159, the feature extraction unit (not shown) includes a TD preprocessing unit 109 responsive to the mono channel (M).
[0179] 6.1 Voice Activity Detection UNCLR classification and XTALK detection use a Voice Activity Detection (VAD) algorithm. In LRTD stereo mode, the VAD algorithm is performed separately on the left channel (L) and the right channel (R). In DFT stereo mode, the VAD algorithm is performed on the downmixed mono channel (M). The output of the VAD algorithm is expressed as a binary flag f VAD VAD flag f VAD is too conservative and has a long hysteresis, making it unsuitable for UNCLR classification and XTALK detection. This prevents fast switching between LRTD and DFT stereo modes, for example, at the end of an intense conversation or during a short pause in the middle of speech. Also, the VAD flag f VAD is sensitive to small changes in the input stereo sound signal 190. This leads to false alarms in crosstalk detection and inaccurate selection of stereo mode. Therefore, UNCLR classification and XTALK detection use an alternative measure of voice activity detection based on changes in relative frame energy. For details about the VAD algorithm, see [1].
[0180] 6.1.1 Relative Frame Energy The UNCLR classification and XTALK detection are performed using the absolute energy E L and the absolute energy E of the right channel (R) R The maximum average energy of the input stereo sound signal can be calculated in the logarithmic domain, for example, using the following relationship:
[0181]
number
[0182] where the index n is added to indicate the current frame, and N=160 is the length of the current frame (160 samples long). The value of the maximum average energy in the logarithmic domain, E ave (n) is restricted to the interval <0; ∞>.
[0183] Then, the relative frame energies of the input stereo sound signal are calculated using, for example, the following relationship: ave It can be computed by linearly mapping (n) to the interval <0; 0,9>.
[0184]
number
[0185] where E up (n) is the relative frame energy E rl (n), and E dn (n) is the relative frame energy E rl (n), where the index n indicates the current frame.
[0186] Relative frame energy E rl The boundary of (n) is determined by the noise update counter a En Counter a is updated at each frame based on (n). For more information about this counter, see [1]. En The purpose of (n) is to signal that the background noise level in each channel in the current frame can be updated. This situation is indicated by the counter a En This occurs when the value of (n) is zero. As a non-limiting example, the counter a En (n) is initialized to 6 and increments or decrements every frame with a lower threshold of 0 and an upper threshold of 6.
[0187] In LRTD stereo mode, noise estimation is performed independently for the left (L) and right (R) channels. Two noise update counters are assigned to the left (L) and right (R) channels, respectively. En,L (n) and a En,R(n). The two counters can then be combined into a single binary parameter with the following relationship:
[0188]
number
[0189] In DFT stereo mode, noise estimation is performed on the downmixed mono channel (M). The noise update counter for the mono channel is a En,M (n). The binary output parameters are calculated using the following relationship:
[0190]
number
[0191] UNCLR classification and XTALK detection are based on the relative frame energy E rl Lower bound E of (n) dn (n) or upper bound E up To allow updating of (n), the binary parameter f En (n) is used. Parameter f En When (n) is equal to zero, the lower bound E dn (n) is updated. Parameter f En When (n) is equal to 1, the upper bound E up (n) is updated.
[0192] Relative frame energy E rl Upper bound E of (n) up (n) is the parameter f En It is updated in frames where (n) equals 1.
[0193]
number
[0194] Here, the index n represents the current frame and the index n-1 represents the previous frame.
[0195] The first and second rows in relation (71) represent the slower and faster updates, respectively. Therefore, by using relation (71), the upper bound E up (n) is updated more quickly when energy increases.
[0196] Relative frame energy E rl Lower bound E of (n) dn (n) is the parameter f En It is updated in frames where (n) equals 0. E dn (n)=0.9E dn (n-1)+0.1E ave (n) (72) Here, the lower threshold is 30.0. up The value of (n) is the lower bound E dn If it gets too close to (n), it will be changed as shown below, for example. E up (n)=E dn (n)+20.0, if E up (n) <E dn (n)+20.0 (73)
[0197] 6.1.2 Alternative VAD Flag Estimation The UNCLR classification and XTALK detection are based on the relative frame energy E calculated in relation (71) as the basis for calculating the alternative VAD flag. rl Use a variant of (n). Set the alternative VAD flag for the current frame to f xVAD (n) Alternative VAD flag f xVAD (n) is the VAD flag generated in the noise estimation module of the TD preprocessor 103 / 104 in the LRTD stereo mode, or the VAD flag f generated in the TD preprocessor 109 in the DFT stereo mode. VAD , the relative frame energy E rl An auxiliary binary parameter f reflecting the change in (n) ErlIt is calculated by combining with (n).
[0198] First, the relative frame energy E rl (n) is averaged over the division of the 10 previous frames, for example using the following relationship:
[0199]
number
[0200] where p is the exponent of the mean. The auxiliary binary parameters are set, for example, according to the following logic:
[0201]
number
[0202] In LRTD stereo mode, the alternative VAD flag f xVAD (n) is the VAD flag f in the left channel (L), for example, using the following relationship: VAD,L (n) and the VAD flag f in the right channel (R) VAD,R (n) and the auxiliary binary parameter f Erl It is calculated using the logical combination with (n). f xVAD (n)=(f VAD,L (n) OR f VAD,R (n)) AND f Erl (n) (76)
[0203] In DFT stereo mode, the alternative VAD flag f xVAD (n) is the VAD flag f in the downmixed mono channel (M), for example, using the following relationship: VAD,M (n) and the auxiliary binary parameter f Erl It is calculated using the logical combination with (n). f xVAD (n)=f VAD,M (n) AND f Erl (n) (77)
[0204] 6.2 Stereo Silence Flag In DFT stereo mode, it is also convenient to calculate a discrete parameter reflecting the low level of the downmixed mono channel (M). Such a parameter, called a stereo silence flag, can be calculated, for example, by comparing the average level of the active signal with a certain predetermined threshold. For example, the long-term active speech level calculated within the VAD algorithm of the TD preprocessor 109.
[0205]
number
[0206] can be used as the basis for computing the stereo silence flag.
[0207] For details about the VAD algorithm, see [1].
[0208] The stereo silence flag may then be calculated using the following relationship:
[0209]
number
[0210] where E M (n) is the absolute energy of the downmixed mono channel (M) in the current frame. The stereo silence flag f sil (n) is restricted to the interval <0; ∞>.
[0211] 7. Classification of Uncorrelated Stereo Content (UNCLR) UNCLR classification in LRTD and DFT stereo modes is based on a logistic regression (LogReg) model (see reference [9]). The LogReg model is trained separately for LRTD and DFT stereo modes on a large labeled database of correlated and uncorrelated stereo signal samples. Uncorrelated stereo training samples are artificially created by combining randomly selected mono samples. The following stereo situations are simulated with such an artificial mix of mono samples: - Speaker A in the left channel and Speaker B in the right channel (or vice versa). - Speaker A in the left channel and music sounds in the right channel (or vice versa). - Speaker A in the left channel and noise in the right channel (or vice versa). - Talker A in the left or right channel and background noise in both channels. - Speaker A on the left or right channel and background music on both channels.
[0212] In a non-limiting implementation, mono samples are selected from the AT&T mono clean speech database sampled at 16 kHz. Only the active segments are extracted from the mono samples using any convenient VAD algorithm, such as the VAD algorithm of the 3GPP EVS codec described in Reference [1]. The total size of the stereo training database with uncorrelated content is approximately 240 MB. No level adjustment is applied to the mono signals before they are combined to form the stereo sound signal. Level adjustment is only applied after this purpose. The level of each stereo sample is normalized to -26 dBov based on a passive mono downmix. Therefore, the inter-channel level difference is not changed and remains the primary factor determining the location of the dominant speaker in the stereo context.
[0213] The correlated stereo training samples are obtained from various real recordings of stereo sound signals. The total size of the training database with correlated stereo content is approximately 220 MB. In a non-limiting implementation, the correlated stereo training samples include samples from the following situations shown in Figure 4, which shows a top-down view of the stereo situation setup for the real recording: - Speaker S1 at position P1 closer to microphone M1 and speaker S2 at position P2 closer to microphone M6. - Speaker S1 at position P4 closer to microphone M3 and speaker S2 at position P3 closer to microphone M4. - Speaker S1 at position P6, closer to microphone M1, and speaker S2 at position P5, closer to microphone M2. - In the stereo recording of M1-M2, only speaker S1 at position P4. - In the stereo recording of M3-M4, only speaker S1 at position P4.
[0214] The total size of the training database is given as follows: N T =N UNC +N CORR (79) where N UNC is the size of the set of decorrelated stereo training samples, and N CORR is the size of the set of correlated stereo training samples. Labels are assigned manually, for example, using the following simple rule:
[0215]
number
[0216] where Ω UNC is the set of features of the entire uncorrelated training database, and Ω CORR is the set of all features in the correlation training database. In this example, non-limiting implementation, inactive frames (VAD=0) are discarded from the training database.
[0217] Each frame in the uncorrelated training database is labeled with a "1," and each frame in the correlated training database is labeled with a "0." Inactive frames with VAD=0 are ignored during the training process.
[0218] 7.1 UNCLR Classification in LRTD Stereo Mode In LRTD stereo mode, the method 150 for coding a stereo sound signal 190 includes an operation of uncorrelated stereo content (UNCLR) classification 161. To perform operation 161, the device 100 for coding a stereo sound signal 190 comprises an UNCLR classifier 111.
[0219] The UNCLR classification operation 161 in LRTD stereo mode is based on a logistic regression (LogReg) model. The following features are extracted by operating the device 100 for coding stereo sound signals (stereo codec) on both the decorrelated and correlated stereo training databases: - Position k of the maximum of the inter-channel cross-correlation function max (Relationship (11)), - Instantaneous target gain g t (Relationship (13)), - logarithm P of the absolute value of the inter-channel correlation function at zero time lag LR (Relationship (14)), - Side-mono energy ratio r SM (Relationship (15)), - The difference d between the maximum and minimum of the dot product between the left / right channel and the mono signal mmLR (Relationship (19)), - the absolute difference d between the dot product between the left channel (L) and the mono signal (M) and the dot product between the right channel (R) and the mono signal (M) in the logarithmic domain LRM (Relationship (20)), - the zero time lag value R0 of the cross-channel correlation function (relation (5)), and - The inter-channel correlation function RR (relation (21)) is used in the UNCLR classification operation 161.
[0220] In total, the UNCLR classifier 111 uses a number F=8 features.
[0221] Prior to the training process, the UNCLR classifier 111 includes a normalizer (not shown) that performs a sub-operation (not shown) of normalizing the set of features by removing the mean of the set and scaling it to unit variance. For this purpose, the normalizer (not shown) uses, for example, the following relationship:
[0222]
number
[0223] where f i,raw denotes the i-th feature of the set, and f i denotes the normalized i-th feature,
[0224]
number
[0225] denotes the overall mean of the i-th feature across the training database, and σ fi is the total variation of the ith feature across the training database.
[0226] The LogReg model used by the UNCLR classifier 111 takes real-valued features as input vectors and makes a prediction about the likelihood of the input belonging to an uncorrelated class (Class 0), which indicates uncorrelated stereo content (UNCLR). To that end, the UNCLR classifier 111 includes a score calculator (not shown) that performs sub-operations (not shown) to calculate a score representative of the uncorrelated stereo content in the input stereo sound signal 190. The score calculator (not shown) calculates the real-valued output of the LogReg model in the form of a linear regression of the extracted features, which can be expressed using the following relationship: y p =b0+b i f i +...+b F f F (82) where b i denotes the coefficient of the LogReg model, and f i indicates the individual features, and then the real-valued output y p is converted to a probability using, for example, the following logistic function:
[0227]
number
[0228] The probability p(class=0) takes a real value between 0 and 1. Intuitively, a probability closer to 1 means that the current frame is highly stereo decorrelated, i.e., has decorrelated stereo content.
[0229] The goal of the learning process is to find the coefficients b i, i=1,.., the goal is to find the best value for F. The coefficients are found iteratively by minimizing the difference between the predicted output p (class=0) and the true output y based on a training database. The UNCLR classifier 111 in LRTD stereo mode is trained using an iterative method of stochastic gradient descent (SGD), for example, as described in Reference
[10] , the entire contents of which are incorporated herein by reference.
[0230] Binary classification can be performed by comparing the probabilistic output p (class=0) with a fixed threshold such as 0.5. However, for the purposes of UNCLR classification in LRTD stereo mode, the probabilistic output p (class=0) is not used. Instead, the raw output y of the LogReg model is used. p is further processed as shown below.
[0231] The score calculator (not shown) of the UNCLR classifier 111 calculates the raw output y of the LogReg model using, for example, a function such as that shown in FIG. p We first normalize the normalization function applied to the raw output of the LogReg model in UNCLR classification in LRTD stereo mode.
[0232] The normalization function in FIG. 5 can be written mathematically as follows:
[0233]
number
[0234] 7.1.1 LogReg Output Weighting Based on Relative Frame Energy A score calculation unit (not shown) of the UNCLR classifier 111 then calculates the normalized output y of the LogReg model using, for example, the following relationship: pn (n) is weighted by the relative frame energy. scr UNCLR (n)=ypn (n)·E rl (n) (85) where E rl (n) is the relative frame energy described by the relation (69). The normalized and weighted output of the LogReg model, scr UNCLR (n) is referred to as the "score" above, which represents or is uncorrelated with the stereo content in the input stereo sound signal 190.
[0235] 7.1.2 Rising Edge Detection Score scr UNCLR Since (n) contains occasional short-term "peaks" resulting from an imperfect statistical model, it cannot be directly used by the UNCLR classifier 111 for UNCLR classification. These peaks can be filtered out by a simple averaging filter, such as a first-order IIR filter. Unfortunately, applying such an averaging filter typically results in blurring of rising edges that represent the transition between stereo-correlated and stereo-uncorrelated content in the input stereo sound signal 190. To preserve rising edges, the smoothing process (application of the averaging IIR filter) is slowed down or even stopped when a rising edge is detected in the input stereo sound signal 190. Detection of a rising edge in the input stereo sound signal 190 is determined by the relative frame energy E rl This is done by analyzing the opening of (n).
[0236] Relative frame energy E rl The rising edge of (n) is found by filtering the relative frame energy with a cascade of P=20 identical first-order resistor-capacitor (RC) filters, each having, for example, the following form:
[0237]
number
[0238] The constants a0, a1, and b1 are selected so that the following relationship holds:
[0239]
number
[0240] Therefore, a single parameter τ edge is used to control the time constant of each RC filter. Experimentally, good results have been obtained with τ edge It is known that the relative frame energy E rl The filtration of (n) can be carried out as follows.
[0241]
number
[0242] where the superscripts p=0, 1,..., P-1 are added to indicate the stages in the cascade of RC filters. The output of the cascade of RC filters is equal to the output from the last stage, i.e.,
[0243]
number
[0244] The reason for using a cascade of first-order RC filters instead of a single higher-order RC filter is to reduce computational complexity. A cascade of multiple first-order RC filters acts as a low-pass filter with a relatively sharp step function. A cascade of multiple first-order RC filters reduces the relative frame energy E rl When used in (n), it attempts to blur occasional short-term spikes while preserving slower but important transitions such as onset and offset. Relative frame energy E rl The rising edge of (n) can be quantified by calculating the difference between the relative frame energy and the filtered output, for example, using the following relationship: f edge (n)=0.95-0.05(E rl (n)-E f (n)) (90)
[0245] term f edge (n) is restricted to the interval <0,9; 0,95>. The score calculation unit (not shown) of the UNCLR classifier 111 calculates f using, for example, the following relationship to generate the normalized, weighted, and smoothed score (the output of the LogReg model): edge The normalized and weighted output scr of the LogReg model is an IIR filter that uses (n) as a forgetting factor. UNCLR (n) is smoothed. wscr UNCLR (n)=f edge (n)·wscr UNCLR (n-1)+(1-f edge (n))·scr UNCLR (n) (91)
[0246] 7.2 UNCLR Classification in DFT Stereo Mode In DFT stereo mode, the method 150 for coding a stereo sound signal 190 includes an operation of uncorrelated stereo content (UNCLR) classification 163. To perform operation 163, the device 100 for coding a stereo sound signal 190 comprises an UNCLR classifier 113.
[0247] The UNCLR classification in DFT stereo mode is performed similarly to the UNCLR classification in LRTD stereo mode as described above. Specifically, the UNCLR classification in DFT stereo mode is also based on a logistic regression (LogReg) model. For simplicity, the symbols / names denoting specific parameters and associated mathematical symbols from the UNCLR classification in LRTD stereo mode are also used for DFT stereo mode. Subscripts are added to avoid ambiguity when referring to the same parameters from multiple parts simultaneously.
[0248] The following features are extracted by operating the device 100 for coding stereo sound signals (stereo codec) on both the stereo decorrelated training database and the stereo correlated training database: - ILD gain g ILD (Relation 43) - IPD gain g IPD (Relationship 48) - IPD rotation angle φ rot (Relation 49) - Prediction gain g pred (Relationship 52) - mean energy of inter-channel coherence E coh (Relationship 55) - Ratio r between the maximum inter-channel amplitude product and the minimum inter-channel amplitude product PP (Relation 57) - total inter-channel spectral magnitude f X (Relation 41)), and - Maximum value G of GCC-PHAT functions ITD (Relation 61) is used by the UNCLR classifier 113 for UNCLR classification in DFT stereo mode.
[0249] In total, the UNCLR classifier 113 uses a number F=8 features.
[0250] Prior to the training process, the UNCLR classifier 113 includes a normalizer (not shown) that performs a sub-operation (not shown) of normalizing the set of features by removing the mean of the set and scaling it to unit variance. For this purpose, the normalizer (not shown) uses, for example, the following relationship:
[0251]
number
[0252] where f i,raw denotes the i-th feature of the set,
[0253]
number
[0254] denotes the overall mean of the i-th feature across the entire training database, and σ fi is the overall variation of the ith feature across the entire training database.
[0255] The overall mean used in relation (92)
[0256]
number
[0257] and the overall change σ fi It should be noted that is different from the same parameters used in relation (81).
[0258] The LogReg model used in DFT stereo mode is similar to the LogReg model used in LRTD stereo mode. The output of the LogReg model, y P is described by relation (82), and the probability that the current frame has uncorrelated stereo content (class=0) is given by relation (83). The classifier training process and the procedure for finding the optimal decision threshold have been described earlier in this specification. Again, to that end, the UNCLR classifier 113 comprises a score calculation unit (not shown) that performs sub-operations (not shown) to calculate a score representing uncorrelated stereo content in the input stereo sound signal 190.
[0259] The score calculation unit (not shown) of the UNCLR classifier 113 calculates the raw output y of the LogReg model according to the same function as in LRTD stereo mode, as shown in FIG. p is first normalized. Normalization can be written mathematically as:
[0260]
number
[0261] 7.2.1 LogReg Output Weighting Based on Relative Frame Energy A score calculation unit (not shown) of the UNCLR classifier 113 then calculates the normalized output y of the LogReg model using, for example, the following relationship: pn (n) is the relative frame energy E rl Weighted by (n). scr UNCLR (n)=y pn (n)·E rl (n) (94) where E rl (n) is the relative frame energy described by the relationship (69).
[0262] The normalized and weighted output of the LogReg model is called the "score" and represents the same quantity as in the LRTD stereo mode described earlier. In DFT stereo mode, the score scr UNCLR (n) is the alternative VAD flag f xVAD When (n) (relationship (77)) is set to 0, it is reset to 0. This is expressed by the following relation: scr UNCLR (n)=0, f xVAD (n)=0 (95)
[0263] 7.2.2 Rising Edge Detection in DFT Stereo Mode Finally, a score calculation unit (not shown) of the UNCLR classifier 113 calculates the score scr in DFT stereo mode using the rising edge detection mechanism described above for UNCLR classification in LRTD stereo mode. UNCLR (n) is smoothed with an IIR filter. To that end, the UNCLR classifier 113 uses the following relationship: wscr UNCLR (n)=f edge (n)·wscr UNCLR(n-1)+(1-f edge (n))·scr UNCLR (n) (96) This is the same as relation (91).
[0264] 7.3 Binary UNCLR determination The final output of the UNCLR classifier 111 / 113 is a binary state. UNCLR (n) indicates the binary state of the UNCLR classifier 111 / 113. UNCLR (n) has a value of "1" to indicate an uncorrelated stereo content class or a value of "0" to indicate a correlated stereo content class. The binary state at the output of the UNCLR classifier 111 / 113 is variable. The binary state is initialized to "0". The state of the UNCLR classifier 111 / 113 changes from the current class to another class in frames where certain conditions are met.
[0265] The mechanism used in the UNCLR classifier 111 / 113 for switching between stereo content classes is depicted in FIG. 6 in the form of a state machine.
[0266] Referring to Figure 6, - (a) Binary state c of the previous frame UNCLR (n-1) is "1" (601), (b) the smoothed score of the current frame, wscr UNCLR (n) is smaller than "-0.07" (602), (c) the variable cnt of the previous frame sw If (n-1) is greater than '0' (603), the binary state of the current frame c UNCLR (n) is switched to "0" (604). - (a) Binary state c of the previous frame UNCLR (n-1) is "1" (601), (b) the smoothed score of the current frame, wscr UNCLR If (n) is not less than "-0.07" (602), the binary state c UNCLR There is no switching of (n). - (a) Binary state c of the previous frame UNCLR(n-1) is "1" (601), (b) the smoothed score of the current frame, wscr UNCLR (n) is smaller than "-0.07" (602), (c) the variable cnt of the previous frame sw If (n-1) is not greater than "0" (603), then the binary state c in the current frame UNCLR There is no switching of (n).
[0267] In the same manner, and with reference to FIG. - (a) Binary state c of the previous frame UNCLR (n-1) is "0" (601), (b) the smoothed score of the current frame, wscr UNCLR (n) is greater than "0.1" (605), (c) the variable cnt of the previous frame sw If (n-1) is greater than '0' (606), the binary state c of the current frame UNCLR (n) is switched to "1" (607). - (a) Binary state c of the previous frame UNCLR (n-1) is "0" (601), (b) the smoothed score of the current frame, wscr UNCLR If (n) is not greater than 0.1 (605), then the binary state c UNCLR There is no switching of (n). - (a) Binary state c of the previous frame UNCLR (n-1) is "0" (601), (b) the smoothed score of the current frame, wscr UNCLR (n) is greater than "0.1" (605), (c) the variable cnt of the previous frame sw If (n-1) is not greater than "0" (606), then the binary state c in the current frame UNCLR There is no switching of (n).
[0268] Finally, the variable cnt in the current frame sw (n) is updated (608) and the procedure is repeated for the next frame (609).
[0269] Variable cnt sw(n) is a frame counter for the UNCLR classifier 111 / 113, which can switch between LRTD stereo mode and DFT stereo mode. This counter is initialized to zero and updated at each frame (608), for example, using the following logic:
[0270]
number
[0271] Counter cnt sw (n) has an upper limit of 100. Variable c type indicates the type of the current frame in the device 100 for coding a stereo audio signal. The frame type is usually determined in a preprocessing operation of the device 100 for coding a stereo audio signal (stereo audio codec), explicitly in the preprocessing unit 103 / 104 / 109. The type of the current frame depends on the following characteristics of the input stereo audio signal 190: - Pitch Period - Vocalization - Spectral tilt - Zero crossing rate - Frame energy difference (short term, long term) is usually selected based on
[0272] As a non-limiting example, a frame type from the 3GPP EVS codec as described in reference [1] is type The frame types in the 3GPP EVS codec are selected from the following set of classes:
[0273]
number
[0274] The parameter VAD0 in relation (97) is the VAD flag without hangover addition. The VAD flag without hangover addition is often calculated in a pre-processing operation of the device for coding a stereo audio signal 100 (stereo audio codec), specifically in the TD pre-processing unit 103 / 104 / 109. As a non-limiting example, the VAD flag without hangover addition from the 3GPP EVS codec as described in reference [1] can be used in the UNCLR classifier 111 / 113 as the parameter VAD0.
[0275] Output binary state c of UNCLR classifier 111 / 113 UNCLR (n) can be changed if the type of the current frame is general, unvoiced, or inactive, or if the VAD flag without hangover addition indicates inactive (VAD0=0) in the input stereo signal. Such frames are generally suitable for switching between LRTD and DFT stereo modes, since they are located either in stable segments or in segments with little perceptual impact on quality. The goal is to minimize the risk of switching artifacts.
[0276] 8. Crosstalk (XTALK) detection XTALK detection is based on LogReg models that are trained separately for LRTD and DFT stereo modes. Both statistical models are trained with features collected from a large database of real stereo recordings and artificially prepared stereo samples. In the training database, each frame is labeled as either single talk or crosstalk. Labeling is done either manually for real stereo recordings or semi-automatically for artificially prepared samples. Manual labeling is done by identifying short, compact segments with crosstalk characteristics. Semi-automatic labeling is done using the VAD outputs from the mono signal before they are mixed into the stereo signal. More details are provided at the end of Section 8.
[0277] In a non-limiting example of the implementation described in this disclosure, actual stereo recordings were sampled at 32 kHz. The total size of these actual stereo recordings is approximately 263 MB, corresponding to approximately 30 minutes. The artificially prepared stereo samples are created by mixing randomly selected speakers from a monophonic clean speech database using an ITU-T G.191 reverberation instrument. The artificially prepared stereo samples are prepared by simulating conditions in a large conference room with an AB microphone setup as shown in FIG. 7, which is a schematic floor plan of a large conference room with an AB microphone setup in which conditions are simulated for XTALK detection.
[0278] Two types of rooms are considered: reverberant (LEAB) and anechoic (LAAB). Referring to FIG. 7, for each type of room, a first speaker S1 may appear at position P4, P5, or P6, and a second speaker S2 may appear at positions P10, P11, and P12. The positions of each speaker S1 and S2 are randomly selected during the preparation of training samples. Thus, speaker S1 is always close to the first simulated microphone M1, and speaker S2 is always close to the second simulated microphone M2. Microphones M1 and M2 are omnidirectional in the illustrated non-limiting implementation of FIG. 7. The pair of microphones M1 and M2 constitutes a simulated AB microphone setup. Mono samples are randomly selected from the training database, downsampled to 32 kHz, and normalized to -26 dBov (dB(overload) - the amplitude of the acoustic signal compared to the maximum that the device can handle before clipping occurs) before further processing. The ITU-T G.191 reverberation instrument contains a database of actual measurements of the room impulse response (RIR) for each talker / microphone pair.
[0279] Randomly selected mono samples from speakers S1 and S2 are then convolved with the room impulse response (RIR) corresponding to the given speaker / microphone, thereby simulating a real AB microphone capture. The contributions from both speakers S1 and S2 at each microphone M1 and M2 are added together. A randomly selected offset ranging from 4 to 4.5 seconds is added to one of the speaker samples before convolution. This ensures that in all training sentences, there is always some period of single talk speech followed by a short period of crosstalk speech and other periods of single talk speech. After RIR convolution and mixing, the samples are renormalized to -26 dBov, and this time is applied to the passive mono downmix.
[0280] Labels are generated semi-automatically using a conventional VAD algorithm, such as the VAD algorithm of the 3GPP EVS codec described in Reference [1]. The VAD algorithm is applied separately to the first speaker (S1) file and the second speaker (S2) file. Then, both binary VAD decisions are combined using a logical "and". This results in a label file. The section where the combined output is equal to "1" determines the crosstalk section. This is shown in FIG. 8, which shows a graph illustrating the automatic labeling of crosstalk samples using VAD. In FIG. 8, the first line shows the speech samples from speaker S1, the second line shows the binary VAD decision for the speech samples from speaker S1, the third line shows the speech samples from speaker S2, the fourth line shows the binary VAD decision for the speech samples from speaker S2, and the fifth line shows the location of the crosstalk section.
[0281] The training set is unbalanced: the ratio of crosstalk to single-talk frames is roughly 1:5, meaning that only about 21% of the training data belongs to the crosstalk class. This is compensated for during the LogReg training process by applying class weights as described in Reference [6], the entire contents of which are incorporated herein by reference.
[0282] The training samples are concatenated and used as input to a device 100 for coding a stereo audio signal (stereo audio codec). Features are collected individually in separate files during the encoding process for each 20 ms frame. This constitutes the training feature set. The total number of frames in the training feature set is, for example, N T =N XTALK +N NORMAL (98) where N XTALK is the total number of crosstalk frames, and N NORMAL is the total number of single talk frames.
[0283] Also, the corresponding binary labels are shown as follows:
[0284]
number
[0285] where Ω XTALK is a superset of all crosstalk frames, and Ω NORMAL is a superset of all single talk frames. Inactive frames (VAD=0) are removed from the training database.
[0286] XTALK detection in 8.1 LRTD stereo mode In LRTD stereo mode, the method 150 for encoding a stereo sound signal includes an operation of detecting crosstalk (XTALK) 160. To perform operation 160, the device 100 for encoding a stereo sound signal comprises an XTALK detection unit 110.
[0287] The operation 160 of detecting crosstalk (XTALK) in LRTD stereo mode is performed similarly to the UNCLR classification in LRTD stereo mode described above. The XTALK detection device 110 is based on a logistic regression (LogReg) model. For simplicity, the names of parameters and associated mathematical symbols from the UNCLR classification are also used in this section. Subscripts are added to avoid ambiguity when referring to the same parameter names from different sections.
[0288] The following characteristics: - L / R class difference class (Relationship (32)), - Maximum autocorrelation L / R difference d v (Relationship (25)), - LSF total L / R difference d LSF (Relationship (23)), - L / R difference d of residual error energy LPC13 (Relationship (22)), - L / R difference of correlation map cmap (Relationship (27)), - L / R difference of noise characteristics d nchar (Relationship (29)), - Non-stationary L / R difference d sta (Relationship (26)), - Spectral diversity L / R difference d sdiv (Relationship (28)), - the non-normalized value P of the inter-channel correlation function at zero time lag LR (Relationship (14)), - Side-mono energy ratio r SM (Relationship (15)), - the difference d between the maximum and minimum of the dot products between the left channel and the mono signal and between the right channel and the mono signalmmLR (Relationship (19)), - the zero time lag value R0 of the cross-channel correlation function (relation (5)), - the cross-correlation function between channels, RR (relation (21)), - Position k of the maximum inter-channel cross-correlation function max (Relationship (11)), - Maximum R of inter-channel correlation function max (Relationship (10)), - The difference Δ between the dot product of L / M and the dot product of R / M LRM (Relationship (20)), and - smoothed ratio between the energy of the side signal and the energy of the mono signal
[0289]
number
[0290] (Relationship (16)) is used by the XTALK detector 110.
[0291] Therefore, the XTALK detector 110 uses a total number of features F=17.
[0292] Before the training process, the XTALK detector 110 is configured with 17 features f i The method further comprises a normalizer (not shown) that performs a sub-operation (not shown) of normalizing the set by removing the mean of the set and scaling it to unit variance. The normalizer (not shown) uses, for example, the following relationship:
[0293]
number
[0294] where f i,raw denotes the i-th feature of the set.
[0295]
number
[0296] is the overall mean of the i-th feature across the training database, and σ fi is the total variation of the ith feature across the training database.
[0297] where the parameters used in relation (100)
[0298]
number
[0299] and σ fi is different from the same parameters used in relation (81).
[0300] LogReg model output y P is described by relation (82), and the probability p(class=0) that the current frame belongs to the crosstalk classification class (class 0) is given by relation (83). Details of the training process and the procedure for finding the optimal decision threshold have been provided above in the description of UNCLR classification in LRTD stereo mode. As mentioned above, for that purpose, the XTALK detection device 110 comprises a score calculation device (not shown) that performs sub-operations (not shown) to calculate a score representing the uncorrelated stereo content in the input stereo sound signal 190.
[0301] The score calculation unit (not shown) of the XTALK detection unit 110 calculates the raw output y of the LogReg model using a function such as that shown in FIG. p is normalized and further processed. Figure 9 shows a graph of a function for scaling the raw output of the LogReg model for XTALK detection in LRTD stereo mode. Such normalization can be mathematically written as:
[0302]
number
[0303] Normalized output y of the LogReg model pn If the previous frame is coded in DFT stereo mode and the current frame is coded in LRTD stereo mode, then (n) is set to 0. Such a procedure prevents switching artifacts.
[0304] 8.1.1 LogReg Output Weighting Based on Relative Frame Energy The score calculation unit (not shown) of the XTALK detection unit 110 calculates the relative frame energy E rl Based on (n), the normalized output y of the LogReg model pn (n). The weighting scheme applied in the XTALK detector 110 in LRTD stereo mode is similar to the weighting scheme applied in the UNCLR classifier 111 in LRTD stereo mode, as described earlier in this specification. The main difference is that the relative frame energy E rl (n) is not used directly as a multiplicative factor as in relation (85). Instead, the score calculation unit (not shown) of the XTALK detection unit 110 calculates the relative frame energy E rl (n) is linearly mapped inversely. This mapping can be done, for example, using the following relationship: w relE (n)=-2.375E rl (n)+2.1375 (102)
[0305] Thus, frames with greater relative energy will have weights closer to 0, while frames with less energy will have weights closer to 0.95. A score calculation unit (not shown) of the XTALK detection unit 110 then calculates the normalized output y of the LogReg model using, for example, the following relationship: pn To filter (n), we use weights w relE Use (n). scr XTALK (n)=wrelE scr XTALK (n-1)+(1-w relE )y pn (n) (103) Here, the index n represents the current frame, and the index n-1 represents the previous frame.
[0306] The normalized and weighted output scr from the XTALK detector 110 XTALK (n) is called the “XTALK score” which represents the crosstalk in the input stereo sound signal 190 .
[0307] 8.1.2 Rising Edge Detection In a manner similar to UNCLR classification in LRTD stereo mode, the score calculation unit (not shown) of the XTALK detector 110 calculates the normalized and weighted output scr of the LogReg model. XTALK (n) is smoothed to smear out occasional short-term "peaks" and "dip" that would otherwise result in false alarms or errors. The smoothing is designed to preserve the rising edges of the LogReg output, since these rising edges may represent important transitions between crosstalk and singletalk segments in the input stereo sound signal 190. The mechanism for detecting rising edges in the XTALK detector 110 in LRTD stereo mode differs from the mechanism for detecting rising edges described above for UNCLR classification in LRTD stereo mode.
[0308] In the XTALK detector 110, the rising edge detection algorithm analyzes the LogReg output values from the previous frame and compares them to a set of pre-calculated "ideal" rising edges with different slopes. The "ideal" rising edges are expressed as a linear function of the frame index n. Figure 10 is a graph illustrating the mechanism for detecting rising edges in the XTALK detector 110 in LRTD stereo mode. Referring to Figure 10, the x-axis contains the index n of the frame previous to the current frame 0. The small grey rectangles represent the XTALK score scr over a period of six frames before the current frame. XTALK As can be seen from Figure 10, the XTALK score scr starting three frames before the current frame is XTALK There is a rising edge at (n). The dotted lines depict a set of four "ideal" rising edges at segments of different lengths.
[0309] For each "ideal" rising edge, the rising edge detection algorithm generates a dotted line and an XTALK score scr. XTALK The output of the rising edge detection algorithm is the minimum mean square error between the tested "ideal" rising edges. The dotted linear functions are min and scr max This is shown in Figure 10 by the large light grey rectangle. The slope of the linear function for each "ideal" rising edge depends on the minimum and maximum thresholds and the length of the segment.
[0310] Rising edge detection is performed by XTALK detector 110 only in frames that meet the following criteria:
[0311]
number
[0312] where K=4 is the maximum length of the rising edge tested.
[0313] The output value of the rising edge detection algorithm is ε 0_1 The use of the "0_1" subscript emphasizes the fact that the output value of the rising edge detection is bounded in the interval <0; 1>. For frames that do not satisfy the criteria in relation (104), the output value of the rising edge detection is directly set to 0, i.e., ε 0_1 =0 (105)
[0314] The set of linear functions that represent the tested "ideal" rising edge can be mathematically expressed by the following relationships:
[0315]
number
[0316] where the index l denotes the length of the rising edge being tested, and the index nk denotes the frame index. The slope of each linear function depends on three parameters: the length l of the rising edge being tested, the minimum threshold scr min , and the maximum threshold scr max For the purposes of the XTALK detector 110 in LRTD stereo mode, the threshold is determined by max =1.0 and scr min = -0.2. These threshold values were found experimentally.
[0317] For each length of the rising edge tested, the rising edge detection algorithm calculates the linear function t (relationship (106)) and the XTALK score scr using, for example, the following relationship: XTALK Calculate the mean square error between
[0318]
number
[0319] where ε is the initial error given by the relation: ε0=|scr XTALK (n)-scr max | 2 (108)
[0320] The minimum mean square error is calculated by the XTALK detector 110 using the following relationship:
[0321]
number
[0322] The smaller the minimum mean square error, the stronger the detected rising edge. In a non-limiting implementation, if the minimum mean square error is greater than 0.3, the output of the rising edge detection is set to 0, i.e., ε 0_1 > if ε min > 0.3 (110) and the rising edge detection algorithm terminates. In all other cases, the minimum mean square error can be mapped linearly in the interval <0; 1>, for example using the following relationship: ε 0_1 =1-2.5ε min (111)
[0323] In the above example, the relationship between the power of the rising edge detection and the minimum mean square error is inversely proportional.
[0324] The XTALK detector 110 normalizes the output of the rising edge detection in the interval <0,5; 0,9> to produce an edge sharpening parameter that is calculated, for example, using the following relationship: f edge (n)=0.9-0.4ε 0_1 (112) 0,5 and 0,9 are used as lower and upper bounds respectively.
[0325] Finally, the score calculation unit (not shown) of the XTALK detection unit 110 calculates f edge Using the IIR filter of the XTALK detector 110, where (n) is used instead of the forgetting factor, the LogReg model scr XTALK 2. Smoothing the normalized and weighted output of (n). Such smoothing may be done using, for example, the following relationship: wscr XTALK (n)=f edge (n)·wscr XTALK (n-1)+(1-f edge (n))·scr XTALK (n) (113)
[0326] Smoothed output wscr XTALK (n) (XTALK score) is reset to 0 in frames where the alternative VAD flag calculated in relation (77) is zero. That is, wscr XTALK (n)=0, if f xVAD (n)=0 (114)
[0327] 8.2 Crosstalk detection in DFT stereo mode In DFT stereo mode, the method 150 for encoding a stereo sound signal 190 includes an operation 162 of detecting crosstalk (XTALK). To perform operation 162, the device 100 for encoding a stereo sound signal 190 comprises an XTALK detection unit 112.
[0328] XTALK detection in DFT stereo mode is performed similarly to XTALK detection in LRTD stereo mode. A logistic regression (LogReg) model is used for binary classification of the input feature vector. For brevity, the names and associated mathematical symbols of specific parameters from XTALK detection in LRTD stereo mode are also used in this section. Subscripts are added to avoid ambiguity when referring to the same parameters from the two parts simultaneously.
[0329] The following characteristics: - ILD gain g ILD (Relation 43) - IPD gain g IPD (Relationship 48) - IPD rotation angle φ rot (Relation 49) - Prediction gain g pred (Relationship 52) - mean energy of inter-channel coherence E coh (Relationship 55) - Ratio r between the maximum inter-channel amplitude product and the minimum inter-channel amplitude product PP (Relation 57) - total inter-channel spectral magnitude f X (Relationship 41) - Maximum value G of GCC-PHAT functions ITD (Relationship 61) - The relationship between the amplitude of the first highest peak and the amplitude of the second highest peak of the GCC-PHAT function GITD12 (Relation 64) - Amplitude m of the second highest peak of GCC-PHAT ITD2 (Relation 66)), and - the difference Δ between the position of the second highest peak in the current frame and the position of the second highest peak in the previous frame ITD2 (Relation 67) are extracted from the device 100 for coding a stereo sound signal 190 by operating the DFT stereo mode on both the single-talk training database and the cross-talk training database.
[0330] In total, the XTALK detector 112 uses a number F=11 features.
[0331] Prior to the training process, the XTALK detector 112 includes a normalizer (not shown) that performs a sub-operation (not shown) of normalizing the set of extracted features by removing the overall mean of the set and scaling it to unit variance, for example, using the following relationship:
[0332]
number
[0333] where f i,raw denotes the i-th feature of the set, and f i denotes the normalized i-th feature,
[0334]
number
[0335] denotes the overall mean of the i-th feature across the training database, and σ fi is the total variation of the ith feature across the training database, where the parameters used in relation (115)
[0336]
number
[0337] and σ fi is different from that used in relation (81).
[0338] The output of the LogReg model is fully described by relation (82), and the probability that the current frame belongs to the crosstalk classification class (class 0) is given by relation (83). Details of the training process and the procedure for finding the optimal decision threshold are provided above in the section on UNCLR classification in LRTD stereo mode. Again, for that purpose, the XTALK detection unit 112 comprises a score calculation unit (not shown) that performs sub-operations (not shown) to calculate a score representative of XTALK detection in the input stereo sound signal 190.
[0339] The score calculator (not shown) of the XTALK detector 112 calculates the raw output y of the LogReg model using a function such as that shown in FIG. pis normalized for further processing. The normalized output of the LogReg model is y pn In DFT stereo mode, weighting based on relative frame energy is not used, so the normalized and weighted output of the LogReg model, specifically the XTALK score scr XTALK (n) is given by the following relation: scr XTALK (n)=y pn (116)
[0340] XTALK score scr XTALK (n) is the alternative VAD flag f xVAD When (n) is set to 0, it is reset to 0. This can be expressed as the following relationship: scr XTALK (n)=0, if f xVAD (n)=0 (117)
[0341] 8.2.1 Rising Edge Detection As in the case of XTALK detection in LRTD stereo mode, a score calculation unit (not shown) of the XTALK detector 112 calculates the XTALK score scr to remove short-term peaks. XTALK (n). Such smoothing is performed using IIR filtering using a rising edge detection mechanism as described for the XTALK detector 110 in LRTD stereo mode. The XTALK score scr XTALK (n) is smoothed with an IIR filter, for example using the following relationship: wscr XTALK (n)=f edge (n)·wscr XTALK (n-1)+(1-f edge (n))·scr XTALK (n) (118) where f edge (n) is the edge sharpening parameter calculated by relation (112).
[0342] 8.3 Binary XTALK determination The final output of the XTALK detector 110 / 112 is binary. XTALK (n) indicates the output of the XTALK detector 110 / 112, with "1" representing crosstalk and "0" representing single talk class. XTALK (n) can also be considered as a state variable. Output c XTALK (n) is initialized to 0. The state variables are changed from the current class to another class only in frames where certain conditions are met. The mechanism for crosstalk class switching is similar to the mechanism for class switching in uncorrelated stereo content, detailed above in Section 7.3. However, there are differences for both LRTD and DFT stereo modes. These differences are detailed below.
[0343] In LRTD stereo mode, the XTALK detector 110 uses a crosstalk switching mechanism as shown in Figure 11. Referring to Figure 11, the following is true. - Output c of the UNCLR classifier 111 for the current frame n UNCLR If (n) is equal to "1" (1101), the output c of the XTALK detection device 110 in the current frame n XTALK There is no switching of (n). (a) Output c of the UNCLR classifier 111 for the current frame n UNCLR (n) is equal to "0" (1101), and (b) the output c of the XTALK detection device 110 in the previous frame n-1 XTALK If (n-1) is equal to "1" (1102), the output c of the XTALK detection device 110 in the current frame n XTALK There is no switching of (n). (a) Output c of the UNCLR classifier 111 for the current frame n UNCLR (n) is equal to "0" (1101), and (b) the output c of the XTALK detection device 110 in the previous frame n-1 XTALK (n-1) is equal to "0" (1102), (c) the smoothed XTALK score wscr in the current frame n XTALKIf (n) is not greater than 0.03 (1104), the output c of the XTALK detector 110 in the current frame n XTALK There is no switching of (n). (a) Output c of the UNCLR classifier 111 for the current frame n UNCLR (n) is equal to "0" (1101), and (b) the output c of the XTALK detection device 110 in the previous frame n-1 XTALK (n-1) is equal to "0" (1102), (c) the smoothed XTALK score wscr in the current frame n XTALK (n) is greater than 0.03 (1104), and (d) the counter cnt in the previous frame n-1 sw If (n-1) is not greater than "0" (1105), the output c of the XTALK detection device 110 in the current frame n XTALK There is no switching of (n). (a) Output c of the UNCLR classifier 111 for the current frame n UNCLR (n) is equal to "0" (1101), and (b) the output c of the XTALK detection device 110 in the previous frame n-1 XTALK (n-1) is equal to "0" (1102), (c) the smoothed XTALK score wscr in the current frame n XTALK (n) is greater than 0.03 (1104), and (d) the counter cnt in the previous frame n-1 sw If (n-1) is greater than "0" (1105), the output c of the XTALK detection device 110 in the current frame n XTALK (n) is switched to "1" (1106).
[0344] Finally, the counter cnt for the current frame n sw (n) is updated (1107) and the procedure is repeated for the next frame (1108).
[0345] Counter cnt sw (n) is common to the UNCLR classifier 111 and the XTALK detector 110 and is defined in relation (97). sw Positive values of (n) are the state variables c XTALK(n) (output c of XTALK detection device 110) XTALK As can be seen in Figure 11, the switching logic indicates that the output c(n) of the UNCLR classifier 111 in the current frame is allowed to be switched. UNCLR (n)(1101). Therefore, it is assumed that the UNCLR classifier 111 is operated before the XTALK detector 110, since the XTALK detector 110 uses the output of the UNCLR classifier 111. Also, the state switching logic of FIG. 11 uses the output c of the XTALK detector 110. XTALK It is unidirectional in the sense that (n) can only change from "0" (single talk) to "1" (cross talk). The state switching logic for the opposite direction, i.e., from "1" (cross talk) to "0" (single talk), is part of the DFT / LRTD stereo mode switching logic, which is explained later in this disclosure.
[0346] In DFT stereo mode, the XTALK detector 112 includes an auxiliary parameter calculator (not shown) that performs sub-operations (not shown) to calculate the following auxiliary parameters: XTALK (n) and the following auxiliary parameters: - Voice Activity Detection (VAD) flag for the current frame (f VAD ), - Amplitude G of the first and second highest peaks of the GCC-PHAT function ITD , m ITD2 (relationships (61) and (66) respectively), - Positions (ITD values) d corresponding to the amplitudes of the first and second highest peaks of the GCC-PHAT function ITD , d ITD2 (Relations (60) and (Paragraph
[0170] (Original paragraph
[0111] ))), and - DFT stereo silence flag f sil (Relationship (78)) Use and.
[0347] In DFT stereo mode, the XTALK detector 112 uses a crosstalk switching mechanism as shown in Figure 12. Referring to Figure 12, the following is true. -d ITD If (n) is equal to "0" (1201), then c XTALK (n) is switched to "0" (1217). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK If (n-1) is not equal to "0" (1202), (c)c XTALK If (n-1) is not equal to "1" (1215), then c XTALK There is no switching of (n). (c)c XTALK (n-1) is equal to "1" (1215), (d) wscr XTALK If (n) is not less than "0.0" (1216), then c XTALK There is no switching of (n). (c)c XTALK (n-1) is equal to "1" (1215), (d) wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is not equal to "1" (1203), (d)c XTALK If (n-1) is not equal to "1" (1215), then c XTALK There is no switching of (n). (d)c XTALK (n-1) is equal to "1" (1215), (e)wscr XTALK If (n) is not less than "0.0" (1216), then c XTALK There is no switching of (n). (d)c XTALK (n-1) is equal to "1" (1215), (e)wscr XTALKIf (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), (d) 0.8G ITD (n) is m ITD2 (n) smaller than (1204), (e) 0.8G ITD (n-1) is m ITD2 (n-1) is smaller than (1205), (f)d ITD2 (n)-d ITD2 (n-1) is smaller than "4.0" (1206), (g)G ITD (n) is greater than "0.15" (1207), (h) G ITD If (n-1) is greater than 0.15 (1208), then c XTALK (n) is switched to "1" (1218). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), and (d) any of the tests 1204 to 1208 is false; (e)wscr XTALK If (n) is greater than "0.8" (1209), then c XTALK (n) is switched to "1" (1218). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), (d) any of the tests 1204 to 1208 is false, and (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil If (n) is not equal to "1" (1210), (g)c XTALK If (n-1) is not equal to "1" (1215), then c XTALKThere is no switching of (n). (g)c XTALK (n-1) is equal to "1" (1215), (h)wscr XTALK If (n) is not less than "0.0" (1216), then c XTALK There is no switching of (n). (g)c XTALK (n-1) is equal to "1" (1215), (h)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), (d) any of the tests 1204 to 1208 is false, and (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil (n) is equal to "1" (1210), (g) d ITD (n) is greater than "8.0" (1211), (h)d ITD If (n-1) is less than "-8.0", then c XTALK (n) is switched to "1" (1218). - (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), (d) any of the tests 1204 to 1208 is false, and (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil (n) is equal to "1" (1210), (g) either test 1211 or 1212 is false, (h) d ITD (n-1) is greater than "8.0" (1213), (i)d ITD If (n) is less than "-8.0" (1214), then c XTALK (n) is switched to "1" (1218). - (a)d ITD(n) is not equal to "0" (1201), (b)c XTALK (n-1) is equal to "0" (1202), (c)f VAD is equal to "1" (1203), (d) any of the tests 1204 to 1208 is false, and (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil If (n) equals "1" (1210), (g) either test 1211 or 1212 is false, and (h) either test 1213 or 1214 is false, · (I C XTALK If (n-1) is not equal to "1" (1215), then c XTALK There is no switching of (n). · (I C XTALK (n-1) is equal to "1" (1215), (j)wscr XTALK If (n) is not less than "0.0" (1216), then c XTALK There is no switching of (n). · (I C XTALK (n-1) is equal to "1" (1215), (j)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219).
[0348] Finally, the counter cnt for the current frame n sw (n) is updated (1220) and the procedure is repeated for the next frame (1221).
[0349] Variable cnt sw (n) is the counter of frames in which it is possible to switch between LRTD stereo mode and DFT stereo mode. sw (n) is common to the UNCLR classifier 113 and the XTALK detector 112. Counter cnt sw (n) is initialized to zero and updated at each frame according to the relationship (97).
[0350] 9. DFT / LRTD stereo mode selection The method 150 for coding a stereo sound signal 190 includes an operation 164 of selecting an LRTD stereo mode or a DFT stereo mode. To perform operation 164, the device 100 for coding a stereo sound signal 190 includes an LRTD / DFT stereo mode selection unit 114 that is delayed by one frame (191) and receives an XTALK decision from an XTALK detection unit 110, an UNCLR decision from an UNCLR classifier 111, an XTALK decision from an XTALK detection unit 112, and an UNCLR decision from an UNCLR classifier 113.
[0351] The LRTD / DFT stereo mode selector 114 selects the binary output c of the UNCLR classifier 111 / 113. UNCLR (n) and XTALK detector 110 / 112 binary output c XTALK (n), the LRTD / DFT stereo mode selector 114 also considers some auxiliary parameters, which are primarily used to prevent stereo mode switching in perceptually sensitive segments or to prevent frequent switching in segments where both the UNCLR classifier 111 / 113 and the XTALK detector 110 / 112 do not provide accurate outputs.
[0352] Operation 164 of selecting LRTD or DFT stereo mode is performed before downmixing and encoding of input stereo sound signal 190. Consequently, operation 164 uses the outputs from UNCLR classifier 111 / 113 and XTALK detector 110 / 112 from the previous frame, as indicated by reference numeral 191 in Figure 1. Operation 164 of selecting LRTD or DFT stereo mode is further depicted in the schematic block diagram of Figure 13.
[0353] As will be explained in the following description, the DFT / LRTD stereo mode selection mechanism used in operation 164 comprises the following sub-operations: - First DFT / LRTD stereo mode selection, - Switching from LRTD stereo mode to DFT stereo mode by detecting crosstalk content Includes:
[0354] 9.1 Initial DFT / LRTD Stereo Mode Selection The DFT stereo mode is the preferred mode for encoding single talk speech with large inter-channel correlation between the left channel (L) and the right channel (R) of the input stereo sound signal 190 .
[0355] The LRTD / DFT stereo mode selection unit 114 begins the initial selection of a stereo mode by determining whether the previous processed frame is a "likely speech frame." This can be done, for example, by examining the log-likelihood ratio between the "speech" class and the "music" class. The log-likelihood ratio is defined as the absolute difference between the log-likelihood of the input stereo sound signal frame generated by a "music" source and the log-likelihood of the input stereo sound signal frame generated by a "speech" source. The following relationship can be used to calculate the log-likelihood ratio: dL SM (n)=L M (n)-L S (n) (119) where L S (n) is the log-likelihood of the "voice" class, and L M (n) is the log-likelihood of the "music" class.
[0356] As an example, a Gaussian Mixture Model (GMM) from the 3GPP EVS codec as described in reference [7], the entire contents of which are incorporated herein by reference, is used to calculate the log-likelihood L of the "voice" class. S (n) and the log-likelihood of the "music" class L M (n) can be used to estimate the log-likelihood ratio (differential score) dL. SM It can also be used to calculate (n).
[0357] Log likelihood ratio dLSM (n) is smoothed with two IIR filters with different forgetting factors, for example using the following relationship:
[0358]
number
[0359] where the superscript (1) designates the first IIR filter and the superscript (2) designates the second IIR filter, respectively.
[0360] Then, the smoothed
[0361]
number
[0362] and
[0363]
number
[0364] is compared to a predetermined threshold.
[0365] If the following combined conditions are met, for example, a new binary flag f SM (n) is set to 1.
[0366]
number
[0367] Flag F SM (n)=1 indicates that the previous frame is likely to be a speech frame. A threshold of 1.0 was found experimentally.
[0368] Next, the first DFT / LRTD stereo mode selection mechanism selects the binary output c of the UNCLR classifier 111 / 113 in the previous frame n-1. UNCLR(n-1) or binary output c of XTALK detector 110 / 112 XTALK If (n-1) is set to 1 and the previous frame is a possible voice frame, a new binary flag f UX Set (n) to 1. This is expressed by the following relationship:
[0369]
number
[0370] M SMODE Let (n)∈(LRTD, DFT) be a discrete variable indicating the selected stereo mode in the current frame n. The stereo mode is initialized in each frame with the value from the previous frame n-1, i.e., M SMODE (n)=M SMODE (n-1) (123)
[0371] Flag F UX If (n) is set to 1, then the LRTD stereo mode is selected for encoding in the current frame n, which can be expressed as the following relationship: M SMODE (n)←LRTD, if, f UX =1 (124)
[0372] Flag F UX If (n) is set to 0 in the current frame n and the stereo mode in the previous frame n-1 is the LRTD stereo mode, the auxiliary stereo mode switching flag f TDM (n-1) is analyzed to select the stereo mode in the current frame n, for example using the following relationship:
[0373]
number
[0374] Auxiliary stereo mode switching flag f TDM (n) is updated every frame in LRTD mode only. Parameter f TDM The update of (n) is explained in the following description.
[0375] As shown in FIG. 13, the LRTD / DFT stereo mode selection unit 114 uses an auxiliary parameter f TDM (n), c LRTD (n), c DFT (n), and m TD (n) is generated by the LRTD energy analysis processor 1301.
[0376] Flag F UX If (n) is set to 0 in the current frame n and the stereo mode in the previous frame n-1 was the DFT stereo mode, then no stereo mode switching is performed and the DFT stereo mode is also selected for the current frame n.
[0377] 9.2 Switching from LRTD Stereo Mode to DFT Stereo Mode in XTALK Detection The XTALK detector 110 in LRTD mode is described in the previous text. As can be seen from FIG. 11, the binary output c of the XTALK detector 110 XTALK (n) can be set to 1 only when crosstalk content is detected in the current frame. As a result, the initial stereo mode selection logic as described above cannot select DFT stereo mode when the XTALK detection device 110 indicates single-talk content. This results in an undesirable extension of LRTD stereo mode in situations when a crosstalk stereo signal segment is followed by a single-talk stereo signal segment. Therefore, an additional mechanism has been implemented to switch from LRTD stereo mode back to DFT stereo mode upon detection of single-talk content. This mechanism is described in the following description.
[0378] If the LRTD / DFT stereo mode selection device 114 selected the LRTD stereo mode in the previous frame n-1, and the initial stereo mode selection selected the LRTD mode in the current frame n, then simultaneously, the binary output c of the XTALK detection device 110 XTALK If (n-1) is 1, the stereo mode can be changed from LRTD stereo mode to DFT stereo mode. This change is allowed, for example, when the following listed conditions are met:
[0379]
number
[0380] The set of conditions defined above includes references to the class and brate parameters. The brate parameter is a high-level constant that contains the overall bit rate used by device 100 for encoding stereo audio signals (stereo codec). The brate parameter is set during initialization of the stereo codec and is left unchanged during the encoding process.
[0381] The clas parameter is a discrete variable that contains information about the type of frame. The clas parameter is usually estimated as part of the signal pre-processing of a stereo codec. As a non-limiting example, the clas parameter from the frame erasure concealment (FEC) module of the 3GPP EVS codec as described in reference [1] may be used in the DFT / LRTD stereo mode selection mechanism. The clas parameter from the FEC module of the 3GPP EVS codec is selected taking into account the frame erasure concealment and decoder recovery strategy. The clas parameter is selected from the following set of predefined classes:
[0382]
number
[0383] It is within the scope of this disclosure to implement the DFT / LRTD stereo mode selection mechanism with other means of frame type classification.
[0384] In the set of conditions (126) defined above, the condition
[0385]
number
[0386] refers to the class parameters calculated during pre-processing of the downmixed mono (M) channel when the device 100 for coding a stereo sound signal operates in DFT stereo mode.
[0387] When the device 100 for coding a stereo sound signal is in the LRTD stereo mode, the condition is replaced by the following relation:
[0388]
number
[0389] Here, the indices "L" and "R" refer to the class parameters calculated in the pre-processing modules for the left (L) and right (R) channels, respectively.
[0390] Parameter c LRTD (n) and c DFT (n) are counters for the LRTD frame and the DFT frame, respectively. These counters are updated every frame as part of the LRTD energy analysis processor 1301. The two counters c LRTD (n) and c DFT The update of (n) is described in detail in the next section.
[0391] 9.3 Auxiliary parameters calculated in the LRTD energy analysis module When the device for coding a stereo sound signal 100 is operated in the LRTD stereo mode, the LRTD / DFT stereo mode selection unit 114 calculates or updates some auxiliary parameters to improve the stability of the DFT / LRTD stereo mode selection mechanism.
[0392] For certain special types of frames, the LRTD stereo mode operates in the so-called "TD submode". The TD submode is usually applied during a short transition period before switching from the LRTD stereo mode to the DFT stereo mode. Whether the LRTD stereo mode operates in the TD submode is determined by the binary submode flag m TD (n) indicates the binary flag m TD (n) is one of the auxiliary parameters, which can be initialized at each frame as follows: m TD (n)=f TDM (n-1) (127) where f TDM is the aforementioned auxiliary switching flag, described later in this section.
[0393] Binary lower mode flag m TD (n) is f UX It is reset to 0 or 1 in the frame where (n)=1. TD The condition for resetting (n) is determined as follows, for example.
[0394]
number
[0395] f UX If (n)=0, the binary lower mode flag m TD (n) remains unchanged.
[0396] The LRTD energy analysis processing unit 1301 calculates the above two counters c LRTD (n) and c DFT (n) Counter c LRTD(n) is one of the auxiliary parameters and counts the number of consecutive LRTD frames. This counter is set to 0 in every frame in which the DFT stereo mode is selected in the device 100 for coding a stereo sound signal, and is incremented by 1 in every frame in which the LRTD stereo mode is selected. This can be expressed as the following relationship:
[0397]
number
[0398] Basically, the counter c LRTD (n) contains the number of frames since the last DFT->LRTD switch point. Counter c LRTD (n) is limited by a threshold of 100. The counter c DFT (n) counts the number of consecutive DFT frames. Counter c DFT (n) is one of the auxiliary parameters, which is set to 0 in every frame in which the LRTD stereo mode is selected in the device 100 for coding a stereo sound signal, and is incremented by 1 in every frame in which the DFT stereo mode is selected, which can be expressed as the following relationship:
[0399]
number
[0400] Basically, the counter c DFT (n) contains the number of frames since the last LRTD->DFT switch point. Counter c DFT (n) is limited by a threshold of 100.
[0401] The final auxiliary parameter calculated by the LRTD energy analysis processor 1301 is the auxiliary stereo mode switching flag f TDM (n). This parameter is a binary flag f UX(n) is initialized every frame. f TDM (n)=f UX (n) (131)
[0402] Auxiliary stereo mode switching flag f TDM (n) is set to 0 when the left (L) and right (R) channels of the input stereo sound signal 190 are out-of-phase (OOP). An exemplary method for OOP detection can be found, for example, in reference [8], the entire contents of which are incorporated herein by reference. If an OOP situation is detected, the binary flag s2m is set to 1 in the current frame, otherwise it is set to zero. The auxiliary stereo mode switch flag f in LRTD stereo mode TDM (n) is set to zero when the binary flag s2m is set to 1. This can be expressed by the relation (132). f TDM (n)←0, if s2m(n)=1 (132)
[0403] If the binary flag s2m(n) is set to zero, the auxiliary switching flag f TDM (n) may be reset to zero based on, for example, the following set of conditions:
[0404]
number
[0405] Of course, the DFT / LRTD stereo mode switching mechanism can be implemented with other methods for OOP detection.
[0406] Auxiliary stereo mode switching flag f TDM (n) can also be reset to 0 based on the following set of conditions:
[0407]
number
[0408] In the two sets of conditions as set out above, clas(n-1)=UNVOICED_CLAS refers to the class parameters calculated during pre-processing of the downmixed mono (M) channel when the device 100 for coding a stereo sound signal operates in DFT stereo mode.
[0409] When the device 100 for coding a stereo sound signal is in the LRTD stereo mode, the condition is replaced by the following relation: Class L (n-1)=UNVOICED_CLAS and class R (n-1)=UNVOICED_CLAS Here, the indices "L" and "R" refer to the class parameters calculated during pre-processing of the left (L) and right (R) channels, respectively.
[0410] 10. Core Encoder The method 150 for coding a stereo sound signal includes an operation 165 of core coding a left channel (L) of the stereo sound signal 190 in LRTD stereo mode, an operation 166 of core coding a right channel (R) of the stereo sound signal 190 in LRTD stereo mode, and an operation 167 of core coding a downmixed mono (M) channel of the stereo sound signal 190 in DFT stereo mode.
[0411] To perform operation 165, device 100 for coding a stereo audio signal comprises a core encoder 115, e.g., a mono core encoder. To perform operation 166, device 100 comprises a core encoder 116, e.g., a mono core encoder. Finally, to perform operation 167, device 100 for coding a stereo audio signal comprises a core encoder 117 capable of operating in a DFT stereo mode to code a downmixed mono (M) channel of stereo audio signal 190.
[0412] It is believed to be within the knowledge of one skilled in the art to select appropriate core encoders 115, 116, and 117. Therefore, these encoders will not be further described in this disclosure.
[0413] 11. Hardware Implementation FIG. 14 is a simplified block diagram of an example configuration of hardware components forming the above-described device 100 and method 150 for encoding a stereo sound signal.
[0414] The device 100 for encoding a stereo sound signal may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The device 100 (identified in FIG. 14 as 1400) comprises an input unit 1402, an output unit 1404, a processing unit 1406, and a storage unit 1408.
[0415] The input unit 1402 is configured to receive the input stereo audio signal 190 of Figure 1 in digital or analog form. The output unit 1404 is configured to provide an output coded stereo audio signal. The input unit 1402 and the output unit 1404 may be implemented in a common module, for example a serial input / output device.
[0416] The processing unit 1406 is operatively connected to the input 1402, the output 1404, and the storage device 1408. The processing unit 1406 is realized as one or more processing units for executing code instructions in support of the functions of the various components of the device 100 for encoding a stereo sound signal as shown in FIG.
[0417] The storage device 1408 may comprise non-transitory storage for storing code instructions executable by the processing device 1406, and specifically may comprise processor-readable storage that comprises / stores non-transitory instructions that, when executed, cause the processing device to perform the operations and components of the method 150 and device 100 for encoding a stereo sound signal as described in this disclosure. The storage device 1408 may also comprise random access memory or buffers for storing intermediate processed data from the various functions performed by the processing device 1406.
[0418] Those skilled in the art will understand that the device 100 and method 150 for encoding stereo sound signals are merely exemplary and are not intended to be limiting in any way. Other embodiments will readily occur to those skilled in the art having the benefit of this disclosure. Furthermore, the disclosed device 100 and method 150 for encoding stereo sound signals may be customized to provide valuable solutions to the needs and problems that exist in encoding and decoding sound.
[0419] For clarity, not all of the routine features of an implementation of the device 100 and method 150 for encoding a stereo sound signal are shown and described. It will, of course, be understood that in developing any such actual implementation of the device 100 and method 150 for encoding a stereo sound signal, numerous implementation-specific decisions may need to be made to achieve the developer's particular goals, such as compatibility with application, system, network, and business-related constraints, and that these particular goals will vary from implementation to implementation and from developer to developer. Furthermore, it will be understood that the development effort may be complex and time-consuming, but would be a routine engineering undertaking for one of ordinary skill in the art of sound processing having the benefit of this disclosure.
[0420] In accordance with this disclosure, the components / processing devices / modules that process the operations and / or data structures described herein can be implemented using various types of operating systems, computer platforms, network devices, computer programs, and / or general-purpose machines. Those skilled in the art will also recognize that devices of a less general-purpose nature, such as hardwired devices, field programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs), can also be used. When a device comprising a series of operations and sub-operations is performed by a processing device, computer, or machine, and the operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processing device, computer, or machine, the operations and sub-operations can be stored on a tangible and / or non-transitory medium.
[0421] The device 100 and method 150 for encoding a stereo sound signal as described herein may use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0422] In the device 100 and method 150 for encoding a stereo sound signal as described herein, various operations and sub-operations may be performed in various orders, and some of the operations and sub-operations may be optional.
[0423] Although the present disclosure has been described above using non-limiting exemplary embodiments thereof, these embodiments can be freely modified within the scope of the appended claims without departing from the spirit and nature of the present disclosure.
[0424] 12. References This disclosure refers to the following references, the entire contents of which are incorporated herein by reference: [1] 3GPP TS 26.445, v.12.0.0, “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”, Sep 2014. [2] M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robillard, J. Lecompte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefevbre, P. Gournay, et al., “The ISO / MPEG Unified Speech and Audio Coding Standard - Consistent High Quality for All Content Types and at All Bit Rates”, J. Audio Eng. Soc., vol. 61, no. 12, pp. 956-977, Dec. 2013. [3] F. Baumgarte, C. Faller, "Binaural cue coding - Part I: Psychoacoustic fundamentals and design principles," IEEE Trans. Speech Audio Processing, vol. 11, pp. 509-519, Nov. 2003. [4] Tommy Vaillancourt, “Method and system using a long-term correlation difference between left and right channels for time domain down mixing a stereo sound signal into primary and secondary channels,” US Patent 10,325,606 B2. [5] 3GPP SA4 contribution S4-170749 “New WID on EVS Codec Extension for Immersive Voice and Audio Services”, SA4 meeting #94, June 26-30, 2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip [6] I. Mani, J. Zhang. “kNN approach to unbalanced data distributions: A case study involving information extraction,” In Proceedings of the Workshop on Learning from Imbalanced Data Sets, pp. 1-7, 2003.KNN [7] V. Malenovsky, T. Vaillancourt, W. Zhe, K. Choo and V. Atti, "Two-stage speech / music classifier with decision smoothing and sharpening in the EVS codec," 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, 2015, pp. 5718-5722. [8] Vaillancourt, T., “Method and system for time-domain down mixing a stereo sound signal into primary and secondary channels using detecting an out-of-phase condition on the left and right channels,” United States Patent US 10,522,157. [9] Maalouf, Maher. “Logistic regression in data analysis: An overview”, 2011 International Journal of Data Analysis Techniques and Strategies. 3. 281-299. 10.1504 / IJDATS.2011.041335.
[10] Ruder, S., “An overview of gradient descent optimization algorithms”. 2016. ArXiv Preprint ArXiv:1609.04747. [Explanation of symbols]
[0425] 100 Stereo sound signal coding device 101, 102 Analyzer 103, 104 Time domain preprocessor 105 FFT Transformation Calculation Device 106 DFT Stereo Parameter Calculation Device 110, 112 XTALK detector 111, 113 UNCLR classifier 114 LRTD / DFT stereo mode selector 150 Stereo sound signal coding method 151 Operation of inter-channel correlation analysis in LRTD stereo mode 152 Inter-channel correlation analysis in DFT stereo mode 153 Operations for time-domain preprocessing of the left channel 154 Operations for time-domain preprocessing of the right channel 155 Operation to calculate the Fast Fourier Transform (FFT) 156 Operation to calculate DFT stereo parameters 157 Downmixing the left and right channels 158 Operations to compute IFFT transforms 159 TD pre-processing operation 160, 162 Crosstalk (XTALK) detection operation 161, 163 Uncorrelated Stereo Content (UNCLR) Classification Behavior 164 Selecting RTD Stereo Mode or DFT Stereo Mode 165 Core encoding of left channel (L) 166 Core coding of the right channel (L) 167 Core encoding of mono (M) channel 190 Stereo Sound Signal 1301 LRTD Energy Analysis Processing Unit 1400 devices 1402 Input section 1404 Output section 1406 Processing equipment 1408 Storage device P1, P2, P3, P4, P5, P6, P7, P8, P9, P10, P11, P12 position S1, S2 Speaker, Speaker M1, M2, M3, M4, M5, M6 microphones
Claims
1. 1. A device for selecting one of a first stereo mode and a second stereo mode for encoding a stereo sound signal including a left channel and a right channel, comprising: a classifier for generating a first output indicative of the presence or absence of uncorrelated stereo content in the stereo sound signal; a detector for generating a second output indicative of the presence or absence of crosstalk in the stereo sound signal caused by two speakers speaking simultaneously; an analysis and processing unit for calculating auxiliary parameters for use in selecting the stereo mode for encoding the stereo sound signal; a stereo mode selection device for selecting the stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameter; Equipped with The stereo mode selection device includes: performing an initial selection of the stereo mode for encoding the stereo sound signal between the first stereo mode and the second stereo mode; selecting the second stereo mode for encoding the stereo sound signal if certain given conditions are met following the first selection of the stereo mode. It is configured as follows: device.
2. 2. The stereo mode selection device of claim 1, wherein the first stereo mode is a time-domain stereo mode in which the left channel and the right channel are coded separately, and the second stereo mode is a frequency-domain stereo mode.
3. 3. The stereo mode selection device of claim 1, wherein in a current frame of the stereo sound signal, the stereo mode selection apparatus uses the first output from a previous frame of the stereo sound signal and the second output from the previous frame.
4. 2. The stereo mode selection device of claim 1, wherein the stereo mode selection apparatus determines whether a previous frame of the stereo sound signal is an audio frame to perform the initial selection of the stereo mode for coding the stereo sound signal.
5. 5. The stereo mode selection device of claim 4, wherein, in the initial selection of the stereo mode, the stereo mode selection device selects the first stereo mode for coding the stereo sound signal if (a) the previous frame is determined to be an audio frame, and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame.
6. 6. The stereo mode selection device of claim 5, wherein, in the initial selection of the stereo mode for encoding the stereo sound signal, if (i) condition (a), condition (b), or both conditions (a) and (b) are not satisfied, and (ii) the stereo mode selected in the previous frame is the second stereo mode, the stereo mode selection apparatus selects the second stereo mode for encoding the stereo sound signal.
7. 7. The stereo mode selection device of claim 5 or 6, wherein the stereo mode selection apparatus selects the stereo mode for coding the stereo sound signal with respect to one of the auxiliary parameters if, in the initial selection of the stereo mode, (i) condition (a), condition (b), or both conditions (a) and (b) are not satisfied, and (ii) the stereo mode selected in the previous frame is the first stereo mode.
8. The stereo mode selection device of claim 7, wherein the one of the auxiliary parameters is an auxiliary stereo mode switch flag.
9. The given condition is selected from the following group: - the first stereo mode is selected in a previous frame of the stereo sound signal; - the first stereo mode is initially selected in a current frame of the stereo sound signal; - the second output of the detection device in the current frame indicates the presence of crosstalk in the stereo sound signal; - (i) the previous frame is determined to be an audio frame, and (ii) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame; - in the previous frame, a counter of a number of consecutive frames using the first stereo mode is greater than a first value; - in the previous frame, a counter of a number of consecutive frames using the second stereo mode is greater than a second value; - in the previous frame, the class of the stereo sound signal is within a set of predetermined classes; and (i) a total bit rate used to encode the stereo sound signal is greater than or equal to a third value, or (ii) a score representing crosstalk in the stereo sound signal from the detection device is less than a fourth value in the previous frame.
2. The stereo mode selection device of claim 1, wherein the stereo mode selection device is selected from:
10. 2. The stereo mode selection device of claim 1, wherein the analysis processing unit calculates as one of the auxiliary parameters an auxiliary lower mode flag that indicates the first stereo mode to operate in a lower mode that is applied over a short transition before switching from the first stereo mode to the second stereo mode.
11. 11. The stereo mode selection device of claim 10, wherein the analysis processing unit resets the auxiliary lower mode flag in a frame of the stereo sound signal if (a) the previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame.
12. 12. The stereo mode selection device of claim 11, wherein the analysis processing device resets the auxiliary lower mode flag to 1 in a frame of the stereo sound signal if (1) an auxiliary stereo mode switching flag calculated by the analysis processing device as an auxiliary parameter is equal to 1, (2) the stereo mode of the previous frame is not the first stereo mode, or (3) a counter of frames using the first stereo mode is less than a given value.
13. 13. The stereo mode selection device of claim 12, wherein the analysis processing unit resets the auxiliary lower mode flag to 0 in a frame of the stereo sound signal if none of conditions (1) to (3) is satisfied.
14. 11. The stereo mode selection device of claim 10, wherein the analysis processing unit does not change the auxiliary lower mode flag in a frame of the stereo sound signal if at least one of the following conditions is not met: (a) the previous frame of the stereo sound signal is determined to be an audio frame; and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame.
15. The stereo mode selection device of claim 1 , wherein the analysis processor calculates a counter of a number of consecutive frames using the first stereo mode as one of the auxiliary parameters.
16. 16. The stereo mode selection device of claim 15, wherein the analysis processing unit increments the counter of a number of consecutive frames using the first stereo mode if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame.
17. 17. The stereo mode selection device of claim 15 or 16, wherein the analysis processing unit resets the counter of several consecutive frames using the first stereo mode to zero if the second stereo mode is selected by the stereo mode selection unit in the current frame.
18. 18. The stereo mode selection device according to claim 1, wherein the analysis processing unit includes as one of the auxiliary parameters a counter of a number of consecutive frames using the second stereo mode.
19. The stereo mode selection device of claim 1 , wherein the analysis processing unit generates an auxiliary stereo mode switching flag as one of the auxiliary parameters.
20. 20. The stereo mode selection device of claim 19, wherein the analysis processing device initializes the auxiliary stereo mode switching flag to 1 in the current frame if (a) the previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, and (ii) to 0 when condition (a), condition (b), or both conditions (a) and (b) are not satisfied.
21. 21. The stereo mode selection device of claim 19, wherein the analysis processing unit sets the auxiliary stereo mode switch flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.
22. 13. The stereo mode selection device according to claim 8, wherein the analysis processing unit generates the auxiliary stereo mode switching flag as one of the auxiliary parameters.
23. 23. The stereo mode selection device of claim 22, wherein the analysis processing device initializes the auxiliary stereo mode switching flag to 1 in the current frame if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output from the classifier indicates the presence of uncorrelated stereo content in the previous frame or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, and (ii) to 0 when condition (a), condition (b), or both conditions (a) and (b) are not satisfied.
24. 24. The stereo mode selection device of claim 22 or 23, wherein the analysis processing unit sets the auxiliary stereo mode switch flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.
25. 1. A method for selecting one of a first stereo mode and a second stereo mode for encoding a stereo sound signal including a left channel and a right channel, comprising: generating a first output indicative of the presence or absence of uncorrelated stereo content in the stereo sound signal; generating a second output indicating the presence or absence of crosstalk in the stereo sound signal caused by two speakers speaking simultaneously; calculating auxiliary parameters for use in selecting the stereo mode for encoding the stereo sound signal; selecting the stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameter; Including, The step of selecting a stereo mode includes: performing an initial selection of the stereo mode for encoding the stereo sound signal between the first stereo mode and the second stereo mode; and selecting the second stereo mode for encoding the stereo sound signal if certain given conditions are met following the first selection of the stereo mode. method.
26. 26. The stereo mode selection method of claim 25, wherein the first stereo mode is a time-domain stereo mode in which the left and right channels are coded separately, and the second stereo mode is a frequency-domain stereo mode.
27. 27. The stereo mode selection method of claim 25 or 26, wherein selecting the stereo mode in a current frame of the stereo sound signal includes using the first output from a previous frame of the stereo sound signal and the second output from the previous frame.
28. 26. The stereo mode selection method of claim 25, wherein the step of selecting the stereo mode includes determining whether a previous frame of the stereo sound signal is an audio frame to implement the initial selection of the stereo mode for coding the stereo sound signal.
29. 29. The stereo mode selection method of claim 28, wherein the step of selecting the stereo mode includes selecting the first stereo mode for coding the stereo sound signal if, in the initial selection of the stereo mode, (a) the previous frame is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal at the previous frame.
30. 30. The stereo mode selection method of claim 29, wherein the step of selecting the stereo mode includes selecting the second stereo mode for coding the stereo sound signal if, in the initial selection of the stereo mode for coding the stereo sound signal, (i) condition (a), condition (b), or both conditions (a) and (b) are not satisfied, and (ii) the stereo mode selected in the previous frame was the second stereo mode.
31. 31. The stereo mode selection method of claim 29 or 30, wherein the step of selecting the stereo mode includes selecting the stereo mode for coding the stereo sound signal with respect to one of the auxiliary parameters if, in the initial selection of the stereo mode, (i) condition (a), condition (b), or both conditions (a) and (b) are not satisfied, and (ii) the stereo mode selected in the previous frame is the first stereo mode.
32. 32. The stereo mode selection method of claim 31, wherein said one of said auxiliary parameters is an auxiliary stereo mode switch flag.
33. The given condition is selected from the following group of conditions: - the first stereo mode is selected in a previous frame of the stereo sound signal; - the first stereo mode is initially selected in a current frame of the stereo sound signal; - the second output in the current frame indicates the presence of crosstalk in the stereo sound signal; - (i) the previous frame is determined as an audio frame, and (ii) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal at the previous frame; - in the previous frame, a counter of a number of consecutive frames using the first stereo mode is greater than a first value; - in the previous frame, a counter of a number of consecutive frames using the second stereo mode is greater than a second value; - in the previous frame, the class of the stereo sound signal is within a set of predetermined classes; and (i) a total bit rate used to code the stereo sound signal is greater than or equal to a third value, or (ii) a score representing crosstalk in the stereo sound signal is less than a fourth value in the previous frame.
26. The stereo mode selection method of claim 25, wherein the stereo mode is selected from:
34. 26. The stereo mode selection method of claim 25, wherein the step of calculating the auxiliary parameters includes calculating as one of the auxiliary parameters an auxiliary lower mode flag that indicates the first stereo mode to operate in a lower mode that applies over a short transition before switching from the first stereo mode to the second stereo mode.
35. 35. The stereo mode selection method of claim 34, wherein the step of calculating the auxiliary parameters includes resetting the auxiliary lower mode flag in the frame of the stereo sound signal if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal in the previous frame.
36. 36. The stereo mode selection method of claim 35, wherein the step of calculating the auxiliary parameter includes resetting the auxiliary lower mode flag to 1 in the frame of the stereo sound signal if (1) an auxiliary stereo mode switching flag calculated as the auxiliary parameter is equal to 1, (2) the stereo mode of the previous frame is not the first stereo mode, or (3) a counter of frames using the first stereo mode is less than a given value.
37. 37. The stereo mode selection method of claim 36, wherein the step of calculating the auxiliary parameters includes resetting the auxiliary sub-mode flag to 0 in a frame of the stereo sound signal if none of conditions (1) to (3) are met.
38. 35. The stereo mode selection method of claim 34, wherein in the step of calculating the auxiliary parameters, the auxiliary sub-mode flag is not changed in the frame of the stereo sound signal if at least one of the following conditions is not met: (a) the previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal in the previous frame.
39. 26. The stereo mode selection method of claim 25, wherein the step of calculating the auxiliary parameters includes calculating a counter of a number of consecutive frames using the first stereo mode as one of the auxiliary parameters.
40. 40. The stereo mode selection method of claim 39, wherein the step of calculating the auxiliary parameter includes incrementing the counter of a number of consecutive frames using the first stereo mode if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal in the previous frame.
41. 41. The stereo mode selection method of claim 39 or 40, wherein the step of calculating the auxiliary parameter includes resetting the counter of several consecutive frames using the first stereo mode to zero if the second stereo mode is selected in the current frame.
42. 42. The stereo mode selection method of claim 25, wherein the step of calculating the auxiliary parameters comprises calculating a counter of a number of consecutive frames using the second stereo mode as one of the auxiliary parameters.
43. 26. The stereo mode selection method of claim 25, wherein the step of calculating the auxiliary parameters includes generating an auxiliary stereo mode switch flag as one of the auxiliary parameters.
44. 44. The stereo mode selection method of claim 43, wherein the step of calculating the auxiliary parameter includes: (i) initializing the auxiliary stereo mode switch flag to 1 in the current frame if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal in the previous frame, and (ii) initializing the auxiliary stereo mode switch flag to 0 when condition (a), condition (b), or both conditions (a) and (b) are not satisfied.
45. 45. The stereo mode selection method of claim 43 or 44, wherein the step of calculating the auxiliary parameters includes setting the auxiliary stereo mode switch flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.
46. 37. The stereo mode selection method of claim 32 or 36, wherein the step of calculating the auxiliary parameters includes generating the auxiliary stereo mode switch flag as one of the auxiliary parameters.
47. 47. The stereo mode selection method of claim 46, wherein the step of calculating the auxiliary parameter includes: (i) initializing the auxiliary stereo mode switch flag to 1 in the current frame if (a) a previous frame of the stereo sound signal is determined to be an audio frame, and (b) the first output indicates the presence of uncorrelated stereo content in the previous frame or the second output indicates the presence of crosstalk in the stereo sound signal in the previous frame; and (ii) initializing the auxiliary stereo mode switch flag to 0 when condition (a), condition (b), or both conditions (a) and (b) are not satisfied.
48. 48. The stereo mode selection method of claim 46 or 47, wherein the step of calculating the auxiliary parameters includes setting the auxiliary stereo mode switch flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.
Citation Information
Patent Citations
Stereo sound encoding / Decoding system
JP1994236200A
periodic speech coding
JP2003522965A
Encoding and decoding of multi-channel signals
JP2004509366A
ADAPTIVE TIME / FREQUENCY BASED CODING MODE DECISION APPARATUS AND CODING MODE DECISION METHOD THEREFOR - Patent application
JP2009524846A
Method and device for determining encoding method
JP2011527762A