Methods and apparatus for classifying unrelated stereo content, detecting crosstalk, and selecting stereo modes in audio codecs.

By performing unrelated stereo content classification and crosstalk detection on stereo audio signals, and combining LRTD and DFT stereo mode selection, the problem of bit rate doubling and sound quality degradation of stereo signals in multi-speaker scenarios is solved, achieving high-quality audio coding at low bit rates.

CN116438811BActive Publication Date: 2026-06-02VOICEAGE CORPORATION

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VOICEAGE CORPORATION
Filing Date
2021-09-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from doubled bit rate and degraded sound quality when transmitting stereo information. In particular, in multi-speaker scenarios, it is difficult to effectively utilize the redundancy of stereo signals, which affects audio quality.

Method used

By classifying unrelated stereo content and detecting crosstalk in stereo audio signals, and combining LRTD and DFT stereo mode selection, the encoding mode is dynamically switched to optimize bit rate utilization. A logistic regression model is used to extract features and perform mode selection.

Benefits of technology

It improves the audio quality of stereo signals at low bit rates, adapts to complex audio scenarios, achieves effective encoding in multi-speaker scenarios, and reduces bit rate requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116438811B_ABST
    Figure CN116438811B_ABST
Patent Text Reader

Abstract

This disclosure describes the classification (hereinafter referred to as "UNCLR classification") and crosstalk detection (hereinafter referred to as "XTALK detection") of irrelevant stereo content in an input stereo audio signal. This disclosure also describes stereo mode selection, such as automatic LRTD / DFT stereo mode selection. Furthermore, this disclosure uses the classification to select one of a first stereo mode and a second stereo mode for encoding a stereo audio signal including the left and right channels; to detect crosstalk in the stereo audio signal including the left and right channels in response to features extracted from the stereo audio signal including the left and right channels; or to classify irrelevant stereo content in the stereo audio signal including the left and right channels in response to features extracted from the stereo audio signal including the left and right channels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to sound coding, and more specifically, but not exclusively, to the classification of unrelated stereo content, crosstalk detection, and stereo mode selection in, for example, multichannel sound codecs that can produce good sound quality in complex audio scenarios with low bit rate and low latency.

[0002] In this disclosure and the appended claims:

[0003] - The term "sound" can be related to speech, audio, and any other sound;

[0004] - The term "stereo" is an abbreviation of "stereophonic"; and

[0005] - The term "mono" is an abbreviation of "monophonic". Background Technology

[0006] Historically, conversational telephones were implemented using handheld devices with only one transducer, outputting sound to only one of the user's ears. In the last decade, users have begun using their portable handheld devices in conjunction with headsets to receive sound through both ears, primarily for listening to music, but sometimes for listening to voice. However, when portable handheld devices are used to send and receive conversational voice, the content remains mono, but is presented to both of the user's ears when headsets are used.

[0007] The quality of encoded sound (e.g., speech and / or audio) transmitted and received via portable handheld devices has been significantly improved using the latest 3GPP voice coding standard EVS (Enhanced Voice Service) as described in reference [1] (the entire contents of which are incorporated herein by reference). The next natural step is to transmit stereo information so that the receiver is as close as possible to the real-life audio scene captured at the other end of the communication link.

[0008] In audio codecs, the transmission of stereo information is often used, for example as described in reference [2] (the entire contents of which are incorporated herein by reference).

[0009] For conversational speech codecs, mono signals are the norm. When stereo audio signals are transmitted, the bit rate is often doubled because both the left and right channels of the stereo audio signal are encoded using a mono codec. This works well in most cases, but it has the drawback of doubling the bit rate and not utilizing any potential redundancy between the two channels (the left and right channels of the stereo audio signal). Furthermore, to keep the overall bit rate at a reasonable level, a very low bit rate is used for each of the left and right channels, thus affecting the overall sound quality. To reduce the bit rate, efficient stereo coding techniques have been developed and used. As non-limiting examples, two stereo coding techniques that can be used efficiently at low bit rates are discussed in the following paragraphs.

[0010] The first stereo coding technique is called parametric stereo. Parametric stereo uses a common mono codec to encode two inputs (left and right channels) into a mono signal plus a certain amount of stereo side information (corresponding to stereo parameters) representing the stereo image. The two input left and right channels are down-mixed into a mono signal, and the stereo parameters are then calculated. This is typically performed in the frequency domain (FD) (e.g., in the Discrete Fourier Transform (DFT) domain). Stereo parameters relate to so-called binaural or interchannel cues. Binaural cues (see, for example, reference [3], the entire contents of which are incorporated herein by reference) include interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC). Depending on the characteristics of the audio signal, the stereo scene configuration, etc., some or all of the binaural cues are encoded and sent to the decoder. Information about which binaural cues are encoded and sent is transmitted as signaling information, which is usually part of the stereo side information. Furthermore, a given binaural cue can be quantized using different encoding techniques, resulting in a variable number of bits being used. Subsequently, in addition to the quantized binaural cue, stereo side information can typically include a quantized residual signal generated by downmixing at mid-bit rates and higher bit rates. The residual signal can be encoded using entropy coding techniques (e.g., arithmetic encoders). In the remainder of this disclosure, parametric stereo will be referred to as "DFT stereo" because parametric stereo coding techniques typically operate in the frequency domain, and this disclosure will describe non-limiting embodiments using DFT.

[0011] Another stereo coding technique operates in the time domain. This stereo coding technique mixes two inputs (left and right channels) into so-called master and consonant channels. For example, following the method described in reference [4] (the entire contents of which are incorporated herein by reference), time-domain mixing can be based on a mixing ratio that determines the respective contributions of the two inputs (left and right channels) in producing the master and consonant channels. The mixing ratio is derived from several metrics, such as the normalized correlation of the two inputs (left and right channels) relative to a mono signal or the long-term correlation difference between the two inputs (left and right channels). The master channel can be encoded by a common mono codec, while the consonant channel can be encoded by a lower bit-rate codec. The encoding of the consonant channel can take advantage of the coherence between the master and consonant channels and can be reused for certain parameters of the master channel. In some sounds where the left and right channels exhibit very little correlation, it is preferable to encode the left and right channels of the stereo input signal individually in the time domain or with minimal inter-channel parameterization. This approach in the encoder is a special case of time-domain TD stereo and will be referred to throughout this disclosure as "LRTD stereo".

[0012] Furthermore, in recent years, the generation, recording, representation, encoding, transmission, and reproduction of audio have been evolving towards enhanced, interactive, and immersive listener experiences. For example, an immersive experience can be described as a state of deep engagement or immersion in a sound scene when sound comes from all directions. In immersive audio (also known as 3D audio), the sound image is reproduced in all three dimensions surrounding the listener, taking into account various sound characteristics such as timbre, directionality, reverberation, transparency, and (auditory) spatial accuracy. Immersive audio is produced for specific sound playback or reproduction systems, such as speaker-based systems, integrated reproduction systems (soundboards), or headphones. Subsequently, the interactivity of the sound reproduction system can include, for example, the ability to adjust sound levels, change the localization of sound, or select different languages ​​for reproduction.

[0013] There are three basic ways to achieve an immersive experience.

[0014] The first approach to achieving an immersive experience is channel-based audio, which uses multiple spaced-apart microphones to capture sound from different directions, with one microphone corresponding to an audio channel in a specific speaker layout. Each recorded channel is then provided to a speaker in a given location. Examples of channel-based audio are, for instance, stereo, 5.1 surround, 5.1+4, etc.

[0015] The second approach to achieving immersive experiences is scene-based audio, which represents the desired sound field in local space as a function of time through the combination of dimensional components. Scene-based audio represents a sound signal independent of the audio source's localization, while the sound field is transformed to fit the selected speaker layout at the presenter. An example of scene-based audio is ambient stereo (“ambisonics”).

[0016] The third approach to achieving immersive experiences is object-based audio, which represents the auditory scene as a collection of individual audio elements (e.g., singer, drums, guitar, etc.) accompanied by information such as their location, allowing them to be presented by a sound reproduction system at their intended positions. This gives object-based audio a great deal of flexibility and interactivity, as each object remains discrete and can be manipulated individually.

[0017] Each of the audio methods described above for achieving immersive experiences presents both advantages and disadvantages. Therefore, it is often not a single audio method, but rather a combination of several audio methods within a complex audio system to create an immersive auditory scene. An example could be an audio system that combines scene-based or channel-based audio with object-based audio, such as ambient stereo with several discrete audio objects.

[0018] In recent years, 3GPP (3rd Generation Partnership Project) has been working on developing a 3D (three-dimensional) sound codec for immersive services called IVAS (Immersive Voice and Audio Services) based on the EVS codec (see reference [5], the entire contents of which are incorporated herein by reference).

[0019] The DFT stereo mode is efficient for encoding single-talk sounds. In cases with two or more speakers, parametric stereo techniques struggle to fully describe the spatial nature of the scene. This problem is particularly pronounced when two speakers are talking simultaneously (crosstalk scenarios) and when the signals in the left and right channels of the stereo input signal are weakly correlated or completely uncorrelated. In such cases, it is preferable to use the LRTD stereo mode to encode the left and right channels of the stereo input signal individually in the time domain, or with minimal inter-channel parameterization. As the scene captured in the stereo input signal evolves, it is desirable to switch between the DFT and LRTD stereo modes based on stereo scene classification. Summary of the Invention

[0020] According to a first aspect, this disclosure relates to a method for classifying irrelevant stereo content in a stereo sound signal including a left channel and a right channel in response to features extracted from the stereo sound signal including a left channel and a right channel, comprising: calculating a score representing irrelevant stereo content in the stereo sound signal in response to the extracted features; and switching between a first category indicating one of irrelevant stereo content and relevant stereo content in the stereo sound signal and a second category indicating the other of irrelevant stereo content and relevant stereo content in response to the score.

[0021] According to a second aspect, this disclosure provides a classifier for unrelated stereo content in a stereo sound signal including the left and right channels, responsive to features extracted from a stereo sound signal including the left and right channels, comprising: a calculator representing a score of unrelated stereo content in the stereo sound signal in response to the extracted features; and a category switching mechanism, responsive to the score, for switching between a first category indicating one of unrelated stereo content and related stereo content in the stereo sound signal and a second category indicating the other of unrelated stereo content and related stereo content.

[0022] This disclosure also relates to a method for detecting crosstalk in a stereo sound signal including a left channel and a right channel in response to features extracted from a stereo sound signal including a left channel and a right channel, comprising: calculating a score representing crosstalk in the stereo sound signal in response to the extracted features; calculating auxiliary parameters for detecting crosstalk in the stereo sound signal; and switching between a first category indicating the presence of crosstalk in the stereo sound signal and a second category indicating the absence of crosstalk in the stereo sound signal in response to the crosstalk score and the auxiliary parameters.

[0023] According to another aspect, this disclosure provides a crosstalk detector in a stereo sound signal including the left and right channels, responsive to features extracted from a stereo sound signal including the left and right channels, comprising: a calculator representing a score of crosstalk in the stereo sound signal in response to the extracted features; a calculator for auxiliary parameters for detecting crosstalk in the stereo sound signal; and a category switching mechanism for switching between a first category indicating the presence of crosstalk in the stereo sound signal and a second category indicating the absence of crosstalk in the stereo sound signal, in response to the crosstalk score and the auxiliary parameters.

[0024] This disclosure also relates to a method for selecting one of a first stereo mode and a second stereo mode for encoding a stereo sound signal including a left channel and a right channel, comprising: generating a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; generating a second output indicating the presence or absence of crosstalk in the stereo sound signal; calculating auxiliary parameters for selecting a stereo mode for encoding the stereo sound signal; and selecting the stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters.

[0025] According to another aspect, this disclosure provides an apparatus for selecting one of a first stereo mode and a second stereo mode for encoding a stereo sound signal including a left channel and a right channel, comprising: a classifier for generating a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; a detector for generating a second output indicating the presence or absence of crosstalk in the stereo sound signal; an analysis processor for calculating auxiliary parameters for selecting a stereo mode for encoding the stereo sound signal; and a stereo mode selector for selecting a stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters.

[0026] The foregoing and other objects, advantages and features of the unrelated stereo content classifier and classification method, the crosstalk detector and detection method, and the stereo mode selection device and method will become more apparent upon reading the following non-limiting description of illustrative embodiments thereof, which are given by way of example only with reference to the accompanying drawings. Attached Figure Description

[0027] In the attached diagram:

[0028] Figure 1 It is a schematic block diagram that simultaneously illustrates the device for encoding stereo sound signals and the corresponding method for encoding stereo sound signals;

[0029] Figure 2 This is a schematic diagram of a crosstalk scenario, in which two opposing speakers are captured by a pair of supercardioid microphones;

[0030] Figure 3 This is a graph showing the location of the peaks in the GCC-PHAT function;

[0031] Figure 4 It is a top-down view used for realistic recording of 3D scene settings;

[0032] Figure 5This is a graph illustrating the normalization function applied to the output of the LogReg model in the classification of unrelated stereo content in LRTD stereo mode;

[0033] Figure 6 It shows the formation Figure 1 A state machine diagram of the switching mechanism between stereo content categories in a classifier for non-relevant stereo content in a device used to encode stereo sound signals;

[0034] Figure 7 It is a schematic floor plan of a large conference room with an AB microphone setup, the conditions of which are simulated for crosstalk detection, where the AB microphone consists of a pair of separately placed cardioid or omnidirectional microphones, positioned such that they cover the space without causing phase problems between them.

[0035] Figure 8 This is a diagram illustrating the automatic labeling of crosstalk samples using VAD (Voice Activity Detection);

[0036] Figure 9 This is a graph representing the function used to scale the raw output of the LogReg model in crosstalk detection in LRTD stereo mode;

[0037] Figure 10 The diagram illustrates the formation of a system for encoding stereo audio signals in LRTD stereo mode. Figure 1 A diagram illustrating the mechanism for detecting rising edges in the crosstalk detector of a device.

[0038] Figure 11 This is a logic diagram illustrating the mechanism for switching between the states of the crosstalk detector output in LRTD stereo mode;

[0039] Figure 12 This is a logic diagram illustrating the mechanism for switching between the states of the crosstalk detector output in DFT stereo mode;

[0040] Figure 13 This is a schematic block diagram illustrating the mechanism for selecting between LRTD and DFT stereo modes; and

[0041] Figure 14 This is a simplified block diagram of an example configuration of hardware components for implementing methods and devices for encoding stereo audio signals. Detailed Implementation

[0042] This disclosure describes the classification of unrelated stereo content in an input stereo audio signal (hereinafter referred to as "UNCLR classification") and crosstalk detection (hereinafter referred to as "XTALK detection"). This disclosure also describes stereo mode selection, such as automatic LRTD / DFT stereo mode selection.

[0043] Figure 1 This is a schematic block diagram that simultaneously illustrates a device 100 for encoding a stereo sound signal 190 and a corresponding method 150 for encoding the stereo sound signal 190.

[0044] Specifically, Figure 1 This illustrates how UNCLR classification, XTALK detection, and stereo mode selection are integrated into stereo audio signal encoding method 150 and device 100.

[0045] UNCLR classification and XTALK detection form two independent techniques. However, they are based on the same statistical model and share certain features and parameters. Furthermore, both UNCLR classification and XTALK detection are designed and trained separately for LRTD stereo patterns and DFT stereo patterns. In this disclosure, LRTD stereo patterns are given as a non-limiting example of time-domain stereo patterns, and DFT stereo patterns are given as a non-limiting example of frequency-domain stereo patterns. Implementations of other time-domain and frequency-domain stereo patterns are also within the scope of this disclosure.

[0046] UNCLR classification analyzes features extracted from the left and right channels of the stereo audio signal 190 and detects weak or zero correlation between the left and right channels. XTALK detection, on the other hand, detects the presence of two speakers speaking simultaneously in a stereo scene. For example, both UNCLR classification and XTALK detection provide binary outputs. These binary outputs are combined in the stereo mode selection logic. As a non-restrictive general rule, stereo mode selection chooses the LRTD stereo mode when UNCLR classification and XTALK detection indicate the presence of two speakers standing on opposite sides of the capture device (e.g., a microphone). This typically results in a weak correlation between the left and right channels of the stereo audio signal 190. The selection of the LRTD stereo mode or DFT stereo mode is performed on a frame-by-frame basis (as is known in the art, the stereo audio signal 190 is sampled at a given sampling rate and processed according to these sample groups, called "frames," which are divided into multiple "subframes"). In addition, the stereo mode selection logic is designed to avoid frequent switching between LRTD and DFT stereo modes, as well as stereo mode switching within perceptually important signal segments.

[0047] In this disclosure, non-limiting, illustrative embodiments of UNCLR classification, XTALK detection, and stereo mode selection will be described by way of example only, with reference to the IVAS coding architecture known as the IVAS codec (or IVAS sound codec). However, incorporating such classification, detection, and selection into any other sound codec is also within the scope of this disclosure.

[0048] 1. Feature Extraction

[0049] UNCLR classification is based on a logistic regression (LogReg) model, as described, for example, in reference [9] (the entire contents of which are incorporated herein by reference). The LogReg model is trained separately for LRTD stereo mode and DFT stereo mode. Training is performed using a large database of features extracted from stereo audio signal encoding device 100 (stereo codec). Similarly, XTALK detection is based on a LogReg model, which is trained separately for LRTD stereo mode and DFT stereo mode. The features used in XTALK detection are different from those used in UNCLR classification. However, some features are shared between the two techniques.

[0050] The features used in UNCLR classification and the features used in XTALK detection are extracted from the following operations:

[0051] - Inter-channel correlation analysis;

[0052] -TD preprocessing; and

[0053] -DFT stereo parameterization.

[0054] The method 150 for encoding stereo sound signals includes an operation (not shown) for extracting the aforementioned features. To perform the feature extraction operation, the apparatus 100 for encoding stereo sound signals includes a feature extractor (not shown).

[0055] 2. Inter-channel correlation analysis

[0056] The feature extraction operations (not shown) include operation 151 for inter-channel correlation analysis of the LRTD stereo mode and operation 152 for inter-channel correlation analysis of the DFT stereo mode. To perform operations 151 and 152, the feature extractor (not shown) includes inter-channel correlation analyzer 101 and inter-channel correlation analyzer 102, respectively. Operations 151 and 152, as well as analyzers 101 and 102, are similar and will be described concurrently.

[0057] Analyzers 101 / 102 receive the left and right channels of the current stereo audio signal frame as input. The left and right channels are first downsampled to 8kHz. For example, let the downsampled left and right channels be labeled as:

[0058] X L (n),X R (n), n=0,..,N-1 (1)

[0059] Where n is the sample index in the current frame, and N = 160 is the length of the current frame (the length of 160 samples). The downsampled left and right channels are used to calculate the interchannel correlation function. First, the absolute energy of the left and right channels is calculated using, for example, the following relationship:

[0060]

[0061] Analyzers 101 / 102 calculate the numerator of the interchannel correlation function based on the dot product between the left and right channels over a hysteresis range of <-40, 40>. For negative hysteresis, the dot product between the left and right channels is calculated, for example, using the following relationship:

[0062]

[0063] Furthermore, for positive lag, the dot product is given, for example, through the following relationship:

[0064]

[0065] Analyzers 101 / 102 then use, for example, the following relationship to calculate the inter-channel correlation function:

[0066]

[0067] The superscript [-1] indicates a reference to the previous frame. The passive mono signal is calculated by averaging the left and right channels:

[0068]

[0069] As a non-restrictive example, using the following relationship, the side signal is calculated as the difference between the left and right channels:

[0070]

[0071] Finally, it is equally useful to define the per-sample product of the left and right channels as:

[0072] X P (n)=X L (n)·X R (n), n=0,..,N-1 (8)

[0073] Analyzers 101 / 102 include an infinite impulse response (IIR) filter (not shown) for smoothing inter-channel correlation functions using, for example, the following relationship:

[0074]

[0075] The superscript [n] indicates the current frame, the superscript [n-1] indicates the previous frame, and α ICA It is a smoothing factor.

[0076] Smoothing factor α ICA It is adaptively set within the interchannel correlation analysis (ICA) module (reference [1]) of the stereo audio signal encoding device 100 (stereo codec). The interchannel correlation function is then weighted at the location of the predicted peak. The mechanism for peak finding and local windowing is implemented within the ICA module and will not be described in this document; additional information about the ICA module can be found in reference [1]. The interchannel correlation function after ICA weighting is denoted as R. w (k), where k∈<-40,40>.

[0077] The location of the maximum value of the interchannel correlation function is an important indicator of the direction in which the dominant sound arrives at the capture point, and is used as a feature by UNCLR classification and XTALK detection in LRTD stereo mode. Analyzers 101 / 102 use, for example, the following relationship to calculate the maximum value of the interchannel correlation function, which is also used as a feature by XTALK detection in LRTD stereo mode:

[0078]

[0079] Furthermore, as a non-limiting embodiment, the following relationship is used to calculate the location of the maximum value:

[0080]

[0081] When the maximum value R of the inter-channel correlation function max When negative, it is set to 0. The maximum value R between the current frame and the previous frame. max The difference between them is calculated, for example, as:

[0082]

[0083] The superscript [-1] indicates a reference to a previous frame.

[0084] The location of the maximum value of the inter-channel correlation function determines which channel becomes the "reference" channel (REF) and the "target" channel (TAR) in the ICA module. If the location k maxIf k ≥ 0, then the left channel (L) is the reference channel (REF) and the right channel (R) is the target channel (TAR). max If < 0, then the right channel (R) is the reference channel (REF), and the left channel (L) is the target channel (TAR). The target channel (TAR) is then shifted to compensate for its delay relative to the reference channel (REF). The number of samples used to shift the target channel (TAR) can, for example, be directly set to |k|. max However, in order to eliminate the location k caused by consecutive frames... max Artifacts caused by sudden changes can be smoothed out by using appropriate filters within the ICA module to reduce the number of samples used for the Shift Target Channel (TAR).

[0085] Let the number of samples used for the target anatomical resonant (TAR) shift be denoted as k. shift , where k shift >0. Let the reference channel signal be labeled X. ref (n), and the target channel signal is labeled as X. tar (n). Instantaneous target gain reflects the energy ratio between the reference channel (REF) and the shifted target channel (TAR). Instantaneous target gain can be calculated, for example, using the following relationship:

[0086]

[0087] Where N is the frame length. Instantaneous target gain is used as a feature by UNCLR classification in LRTD stereo mode.

[0088] 2.1 Inter-channel characteristics

[0089] Analyzers 101 / 102 directly derive the first feature set used in UNCLR classification and XTALK detection from inter-channel analysis. The value of the inter-channel correlation function R(0) at zero hysteresis is itself used as a feature for UNCLR classification and XTALK detection in LRTD stereo mode. Another feature used by UNCLR classification and XTALK detection in LRTD stereo mode is obtained by calculating the logarithm of the absolute value of C(0), as follows:

[0090]

[0091] The energy ratio of the side signal to the mono signal was also used as a feature by UNCLR classification and XTALK detection in the LRTD stereo mode. This ratio was calculated, for example, using the following relationship:

[0092]

[0093] The energy ratio of relation (15) is smoothed over time, for example, as shown below:

[0094]

[0095] Among them, c hang The VAD hangover frame count is calculated as part of the VAD (voice activity detection) module of the stereo audio signal encoding device 100 (stereo codec) (see, for example, reference [1]). The smoothed ratio of relation (16) is used as a feature by XTALK detection in LRTD stereo mode.

[0096] Analyzers 101 / 102 derive the following dot products from the left channel and mono signals, and from the right channel and mono signals. First, the dot product between the left channel and mono signals is expressed, for example, as:

[0097]

[0098] And the dot product between the right channel and the mono signal is, for example:

[0099]

[0100] Both dot products are positive, with a lower bound of 0. A metric based on the difference between the maximum and minimum values ​​of these two dot products is used as a feature for UNCLR classification and XTALK detection in LRTD stereo mode. It can be calculated using the following relationship:

[0101] d mmLR =max[C LM C RM ]-min[C LM C RM (19)

[0102] The similarity measure used as an independent feature by UNCLR classification and XTALK detection in LRTD stereo mode is directly based on the absolute difference between two dot products in the linear and logarithmic domains, which is calculated, for example, using the following relationship:

[0103]

[0104] The final feature used by the UNCLR classification and XTALK detection in the LRTD stereo mode is calculated as part of inter-channel correlation analysis operation 151 / 152 and reflects the evolution of the inter-channel correlation function. It can be calculated as follows:

[0105]

[0106] The superscript [-2] indicates a reference to the second frame preceding the current frame.

[0107] 3. Time Domain (TD) Preprocessing

[0108] In LRTD stereo mode, there is no mono downmixing, and both the left and right channels of the input stereo audio signal 190 are analyzed to extract features in corresponding temporal preprocessing operations (i.e., operation 153 for temporal preprocessing of the left channel of the stereo audio signal 190 and operation 154 for temporal preprocessing of the right channel of the stereo audio signal 190). To perform operations 153 and 154, the feature extractor (not shown) includes, for example... Figure 1 The corresponding time-domain preprocessors 103 and 104 are shown. Operations 153 and 154, as well as the corresponding preprocessors 103 and 104, are similar and will be described simultaneously.

[0109] Temporal preprocessing operations 153 / 154 perform multiple sub-operations to produce certain parameters that are used as extracted features for UNCLR classification and XTALK detection. These sub-operations may include:

[0110] - Spectrum analysis;

[0111] - Linear predictive analysis;

[0112] - Open-loop pitch estimation;

[0113] -Voice Activity Detection (VAD);

[0114] - Background noise estimation; and

[0115] - Frame Error Hiding (FEC) category.

[0116] The time-domain preprocessors 103 / 104 perform linear predictive analysis using the Levinson-Durbin algorithm. The output of the Levinson-Durbin algorithm is a set of linear prediction coefficients (LPCs). The Levinson-Durbin algorithm is an iterative method, and the total number of iterations in the Levinson-Durbin algorithm can be denoted as M. In each i-th iteration, where i = 1, ..., M, the residual energy... The calculation is performed. In this disclosure, as a non-limiting illustrative implementation, it is assumed that the Levinson-Durbin algorithm runs iteratively with M=16. The difference in residual energy between the left and right channels of the input stereo audio signal 190 is used as a feature for XTALK detection in the LRTD stereo mode. The difference in residual energy can be calculated as follows:

[0117]

[0118] Here, subscripts L and R are added to represent the left and right channels of the input stereo audio signal 190, respectively. In this non-limiting embodiment, the feature is that the residual energy from the 14th iteration, rather than the last iteration, is used to calculate (difference). LPC13 This is because experiments have shown that this iteration has the highest discriminative potential for UNCLR classification. More information about the Levinson-Durbin algorithm and details about residual energy calculation can be found, for example, in reference [1].

[0119] The LPC coefficients estimated using the Levinson-Durbin algorithm are converted into line spectral frequencies, LSF(i), i = 0, ..., M-1. The sum of the LSF values ​​can be used as an estimate of the gravity point of the envelope of the input stereo audio signal 190°. The difference between the sums of the LSF values ​​in the left and right channels contains information about the similarity between the two channels. Therefore, this difference is used as a feature in XTALK detection in the LRTD stereo mode. The difference between the sums of the LSF values ​​in the left and right channels can be calculated using the following relationship:

[0120]

[0121] Additional information about the LPC to LSF conversion described above can be found, for example, in reference [1].

[0122] The time-domain preprocessors 103 / 104 perform open-loop pitch estimation and use the autocorrelation function to calculate the open-loop pitch difference between the left channel (L) and the right channel (R). The open-loop pitch difference between the left channel (L) and the right channel (R) can be calculated using the following relationship:

[0123]

[0124] Where T [k] This is the open-loop pitch estimate in the k-th segment of the current frame. In this disclosure, as a non-limiting illustrative example, it is assumed that the open-loop pitch analysis is performed in three adjacent half-frames (segments) with indices k = 1, 2, 3, where two segments are in the current frame and one segment is in the second half of the previous frame. Different numbers of segments, as well as different segment lengths and overlaps, can be used. Additional information regarding open-loop pitch estimation can be found, for example, in reference [1].

[0125] The difference between the maximum autocorrelation values ​​(sound output) of the left and right channels of the input stereo audio signal 190 (determined by the autocorrelation function described above) is also used as a feature by XTALK detection in the LRTD stereo mode. The difference between the maximum autocorrelation values ​​of the left and right channels can be calculated using the following relationship:

[0126]

[0127] Where v [k] This represents the maximum autocorrelation value of the left (L) channel and the right (R) channel in the k-th half-frame.

[0128] Background noise estimation is part of the Voice Activity Detection (VAD) algorithm (see reference [1]). Specifically, background noise estimation uses an active / inactive signal detector (not shown) that depends on a feature set, some of which are used by UNCLR classification and XTALK detection. For example, the active / inactive signal detector (not shown) generates nonstationary parameters f for the left channel (L) and right channel (R). sta This is used as a measure of spectral stability. The difference in non-stationarity between the left and right channels of the input stereo audio signal 190 is used as a feature by XTALK detection in the LRTD stereo mode. The difference in non-stationarity between the left (L) channel and the right (R) channel can be calculated using the following relationship:

[0129] d sta =|f sta,L -f sta,R | (26)

[0130] The active / inactive signal detector (not shown) depends on the correlation graph parameter C. map Harmonic analysis was performed. The correlation plot is a measurement of the pitch stability of the input stereo audio signal 190, and it is used for UNCLR classification and XTALK detection. The difference between the correlation plots of the left (L) channel and the right (R) channel is used as a feature by XTALK detection in the LRTD stereo mode, and is calculated using, for example, the following relationship:

[0131] d cmap =|C map,L -C map,R | (27)

[0132] Finally, an active / inactive signal detector (not shown) periodically measures the spectral diversity and noise characteristics in each frame. These two parameters are also used as features by UNCLR classification and XTALK detection in the LRTD stereo mode. Specifically, (a) the difference in spectral diversity between the left channel (L) and the right channel (R) can be calculated as follows:

[0133] d sdiv =|log(S) div,L )-log(S div,R (28)

[0134] Where S divThis represents a measurement of spectral diversity in the current frame, and (b) the difference in noise characteristics between the left channel (L) and the right channel (R) can be calculated as follows:

[0135] d nchar =|log(n) char,L )-log(n char,R (29)

[0136] Where n char This represents a measurement of the noise characteristics in the current frame. For details on the calculation of correlation plots, nonstationarity, spectral diversity, and noise characteristic parameters, please refer to [1].

[0137] As described in reference [1], the ACELP (Algebraic Digital Excited Linear Prediction) core encoder, which is part of the stereo audio signal encoding device 100, includes specific settings for encoding unvoiced sounds. The use of these settings is adjusted by several factors, including the measurement of sudden energy increases in short segments within the current frame. The settings for encoding unvoiced sounds in the ACELP core encoder are applied only when there are no sudden energy increases within the current frame. The start of crosstalk segments can be located by comparing the measurements of sudden energy increases in the left and right channels. Sudden energy increases can be similar to those described in the 3GPP EVS codec (reference [1]). d The difference in sudden energy increase between the left channel (L) and the right channel (R) can be calculated using the following relationship:

[0138] d dE =|log(E) d,L )-log(E d,R (30)

[0139] The subscripts L and R are added to represent the left and right channels of the input stereo audio signal 190, respectively.

[0140] The temporal preprocessors 103 / 104 and preprocessing operations 153 / 154 use an FEC classification module containing a state machine for FEC technology. The FEC category in each frame is selected from predefined categories based on an evaluation function. The difference between the FEC categories selected for the left channel (L) and right channel (R) in the current frame is used as a feature by XTALK detection in the LRTD stereo mode. However, for such classification and detection purposes, the FEC categories may be subject to the following limitations:

[0141]

[0142] Where t classThe FEC category is selected in the current frame. Therefore, the FEC category is limited to voiced (“VOICED”) and unvoiced (“UNVOICED”). The difference between the classes in the left channel (L) and the right channel (R) can be calculated as follows:

[0143] d class =|t class,L -t class,R |(32)

[0144] For additional details regarding the FEC classification, please refer to [1].

[0145] The temporal preprocessors 103 / 104 and preprocessing operations 153 / 154 implement speech / music classification and the corresponding speech / music classifier. This speech / music classification performs a binary decision in each frame based on power spectral divergence and power spectral stability. The difference in power spectral divergence between the left channel (L) and the right channel (R) is calculated, for example, using the following relationship:

[0146] d Pdiff =|P diff,L -P diff,R |(33)

[0147] Among them, P diff This represents the power spectral divergence in the left (L) and right (R) channels of the current frame, and the difference in power spectral stability between the left (L) and right (R) channels is calculated, for example, using the following relationship:

[0148] d Psta =|P sta,L -P sta,R |(34)

[0149] Where P sta This indicates the power spectral stability of the left (L) and right (R) channels in the current frame.

[0150] Reference [1] describes the details of power spectral divergence and power spectral stability calculated in speech / music classification.

[0151] 4. DFT stereo parameters

[0152] The method 150 for encoding stereo audio signal 190 includes an operation 155 of calculating the Fast Fourier Transform (FFT) of the left channel (L) and the right channel (R). To perform operation 155, the apparatus 100 for encoding stereo audio signal 190 includes an FFT calculator 105.

[0153] The feature extraction operation (not shown) includes operation 156 for calculating DFT stereo parameters. To perform operation 156, the feature extractor (not shown) includes a calculator 106 for DFT stereo parameters.

[0154] In DFT stereo mode, the transform calculator 105 transforms the left channel (L) and right channel (R) of the input stereo sound signal 190 to the frequency domain through FFT transformation.

[0155] Let the complex spectrum of the left channel (L) be labeled as Furthermore, the complex spectrum of the right channel (R) is denoted as Where k = 0, ..., N FFT -1 is the index of the frequency element, and N FFT This is the length of the FFT transform. For example, when the sampling rate of the input stereo audio signal is 32kHz, the DFT stereo parameter calculator 106 calculates the complex spectrum over a 40ms window, obtaining N. FFT = 1280 samples. Subsequently, as a non-limiting embodiment, the following relationship can be used to calculate the multi-channel spectrum:

[0156]

[0157] The asterisk superscript indicates complex conjugation. The complex cross-channel spectrum can be decomposed into real and imaginary parts using the following relationship:

[0158]

[0159] Using the real and imaginary parts, the absolute amplitude of the complex cross-channel spectrum can be expressed as:

[0160]

[0161] The total absolute amplitude of the complex cross-channel spectrum is obtained by summing the absolute amplitudes of the spectrum at the frequency elements using the following relationship:

[0162]

[0163] The energy spectrum of the left channel (L) and the energy spectrum of the right channel (R) can be expressed as:

[0164]

[0165] The total energy of the left channel (L) and the right channel (R) can be obtained by summing the energy spectra of the left channel (L) and the right channel (R) over the frequency elements using the following relationship:

[0166]

[0167] UNCLR classification and XTALK detection in DFT stereo mode use the total absolute amplitude of the multi-channel spectrum as one of their features, but not in the direct form defined above, but in an energy-normalized form, and in the logarithmic domain, as expressed using, for example, the following relation:

[0168]

[0169] The DFT stereo parameter calculator 106 can calculate the mixing energy in mono using, for example, the following relationship:

[0170] E M =E L +E R +2X LR |(42)

[0171] Inter-channel level difference (ILD) is a feature used for UNCLR classification and XTALK detection in DFT stereo mode because it contains information about the angle from which the main sound originates. For UNCLR classification and XTALK detection purposes, inter-channel level difference (ILD) can be expressed as a gain factor. The DFT stereo parameter calculator 106 uses, for example, the following relationship to calculate the inter-channel level difference (ILD) gain:

[0172]

[0173] Inter-channel phase difference (IPD) contains information that a listener can infer from it the direction of the incoming sound signal. The DFT stereo parameter calculator 106 uses, for example, the following relationship to calculate the inter-channel phase difference (IPD):

[0174]

[0175] in:

[0176]

[0177] The inter-channel phase difference (IPD) relative to the previous frame is calculated, for example, using the following relationship:

[0178] d IPD =|IPD [n] -IPD [n-1] |(46)

[0179] The superscript n is used to indicate the current frame, and the superscript n-1 is used to indicate the previous frame. Finally, the calculator 106 can calculate the IPD gain as the energy E of the phase-aligned (IPD=0) under-mixed energy (the numerator of relation (47)) and the mono under-mixed energy. M The ratio between:

[0180]

[0181] IPD gain g IPD_lin The value is limited to the interval <0,1>. If the value exceeds the upper threshold of 1.0, it is replaced with the IPD gain value from the previous frame. UNCLR classification and XTALK detection in DFT stereo mode use IPD gain in the logarithmic domain as a feature. Calculator 106 uses, for example, the following relationship to determine IPD gain in the logarithmic domain:

[0182] g IPD =log(1-g IPD_lin (48)

[0183] Inter-channel phase difference (IPD) can also be expressed as an angle characterized by UNCLR classification and XTALK detection in the DFT stereo mode, and is calculated, for example, as follows:

[0184]

[0185] Side channels can be calculated as the difference between the left channel (L) and the right channel (R). This difference (E) can be calculated using the following relationship. L –E R The absolute value of the energy relative to the mixing energy E in mono. M The gain of the side channels is expressed as a ratio:

[0186]

[0187] Gain g side The higher the gain g, the greater the energy difference between the left channel (L) and the right channel (R). side Values ​​are restricted to the range <0.01, 0.99>. Values ​​outside this range are restricted.

[0188] The phase difference between the left channel (L) and right channel (R) of the input stereo audio signal 190 can also be analyzed based on the prediction gain calculated using, for example, the following relationship:

[0189] g pred_lin =(1-g side E L +(1+g side E R -2|X LR | (51)

[0190] Where the prediction gain g pred_lin The value of g is restricted to the interval <0,∞>, i.e., positive values. pred_lin The above expression captures the cross-channel spectrum (X).LR Energy and mixed energy under mono channel E M =E L +E R +2|X LR The difference between |. Calculator 106 uses, for example, relation (52) to calculate this gain g. pred_lin Transform to the logarithmic domain for use as features in DFT stereo mode UNCLR classification and XTALK detection:

[0191] g pred =log(g pred_lin +1) (52)

[0192] Calculator 106 also uses the per-bin channel energy of relation (39) to calculate the mean energy of inter-channel coherence (ICC), which forms a clue for determining the difference between the left channel (L) and the right channel (R) that are not captured by inter-channel time difference (ITD) (described below) and inter-channel phase difference (IPD). First, calculator 106 uses, for example, the following relation to calculate the total energy of the cross-channel spectrum:

[0193] E X =Re(X) LR ) 2 +Im(X LR ) 2 (53)

[0194] To express the mean energy of inter-channel coherence (ICC), it is useful to calculate the following parameters:

[0195]

[0196] Subsequently, the mean energy of inter-channel coherence (ICC) was used as a feature by UNCLR classification and XTALK detection in the DFT stereo mode, and can be expressed as

[0197]

[0198] If the inner term is less than 1.0, then the mean energy E coh The value is set to 0. Another possible interpretation of inter-channel coherence (ICC) is calculated as the side-to-mono energy ratio as follows:

[0199]

[0200] Finally, calculator 106 determines the ratio r of the maximum to minimum intra-channel amplitude product used in UNCLR classification and XTALK detection. ppThe feature used as a feature by UNCLR classification and XTALK detection in DFT stereo mode is calculated, for example, using the following relation:

[0201]

[0202] The amplitude product within the vocal tract is defined as follows:

[0203]

[0204] One parameter used in stereo sound signal reproduction is the inter-channel time difference (ITD). In DFT stereo mode, the DFT stereo parameter calculator 106 estimates the inter-channel time difference (ITD) based on the generalized cross-channel correlation (GCC-PHAT) function with phase difference. The inter-channel time difference (ITD) corresponds to the time delay of arrival (TDOA) estimate. The GCC-PHAT function is a robust method for estimating the inter-channel time difference (ITD) on reverberant signals. The GCC-PHAT is calculated, for example, using the following relationship:

[0205]

[0206] IFFT stands for Inverse Fast Fourier Transform.

[0207] The inter-channel time difference (ITD) is then estimated using, for example, the following relationship according to the GCC-PHAT function:

[0208]

[0209] Where d represents the time lag in the sample corresponding to a time delay ranging from -5ms to +5ms. ITD The maximum value of the GCC-PHAT function is used as a feature by UNCLR classification and XTALK detection in DFT stereo mode, and can be retrieved using the following relation:

[0210]

[0211] In monophonic scenarios, a single dominant peak is typically present in the GCC-PHAT function corresponding to the inter-channel time difference (ITD). However, in crosstalk situations where two speakers are located on opposite sides of the capture microphone, two dominant peaks are typically present and located separately from each other. Figure 2 The diagram illustrates this situation. Specifically, based on a non-limiting illustrative example, Figure 2 It is a plan view of a crosstalk scene with two opposing speakers, S1 and S2, captured by a pair of supercardioid microphones M1 and M2, and Figure 3 This is a graph showing the locations of the two dominant peaks in the GCC-PHAT function.

[0212] The amplitude of the first peak G ITD It is calculated using relation (61), and its location d ITD It is calculated using relation (60). The amplitude of the second peak is located by searching for the second maximum value of the GCC-PHAT function in the opposite direction to the first peak. More specifically, the search direction for the second peak is s ITD The location of the first peak d ITD The sign is determined by:

[0213] s ITD =sgn(d ITD (62)

[0214] Where sgn(.) is the sign function.

[0215] The DFT stereo parameter calculator 106 can then use, for example, the following relationship to retrieve the values ​​in direction s. ITD The second maximum value (second highest peak value) of the GCC-PHAT function on the above:

[0216]

[0217] As a non-limiting embodiment, the threshold thr xt =8 Ensure the second peak of the GCC-PHAT function is at a distance from the start (d ITD =0) The search is conducted at a distance of at least 8 samples. In terms of crosstalk (XTALK) detection, this means that any potential secondary speaker in the scene must exist at least a certain minimum distance from both the first “dominant” speaker and the intermediate point (d=0).

[0218] The location of the second highest peak of the GCC-PHAT function is calculated by substituting the arg max(.) function for the max(.) function and using relation (63). The location of the second highest peak of the GCC-PHAT function will be denoted as d. ITD2 .

[0219] The relationship between the amplitudes of the first peak and the second highest peak of the GCC-PHAT function is used as a feature by XTALK detection in the DFT stereo mode, and can be evaluated using the following ratios:

[0220]

[0221] ratio r GITD12It has high discriminative potential, but in order to use it as a feature, XTALK detection eliminates accidental false alarms caused by the limited time resolution applied during frequency transformations in DFT stereo mode. This can be achieved by using, for example, the following relationship to measure the ratio r in the current frame. GITD12 The value is multiplied by the value from the previous frame at the same ratio to complete the process:

[0222] r GITD12 ←r GITD12 (n)·r GITD12 (n-1) (65)

[0223] An index n is added to identify the current frame, and an index n-1 is added to identify the previous frame. For simplicity, the parameter name is r. GITD12 It is reused to identify output parameters.

[0224] The amplitude of the second highest peak alone serves as an indicator of the intensity of the secondary speaker in the scene. Similar to the ratio r... GITD12 Use, for example, the following relation (66) to reduce the value G. ITD2 The random "spikes" are used to obtain another feature employed by XTALK detection in the DFT stereo mode:

[0225] m ITD2 =G ITD2 (n)·G ITD2 (n-1) (66)

[0226] Another feature used in XTALK detection in DFT stereo mode is the localization of the second highest peak in the current frame. ITD2 (n) is the difference relative to the previous frame, which is calculated using, for example, the following relationship:

[0227] Δ ITD2 =|d ITD2 (n)-d ITD2 (n-1)| (67)

[0228] 5. Lower Mixture and Inverse Fast Fourier Transform (IFFT)

[0229] In DFT stereo mode, the method 150 for encoding a stereo audio signal includes an operation 157 of downmixing the left channel (L) and right channel (R) of the stereo audio signal 190 and an operation 158 of calculating the IFFT transform of the downmixed signal. To perform operations 157 and 158, the device 100 for encoding the stereo audio signal 190 includes a downmixer 107 and an IFFT transform calculator 108.

[0230] For example, as described in reference [6] (the entire contents of which are incorporated herein by reference), the downmixer 107 downmixes the left channel (L) and right channel (R) of the stereo audio signal into a mono channel (M) and a side channel (S).

[0231] The IFFT transform calculator 108 then calculates the IFFT transform of the downmixed mono (M) from the downmixer 107 to produce the time-domain mono (M) to be processed in the TD preprocessor 109. The IFFT transform used in calculator 108 is the inverse of the FFT transform used in calculator 105.

[0232] 6. TD preprocessing in DFT stereo mode

[0233] In DFT stereo mode, the feature extraction operation (not shown) includes a TD preprocessing operation 159 for extracting features used in UNCLR classification and XTALK detection. To perform operation 159, the feature extractor (not shown) includes a TD preprocessor 109 responsive to mono (M) audio.

[0234] 6.1 Voice Activity Detection

[0235] UNCLR classification and XTALK detection use the Voice Activity Detection (VAD) algorithm. In LRTD stereo mode, the VAD algorithm runs separately on the left (L) and right (R) channels. In DFT stereo mode, the VAD algorithm runs on the lower-mixed mono channel (M). The output of the VAD algorithm is a binary flag f. VAD VAD logo f VAD It is not suitable for UNCLR classification and XTALK detection because it is too conservative and has a long hysteresis. This prevents rapid switching between LRTD stereo mode and DFT stereo mode (e.g., at the end of a conversation burst or during a short pause in the middle of a speech). Furthermore, the VAD flag f VAD It is sensitive to small changes in the input stereo audio signal 190. This can lead to false alarms and incorrect selection of stereo modes in crosstalk detection. Therefore, UNCLR classification and XTALK detection use alternative measures for voice activity detection based on changes in relative frame energy. For details on the VAD algorithm, see [1].

[0236] 6.1.1 Relative Frame Energy

[0237] UNCLR classification and XTALK detection use the absolute energy E of the left channel (L) obtained by utilizing relation (2). L The absolute energy E of the right channel (R) R The maximum average energy of the input stereo sound signal can be calculated in the logarithmic domain using, for example, the following relationship:

[0238]

[0239] An index n is added to identify the current frame, and N = 160 is the length of the current frame (the length of 160 samples). The maximum average energy E in the logarithmic domain. ave The value of (n) is restricted to the interval <0; ∞>.

[0240] The maximum average energy E can then be calculated using, for example, the following relationship: ave (n) The relative frame energy of the input stereo audio signal is calculated by linearly mapping to the interval <0; 0.9>:

[0241]

[0242] Where E up (n) indicates the relative frame energy E rl The upper bound of (n), E dn (n) indicates the relative frame energy E rl The lower bound of (n) is determined, and the index n indicates the current frame.

[0243] Relative frame energy E rl The bound of (n) is updated based on the noise-based count value a in each frame. En (n) (which is part of the noise estimation module of TD preprocessors 103, 104 and 109) is updated. For additional information about this count value, see [1]. Count value a En The purpose of (n) is to signal that the background noise level of each channel in the current frame can be updated. This occurs when the count value a En When (n) has a value of zero. As a non-restrictive example, the count value a in each channel. En (n) is initialized to 6 and increments or decrements in each frame, having a lower threshold of 0 and an upper threshold of 6.

[0244] In LRTD stereo mode, noise estimation is performed independently on the left channel (L) and right channel (R). The two noise update count values ​​are represented as a for the left channel (L) and right channel (R), respectively. En,L (n) and a En,R (n). The two count values ​​can then be combined into a single binary parameter using the following relationship:

[0245]

[0246] In DFT stereo mode, noise estimation is performed on the lower-mixed mono channel (M). The noise update count value in the mono channel is denoted as a. En,M(n). The binary output parameters are calculated using the following relationship:

[0247]

[0248] UNCLR classification and XTALK detection use a binary parameter f En (n) to achieve relative frame energy E rl E is the lower bound of (n). dn (n) or upper bound E up Update (n). When parameter f En When (n) equals zero, the lower bound E dn (n) is updated. When parameter f En When (n) equals 1, the upper bound E up (n) is updated.

[0249] Relative frame energy E rl E is the upper bound of (n). up (n) is where the parameter f is... En Frames where (n) equals 1 are updated using, for example, the following relationship:

[0250]

[0251] Where index n represents the current frame, and index n-1 represents the previous frame.

[0252] The first and second rows in relation (71) represent slower and faster updates, respectively. Therefore, using relation (71), the upper bound E increases as energy increases. up (n) is updated more quickly.

[0253] Relative frame energy E rl E is the lower bound of (n). dn (n) is where the parameter f is... En Frames where (n) equals 0 are updated using, for example, the following relationship:

[0254] E dn (n) = 0.9E dn (n-1)+0.1E ave (n) (72)

[0255] The lower threshold is 30.0. If the upper bound E... up The value of (n) and the lower bound E dn If (n) is too close, it is modified, for example, as shown below:

[0256]

[0257] 6.1.2 Alternative VAD Marker Estimation

[0258] UNCLR classification and XTALK detection use the relative frame energy E calculated in relation (71). rl The change of (n) serves as the basis for calculating the alternative VAD flag. Let the alternative VAD flag in the current frame be denoted as f. xVAD (n). Replacement VAD flag f xVAD (n) is the VAD flag generated in the noise estimation module of TD preprocessor 103 / 104 in LRTD stereo mode or in TD preprocessor 109 in DFT stereo mode. VAD Reflecting relative frame energy E rl The auxiliary binary parameter f of the change of (n) Erl (n) is calculated by combining them.

[0259] First, the relative frame energy E is calculated over segments from 10 previous frames using, for example, the following relationship. rl (n) Calculate the average:

[0260]

[0261] Where p is the average index. The auxiliary binary parameters are set according to, for example, the following logic:

[0262]

[0263] In LRTD stereo mode, the VAD flag is replaced by f. xVAD (n) is through the VAD flag in the left channel (L). VAD,L (n), VAD flag in the right channel (R) VAD,R (n), and auxiliary binary parameter f Erl Logical combinations of (n) are calculated using, for example, the following relations:

[0264] f xVAD (n)=(f VAD,L (n)ORf VAD,R (n))ANDf Erl (n) (76)

[0265] In DFT stereo mode, the VAD flag is replaced by f. xVAD (n) is the VAD flag f in the lower-mix mono channel (M). VAD,M (n), and auxiliary binary parameter f Erl Logical combinations of (n) are computed using, for example, the following relations.

[0266] f xVAD (n)=f VAD,M (n)ANDf Erl(n) (77)

[0267] 6.2 Stereo Mute Indicator

[0268] In DFT stereo mode, it is also convenient to calculate discrete parameters reflecting the low levels of the mixed mono (M) channels. Such parameters, referred to as stereo mute flags, can be calculated, for example, by comparing the average level of the active signal with a specific predefined threshold. As an example, the long-term active speech level is calculated within the VAD algorithm of the TD preprocessor 109. It can be used as the basis for calculating stereo mute flags. For details on the VAD algorithm, see [1].

[0269] The stereo mute indicator can then be calculated using the following relationship:

[0270]

[0271] Where E M (n) is the absolute energy of the lower-mixed mono channel (M) in the current frame. Stereo mute flag f sil (n) is restricted to the interval <0; ∞>.

[0272] 7. Classification of Unrelated Stereo Content (UNCLR)

[0273] The UNCLR classification in LRTD stereo mode and DFT stereo mode is based on a logistic regression (LogReg) model (see reference [9]). The LogReg model is trained separately for LRTD stereo mode and DFT stereo mode on a large labeled database consisting of correlated and uncorrelated stereo sound signal samples. Uncorrelated stereo training samples are artificially created by combining randomly selected mono samples. The following stereo scenes can be simulated using this artificial mixing of mono samples:

[0274] Speaker A is in the left channel, and speaker B is in the right channel (or vice versa);

[0275] - Speaker A is in the left channel, and the music is in the right channel (or vice versa);

[0276] - Speaker A is in the left channel, and noise is in the right channel (or vice versa);

[0277] - Speaker A is in the left or right channel, and background noise is in both channels;

[0278] Speaker A is in either the left or right channel, and the background music is in both channels.

[0279] In a non-limiting implementation, the mono samples were selected from an AT&T mono clean speech database sampled at 16 kHz. Active segments were extracted from the mono samples using only any convenient VAD algorithm (e.g., the VAD algorithm for the 3GPP EVS codec described in reference [1]). The total size of the stereo training database with irrelevant content was approximately 240 MB. No level adjustment was applied to the mono signals before they were combined to form the stereo sound signal. Level adjustment was applied only after this process. Based on passive mono downmixing, the level of each stereo sample was normalized to -26 dBov. Therefore, the inter-channel level difference remained unchanged and was still the primary factor in determining the dominant speaker's location in the stereo scene.

[0280] The relevant stereo training samples are obtained from various real records of stereo sound signals. The total size of the training database with relevant stereo content is approximately 220 MB. In a non-limiting embodiment, the relevant stereo training samples include data from... Figure 4 The following are examples of scenarios shown. Figure 4 A top view of the stereo scene setup used for actual recording is shown:

[0281] - Speaker S1 is positioned at point P1, close to microphone M1; speaker S2 is positioned at point P2, close to microphone M6.

[0282] - Speaker S1 is positioned at P4, close to microphone M3; Speaker S2 is positioned at P3, close to microphone M4.

[0283] - Speaker S1 is positioned at P6, close to microphone M1; speaker S2 is positioned at P5, close to microphone M2.

[0284] - Only speaker S1 is positioned at P4, in the M1-M2 stereo recording;

[0285] - Only speaker S1 is positioned at P4, in the M3-M4 stereo recording;

[0286] Let the total size of the training database be denoted as:

[0287] N T =N UNC +N CORR (79)

[0288] Where N UNC N is the size of the uncorrelated stereo training sample set. CORR This is the size of the relevant stereo training sample set. Labels are manually assigned using simple rules such as the following:

[0289]

[0290] Where Ω UNC It is the entire feature set of the unrelated training database, and Ω CORR This is the entire feature set of the relevant training database. In this illustrative, non-limiting implementation, inactive frames (VAD=0) are discarded from the training database.

[0291] Each frame in the unrelated training database is labeled "1", and each frame in the related training database is labeled "0". Inactive frames with VAD=0 are ignored during training.

[0292] 7.1LRTD Stereo Mode UNCLR Classification

[0293] In LRTD stereo mode, the method 150 for encoding the stereo audio signal 190 includes an operation 161 of classifying unrelated stereo content (UNCLR). To perform operation 161, the device 100 for encoding the stereo audio signal 190 includes an UNCLR classifier 111.

[0294] The UNCLR classification operation in LRTD stereo mode is based on a logistic regression (LogReg) model. The following features, extracted by running a device 100 (stereo codec) for encoding stereo sound signals on both uncorrelated and correlated stereo training databases, are used for the UNCLR classification operation:

[0295] -Location of the maximum value of the inter-channel cross-correlation function, k max (Relationship(11));

[0296] -Instantaneous target gain, g t (Relationship(13));

[0297] - The logarithm of the absolute value of the inter-channel correlation function at zero hysteresis, p LR (Relationship(14));

[0298] - Side-to-mono energy ratio, r SM (Relationship(15));

[0299] - The difference between the maximum and minimum values ​​of the dot product between the left / right channel and the mono signal, d mmLR (Relationship(19));

[0300] - The absolute difference in the logarithmic domain between the dot product of the left channel (L) and the mono signal (M) and the dot product of the right channel and the mono signal (M), d LRM (Relationship(20));

[0301] - The zero hysteresis value of the cross-channel correlation function, R0 (relation (5)); and

[0302] - Evolution of the inter-channel correlation function, RR (relationship (21)).

[0303] The UNCLR classifier 111 uses a total of F = 8 features.

[0304] Prior to the training process, the UNCLR classifier 111 includes a normalizer (not shown), which performs a suboperation (not shown) to normalize the feature set by removing the mean of the features and scaling them to unit variance. The normalizer (not shown) uses, for example, the following relationship for this purpose:

[0305]

[0306] Where f i,raw The i-th feature of the set, f i Indicates the i-th feature of the normalization, Let σ represent the global mean of the i-th feature across the training database. fi It is the global variance of the i-th feature across the training database.

[0307] The LogReg model used by the UNCLR classifier 111 takes real-valued features as an input vector and predicts the probability that the input belongs to the unrelated class (class 0) indicating unrelated stereo content (UNCLR). For this purpose, the UNCLR classifier 111 includes a score calculator (not shown) that performs a suboperation (not shown) to calculate a score representing the unrelated stereo content in the input stereo sound signal 190. The score calculator (not shown) calculates the output of the LogReg model in the form of a linear regression of the extracted features, which can be expressed using the following relationship:

[0308] y p =b0+b1f1+...+b F f F (82)

[0309] Among them, b i Indicates the coefficients of the LogReg model, and f i Individual characteristics are identified. Then, a logical function, such as the following, is used to output the real value y. p Transform into probabilities:

[0310]

[0311] The probability p(class=0) takes a real value between 0 and 1. Intuitively, a probability closer to 1 means that the current frame is highly stereo uncorrelated, that is, it has uncorrelated stereo content.

[0312] The goal of the learning process is to find the coefficient b based on the training data. i The optimal value of F is found by minimizing the difference between the predicted output p (class=0) and the true output y on the training database. The UNCLR classifier 111 in the LRTD stereo mode is trained using an iterative method of stochastic gradient descent (SGD) as described, for example, in reference

[10] (the entire contents of which are incorporated herein by reference).

[0313] Binary classification can be performed by comparing the probability output p(class=0) with a fixed threshold (e.g., 0.5). However, for the purpose of UNCLR classification in LRTD stereo mode, the probability output p(class=0) is not used. Instead, the raw output y of the LogReg model is used. p It will be further processed as shown below.

[0314] The UNCLR classifier 111 score calculator (not shown) first uses, for example... Figure 5 The function shown represents the original output y of the LogReg model. p Normalize. Figure 5 This is a graph showing the normalization function applied to the raw output of the LogReg model in the UNCLR classification in LRTD stereo mode.

[0315] Figure 5 The normalization function can be mathematically described as follows:

[0316]

[0317] 7.1.1 LogReg Output Weighting Based on Relative Frame Energy

[0318] The UNCLR classifier 111's score calculator (not shown) then uses, for example, the following relationship to normalize the LogReg model's output y using relative frame energy. pn (n) weighted:

[0319] scr UNCLR (n)=y pn (n)·E rl (n) (85)

[0320] Where E rl (n) is the relative frame energy described by relation (69). The normalized weighted output scr of the LogReg modelUNCLR (n) is referred to as the “score” or unrelated stereo content in the input stereo sound signal 190.

[0321] 7.1.2 Rising Edge Detection

[0322] For UNCLR classification, the score is scr UNCLR (n) still cannot be directly used by the UNCLR classifier 111 because it contains occasional short-term "peaks" generated by an imperfect statistical model. These peaks can be filtered out by a simple averaging filter such as a first-order IIR filter. Unfortunately, the application of such an averaging filter usually results in the smoothing of the rising edges that represent the transition between stereo correlated and uncorrelated content in the input stereo audio signal 190. To preserve the rising edges, the smoothing process (the application of the averaging IIR filter) is reduced or even stopped when a rising edge is detected in the input stereo audio signal 190. The detection of rising edges in the input stereo audio signal 190 is achieved by analyzing the relative frame energy E. rl It is accomplished through the evolution of (n).

[0323] Relative frame energy E rl The rising edge of (n) is obtained by filtering the relative frame energy using a cascaded array of P = 20 identical first-order resistor-capacitor (RC) filters, each filter having, for example, the following form:

[0324]

[0325] Constants a0, a1, and b1 are chosen such that:

[0326]

[0327] Therefore, a single parameter τ edge It is used to control the time constant of each RC filter. Experimentally, it was found that using τ... edge A value of 0.3 yields good results. Using a cascaded array of P=20 RC filters, the relative frame energy E... rl The filtering of (n) can be performed as follows:

[0328]

[0329] The superscripts p = 0, 1, ..., P–1 are added to denote the stages in a cascaded RC filter. The output of the cascaded RC filter is equal to the output from the last stage, i.e.

[0330]

[0331] The reason for using cascaded first-order RC filters instead of a single higher-order RC filter is to reduce computational complexity. Multiple cascaded first-order RC filters act as low-pass filters with relatively sharp step functions. When the relative frame energy E... rl When used on (n), it tends to smooth out occasional short spikes while retaining slower but important transitions, such as start and offset. Relative frame energy E rl The rising edge of (n) can be quantized by calculating the difference between the relative frame energy and the filtered output using, for example, the following relationship:

[0332] f edge (n) = 0.95 - 0.05(E rl (n)-E f (n)) (90)

[0333] Item f edge (n) is restricted to the interval <0.9; 0.95>. The UNCLR classifier 111's score calculator (not shown) utilizes, for example, the following relationship, using f edge (n) An IIR filter, acting as a forgetting factor, is used to smooth the normalized weighted output scr of the LogReg model. UNCLR (n) is used to produce normalized, weighted, and smoothed scores (the output of the LogReg model):

[0334] wscr UNCLR (n)=f edge (n)·wscr UNCLR (n-1)+(1-f edge (n))·scr UNCLR (n) (91)

[0335] 7.2 UNCLR Classification in DFT Stereo Mode

[0336] In DFT stereo mode, the method 150 for encoding the stereo sound signal 190 includes an operation 163 of classifying unrelated stereo content (UNCLR). To perform operation 163, the device 100 for encoding the stereo sound signal 190 includes an UNCLR classifier 113.

[0337] The UNCLR classification in the DFT stereo mode is performed similarly to the UNCLR classification in the LRTD stereo mode, as described above. Specifically, the UNCLR classification in the DFT stereo mode is also based on a logistic regression (LogReg) model. For simplicity, the symbols / names indicating specific parameters from the UNCLR classification in the LRTD stereo mode and their associated mathematical notations are also used in the DFT stereo mode. When the same parameter is referenced from multiple parts simultaneously, a subscript is added to avoid ambiguity.

[0338] The following features, extracted by running on both stereo uncorrelated and stereo correlated training databases on a device 100 (stereo codec) for encoding stereo sound signals, are used by the UNCLR classifier 113 for UNCLR classification in DFT stereo patterns:

[0339] -ILD gain, g ILD (Relationship(43));

[0340] -IPD gain, g IPD (Relationship(48));

[0341] -IPD rotation angle, (Relationship(49));

[0342] -Predicted gain, g pred (Relationship(52));

[0343] -Mean energy of inter-channel coherence, E coh (Relationship(55));

[0344] - The ratio of the amplitude products within the maximum and minimum channels, r PP (Relationship(57));

[0345] - The overall amplitude of the cross-channel spectrum, f X (relation(41)); and

[0346] The maximum value of the GCC-PHAT function, G ITD (Relationship(61)).

[0347] The UNCLR classifier 113 uses a total of F=8 features.

[0348] Prior to the training process, the UNCLR classifier 113 includes a normalizer (not shown) that performs a suboperation (not shown) to normalize the feature set by removing the mean of the features and scaling them to unit variance. The normalizer (not shown) uses, for example, the following relationship for this purpose:

[0349]

[0350] Among them, f i,raw Indicate the i-th feature of the set, Let σ represent the global mean of the i-th feature across the entire training database, and σi represent the global mean of the i-th feature. fi This is again the global variance of the i-th feature across the entire training database. It should be noted that the global mean used in relation (92) is different. and global variance σ fiUnlike the same parameters used in relation (81).

[0351] The LogReg model used in DFT stereo mode is similar to the LogReg model used in LRTD stereo mode. The output y of the LogReg model... p The probability that the current frame has irrelevant stereo content (category = 0) is given by relation (83), as described by relation (82). The classifier training process and the process of finding the optimal decision threshold have been described above. Again, for this purpose, the UNCLR classifier 113 includes a score calculator (not shown) that performs suboperations (not shown) that calculate a score representing the irrelevant stereo content in the input stereo sound signal 190.

[0352] Similar to LRTD stereo mode and according to such Figure 5 The function shown, the score calculator for UNCLR classifier 113 (not shown), first takes the raw output y of the LogReg model. p Normalization. Normalization can be mathematically described as follows:

[0353]

[0354] 7.2.1 LogReg Output Weighting Based on Relative Frame Energy

[0355] The UNCLR classifier 113 score calculator (not shown) then uses, for example, the following relationship with relative frame energy E rl (n) Normalized output y of the LogReg model pn (n) weighted:

[0356] scr UNCLR (n)=y pn (n)·E rl (n) (94)

[0357] Where E rl (n) is the relative frame energy described by relation (69).

[0358] The weighted normalized output of the LogReg model is called the "score," and it represents the same quantity as in the LRTD stereo mode described above. In the DFT stereo mode, when the VAD flag f is replaced... xVAD (n)(When relation (77) is set to 0, the score is scr UNCLR (n) is reset to 0. This is represented by the following relation:

[0359]

[0360] 7.2.2 Rising Edge Detection in DFT Stereo Mode

[0361] The UNCLR classifier 113's score calculator (not shown) ultimately uses the aforementioned rising edge detection mechanism in the UNCLR classification of the LRTD stereo mode, and utilizes an IIR filter to smooth the score in the DFT stereo mode. UNCLR (n). For this purpose, the UNCLR classifier 113 uses the following relation:

[0362] wscr UNCLR (n)=f edge (n)·wscr UNCLR (n-1)+(1-f edge (n))·scr UNCLR (n) (96)

[0363] This is the same as relation (91).

[0364] 7.3 Binary UNCLR Decision

[0365] The final output of UNCLR classifier 111 / 113 is a binary state. Let c UNCLR (n) represents the binary state of UNCLR classifier 111 / 113. Binary state c UNCLR (n) has a value of "1" to indicate an unrelated stereo content category, or a value of "0" to indicate a related stereo content category. The binary state at the output of UNCLR classifier 111 / 113 is variable. It is initialized to "0". The state of UNCLR classifier 111 / 113 changes from the current category to another category in frames that meet specific conditions.

[0366] The mechanism used in UNCLR classifiers 111 / 113 for switching between stereo content categories is... Figure 6 It is described in the form of a state machine.

[0367] refer to Figure 6 :

[0368] -If (a) the binary state c of the previous frame UNCLR (n–1) is “1” (601), (b) the smoothed score wscr of the current frame UNCLR (n) is less than -0.07 (602), and (c) the variable cnt from the previous frame. sw If (n–1) is greater than “0” (603), then the binary state c of the current frame is... UNCLR (n) is switched to "0" (604);

[0369] -If (a) the binary state c of the previous frame UNCLR(n–1) is “1” (601), and (b) the smoothed score wscr of the current frame UNCLR If (n) is not less than -0.07 (602), then there is no binary state c in the current frame. UNCLR Switching of (n);

[0370] -If (a) the binary state c of the previous frame UNCLR (n–1) is “1” (601), (b) the smoothed score wscr of the current frame UNCLR (n) is less than -0.07 (602), and (c) the variable cnt from the previous frame. sw If (n–1) is not greater than “0” (603), then there is no binary state c in the current frame. UNCLR Switching of (n).

[0371] In the same way, refer to Figure 6 :

[0372] -If (a) the binary state c of the previous frame UNCLR (n–1) is “0” (601), (b) the smoothed score wscr of the current frame UNCLR (n) is greater than “0.1” (605), and (c) the variable cnt from the previous frame. sw If (n–1) is greater than “0” (606), then the binary state c of the current frame is... UNCLR (n) is switched to "1" (607);

[0373] -If (a) the binary state c of the previous frame UNCLR (n–1) is “0” (601), and (b) the smoothed score wscr of the current frame UNCLR If (n) is not greater than "0.1" (605), then there is no binary state c in the current frame. UNCLR Switching of (n);

[0374] -If (a) the binary state c of the previous frame UNCLR (n–1) is “0” (601), (b) the smoothed score wscr of the current frame UNCLR (n) is greater than “0.1” (605), and (c) the variable cnt from the previous frame. sw If (n–1) is not greater than “0” (606), then there is no binary state c in the current frame. UNCLR Switching of (n).

[0375] Finally, the variable cnt in the current frame sw (n) is updated (608), and the process is repeated for the next frame (609).

[0376] variable cntsw (n) is the count value of the frames in the UNCLR classifier 111 / 113 that can switch between LRTD and DFT stereo modes. This count value is initialized to zero and is updated in each frame using, for example, the following logic (608):

[0377]

[0378] Count value cnt sw The upper limit of (n) is 100. Variable c type This indicates the type of the current frame in the device 100 used for encoding stereo audio signals. The frame type is typically determined during preprocessing operations of the device 100 (stereo codec) for encoding stereo audio signals, specifically in preprocessors 103 / 104 / 109. The type of the current frame is typically selected based on the following characteristics of the input stereo audio signal 190:

[0379] -Pitch duration

[0380] -voicing

[0381] - Spectrum tilt

[0382] -Zero crossing rate

[0383] -Frame energy difference (short-term, long-term)

[0384] As a non-limiting example, the frame type from the 3GPP EVS codec described in reference [1] can be used as a parameter c of relation (97) in the UNCLR classifier 111 / 113. type The frame types in the 3GPP EVS codec are selected from the following set of categories:

[0385] c type ∈(INACTIVE,UNVOICED,VOICED,GENERIC,TRANSITION,AUDIO)

[0386] The parameter VAD0 in relation (97) is the VAD flag without any trailing. The VAD flag without trailing is typically calculated during the preprocessing operation of the device 100 (stereo codec) used to encode the stereo audio signal, specifically in the TD preprocessors 103 / 104 / 109. As a non-limiting example, the VAD flag without trailing from the 3GPP EVS codec described in reference [1] can be used as parameter VAD0 in the UNCLR classifiers 111 / 113.

[0387] If the current frame type is general (“GENERIC”), unvoiced (“UNVOICED”), or inactive (“INACTIVE”), or if there is no VAD flag added to indicate inactivity in the input stereo audio signal (VAD0 = 0), then the output binary state c of the UNCLR classifier 111 / 113 is... UNCLR (n) can be changed. Such frames are generally suitable for switching between LRTD and DFT stereo modes because they are located in stable segments or segments that have a low perceptual impact on quality. The goal is to minimize the risk of switching artifacts.

[0388] 8. Crosstalk (XTALK) detection

[0389] XTALK detection is based on the LogReg model, which is trained separately for LRTD stereo mode and DFT stereo mode. Both statistical models are trained on features collected from a large database of real stereo recordings and manually prepared stereo samples. In the training database, each frame is labeled as either a monotone or crosstalk. Labeling is performed manually in the case of real stereo recordings, or semi-automatically in the case of manually prepared samples. Manual labeling is done by identifying short, compact segments with crosstalk characteristics. Semi-automatic labeling is performed using the VAD output from the mono signal before mixing the mono signal into a stereo sound signal. Details are provided at the end of Section 8.

[0390] In a non-limiting example of the implementation described in this disclosure, the real stereo recordings are sampled at 32 kHz. The total size of these real stereo recordings is approximately 263 MB, corresponding to approximately 30 minutes. The artificially prepared stereo samples are created by mixing randomly selected speakers from a mono clean speech database using the ITU-T G.191 reverberation tool. The artificially prepared stereo samples are obtained by simulating... Figure 7 The AB microphone setup shown is prepared for the conditions in a large conference room. Figure 7 It is a schematic floor plan of a large conference room with AB microphone setup, the conditions of which were simulated for XTALK testing.

[0391] Consider two types of rooms: echoing (LEAB) and unaechoing (LAAB). (Reference) Figure 7 For each type of room, the first speaker S1 can appear at positions P4, P5, or P6, and the second speaker S2 can appear at positions P10, P11, and P12. The positions of each speaker S1 and S2 are randomly selected during the preparation of training samples. Therefore, speaker S1 is always close to the first analog microphone M1, while speaker S2 is always close to the second analog microphone M2. Microphones M1 and M2 are located in... Figure 7 The non-limiting implementation shown is omnidirectional. Microphone pairs M1 and M2 constitute an analog AB microphone setup. Before further processing, mono samples are randomly selected from the training database, downsampled to 32 kHz, and normalized to -26 dBov (dB (overload) – the amplitude of the audio signal compared to the maximum value that the device can handle before clipping occurs). The ITU-T G.191 reverberation tool contains a database of real measurements of the room impulse response (RIR) for each speaker / microphone pair.

[0392] Randomly selected mono samples for speakers S1 and S2 are then convolved with the room impulse response (RIR) corresponding to a given speaker / microphone location, simulating real-world AB microphone capture. Contributions from both speakers S1 and S2 in each microphone M1 and M2 are summed. Before convolution, a randomly selected offset within the 4–4.5 second range is added to one of the speaker samples. This ensures that there is always a period of monophonic speech in all training utterances, followed by a short period of crosstalk and another period of monophonic speech. After RIR convolution and mixing, the samples are normalized again to -26 dBov, this time applied to the passive mono mixing.

[0393] The tokens are created semi-automatically using a standard VAD algorithm (e.g., the VAD algorithm for the 3GPP EVS codec described in reference [1]). The VAD algorithm is applied separately to the first speaker (S1) file and the second speaker (S2) file. The binary VAD decisions are then combined using a logical AND. This produces the token file. The segments whose combined output equals "1" determine the crosstalk segments. This is in Figure 8 It is shown in the middle, Figure 8 The diagram illustrates the automatic tagging of crosstalk samples using VAD. Figure 8 In the diagram, the first row shows the speech sample from speaker S1, the second row shows the binary VAD decision for the speech sample from speaker S1, the third row shows the speech sample from speaker S2, the fourth row shows the binary VAD decision for the speech sample from speaker S2, and the fifth row shows the location of the crosstalk segment.

[0394] The training set is imbalanced. The ratio of crosstalk frames to monotone frames is approximately 1 to 5, meaning that only about 21% of the training data belongs to the crosstalk category. This is compensated for during the LogReg training process by applying class weights as described in reference [6] (the entire contents of which are incorporated herein by reference).

[0395] The training samples are concatenated and used as input to device 100 (stereo codec) for encoding stereo audio signals. Features are collected individually for each 20ms frame during the encoding process, in separate files. This constitutes the training feature set. Let the total number of frames in the training feature set be denoted as, for example:

[0396] N T =N XTALK +N NORMAL (98)

[0397] Where N XTALK It is the total number of crosstalk frames, while N NORMAL It represents the total number of single-tone frames.

[0398] Furthermore, let the corresponding binary label be denoted as, for example:

[0399]

[0400] Where Ω XTALK It is a superset of all crosstalk frames, and Ω NORMAL It is a superset of all monotone frames. Inactive frames (VAD=0) are removed from the training database.

[0401] XTALK detection in 8.1LRTD stereo mode

[0402] In LRTD stereo mode, the method 150 for encoding stereo audio signals includes an operation 160 for detecting crosstalk (XTALK). To perform operation 160, the device 100 for encoding stereo audio signals includes an XTALK detector 110.

[0403] The operation 160 for detecting crosstalk (XTALK) in LRTD stereo mode is performed similarly to the UNCLR classification in LRTD stereo mode as described above. The XTALK detector 110 is based on a logistic regression (LogReg) model. For simplicity, parameter names and associated mathematical notation from the UNCLR classification are also used in this section. When referring to the same parameter name from different parts, subscripts are added to the notation to avoid ambiguity.

[0404] The following features are used by the XTALK detector 110:

[0405] -L / R category difference, d class (Relationship(32));

[0406] - Maximum autocorrelation L / R difference, d v (Relationship(25));

[0407] -L / R difference of the sum of LSF, dLSF (Relationship(23));

[0408] -L / R difference of residual energy, d LPC13 (Relationship(22));

[0409] -L / R difference in the correlation plot, d cmap (Relationship(27));

[0410] - The L / R difference in noise characteristics, d nchar (Relationship(29));

[0411] - Non-stationary L / R difference, d sta (Relationship(26));

[0412] - L / R difference of spectral diversity, d sdiv (Relationship(28));

[0413] - The unnormalized value of the inter-channel correlation function at hysteresis 0, p LR (Relationship(14));

[0414] - Side-to-mono energy ratio, r SM (Relationship(15));

[0415] -The difference between the maximum and minimum values ​​of the dot products between the left channel and the mono channel, and between the right channel and the mono signal, d mmLR (Relationship(19));

[0416] - The zero hysteresis value of the cross-channel correlation function, R0(relation(5));

[0417] - Evolution of the cross-correlation function between vocal tracts, RR (relationship (21));

[0418] -Location of the maximum value of the inter-channel cross-correlation function, k max (Relationship(11));

[0419] - The maximum value of the inter-channel correlation function, R max (Relationship(10));

[0420] The difference between the dot products of -L / M and R / M, Δ LRM (relation(20)); and

[0421] - Smoothed energy ratio of side signal and mono signal (Relationship(16)).

[0422] Accordingly, the XTALK detector 110 uses a total of F = 17 features.

[0423] Prior to the training process, the XTALK detector 110 includes a normalizer (not shown), which performs normalization on the 17 features f by removing the mean of the features and scaling them to unit variance. i The set is normalized using a sub-operation (not shown). The normalizer (not shown) uses, for example, the following relation:

[0424]

[0425] Among them, f i,raw Indicate the i-th feature of the set, Let σ represent the global mean of the i-th feature across the training database. fi It is the global variance of the i-th feature across the training database. Here, the parameter used in relation (100) and σ fi Unlike the same parameters used in relation (81).

[0426] The output y of the LogReg model p The probability p (category = 0) of the current frame belonging to the crosstalk segment category (category 0) is given by relation (83). Details of the training process and the process of finding the optimal decision threshold are provided above in the description of UNCLR classification in LRTD stereo mode. As mentioned above, for this purpose, the XTALK detector 110 includes a score calculator (not shown) that performs a suboperation (not shown) to calculate a score representing the score of the unrelated stereo content in the input stereo sound signal 190.

[0427] The XTALK detector 110's score calculator (not shown) utilizes, for example... Figure 9 The function shown, and further processed, outputs the raw y of the LogReg model. p Normalize. Figure 9 This is a graph representing the function used to scale the raw output of the LogReg model in XTALK detection in LRTD stereo mode. This type of normalization can be mathematically described as follows:

[0428]

[0429] If the previous frame was encoded in DFT stereo mode and the current frame is encoded in LRTD stereo mode, then the normalized output y of the LogReg model is... pn (n) is set to 0. This type of process prevents switching artifacts.

[0430] 8.1.1 LogReg Output Weighting Based on Relative Frame Energy

[0431] The score calculator (not shown) for the XTALK detector 110 is based on the relative frame energy E. rl (n) Normalized output y of the LogReg model pn (n) Weighting is performed. The weighting scheme applied in the XTALK detector 110 in LRTD stereo mode is similar to the weighting scheme applied in the UNCLR classifier 111 in LRTD stereo mode (as described above). The main difference lies in the relative frame energy E rl (n) is not directly used as a multiplication factor as in relation (85). Instead, the score calculator (not shown) of the XTALK detector 110 inversely proportionally represents the relative frame energy E. rl (n) linearly maps to the interval <0; 0.95>. This mapping can be accomplished, for example, using the following relation:

[0432] w relE (n) = -2.375E rl (n)+2.1375 (102)

[0433] Therefore, in frames with higher relative energy, the weights will be close to 0, while in frames with lower energy, the weights will be close to 0.95. The score calculator of the XTALK detector 110 (not shown) then uses, for example, the following relationship, with weight w relE (n) is used to normalize the output y of the LogReg model. pn (n) Perform filtering:

[0434] scr XTALK (n)=w relE scr XTALK (n-1)+(1-w relE )y pn (n) (103)

[0435] Here, index n indicates the current frame, and n-1 indicates the previous frame.

[0436] Normalized weighted output scr from XTALK detector 110 XTALK (n) is called the “XTALK score”, which represents the crosstalk in the input stereo audio signal 190.

[0437] 8.1.2 Rising Edge Detection

[0438] Similar to UNCLR classification in LRTD stereo mode, the score calculator (not shown) of the XTALK detector 110 smooths the normalized weighted output score of the LogReg model. XTALK(n). The reason is to smooth out occasional short-term "peaks" and "drops" that would otherwise lead to false alarms or errors. Smoothing is designed to preserve the rising edges of the LogReg output because these rising edges may represent important transitions between crosstalk and monotone segments in the input stereo audio signal 190. The mechanism for rising edge detection in the XTALK detector in LRTD stereo mode differs from the rising edge detection mechanism described above regarding the UNCLR classification in LRTD stereo mode.

[0439] In the XTALK detector 110, the rising edge detection algorithm analyzes the LogReg output value from the previous frame and compares it with a pre-computed "ideal" rising edge with a different slope. The "ideal" rising edge is represented as a linear function of the frame index n. Figure 10 This diagram illustrates the mechanism for detecting the rising edge in the XTALK detector 110 in LRTD stereo mode. (Reference) Figure 10 The x-axis contains the index n of the frames preceding the current frame (frame 0). The small gray rectangle represents the XTALK score (scr) over the six frames preceding the current frame. XTALK Example output of (n). From Figure 10 It can be seen that in the XTALK score, scr XTALK In (n), rising edges begin to exist starting from three frames prior to the current frame. Dashed lines represent the set of four “ideal” rising edges on segments of different lengths.

[0440] For each "ideal" rising edge, the rising edge detection algorithm calculates the dashed line and the XTALK score. XTALK The mean square error between (n) is the minimum mean square error among the tested "ideal" rising edges. The output of the rising edge detection algorithm is the minimum mean square error among the rising edges. The linear function represented by the dashed line is based on the minimum value of scr. min and maximum value scr max It is pre-calculated using a predefined threshold. This is in Figure 10 The image is represented by a large light gray rectangle. The slope of each "ideal" rising edge of the linear function depends on the minimum and maximum thresholds and the length of the segment.

[0441] Rising edge detection is performed by the XTALK detector 110 only in frames that meet the following criteria:

[0442]

[0443] Where K=4 is the maximum length of the rising edge being tested.

[0444] Let the output value of the rising edge detection algorithm be labeled as ε. 0_1The use of the subscript "0_1" emphasizes that the output value of the rising edge detection is restricted to the interval <0; 1>. For frames that do not satisfy the criteria in relation (104), the output value of the rising edge detection is directly set to 0, i.e.

[0445] ε 0_1 =0 (105)

[0446] The set of linear functions representing the "ideal" rising edge of the test can be mathematically expressed using the following relationship:

[0447]

[0448] Here, index l indicates the length of the tested rising edge, and n–k is the frame index. The slope of each linear function is determined by three parameters: the length l of the tested rising edge, the minimum threshold scr, and so on. min and the maximum threshold scr max For the purpose of the XTALK detector 110 in the LRTD stereo mode, the threshold is set to scr. max =1.0 and scr min = -0.2. These threshold values ​​were found experimentally.

[0449] For each tested rising edge length, the rising edge detection algorithm uses, for example, the following relationship to calculate a linear function t (relation (106)) and an XTALK score scr XTALK Mean square error between:

[0450]

[0451] Where ε0 is the initial error given by the following formula:

[0452] ε0=[scr XTALK (n)-scr max ] 2 (108)

[0453] The minimum mean square error is calculated by the XTALK detector 110 using the following formula:

[0454]

[0455] The lower the minimum mean square error, the stronger the detected rising edge. In the non-restrictive implementation, if the minimum mean square error is higher than 0.3, the rising edge detection output is set to 0, i.e.:

[0456]

[0457] And the rising edge detection algorithm exits. In all other cases, the minimum mean square error can be linearly mapped onto the interval <0; 1> using, for example, the following relationship:

[0458] ε 0_1 =1-2.5ε min (111)

[0459] In the example above, the output of rising edge detection is inversely proportional to the minimum mean square error.

[0460] The XTALK detector 110 normalizes the rising edge detection output in the range <0.5; 0.9> to produce an edge sharpness parameter calculated using, for example, the following relationship:

[0461] f edge (n) = 0.9 - 0.4ε 0_1 (112)

[0462] 0.5 and 0.9 are used as the lower limit and upper limit, respectively.

[0463] Finally, the score calculator (not shown) of the XTALK detector 110 smooths the normalized weighted output score of the LogReg model using the IIR filter of the XTALK detector 110. XTALK (n), where f edge (n) is used to replace the forgetting factor. Such smoothing uses, for example, the following relationship:

[0464] wscr XTALK (n)=f edge (n)·wscr XTALK (n-1)+(1-f edge (n))·scr XTALK (n) (113)

[0465] In frames where the alternative VAD flag calculated in relation (77) is zero, the smoothed output wscr XTALK (n)(XTALK score) is reset to 0. That is:

[0466]

[0467] 8.2 Crosstalk Detection in DFT Stereo Mode

[0468] In DFT stereo mode, the method 150 for encoding the stereo audio signal 190 includes an operation 162 for detecting crosstalk (XTALK). To perform operation 162, the device 100 for encoding the stereo audio signal 190 includes an XTALK detector 112.

[0469] XTALK detection in the DFT stereo mode is performed similarly to XTALK detection in the LRTD stereo mode. A logistic regression (LogReg) model is used for binary classification of the input feature vector. For simplicity, the names of the specific parameters from the LRTD stereo mode XTALK detection and their associated mathematical notation are also used in this section. Subscripts are added to avoid ambiguity when referring to the same parameters from both parts.

[0470] By running the DFT stereo pattern on both the monotone and crosstone training databases, the following features are extracted from the device 100 used to encode the stereo sound signal 190:

[0471] -ILD gain, g ILD (Relationship(43));

[0472] -IPD gain, g IPD (Relationship(48));

[0473] -IPD rotation angle, (Relationship(49));

[0474] -Predicted gain, g pred (Relationship(52));

[0475] -Mean energy of inter-channel coherence, E coh (Relationship(55));

[0476] - The ratio of the amplitude products within the maximum and minimum channels, r PP (Relationship(57));

[0477] - The overall amplitude of the cross-channel spectrum, f X (Relationship(41));

[0478] The maximum value of the GCC-PHAT function, G ITD (Relationship(61));

[0479] The relationship between the amplitudes of the first and second highest peaks of the -GCC-PHAT function, r GITD12 (Relationship(64));

[0480] -The amplitude of the second highest peak of GCC-PHAT, m ITD2 (relationship(66)); and

[0481] - The difference between the location of the second highest peak in the current frame and the location of the second highest peak in the previous frame, Δ ITD2 (Relationship(67)).

[0482] The XTALK detector 112 uses a total of F = 11 features.

[0483] Prior to the training process, the XTALK detector 112 includes a normalizer (not shown) that performs a suboperation (not shown) to normalize the set of extracted features by removing the global mean of the features and scaling them to unit variance using, for example, the following relation.

[0484]

[0485] Where f i,raw The i-th feature of the set, f i Indicates the i-th feature of the normalization, It is the global mean of the i-th feature across the training database, and σ fi It is the global variance of the i-th feature across the training database. The parameters used in relation (115) and σ fi Unlike the parameters used in relation (81).

[0486] The output of the LogReg model is entirely described by relation (82), and the probability that the current frame belongs to the crosstalk segment category (category 0) is given by relation (83). Details of the training process and the process of finding the optimal decision threshold are provided above in the section on UNCLR classification in LRTD stereo mode. Again, for this purpose, the XTALK detector 112 includes a score calculator (not shown) that performs suboperations (not shown) that calculate the score of the XTALK detection in the input stereo sound signal 190.

[0487] The score calculator (not shown) for the XTALK detector 112 is used. Figure 5 The function shown, and further processed, outputs the raw y of the LogReg model. p Normalization is performed. The normalized output of the LogReg model is denoted as y. pn In DFT stereo mode, weighting based on relative frame energy is not used. Therefore, the normalized weighted output of the LogReg model (specifically, the XTALK score) is... XTALK (n) is given by the following formula:

[0488] scr XTALK (n)=y pn (116)

[0489] When replacing the VAD flag f xVAD When (n) is set to 0, the XTALK score is... XTALK (n) is reset to 0. This can be expressed as follows:

[0490]

[0491] 8.2.1 Rising Edge Detection

[0492] As in the case of XTALK detection in LRTD stereo mode, the score calculator (not shown) of XTALK detector 112 smooths the XTALK score. XTALK (n) to remove short-term peaks. This smoothing is performed using a rising edge detection mechanism, as described regarding the XTALK detector 110 in LRTD stereo mode, via IIR filtering. XTALK score scr XTALK (n) is smoothed using an IIR filter with, for example, the following relationship:

[0493] wscr XTALK (n)=f edge (n)·wscr XTALK (n-1)+(1-f edge (n))·scr XTALK (n) (118)

[0494] Where f edge (n) is the edge sharpness parameter calculated in relation (112).

[0495] 8.3 Binary XTALK Decision Making

[0496] The final output of the XTALK detector 110 / 112 is binary. Let c XTALK (n) indicates the output of XTALK detector 110 / 112, where "1" represents the crosstalk category and "0" represents the monotone category. Output c XTALK (n) can also be viewed as a state variable. It is initialized to 0. The state variable changes from the current class to another class only in frames that meet specific conditions. The mechanism for crosstalk category switching is similar to the mechanism for category switching of unrelated stereo content, which has been described in detail in Section 7.3 above. However, differences exist for both LRTD stereo mode and DFT stereo mode. These differences will be discussed below.

[0497] In LRTD stereo mode, the XTALK detector 110 uses, for example... Figure 11 The crosstalk switching mechanism is shown. (Reference) Figure 11 :

[0498] -If the output c of UNCLR classifier 111 in the current frame n UNCLR If (n) equals "1" (1101), then the output c of XTALK detector 110 does not exist in the current frame n. XTALK Switching of (n).

[0499] -If (a) the output c of UNCLR classifier 111 in the current frame n UNCLR (n) equals “0” (1101), and (b) the output of XTALK detector 110 in the previous frame n–1. XTALK If (n–1) equals “1” (1102), then the output c of XTALK detector 110 does not exist in the current frame n. XTALK Switching of (n).

[0500] -If (a) the output c of UNCLR classifier 111 in the current frame n UNCLR (n) equals "0" (1101), (b) the output of XTALK detector 110 in the previous frame n–1 XTALK (n–1) equals “0” (1102), and (c) the smoothed XTALK score wscr in the current frame n. XTALK If (n) is not greater than 0.03 (1104), then the output c of XTALK detector 110 does not exist in the current frame n. XTALK Switching of (n).

[0501] -If (a) the output c of UNCLR classifier 111 in the current frame n UNCLR (n) equals "0" (1101), (b) the output of XTALK detector 110 in the previous frame n-1. XTALK (n–1) equals “0” (1102), (c) smoothed XTALK score wscr in the current frame n XTALK (n) is greater than 0.03 (1104), and (d) is the count value cnt in the previous frame n–1. sw If (n–1) is not greater than “0” (1105), then the output c of XTALK detector 110 does not exist in the current frame n. XTALK Switching of (n).

[0502] -If (a) the output c of UNCLR classifier 111 in the current frame n UNCLR (n) equals "0" (1101), (b) the output of XTALK detector 110 in the previous frame n–1 XTALK (n–1) equals “0” (1102), (c) smoothed XTALK score wscr in the current frame n XTALK (n) is greater than 0.03 (1104), and (d) is the count value cnt in the previous frame n–1. sw If (n–1) is greater than “0” (1105), then the output c of XTALK detector 110 in the current frame n is... XTALK (n) is switched to "1" (1106).

[0503] Finally, the count value cnt in the current frame n sw (n) is updated (1107), and the process is repeated for the next frame (1108).

[0504] Count value cnt sw (n) is common to both the UNCLR classifier 111 and the XTALK detector 110, and is defined in relation (97). The count value cnt sw A positive value of (n) indicates the state variable c XTALK (n)(output of XTALK detector 110 c) XTALK Switching between (n) is allowed. For example... Figure 11 As can be seen, the switching logic uses the output c of the UNCLR classifier 111 in the current frame. UNCLR (n)(1101). Therefore, it is assumed that the UNCLR classifier 111 runs before the XTALK detector 110 because it uses its output. Furthermore, Figure 11 The state switching logic is unidirectional, meaning that the output c of XTALK detector 110... XTALK (n) can only change from "0" (single tone) to "1" (crosstone). The state switching logic for the opposite direction (i.e., from "1" (crosstone) to "0" (single tone) is part of the DFT / LRTD stereo mode switching logic, which will be described later in this disclosure.

[0505] In DFT stereo mode, the XTALK detector 112 includes an auxiliary parameter calculator (not shown), which performs sub-operations (not shown) to calculate the following auxiliary parameters. Specifically, the crosstalk switching mechanism uses the output wscr of the XTALK detector 112. XTALK (n), and the following auxiliary parameters:

[0506] -Voice Activity Detection (VAD) flag in the current frame (f VAD );

[0507] The amplitudes of the first and second highest peaks of the GCC-PHAT function, G ITD ,m ITD2 (Relations (61) and (66) respectively);

[0508] - Corresponding to the location of the first and second highest peaks of the GCC-PHAT function (ITD values), d ITD ,d ITD2 (respectively, relation (60) and (paragraph

[00111] )); and

[0509] -DFT stereo mute indicator, f sil Relationship (78).

[0510] In DFT stereo mode, the XTALK detector 112 uses, for example... Figure 12 The crosstalk switching mechanism is shown. (Reference) Figure 12 :

[0511] -If d ITD If (n) equals "0" (1201), then c XTALK (n) is switched to "0" (1217);

[0512] -If (a)d ITD (n) is not equal to "0" (1201), and (b)c XTALK (n-1) is not equal to "0" (1202).

[0513] ■If (c)c XTALK If (n-1) is not equal to "1" (1215), then c does not exist. XTALK Switching of (n);

[0514] ■If (c)c XTALK (n-1) equals "1" (1215), and (d)wscr XTALK If (n) is not less than "0.0" (1216), then c does not exist. XTALK Switching of (n);

[0515] ■If (c)c XTALK (n-1) equals "1" (1215), and (d)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219);

[0516] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), and (c)f VAD Not equal to "1" (1203)

[0517] ■If (d)c XTALK If (n-1) is not equal to "1" (1215), then c does not exist. XTALK Switching of (n);

[0518] ■If (d)c XTALK (n-1) equals "1" (1215), and (e)wscr XTALK If (n) is not less than "0.0" (1216), then c does not exist. XTALK Switching of (n);

[0519] ■If (d)c XTALK(n-1) equals "1" (1215), and (e)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219);

[0520] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD Equals "1" (1203), (d)0.8G ITD (n) is less than m ITD2 (n)(1204), (e)0.8G ITD (n-1) is less than m ITD2 (n-1)(1205), (f)d ITD2 (n)-d ITD2 (n-1) is less than "4.0" (1206), (g)G ITD (n) is greater than "0.15" (1207), (h)G ITD If (n-1) is greater than "0.15" (1208), then c XTALK (n) is switched to "1" (1218);

[0521] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD The result is equal to "1" (1203), and (d) tests any of 1204 to 1208 and finds them to be negative.

[0522] ■If (e)wscr XTALK If (n) is greater than "0.8" (1209), then c XTALK (n) is switched to "1" (1218);

[0523] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD (d) Test 1204 through 1208 and it is negative; (e) wscr XTALK (n) is not greater than “0.8” (1209), and (f)f sil (n) is not equal to "1" (1210).

[0524] ■If (g)c XTALKIf (n-1) is not equal to "1" (1215), then c does not exist. XTALK Switching of (n);

[0525] ■If (g)c XTALK (n-1) equals "1" (1215), and (h)wscr XTALK If (n) is not less than "0.0" (1216), then c does not exist. XTALK Switching of (n);

[0526] ■If (g)c XTALK (n-1) equals "1" (1215), and (h)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219);

[0527] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD (d) Test 1204 through 1208 and it is negative; (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil (n) equals "1" (1210), (g)d ITD (n) is greater than “8.0” (1211), and (h)d ITD If (n-1) is less than -8.0, then c XTALK (n) is switched to "1" (1218);

[0528] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD (d) Test 1204 through 1208 and it is negative; (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil (n) equals "1" (1210), (g) tests if either 1211 or 1212 is negative, (h) d ITD (n-1) is greater than “8.0” (1213), and (i)d ITD If (n) is less than -8.0 (1214), then c XTALK (n) is switched to "1" (1218);

[0529] -If (a)d ITD (n) is not equal to "0" (1201), (b)c XTALK (n-1) equals “0” (1202), (c)f VAD (d) Test 1204 through 1208 and it is negative; (e) wscr XTALK (n) is not greater than "0.8" (1209), (f)f sil (n) equals "1" (1210), (g) tests if either 1211 or 1212 is negative, and (h) tests if either 1213 or 1214 is negative.

[0530] ■If (i)c XTALK If (n-1) is not equal to "1" (1215), then c does not exist. XTALK Switching of (n);

[0531] ■If (i)c XTALK (n-1) equals "1" (1215), and (j)wscr XTALK If (n) is not less than "0.0" (1216), then c does not exist. XTALK Switching of (n);

[0532] ■If (i)c XTALK (n-1) equals "1" (1215), and (j)wscr XTALK If (n) is less than "0.0" (1216), then c XTALK (n) is switched to "0" (1219);

[0533] Finally, the count value cnt in the current frame n sw (n) is updated (1220), and the process is repeated for the next frame (1221).

[0534] variable cnt sw (n) is the count of frames that can switch between LRTD and DFT stereo modes. This count value is cnt. sw (n) is common to both the UNCLR classifier 113 and the XTALK detector 112. The count value is cnt. sw (n) is initialized to zero and is updated in each frame according to relation (97).

[0535] 9. DFT / LRTD Stereo Mode Selection

[0536] The method 150 for encoding stereo audio signal 190 includes an operation 164 of selecting an LRTD or DFT stereo mode. To perform operation 164, the apparatus 100 for encoding stereo audio signal 190 includes an LRTD / DFT stereo mode selector 114 that receives an XTALK decision delayed by one frame (191) from an XTALK detector 110, an UNCLR decision 111 from an UNCLR classifier, an XTALK decision from an XTALK detector 112, and an UNCLR decision from an UNCLR classifier 113.

[0537] The LRTD / DFT stereo mode selector 114 is based on the binary output c of the UNCLR classifier 111 / 113. UNCLR (n) and the binary output c of XTALK detector 110 / 112 XTALK (n) is used to select either LRTD or DFT stereo mode. The LRTD / DFT stereo mode selector 114 also considers certain auxiliary parameters. These parameters are mainly used to prevent stereo mode switching in perceptually sensitive segments, or to prevent frequent switching in segments in which neither the UNCLR classifiers 111 / 113 nor the XTALK detectors 110 / 112 provide accurate output.

[0538] The operation 164, which selects the LRTD or DFT stereo mode, is performed before the input stereo audio signal 190 is downmixed and encoded. Therefore, as... Figure 1 As shown in 191, operation 164 uses the outputs of the UNCLR classifier 111 / 113 and the XTALK detector 110 / 112 from the previous frame. Operation 164 selects either LRTD or DFT stereo mode in Figure 13 The schematic block diagram is further described.

[0539] As will be described below, the DFT / LRTD stereo mode selection mechanism used in operation 164 includes the following sub-operations:

[0540] - Initial DFT / LRTD stereo mode selection; and

[0541] - Switching from LRTD to DFT stereo mode when crosstalk content is detected.

[0542] 9.1 Initial DFT / LRTD Stereo Mode Selection

[0543] DFT stereo mode is a preferred mode for encoding monotone speech with high interchannel correlation between the left (L) channel and the right (R) channel of the input stereo sound signal 190.

[0544] The LRTD / DFT stereo mode selector 114 begins the initial selection of a stereo mode by determining whether a previously processed frame "might be a speech frame". This can be done, for example, by checking the log-likelihood ratio between the "speech" category and the "music" category. The log-likelihood ratio is defined as the absolute difference between the log-likelihood of the input stereo sound signal frame generated from the "music" source and the log-likelihood of the input stereo sound signal frame generated from the "speech" source. The following relationship can be used to calculate the log-likelihood ratio:

[0545] dL SM (n)=L M (n)-L S (n)(119)

[0546] Where L S (n) is the log-likelihood of the "speech" category, while L M (n) is the log-likelihood of the "music" category.

[0547] For example, the Gaussian mixture model (GMM) of the 3GPP EVS codec described in reference [7] (the entire contents of which are incorporated herein by reference) can be used to estimate the log-likelihood L of the “speech” category. S (n), and the log-likelihood L of the "music" category. M (n). Other methods for speech / music classification can also be used to calculate the log-likelihood ratio (difference score) dL. SM (n).

[0548] Smoothing the log-likelihood ratio dL can be achieved by using, for example, two IIR filters with different forgetting factors that have the following relationship. SM (n):

[0549]

[0550] Accordingly, superscript (1) indicates the first IIR filter, while superscript (2) indicates the second IIR filter.

[0551] Then the smoothed value and The new binary flag f is compared with a predetermined threshold, and if, for example, the following combination of conditions is met, then... SM (n) is set to 1:

[0552]

[0553] Mark f SM (n) = 1 is an indicator that the previous frame may be a speech frame. The threshold of 1.0 was found experimentally.

[0554] If the binary output c of the UNCLR classifier 111 / 113 in the previous frame n-1 UNCLR (n-1) or the binary output c of the XTALK detector 110 / 112 XTALK (n-1) is set to 1, and if the previous frame was a speech frame, the initial DFT / LRTD stereo mode selection mechanism will set the new binary flag f. UX (n) is set to 1. This can be expressed by the following relation:

[0555]

[0556] Let M SMODE (n)∈(LRTD,DFT) is a discrete variable representing the stereo mode selected in the current frame n. The stereo mode is initialized in each frame using values ​​from the previous frame n-1, i.e.:

[0557] M SMODE (n)=M SMODE (n-1) (123)

[0558] If the flag f UX When (n) is set to 1, the LRTD stereo mode is selected for encoding in the current frame n. This can be expressed as follows:

[0559]

[0560] If flag f is in the current frame n UX If (n) is set to 0, and the stereo mode in the previous frame n-1 is LRTD stereo mode, then the auxiliary stereo mode switching flag f from the LRTD energy analysis processor 1301 of the LRTD / DFT stereo mode selector 114 will be set to 0. TDM (n-1) (described below) is analyzed to select the stereo mode in the current frame n using, for example, the following relationship:

[0561]

[0562] Update the auxiliary stereo mode switching flag f only in each frame in LRTD mode. TDM (n). Parameter f TDM The update of (n) is described below.

[0563] like Figure 13 As shown, the LRTD / DFT stereo mode selector 114 includes an LRTD energy analysis processor 1301 to generate auxiliary parameters f, which will be described in more detail later in this disclosure. TDM (n), c LRTD (n), c DFT(n) and m TD (n).

[0564] If flag f is in the current frame n UX If (n) is set to 0 and the stereo mode in the previous frame n-1 was DFT stereo mode, then no stereo mode switching is performed, and DFT stereo mode is also selected in the current frame n.

[0565] 9.2XTALK test: LRTD to DFT stereo mode switching

[0566] The XTALK detector 110 in LRTD mode has been described in the preceding description. From Figure 11 It can be seen that the binary output c of the XTALK detector 110 XTALK (n) can only be set to 1 when crosstalk content is detected in the current frame. As a result, the initial stereo mode selection logic described above cannot select DFT stereo mode when XTALK detector 110 indicates monotone content. In cases where a crosstalk stereo signal segment is followed by a monotone stereo signal segment, this may lead to an undesirable extension of the LRTD stereo mode. Therefore, an additional mechanism has been implemented for switching back from LRTD stereo mode to DFT stereo mode when monotone content is detected. This mechanism is described in the following description.

[0567] If the LRTD / DFT stereo mode selector 114 selects the LRTD stereo mode in the previous frame n-1 and the initial stereo mode selection selects the LRTD mode in the current frame n, and if the binary output c of the XTALK detector 110 is simultaneously... XTALK If (n-1) is 1, then the stereo mode can be changed from LRTD to DFT stereo mode. The latter change is allowed, for example, when the following conditions are met:

[0568]

[0569] The set of conditions defined above includes references to the `clas` and `brate` parameters. The `brate` parameter is a high-level constant that contains the total bit rate used by device 100 (stereo codec) for encoding the stereo audio signal. It is set during the initialization of the stereo codec and remains unchanged during the encoding process.

[0570] The clas parameter is a discrete variable containing information about the frame type. Estimation of the clas parameter is typically part of the signal preprocessing for a stereo codec. As a non-limiting example, the clas parameter from the Frame Erasure Hiding (FEC) module of the 3GPP EVS codec, as described in reference [1], can be used in the DFT / LRTD stereo mode selection mechanism. The clas parameter from the FEC module of the 3GPP EVS codec is selected considering frame erasure hiding and decoder recovery strategies. The clas parameter is selected from the following predefined set of categories:

[0571]

[0572] DFT / LRTD stereo mode selection mechanisms implemented using other means of frame type classification are within the scope of this disclosure.

[0573] In the condition set (126) defined above, the condition

[0574]

[0575] The clas parameter refers to the parameter calculated during the preprocessing of the downmixed mono (M) channel when the device 100 used to encode stereo audio signals is running in DFT stereo mode.

[0576] When the device 100 used for encoding stereo audio signals is in LRTD stereo mode, the condition should be replaced as follows:

[0577]

[0578] The indices “L” and “R” refer to the clas parameters calculated in the preprocessing modules of the left (L) channel and right (R) channel, respectively.

[0579] Parameter c LRTD (n) and c DFT (n) are the count values ​​for the LRTD and DFT frames, respectively. These count values ​​are updated in each frame as part of the LRTD energy analysis processor 1301. The two count values ​​c LRTD (n) and c DFT The update of (n) is described in detail in the next section.

[0580] 9.3 Auxiliary parameters calculated in the LRTD energy analysis module

[0581] When the device 100 used to encode stereo audio signals is running in LRTD stereo mode, the LRTD / DFT stereo mode selector 114 calculates or updates several auxiliary parameters to improve the stability of the DFT / LRTD stereo mode selection mechanism.

[0582] For certain special frame types, LRTD stereo mode operates in a so-called "TD sub-mode." The TD sub-mode is typically used during a brief transition period before switching from LRTD stereo mode to DFT stereo mode. Whether LRTD stereo mode will operate in the TD sub-mode is determined by the binary sub-mode flag m. TD (n) Indication. Binary symbol m TD (n) is one of the auxiliary parameters, and it can be initialized in each frame as follows:

[0583] m TD (n)=f TDM (n-1) (127)

[0584] Where f TDM (n) refers to the auxiliary switching flag described later in this section.

[0585] In which f UX In frames where (n) = 1, the binary sub-pattern flag m TD (n) is reset to 0 or 1. Used to reset m. TD The conditions for (n) are defined, for example, as follows:

[0586]

[0587] If f UX If (n) = 0, then the binary sub-pattern flag m TD (n) remains unchanged.

[0588] The LRTD energy analysis processor 1301 includes the two count values ​​c mentioned above. LRTD (n) and c DFT (n). Count value c LRTD (n) is one of the auxiliary parameters and counts the number of consecutive LRTD frames. This count value is set to 0 in each frame selected in the device 100 for encoding the stereo sound signal in DFT stereo mode, and incremented by 1 in each frame selected in LRTD stereo mode. This can be expressed as follows:

[0589]

[0590] In essence, the count value c LRTD (n) contains the number of frames since the last DFT->LRTD switch point. The count value is c. LRTD(n) is limited by a threshold of 100. The count value c DFT (n) Counts the number of consecutive DFT frames. The count value is c. DFT (n) is one of the auxiliary parameters, and it is set to 0 in each frame selected in the LRTD stereo mode in the device 100 for encoding stereo sound signals, and incremented by 1 in each frame selected in the DFT stereo mode. This can be expressed as follows:

[0591]

[0592] In essence, the count value c DFT (n) contains the number of frames since the last LRTD->DFT switch point. The count value is c. DFT (n) is limited by the threshold of 100.

[0593] The final auxiliary parameter calculated in the LRTD energy analysis processor 1301 is the auxiliary stereo mode switching flag f. TDM (n). This parameter utilizes the binary flag f in each frame. UX Initialize with (n), as shown below:

[0594] f TDM (n)=f UX (n) (131)

[0595] When the left (L) channel and right (R) channel of the input stereo audio signal 190 are out of phase (OOP), the auxiliary stereo mode switching flag f TDM (n) is set to 0. An exemplary method for OOP detection can be found, for example, in reference [8] (the entire contents of which are incorporated herein by reference). When an OOP situation is detected, the binary flag s2m is set to 1 in the current frame n, otherwise it is set to zero. When the binary flag s2m is set to 1, the auxiliary stereo mode switching flag f in the LRTD stereo mode is set to 1. TDM (n) is set to 0. This can be expressed using relation (132):

[0596]

[0597] If the binary flag s2m(n) is set to zero, then the auxiliary switching flag f TDM (n) can be reset to zero, for example, based on the following set of conditions:

[0598] (133)

[0600] Of course, the DFT / LRTD stereo mode switching mechanism can be implemented using other methods for OOP detection.

[0601] Auxiliary stereo mode switching indicator f TDM (n) can also be reset to 0 based on the following set of conditions:

[0602]

[0603] In the two sets of conditions defined above, the conditions

[0604] clas(n-1) = UNVOICED_CLAS

[0605] The clas parameter refers to the parameter calculated during the preprocessing of the downmixed mono (M) channel when the device 100 used to encode stereo audio signals is running in DFT stereo mode.

[0606] When the device 100 used for encoding stereo audio signals is in LRTD stereo mode, the condition should be replaced as follows:

[0607]

[0608] The indices “L” and “R” refer to the clas parameters calculated during the preprocessing of the left (L) channel and the right (R) channel, respectively.

[0609] 10. Core Encoder

[0610] Method 150 for encoding stereo audio signals includes operations 115 of core encoding the left channel (L) of stereo audio signal 190 in LRTD stereo mode, operations 116 of core encoding the right channel (R) of stereo audio signal 190 in LRTD stereo mode, and operations 117 of core encoding the lower mixed mono channel (M) of stereo audio signal 190 in DFT stereo mode.

[0611] To perform operation 115, the device 100 for encoding the stereo audio signal includes a core encoder 115, such as a mono core encoder. To perform operation 116, the device 100 includes a core encoder 116, such as a mono core encoder. Finally, to perform operation 167, the device 100 for encoding the stereo audio signal includes a core encoder 117, which is capable of operating in DFT stereo mode to encode the downmixed mono (M) channel of the stereo audio signal 190.

[0612] The selection of suitable core encoders 115, 116, and 117 is considered to be within the knowledge of those skilled in the art. Accordingly, these encoders will not be further described in this disclosure.

[0613] 11. Hardware Implementation

[0614] Figure 14 This is a simplified block diagram of an example configuration of the hardware components that form the device 100 and method 150 for encoding stereo sound signals described above.

[0615] The device 100 for encoding stereo audio signals can be implemented as part of a mobile terminal, a portable media player, or in any similar device. Device 100 (in...) Figure 14 The component (identified as 1400) includes input 1402, output 1404, processor 1406, and memory 1408.

[0616] Input 1402 is configured to receive digital or analog data. Figure 1 The input stereo audio signal is 190. The output 1404 is configured to supply an encoded stereo audio signal. Input 1402 and output 1404 can be implemented in a common module, such as a serial input / output device.

[0617] Processor 1406 is operatively connected to input 1402, output 1404, and memory 1408. Processor 1406 is implemented as one or more processors for executing code instructions that support, for example... Figure 1 The functions of the various components of the device 100 for encoding stereo audio signals are shown.

[0618] Memory 1408 may include non-transitory memory for storing code instructions executable by processor 1406, specifically including / storing processor-readable memory for storing non-transitory instructions that, when executed, cause the processor to implement the operations and components of the method 150 and apparatus 100 for encoding stereo sound signals as described in this disclosure. Memory 1408 may also include random access memory or buffers for storing intermediate processing data from various functions performed by processor 1406.

[0619] Those skilled in the art will recognize that the description of the apparatus 100 and method 150 for encoding stereo sound signals is merely illustrative and not intended to be limiting in any way. Other embodiments will readily conceive of those skilled in the art upon receiving this disclosure. Furthermore, the disclosed apparatus 100 and method 150 for encoding stereo sound signals can be customized to provide valuable solutions to existing needs and problems related to sound encoding and decoding.

[0620] For clarity, not all conventional features of the implementation of the device 100 and method 150 for encoding stereo sound signals are shown or described. It will be understood, of course, that in the development of any such actual implementation of the device 100 and method 150 for encoding stereo sound signals, many implementation-specific decisions may need to be made to achieve the developer’s specific goals, such as compliance with constraints related to the application, system, network, and business, and these specific goals will vary from implementation to implementation and from developer to developer. Furthermore, it should be understood that while the development work may be complex and time-consuming, it remains a routine engineering task for those skilled in the art of sound processing who will benefit from this disclosure.

[0621] According to this disclosure, the components / processors / modules, processing operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. Furthermore, those skilled in the art will recognize that devices with less general-purpose characteristics, such as hardwired devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., can also be used. Where a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and those operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, they can be stored on tangible and / or non-transitory media.

[0622] The apparatus 100 and method 150 for encoding stereo audio signals as described herein may use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0623] In the apparatus 100 and method 150 for encoding stereo audio signals as described herein, various operations and sub-operations may be performed in various orders, and some of the operations and sub-operations may be optional.

[0624] Although the present disclosure has been described above by way of non-limiting, illustrative embodiments, these embodiments may be modified in any way within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0625] 12. References

[0626] This disclosure references the following sources, the entire contents of which are incorporated herein by reference:

[0627] [1] 3GPP TS 26.445, v.12.0.0, “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”, September 2014.

[0628] [2] M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robillard, J. Lecompte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefevbre, P. Gournay et al., “The ISO / MPEG Unified Speech and Audio Coding Standard - Consistent High Quality for All Content Types and at All Bit Rates”, Journal of the Audio Engineering Society, Vol. 61, No. 12, pp. 956-977, December 2013.

[0629] [3] F. Baumgarte, C. Faller, “Binaural cue coding - Part I: Psychoacoustic fundamentals and design principles”, IEEE Transactions on Speech and Audio Processing, Vol. 11, pp. 509-519, November 2003.

[0630] [4] Tommy Vaillancourt, “Method and system using a long-term correlation difference between left and right channels for time domaindown mixing a stereo sound signal into primary and secondary channels,” U.S. Patent 10,325,606B2.

[0631] [5] 3GPP SA4 Document S4-170749 “New WID on EVS Codec Extension for Immersive Voice and Audio Services”, SA4 Conference #94, June 26-30, 2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip

[0632] [6] I. Mani, J. Zhang, “kNN approach to unbalanced data distributions: A case study involving information extraction”, pp. 1-7, 2003, in the notes of a workshop on learning from imbalanced datasets.

[0633] [7] V. Malenovsky, T. Vaillancourt, W. Zhe, K. Choo and V. Atti, “Two-stage speech / music classifier with decision smoothing and sharpening in the EVS codec”, 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Queensland, 2015, pp. 5718-5722.

[0634] [8] Vaillantour, T., “Method and system for time-domain down mixing of stereo sound signal into primary and secondary channels using detecting an out-of-phase condition on the left and right channels”, US Patent 10,522,157.

[0635] [9] Maalouf, Maher. “Logistic regression in data analysis: An overview”, International Journal of Data Analysis Techniques and Strategies, 2011. 3. 281-299. 10. 1504 / IJDATS. 2011.041335.

[0636]

[10] Ruder, S. “An overview of gradient descent optimization algorithms”. 2016. ArXiv preprint ArXiv:1609.04747.

Claims

1. A stereo mode selection device for selecting one of a first stereo mode and a second stereo mode to encode a stereo audio signal including a left channel and a right channel, comprising: A classifier for generating a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; A detector used to generate a second output indicating whether crosstalk caused by two speakers talking simultaneously exists in the stereo sound signal; An analysis processor for calculating auxiliary parameters to select the stereo mode for encoding the stereo sound signal; as well as A stereo mode selector for selecting the stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters. The stereo mode selector is configured as follows: An initial selection of the stereo mode for encoding the stereo sound signal is performed between the first stereo mode and the second stereo mode; and, After the initial selection of the stereo mode, if several given conditions are met, the second stereo mode is selected for encoding the stereo sound signal.

2. The stereo mode selection device as claimed in claim 1, wherein the first stereo mode is a time-domain stereo mode in which the left channel and the right channel are individually encoded, and the second stereo mode is a frequency-domain stereo mode.

3. The stereo mode selection device as claimed in claim 1 or 2, wherein in the current frame of the stereo audio signal, the stereo mode selector uses the first output from a previous frame of the stereo audio signal and the second output from the previous frame.

4. The stereo mode selection device of claim 1, wherein, in order to perform the initial selection of the stereo mode for encoding the stereo sound signal, the stereo mode selector determines whether a previous frame of the stereo sound signal is a speech frame.

5. The stereo mode selection device of claim 4, wherein in the initial selection of the stereo mode for encoding the stereo sound signal, the stereo mode selector initializes the stereo mode for encoding the stereo sound signal to the stereo mode selected in the previous frame in each frame of the stereo sound signal.

6. The stereo mode selection device of claim 4, wherein in the initial selection of the stereo mode, if (a) the previous frame is determined to be a speech frame and (b) the first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, then the stereo mode selector selects the first stereo mode for encoding the stereo sound signal.

7. The stereo mode selection device of claim 6, wherein in the initial selection of the stereo mode for encoding the stereo sound signal, if (i) at least one of conditions (a) and (b) is not satisfied, and (ii) the stereo mode selected in the previous frame is the second stereo mode, then the stereo mode selector selects the second stereo mode for encoding the stereo sound signal.

8. The stereo mode selection device of claim 6, wherein in the initial selection of the stereo mode, if (i) at least one of the conditions (a) and (b) is not satisfied, and (ii) the stereo mode selected in the previous frame is the first stereo mode, then the stereo mode selector selects the stereo mode for encoding the stereo sound with respect to one of the auxiliary parameters.

9. The stereo mode selection device as claimed in claim 8, wherein one of the auxiliary parameters is an auxiliary stereo mode switching flag.

10. The stereo mode selection device of claim 1, wherein the given condition is selected from the group consisting of: - The first stereo mode was selected in a previous frame of the stereo sound signal; - The first stereo mode was initially selected in the current frame of the stereo sound signal; - In the current frame, the second output of the detector indicates the presence of crosstalk in the stereo audio signal; - (i) the previous frame is determined to be a speech frame, and (ii) the first output from the classifier indicates that there is irrelevant stereo content in the previous frame, or the second output from the detector indicates that there is crosstalk in the stereo sound signal in the previous frame; - In the previous frame, the count value of the number of consecutive frames using the first stereo mode is higher than the first value; - In the previous frame, the count value of the number of consecutive frames using the second stereo mode is higher than the second value; - In the previous frame, the category of the stereo sound signal is within a predefined category set; as well as - (i) the total bit rate used to encode the stereo sound signal is equal to or greater than the third value, or (ii) the score from the detector in the previous frame representing crosstalk in the stereo sound signal is less than the fourth value.

11. The stereo mode selection device of claim 1, wherein the analysis processor calculates an auxiliary sub-mode flag as an auxiliary parameter among the auxiliary parameters, the auxiliary sub-mode flag indicating that the first stereo mode operates in a sub-mode, the sub-mode being applied to a brief transition before switching from the first stereo mode to the second stereo mode.

12. The stereo mode selection device of claim 11, wherein the analysis processor resets the auxiliary sub-mode flag in a frame of the stereo audio signal under the following conditions: (a) a previous frame of the stereo audio signal is determined to be a speech frame, and (b) a first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or a second output from the detector indicates the presence of crosstalk in the stereo audio signal in the previous frame.

13. The stereo mode selection device of claim 12, wherein the analysis processor resets the auxiliary sub-mode flag to 1 in frames of the stereo sound signal under the following conditions: (1) the auxiliary stereo mode switching flag calculated by the analysis processor as an auxiliary parameter is equal to 1, (2) the stereo mode of the previous frame is not the first stereo mode, or (3) the count value of the frame using the first stereo mode is less than a given value.

14. The stereo mode selection device as claimed in claim 13, wherein the analysis processor resets the auxiliary submode flag to 0 in frames in which none of the conditions (1) to (3) of the stereo sound signal are satisfied.

15. The stereo mode selection device of claim 11, wherein the analysis processor does not change the auxiliary sub-mode flag in a frame in which at least one of the following conditions of the stereo audio signal is not met: (a) a previous frame of the stereo audio signal is determined to be a speech frame, and (b) a first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or a second output from the detector indicates the presence of crosstalk in the stereo audio signal in the previous frame.

16. The stereo mode selection device of claim 1, wherein the analysis processor includes a count of the number of consecutive frames using the first stereo mode as one of the auxiliary parameters.

17. The stereo mode selection device of claim 16, wherein if (a) a previous frame of the stereo sound signal is determined to be a speech frame, and (b) the first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, the analysis processor increments the count value of the number of consecutive frames using the first stereo mode.

18. The stereo mode selection device of claim 16, wherein if the stereo mode selector selects the second stereo mode in the current frame of the stereo sound signal, the analysis processor resets the count value of the number of consecutive frames using the first stereo mode to zero.

19. The stereo mode selection device of claim 16, wherein the count value of the number of consecutive frames using the first stereo mode is limited to an upper threshold.

20. The stereo mode selection device of claim 1, wherein the analysis processor includes a count of the number of consecutive frames using the second stereo mode as one of the auxiliary parameters.

21. The stereo mode selection device of claim 20, wherein if the second stereo mode is selected in the current frame of the stereo sound signal, the analysis processor increments the count value of the number of consecutive frames using the second stereo mode.

22. The stereo mode selection device of claim 20, wherein if the stereo mode selector selects the first stereo mode in the current frame of the stereo sound signal, the analysis processor resets the count value of the number of consecutive frames using the second stereo mode to zero.

23. The stereo mode selection device of claim 20, wherein the count value of the number of consecutive frames using the second stereo mode is limited to an upper threshold.

24. The stereo mode selection device of claim 1, wherein the analysis processor generates an auxiliary stereo mode switching flag as an auxiliary parameter among the auxiliary parameters.

25. The stereo mode selection device of claim 24, wherein the analysis processor (i) initializes the auxiliary stereo mode switching flag to 1 in the current frame of the stereo sound signal when (a) a previous frame of the stereo sound signal is determined to be a speech frame, and (b) a first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or a second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, and (ii) initializes the auxiliary stereo mode switching flag to 0 in the current frame when at least one of conditions (a) and (b) is not satisfied.

26. The stereo mode selection device of claim 24, wherein the analysis processor sets the auxiliary stereo mode switching flag to 0 when the left and right channels of the stereo sound signal are out of phase.

27. The stereo mode selection device of claim 9 or 13, wherein the analysis processor generates the auxiliary stereo mode switching flag as an auxiliary parameter among the auxiliary parameters.

28. The stereo mode selection device of claim 27, wherein the analysis processor (i) initializes the auxiliary stereo mode switching flag to 1 in the current frame of the stereo sound signal when (a) the previous frame is determined to be a speech frame and (b) the first output from the classifier indicates the presence of irrelevant stereo content in the previous frame, or the second output from the detector indicates the presence of crosstalk in the stereo sound signal in the previous frame, and (ii) initializes the auxiliary stereo mode switching flag to 0 in the current frame when at least one of the conditions (a) and (b) is not satisfied.

29. The stereo mode selection device of claim 27, wherein the analysis processor sets the auxiliary stereo mode switching flag to 0 when the left and right channels of the stereo sound signal are out of phase.

30. An apparatus for selecting one of a first stereo mode and a second stereo mode for encoding a stereo audio signal including a left channel and a right channel, comprising: At least one processor; as well as A memory coupled to the processor and storing non-transitory instructions that, when executed, cause the processor to perform: A classifier for generating a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; A detector used to generate a second output indicating whether crosstalk caused by two speakers talking simultaneously exists in the stereo sound signal; An analysis processor for calculating auxiliary parameters to select the stereo mode for encoding the stereo sound signal; as well as A stereo mode selector for selecting the stereo mode for encoding the stereo sound signal in response to the first output, the second output, and the auxiliary parameters. The stereo mode selector is configured as follows: An initial selection of the stereo mode for encoding the stereo sound signal is performed between the first stereo mode and the second stereo mode; and, After the initial selection of the stereo mode, if several given conditions are met, the second stereo mode is selected for encoding the stereo sound signal.

31. An apparatus for selecting one of a first stereo mode and a second stereo mode for encoding a stereo audio signal including a left channel and a right channel, comprising: At least one processor; as well as A memory coupled to the processor and storing non-transitory instructions that, when executed, cause the processor to: Generate a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; A second output is generated that indicates whether crosstalk caused by two speakers talking simultaneously is present or absent in the stereo sound signal; Calculate auxiliary parameters to select the stereo mode for encoding the stereo sound signal; as well as The stereo mode for encoding the stereo sound signal is selected in response to the first output, the second output, and the auxiliary parameters. The selection of the stereo mode includes: An initial selection of the stereo mode for encoding the stereo sound signal is performed between the first stereo mode and the second stereo mode. as well as, After the initial selection of the stereo mode, if several given conditions are met, the second stereo mode is selected for encoding the stereo sound signal.

32. A stereo mode selection method for selecting one of a first stereo mode and a second stereo mode for encoding a stereo audio signal including a left channel and a right channel, comprising: Generate a first output indicating the presence or absence of irrelevant stereo content in the stereo sound signal; A second output is generated that indicates whether crosstalk caused by two speakers talking simultaneously is present or absent in the stereo sound signal; Calculate auxiliary parameters to select the stereo mode for encoding the stereo sound signal; as well as The stereo mode for encoding the stereo sound signal is selected in response to the first output, the second output, and the auxiliary parameters. The selection of the stereo mode includes: An initial selection of the stereo mode for encoding the stereo sound signal is performed between the first stereo mode and the second stereo mode. as well as, After the initial selection of the stereo mode, if several given conditions are met, the second stereo mode is selected for encoding the stereo sound signal.

33. The stereo mode selection method of claim 32, wherein the first stereo mode is a time-domain stereo mode in which the left channel and the right channel are encoded separately, and the second stereo mode is a frequency-domain stereo mode.

34. The stereo mode selection method of claim 32, wherein selecting the stereo mode in the current frame of the stereo audio signal includes using the first output from a previous frame of the stereo audio signal and the second output from the previous frame.

35. The stereo mode selection method of claim 32, wherein, in order to perform the initial selection of the stereo mode for encoding the stereo sound signal, selecting the stereo mode includes determining whether a previous frame of the stereo sound signal is a speech frame.

36. The stereo mode selection method of claim 35, wherein in the initial selection of the stereo mode for encoding the stereo sound signal, selecting the stereo mode includes initializing the stereo mode for encoding the stereo sound signal to the stereo mode selected in the previous frame in each frame of the stereo sound signal.

37. The stereo mode selection method as described in claim 35, wherein in the initial selection of the stereo mode, selecting the stereo mode includes: If (a) the previous frame is determined to be a speech frame, and (b) the first output indicates that there is irrelevant stereo content in the previous frame, or the second output indicates that there is crosstalk in the stereo audio signal in the previous frame, then the first stereo mode is selected for encoding the stereo audio signal.

38. The stereo mode selection method of claim 37, wherein in the initial selection of the stereo mode for encoding the stereo sound signal, selecting the stereo mode includes: If (i) at least one of conditions (a) and (b) is not satisfied and (ii) the stereo mode selected in the previous frame is the second stereo mode, then the second stereo mode is selected for encoding the stereo sound signal.

39. The stereo mode selection method as described in claim 37, wherein in the initial selection of the stereo mode, selecting the stereo mode includes: If (i) at least one of the conditions (a) and (b) is not met and (ii) the stereo mode selected in the previous frame is the first stereo mode, then the stereo mode for encoding the stereo sound is selected with respect to one of the auxiliary parameters.

40. The stereo mode selection method as described in claim 39, wherein one of the auxiliary parameters is an auxiliary stereo mode switching flag.

41. The stereo mode selection method of claim 32, wherein the given condition is selected from the group consisting of: - The first stereo mode was selected in a previous frame of the stereo sound signal; - The first stereo mode was initially selected in the current frame of the stereo sound signal; - In the current frame, the second output indicates the presence of crosstalk in the stereo audio signal; - (i) the previous frame is determined to be a speech frame, and (ii) the first output indicates that there is unrelated stereo content in the previous frame, or the second output indicates that there is crosstalk in the stereo sound signal in the previous frame; - In the previous frame, the count value of the number of consecutive frames using the first stereo mode is higher than the first value; - In the previous frame, the count value of the number of consecutive frames using the second stereo mode is higher than the second value; - In the previous frame, the category of the stereo sound signal is within a predefined category set; as well as - (i) the total bit rate used to encode the stereo audio signal is equal to or greater than the third value, or (ii) the score of crosstalk in the stereo audio signal in the previous frame is less than the fourth value.

42. The stereo mode selection method of claim 32, wherein calculating the auxiliary parameters includes calculating an auxiliary sub-mode flag as one of the auxiliary parameters, the auxiliary sub-mode flag indicating that the first stereo mode operates in a sub-mode, the sub-mode being applied to a brief transition before switching from the first stereo mode to the second stereo mode.

43. The stereo mode selection method of claim 42, wherein calculating the auxiliary parameter includes resetting the auxiliary sub-mode flag in a frame of the stereo audio signal under the following conditions: (a) a previous frame of the stereo audio signal is determined to be a speech frame, and (b) a first output indicates the presence of irrelevant stereo content in the previous frame, or a second output indicates the presence of crosstalk in the stereo audio signal in the previous frame.

44. The stereo mode selection method of claim 43, wherein calculating the auxiliary parameter includes resetting the auxiliary sub-mode flag to 1 in frames of the stereo sound signal under the following conditions: (1) the auxiliary stereo mode switching flag, which is an auxiliary parameter, is equal to 1; (2) the stereo mode of the previous frame is not the first stereo mode; or (3) the count value of the frame using the first stereo mode is less than a given value.

45. The stereo mode selection method as claimed in claim 44, wherein calculating the auxiliary parameter includes resetting the auxiliary sub-mode flag to 0 in frames in which none of the conditions (1) to (3) of the stereo sound signal are satisfied.

46. ​​The stereo mode selection method of claim 42, wherein calculating the auxiliary parameter includes not changing the auxiliary sub-mode flag in a frame in which at least one of the following conditions of the stereo audio signal is not met: (a) a previous frame of the stereo audio signal is determined to be a speech frame, and (b) a first output indicates the presence of irrelevant stereo content in the previous frame, or a second output indicates the presence of crosstalk in the stereo audio signal in the previous frame.

47. The stereo mode selection method of claim 32, wherein calculating the auxiliary parameters includes calculating a count of the number of consecutive frames using the first stereo mode as one of the auxiliary parameters.

48. The stereo mode selection method as described in claim 47, wherein calculating the auxiliary parameters includes: If (a) a previous frame of the stereo audio signal is determined to be a speech frame, and (b) the first output indicates the presence of unrelated stereo content in the previous frame, or the second output indicates the presence of crosstalk in the stereo audio signal in the previous frame, then the count value of the number of consecutive frames using the first stereo mode is incremented.

49. The stereo mode selection method as described in claim 47, wherein calculating the auxiliary parameters includes: If the second stereo mode is selected in the current frame of the stereo audio signal, the count value of the number of consecutive frames using the first stereo mode is reset to zero.

50. The stereo mode selection method of claim 47, further comprising limiting the count value of the number of consecutive frames using the first stereo mode to an upper threshold.

51. The stereo mode selection method of claim 32, wherein calculating the auxiliary parameters includes calculating a count of the number of consecutive frames using the second stereo mode as one of the auxiliary parameters.

52. The stereo mode selection method as described in claim 51, wherein calculating the auxiliary parameters includes: If the second stereo mode is selected in the current frame of the stereo sound signal, the count value of the number of consecutive frames using the second stereo mode is incremented.

53. The stereo mode selection method as described in claim 51, wherein calculating the auxiliary parameters includes: If the first stereo mode is selected in the current frame of the stereo sound signal, the count value of the number of consecutive frames using the second stereo mode is reset to zero.

54. The stereo mode selection method of claim 51, further comprising limiting the count value of the number of consecutive frames using the second stereo mode to an upper threshold.

55. The stereo mode selection method of claim 32, wherein calculating the auxiliary parameters includes generating an auxiliary stereo mode switching flag as one of the auxiliary parameters.

56. The stereo mode selection method as described in claim 55, wherein calculating the auxiliary parameters includes: (i) In the case that a previous frame of the stereo audio signal in (a) is determined to be a speech frame, and (b) the first output indicates that there is irrelevant stereo content in the previous frame, or the second output indicates that there is crosstalk in the stereo audio signal in the previous frame, the auxiliary stereo mode switching flag is initialized to 1 in the current frame of the stereo audio signal, and (ii) in the case that at least one of the conditions (a) and (b) is not satisfied, the auxiliary stereo mode switching flag is initialized to 0 in the current frame.

57. The stereo mode selection method as described in claim 55, wherein calculating the auxiliary parameter includes setting the auxiliary stereo mode switching flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.

58. The stereo mode selection method as described in claim 40 or 44, wherein calculating the auxiliary parameters includes generating the auxiliary stereo mode switching flag as an auxiliary parameter among the auxiliary parameters.

59. The stereo mode selection method as described in claim 58, wherein calculating the auxiliary parameters includes: (i) In the case that (a) the previous frame is determined to be a speech frame, and (b) the first output indicates the presence of unrelated stereo content in the previous frame, or the second output indicates the presence of crosstalk in the stereo audio signal in the previous frame, the auxiliary stereo mode switching flag is initialized to 1 in the current frame of the stereo audio signal, and (ii) the auxiliary stereo mode switching flag is initialized to 0 in the current frame when at least one of the conditions (a) and (b) is not met.

60. The stereo mode selection method as described in claim 58, wherein calculating the auxiliary parameter includes setting the auxiliary stereo mode switching flag to 0 when the left channel and the right channel of the stereo sound signal are out of phase.