Stereo sound decoding method and stereo sound decoding system

By using a stereo sound decoding method and employing a factor β to mix the main and auxiliary channels in the time domain, the bit rate allocation is optimized, solving the problem of transmitting stereo information at low bit rates in complex audio scenarios and achieving high-quality sound transmission in broadband and ultra-wideband operations.

CN116343802BActive Publication Date: 2026-06-02VOICEAGE CORPORATION

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VOICEAGE CORPORATION
Filing Date
2016-09-22
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively transmit stereo information at low bit rates in complex audio scenarios, leading to either doubling the bit rate or degrading sound quality. In particular, existing stereo technologies struggle to maintain good quality when there is background noise and interference from the speaker.

Method used

A stereo sound decoding method is adopted. By receiving the encoding parameters, multiple encoding models are used to decode the main channel and the auxiliary channel. The left channel and the right channel are generated by mixing in the time domain using the factor β, and the bit rate allocation is optimized to maintain the sound quality.

Benefits of technology

It significantly improves stereo voice quality and clarity at low bit rates in complex audio scenarios, reduces bit rate requirements, and improves sound quality in broadband and ultra-wideband operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343802B_ABST
    Figure CN116343802B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a stereo sound decoding method and a stereo sound decoding system. The stereo sound decoding method includes receiving encoded parameters including encoded parameters of a primary channel and encoded parameters of a secondary channel, wherein the primary channel encoded parameters include LP filter coefficients of the primary channel; decoding the primary channel in response to the primary channel encoded parameters; and decoding the secondary channel using one of a plurality of encoding models, wherein (a) at least one of the encoding models uses the primary channel LP filter coefficients to decode the secondary channel, and (b) at least one of the encoding models uses primary channel encoded parameters other than the LP filter coefficients to decode the secondary channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of the following invention patent application:

[0002] Application Number: 201680062619.2

[0003] Application date: September 22, 2016

[0004] Invention Title: Method and System for Decoding Left and Right Channels of Stereo Audio Signals Technical Field

[0005] This disclosure relates to stereo sound coding, and more particularly, but not exclusively, to stereo speech and / or audio coding capable of producing good stereo quality in complex audio scenarios at low bit rates and low latency. Background Technology

[0006] Historically, telephone conversations have been conducted using a handset with only one transducer, outputting sound to only one ear of the user. In the last decade or so, users have begun using their portable telephone handsets in conjunction with headsets to receive sound across their ears, primarily for listening to music, and sometimes for voice conversations. However, when using a portable telephone handset to transmit and receive voice conversations, the content remains mono, but when using a headset, the content is presented to both ears of the user.

[0007] The quality of encoded sound has been significantly improved by utilizing the latest 3GPP voice coding standard described in reference [1] (the entire contents of which are incorporated herein by reference). For example, voice and / or audio transmitted and received through a portable telephone handset. The next natural step is to transmit stereo information so that the receiver is as close as possible to the real-life audio scene captured on the other side of the communication link.

[0008] In audio codecs, the transmission of stereo information is normally used, for example as described in reference [2] (the entire contents of which are incorporated herein by reference).

[0009] For dialogue voice codecs, mono signals are the norm. When transmitting mono signals, the bit rate typically needs to be doubled because a mono codec is used to encode both the left and right channels. This works well in most cases, but presents the following drawbacks: the bit rate is doubled, and any potential redundancy between the two channels (left and right) cannot be fully utilized. Furthermore, in order to maintain the overall bit rate at a reasonable level, a very low bit rate is used for each channel, thereby affecting the overall sound quality.

[0010] A possible alternative is to use so-called parametric stereo as described in reference [6] (the entire contents of which are incorporated herein by reference). Parametric stereo transmits information such as binaural time difference (ITD) or binaural intensity difference (IID). The latter information is transmitted per frequency band, and at low bit rates, the bit budget associated with stereo transmission is not high enough to allow these parameters to work effectively.

[0011] While panning factors can help create basic stereo effects at low bit rates, this technique cannot preserve the surrounding environment and has inherent limitations. Too rapid an adaptation of the panning factor becomes distracting for the listener, while too slow an adaptation fails to reflect the speaker's true location, making it difficult to achieve good quality when the speaker is distracted or when background noise fluctuations are significant. Currently, encoding decent-quality conversational stereo speech for all possible audio scenarios requires a minimum bit rate of approximately 24 kb / s for wideband (WB) signals; below this bit rate, speech quality begins to degrade.

[0012] With the increasing globalization of the workforce and the fragmentation of work teams worldwide, there is a need for improved communications. For example, participants in a conference call may be in different and distant locations. Some participants may be in their cars, others in large anechoic chambers, or even in their living rooms. In fact, all participants want to feel as if they are having a face-to-face discussion. Achieving stereo voice (or more generally, stereo sound) in portable devices would be a major step in this direction. Summary of the Invention

[0013] According to a first aspect, this disclosure relates to a stereo sound decoding method for decoding the left and right channels of a stereo sound signal, comprising: receiving encoding parameters, the encoding parameters including encoding parameters for a main channel, encoding parameters for a secondary channel, and a factor β, wherein the main channel encoding parameters include LP filter coefficients for the main channel; decoding the main channel in response to the main channel encoding parameters; decoding the secondary channel using one of a plurality of encoding models, wherein at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel; and temporally mixing the decoded main and secondary channels using the factor β to generate the left and right channels of the decoded stereo sound signal, wherein the factor β determines the respective contributions of the main and secondary channels in the generation of the left and right channels.

[0014] According to a second aspect, a stereo sound decoding system for decoding the left and right channels of a stereo sound signal is provided, comprising: a component for receiving encoding parameters including encoding parameters of a main channel, encoding parameters of a secondary channel, and a factor β, wherein the main channel encoding parameters include LP filter coefficients of the main channel; a decoder for the main channel in response to the main channel encoding parameters; a decoder for the secondary channel using one of a plurality of encoding models, wherein at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel; and a temporal mixer for the decoded main and secondary channels using the factor β to generate the left and right channels of the decoded stereo sound signal, wherein the factor β determines the respective contributions of the main and secondary channels in the generation of the left and right channels.

[0015] According to a third aspect, a stereo sound decoding system for decoding the left and right channels of a stereo sound signal is provided, comprising: at least one processor; and a memory coupled to the processor, and including non-transient instructions that, when executed, cause the processor to implement: means for receiving encoding parameters including encoding parameters of a main channel, encoding parameters of a secondary channel, and an encoding parameter of factor β, wherein the main channel encoding parameters include LP filter coefficients of the main channel; a decoder for the main channel in response to the main channel encoding parameters; a decoder for the secondary channel using one of a plurality of encoding models, wherein at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel; and a temporal mixer for the decoded main and secondary channels using factor β to generate the left and right channels of the decoded stereo sound signal, wherein factor β determines the respective contributions of the main and secondary channels in the generation of the left and right channels.

[0016] On the other hand, a stereo sound decoding system for decoding the left and right channels of a stereo sound signal is provided, comprising: at least one processor; and a memory coupled to the processor, and including non-transient instructions that, when executed, cause the processor to: receive encoding parameters including encoding parameters for a main channel, encoding parameters for a secondary channel, and a factor β, wherein the main channel encoding parameters include LP filter coefficients for the main channel; decode the main channel in response to the main channel encoding parameters; decode the secondary channel using one of a plurality of encoding models, wherein at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel; and temporally mix the decoded main and secondary channels using the factor β to generate the left and right channels of the decoded stereo sound signal, wherein the factor β determines the respective contributions of the main and secondary channels in the generation of the left and right channels.

[0017] This disclosure also relates to a processor-readable memory including non-transitory instructions that, when executed, cause the processor to perform the operations described above.

[0018] This disclosure also relates to a stereo sound decoding method, comprising: receiving encoding parameters including encoding parameters of a main channel and encoding parameters of a secondary channel, wherein the main channel encoding parameters include LP filter coefficients of the main channel; decoding the main channel in response to the main channel encoding parameters; and decoding the secondary channel using one of a plurality of encoding models, wherein (a) at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel, and (b) at least one of the encoding models uses main channel encoding parameters other than the LP filter coefficients to decode the secondary channel.

[0019] This disclosure also relates to a stereo sound decoding system, comprising: a component for receiving encoding parameters including encoding parameters of a main channel and encoding parameters of a secondary channel, wherein the main channel encoding parameters include LP filter coefficients of the main channel; a decoder for the main channel in response to the main channel encoding parameters; and a decoder for the secondary channel using one of a plurality of encoding models, wherein (a) at least one of the encoding models uses the main channel LP filter coefficients to decode the secondary channel, and (b) at least one of the encoding models uses main channel encoding parameters other than the LP filter coefficients to decode the secondary channel.

[0020] The foregoing and other objects, advantages and features of the stereo sound decoding method and system for decoding the left and right channels of a stereo sound signal will become clearer by reading the following non-limiting description of its illustrative embodiments, which are given by way of example only with reference to the accompanying drawings. Attached Figure Description

[0021] In the attached diagram:

[0022] Figure 1 This is a schematic block diagram of a stereo sound processing and communication system, which depicts the possible context for the implementation of the stereo sound coding methods and systems disclosed in the following description.

[0023] Figure 2 The concurrent diagram illustrates the block diagram of the stereo sound coding method and system based on the first model (presented as an integrated stereo design);

[0024] Figure 3 The concurrent diagram illustrates the block diagram of the stereo sound coding method and system based on the second model (presented as an embedded model);

[0025] Figure 4 It is concurrently displayed Figure 2 and 3 Sub-operations of temporal mixing operations in stereo sound coding methods, and Figure 2 and 3 A block diagram of the channel mixer module of a stereo sound coding system;

[0026] Figure 5 It is a graph showing how the linearized long-term correlation difference is mapped to the factor β and the energy normalization factor ε;

[0027] Figure 6 It is a multi-plot showing the difference between using the PCA / KLT scheme over the entire frame and using the "cosine" mapping function;

[0028] Figure 7 It shows a multi-curve diagram of the main channel, auxiliary channel, and the spectrum of these main and auxiliary channels, produced by applying temporal downmixing to a stereo sample recorded in a small echo chamber using a binaural microphone setup with background office noise.

[0029] Figure 8 It is a block diagram illustrating a stereo sound coding method and system, showing the possible implementations and optimizations of coding both the main Y and auxiliary X channels of stereo sound signals;

[0030] Figure 9 It's a diagram. Figure 8 The stereo sound coding method and system LP filter coherence analysis operation and the block diagram of the corresponding LP filter coherence analyzer;

[0031] Figure 10 This is a block diagram illustrating the stereo sound decoding method and the stereo sound decoding system.

[0032] Figure 11 It's a diagram. Figure 10 A block diagram of the stereo sound decoding method and additional features of the system;

[0033] Figure 12 This is a simplified block diagram of an example configuration of the hardware components that form the stereo sound encoding system and stereo sound decoder of this disclosure;

[0034] Figure 13 The concurrent illustration shows the use of a pre-adjustment factor to enhance the stability of stereo images. Figure 2 and 3 Sub-operations of temporal mixing operations in stereo sound coding methods, and Figure 2 and 3 Block diagram of another embodiment of the channel mixer module of a stereo sound coding system;

[0035] Figure 14 This is a block diagram illustrating the operation of time delay correction and the modules of the time delay corrector.

[0036] Figure 15 The concurrent diagram illustrates the block diagram of the replacement stereo sound encoding method and system;

[0037] Figure 16 This is a block diagram illustrating the sub-operations of pitch coherence analysis and the modules of the pitch coherence analyzer.

[0038] Figure 17 The concurrent diagram illustrates a block diagram of a stereo coding method and system using time-domain hybridization, capable of operation in both the time and frequency domains; and

[0039] Figure 18 The concurrent diagram illustrates block diagrams of other stereo coding methods and systems that utilize time-domain mixing and have operational capabilities in both the time and frequency domains. Detailed Implementation

[0040] This disclosure relates to the generation and transmission of realistic representations of stereo sound content, such as speech and / or audio content, from specific but not exclusive complex audio scenarios, with low bit rates and low latency. Complex audio scenarios include situations where (a) there is low correlation between sound signals recorded by microphones, (b) there are significant fluctuations in background noise, and / or (c) there is interference from the speaker. Examples of complex audio scenarios include large anechoic conference rooms with A / B microphone configurations, small echo chambers with binaural microphones, and small echo chambers with mono / side microphone configurations. All of these room configurations can include fluctuating background noise and / or interference from the speaker.

[0041] Known stereo audio codecs such as those described in reference [7] (the entire contents of which are incorporated herein by reference) are inefficient for encoding audio that does not closely approximate a mono model (especially at low bit rates). Certain cases are particularly difficult to encode using existing stereo techniques. Such cases include:

[0042] -LAAB (Large echo-free chamber with A / B microphone setup);

[0043] -SEBI (Small Echo Chamber with Binaural Microphone Setup); and

[0044] -SEMS (Small echo chambers with mono / dual microphone setup).

[0045] Adding fluctuating background noise and / or interfering with the speaker makes these audio signals more difficult to encode at low bit rates using techniques specifically designed for stereo (such as parametric stereo). The drawback of encoding such signals is the use of two mono channels, which doubles the bit rate and network bandwidth required.

[0046] The latest 3GPP EVS voice standard provides a bit rate range from 7.2 kb / s to 96 kb / s for wideband (WB) operation and a bit rate range from 9.6 kb / s to 96 kb / s for ultra-wideband (SWB) operation. This means that the three lowest dual-mono bit rates for using EVS are 14.4, 16.0, and 19.2 kb / s for WB operation and 19.2, 26.3, and 32.8 kb / s for SWB operation. Although the voice quality of the deployed 3GPP AMR-WB described in reference [3] (the entire contents of which are incorporated herein by reference) is improved over its predecessor codec, the quality of coded voice at 7.2 kb / s in noisy environments is far from transparent, and therefore the voice quality of dual-mono at 14.4 kb / s can also be expected to be limited. At such low bit rates, bit rate usage is maximized so that the best possible voice quality is obtained as often as possible. Using the stereo audio coding methods and systems disclosed in the following description, the minimum total bit rate for dialogue stereo speech content (even in complex audio scenarios) should be approximately 13 kb / s for WB and approximately 15.0 kb / s for SWB. At a lower bit rate than that used in dual-mono schemes, the quality and intelligibility of stereo speech are significantly improved for complex audio scenarios.

[0047] Figure 1 This is a schematic block diagram of a stereo sound processing and communication system 100, which depicts a possible context for the implementation of the stereo sound coding methods and systems disclosed in the following description.

[0048] Figure 1 The stereo sound processing and communication system 100 supports the transmission of stereo sound signals via a communication link 101. The communication link 101 may include, for example, a cable or fiber optic link. Alternatively, the communication link 101 may include at least a portion of a radio frequency (RF) link. RF links typically support multiple simultaneous communications requiring shared bandwidth resources, such as those available through cellular telephones. Although not shown, the communication link 101 may be replaced by a storage device in a single implementation of the processing and communication system 100 that records and stores the encoded stereo sound signals for later playback.

[0049] Still referencing Figure 1 For example, a pair of microphones 102 and 122 produce left 103 and right 123 channels of a raw analog stereo sound signal, for example, detected in a complex audio scene. As indicated in the description above, the sound signal may specifically, but not exclusively, include speech and / or audio. Microphones 102 and 122 may be arranged according to A / B, binaural, or mono / sideways configurations.

[0050] The left 103 and right 123 channels of the original analog audio signal are supplied to the analog-to-digital (A / D) converter 104 to convert them into the left 105 and right 125 channels of the original digital stereo audio signal. The left 105 and right 125 channels of the original digital stereo audio signal can also be recorded and supplied from a storage device (not shown).

[0051] The stereo audio encoder 106 encodes the left 105 and right 125 channels of the digital stereo audio signal, thereby producing a set of multiplexed encoding parameters in the form of a bitstream 107 passed to the optional error correction encoder 108. Before transmitting the bitstream 111 obtained via the communication link 101, the optional error correction encoder 108 (if present) adds redundancy to the binary representation of the encoding parameters in the bitstream 107.

[0052] On the receiver side, the optional error correction decoder 109 utilizes the aforementioned redundancy information in the received digital bitstream 111 to detect and correct errors that may occur during transmission through the communication link 101, generating a bitstream 112 with the received encoding parameters. The stereo audio decoder 110 converts the received encoding parameters in the bitstream 112 to create the synthesized left 113 and right 133 channels of the digital stereo audio signal. The reconstructed left 113 and right 133 channels of the digital stereo audio signal in the stereo audio decoder 110 are converted into the synthesized left 114 and right 134 channels of the analog stereo audio signal in the digital-to-analog (D / A) converter 115.

[0053] The left 114 and right 134 channels of the synthesized analog stereo sound signal are reproduced in a pair of speaker units 116 and 136, respectively. Alternatively, the left 113 and right 133 channels of the digital stereo sound signal from the stereo sound decoder 110 can also be supplied to a storage device (not shown) and recorded therein.

[0054] Figure 1 The left 105 and right 125 channels of the original digital stereo audio signal correspond to Figure 2 , 3 The left L and right R channels of frequencies 4, 8, 9, 13, 14, 15, 17, and 18. Furthermore, Figure 1 The stereo sound encoder 106 corresponds to Figure 2 , 3 Stereo sound coding systems of 8, 15, 17 and 18.

[0055] The stereo sound coding method and system disclosed herein are two-fold; providing first and second models.

[0056] Figure 2The concurrent diagram illustrates the block diagram of the stereo sound encoding method and system based on the first model (presented as an integrated stereo design based on the EVS kernel).

[0057] refer to Figure 2 The stereo sound coding method based on the first model includes a time-domain mixing operation 201, a main channel coding operation 202, a secondary channel coding operation 203, and a multiplexing operation 204.

[0058] In order to perform the time-domain mixing operation 201, the channel mixer 251 mixes the two input stereo channels (right channel R and left channel L) to produce the main channel Y and the auxiliary channel X.

[0059] To perform the consonant channel encoding operation 203, the consonant channel encoder 253 selects and uses a minimum number of bits (minimum bit rate) to encode consonant channel X using one of the encoding modes defined in the following description, and produces a corresponding consonant channel encoded bitstream 206. The associated bit budget may vary per frame depending on the frame content.

[0060] To implement the main channel encoding operation 202, a main channel encoder 252 is used. The auxiliary channel encoder 253 transmits a signaling message to the main channel encoder 252 indicating the number of bits 208 used to encode the auxiliary channel X in the current frame. Any suitable type of encoder can be used as the main channel encoder 252. As a non-limiting example, the main channel encoder 252 can be a CELP-type encoder. In this illustrative embodiment, the main channel CELP-type encoder is a modified version of a conventional EVS encoder, wherein the EVS encoder is modified to exhibit greater bitrate scalability to allow flexible bitrate allocation between the main and auxiliary channels. In this manner, the modified EVS encoder will be able to use all the bits not used to encode the auxiliary channel X to encode the main channel Y at the corresponding bitrate, producing a bitstream 205 corresponding to the main channel encoding.

[0061] Multiplexer 254 concatenates the main channel bitstream 205 and the auxiliary channel bitstream 206 to form a multiplexed bitstream 207 to complete the multiplexing operation 204.

[0062] In the first model, the number of bits and corresponding bit rate used to encode the consonant channel X (in bitstream 206) are less than the number of bits and corresponding bit rate used to encode the main channel Y (in bitstream 205). This can be viewed as two (2) variable bit rate channels, where the sum of the bit rates of the two channels X and Y represents a constant total bit rate. This scheme can have different flavors, with more or less emphasis on the main channel Y. According to the first example, when maximum emphasis is applied to the main channel Y, the bit budget of the consonant channel X is strongly forced to be minimized. According to the second example, if less emphasis is applied to the main channel Y, the bit budget of the consonant channel X can be made more constant, which means that the average bit rate of the consonant channel X is slightly higher than in the first example.

[0063] It is important to note that the right R and left L channels of the input digital stereo audio signal are processed by consecutive frames of a given duration, which can correspond to the duration of the frames used in EVS processing. Each frame depends on the duration and sampling rate of the given frame being used, and includes multiple samples from the right R and left L channels.

[0064] Figure 3 The concurrent diagram illustrates the block diagram of the stereo sound coding method and system based on the second model (presented as an embedded model).

[0065] refer to Figure 3 The stereo sound coding method based on the second model includes a time-domain mixing operation 301, a main channel coding operation 302, a secondary channel coding operation 303, and a multiplexing operation 304.

[0066] In order to complete the time-domain mixing operation 301, the channel mixer 351 mixes the two input right R and left L channels to form the main channel Y and the auxiliary channel X.

[0067] In the main channel encoding operation 302, the main channel encoder 352 encodes the main channel Y to produce a main channel encoded bitstream 305. Furthermore, any suitable type of encoder can be used as the main channel encoder 352. As a non-limiting example, the main channel encoder 352 can be a CELP type encoder. In this illustrative embodiment, the main channel encoder 352 uses a voice coding standard such as a conventional EVS mono coding mode or an AMR-WB-IO coding mode, meaning that when the bitrate is compatible with such a decoder, the mono portion of the bitstream 305 will operate with a conventional EVS, AMR-WB-IO, or conventional AMR-WB decoder. Depending on the selected coding mode, some adjustments to the main channel Y may be required for processing by the main channel encoder 352.

[0068] In the consonant channel encoding operation 303, the consonant channel encoder 353 encodes the consonant channel X at a lower bit rate using one of the encoding modes defined in the following description. The consonant channel encoder 353 produces a consonant channel encoded bitstream 306.

[0069] To perform multiplexing operation 304, multiplexer 354 links the main channel encoded bitstream 305 and the auxiliary channel encoded bitstream 306 to form a multiplexed bitstream 307. This is called embedding mode because the auxiliary channel encoded bitstream 306, associated with stereo, is added on top of the cooperating bitstream 305. As described here above, the auxiliary channel bitstream 306 can be stripped off at any time from the multiplexed stereo bitstream 307 (linked bitstreams 305 and 306) that makes the bitstream decodeable by conventional codecs, while users of the latest version of the codec can still enjoy full stereo decoding.

[0070] The first and second models described above are actually very similar to each other. The main difference between the two models is that in the first model, dynamic bit allocation between the two channels Y and X can be used, while in the second model, bit allocation is more restricted due to common operational considerations.

[0071] The following description provides examples of implementations and schemes for the first and second models described above.

[0072] 1) Temporal Hybridization

[0073] As described above, known stereo models operating at low bit rates have difficulty encoding speech that is not close to a mono model. Conventional schemes use, for example, the Karhunen-Loève transform (klt), and perform downmixing in the frequency domain (per band) using, for example, correlations associated with principal component analysis (pca) for each band, to obtain two vectors, as described in references [4] and [5], the entire contents of which are merged here by reference. One of these two vectors merges all highly correlated content, while the other vector defines all content that is not very correlated. The best known method for encoding speech at low bit rates uses a time-domain codec, such as the CELP (Code-Excited Linear Prediction) codec, where known frequency-domain schemes cannot be directly applied. For this reason, although the idea behind per-band pca / klt is interesting, when the content is speech, the main channel Y needs to be transformed back to the time domain, and after such a transformation, its content no longer looks like conventional speech, especially in the above configuration using a speech-specific model such as CELP. This has the effect of reducing the performance of the speech codec. Furthermore, at low bit rates, the input to the voice codec should be as close as possible to the codec's internal model expectation.

[0074] Starting with the idea that the input to a low-bit-rate voice codec should be as close as possible to the desired voice signal, a first technique has been developed. This first technique is based on an evolution of the conventional PCA / KLT scheme. While the conventional scheme calculates the PCA / KLT for each frequency band, the first technique calculates it directly over the entire frame in the time domain. This works adequately during active speech segments if there is no background noise or interference from the speaker. The PCA / KLT scheme determines which channel (left L or right R channel) contains the most useful information, and that channel is sent to the main channel encoder. Unfortunately, the frame-based PCA / KLT scheme is unreliable in the presence of background noise or when two or more people are talking to each other. The principle of the PCA / KLT scheme involves the selection of one input channel (R or L) or another, which often results in drastic changes to the content of the main channel to be encoded. For at least the above reasons, the first technique is not reliable enough, and therefore, a second technique is presented here to overcome the shortcomings of the first technique and allow for smoother transitions between input channels. References will follow below. Figure 4-9 To describe this second technology.

[0075] refer to Figure 4 Hybrid 201 / 301 in the time domain Figure 2 and 3 The operations include the following sub-operations: energy analysis sub-operation 401, energy trend analysis sub-operation 402, L and R channel normalized correlation analysis sub-operation 403, long-term (LT) correlation difference calculation sub-operation 404, long-term correlation difference to factor β conversion and quantization sub-operation 405, and time-domain mixing sub-operation 406.

[0076] Bearing in mind the idea that the input to a low bitrate audio (such as voice and / or audio) codec should be as homogeneous as possible, the energy analysis sub-operation 401 is performed by the energy analyzer 451 in the channel mixers 251 / 351 to first determine the RMS (root mean square) energy of each input channel R and L using relation (1) through the frame:

[0077]

[0078] Where the subscripts L and R represent the left and right channels respectively, L(i) represents sample i of channel L, R(i) represents sample i of channel R, N corresponds to the number of samples per frame, and t represents the current frame.

[0079] The energy analyzer 451 then uses relation (2) to determine the long-term RMS value of each channel using the RMS value from relation (1).

[0080]

[0081] Where t represents the current frame and t -1 This indicates the previous frame.

[0082] In order to perform the energy trend analysis sub-operation 402, the energy trend analyzer 452 of the channel mixer 251 / 351 uses long-term RMS values. The trend of energy in L and R of each channel is determined using relation (3).

[0083]

[0084] The trend of long-term RMS values ​​is used to indicate whether the temporal events captured by the microphone are fading out or changing channels. Long-term RMS values ​​and their trends are also used to determine the convergence rate α of the long-term correlation difference, as described later.

[0085] To perform the channel L and R normalized correlation analysis suboperation 403, the L and R normalized correlation analyzer 453 uses relation (4) to calculate the correlation G for each of the left L and right R channels normalized to the mono signal version m(i) in frame t for sound (e.g., speech and / or audio). L|R :

[0086]

[0087] As already mentioned, N corresponds to the number of samples in the frame, and t represents the current frame. In the current embodiment, all normalized correlation and RMS values ​​determined by relations 1 to 4 are computed in the time domain for the entire frame. In another possible configuration, these values ​​can be computed in the frequency domain. For example, the techniques described herein applicable to audio signals with speech characteristics can be part of a larger framework that allows switching between frequency-domain general stereo audio coding methods and the methods described in this disclosure. In this case, computed normalized correlation and RMS values ​​in the frequency domain may present certain advantages in terms of complexity or code reuse.

[0088] In order to calculate the long-term (LT) correlation difference in suboperation 404, calculator 454 uses relation (5) to calculate the smoothed normalized correlation for each channel L and R in the current frame:

[0089]

[0090] Where α is the convergence rate mentioned above. Finally, calculator 454 uses relation (6) to determine the long-term (LT) correlation difference.

[0091]

[0092] In one example embodiment, the convergence rate α can have a value of 0.8 or 0.5, depending on the trend of the long-term energy calculated in relation (2) and relation (3). For example, when the long-term energies of the left L and right R channels evolve in the same direction, the convergence rate α can have a value of 0.8, and the long-term correlation difference at frame t... With frame t -1 Long-term correlation difference The difference between them is low (less than 0.31 in this example embodiment), and at least one of the long-term RMS values ​​of the left L and right R channels is above a certain threshold (2000 in this example embodiment). This means that the two channels L and R are evolving smoothly, there is no rapid change in energy from one channel to the other, and at least one channel contains a meaningful energy level. Otherwise, when the long-term energies of the right R and left L channels evolve in different directions, when the difference between the long-term correlation differences is high, or when both right R and left L channels have low energy, α would be set to 0.5 to increase the long-term correlation difference. Adjustment speed.

[0093] In order to perform the transformation and quantization sub-operation 405, once the long-term correlation difference has been appropriately estimated in calculator 454 The converter and quantizer 455 then convert the difference into a quantization factor β and supply it to (a) the main channel encoder 252. Figure 2 (b) Auxiliary channel encoder 253 / 353 Figure 2 and 3 (c) Multiplexer 254 / 354 Figure 2 and 3 ), used for through such Figure 1 The 101 communication link is transmitted to the decoder in a multiplexed bit stream 207 / 307.

[0094] The factor β represents two aspects of the stereo input combined into a single parameter. First, factor β represents the proportion or contribution of each of the right R channel and left L channel combined to create the main channel Y. Second, it also represents the energy scaling factor applied to the main channel Y to obtain a main channel that will appear as close in energy domain to the mono signal version of the sound. Therefore, in the case of an embedded architecture, it allows the main channel Y to be decoded separately without receiving the secondary bitstream 306 carrying stereo parameters. This energy parameter can also be used to rescale the energy of the secondary channel X before its encoding, so that the global energy of the secondary channel X is closer to the optimal energy range of the secondary channel encoder. Figure 2As shown, energy information, which is essentially present in factor β, can also be used to improve bit allocation between the main and auxiliary channels.

[0095] An index can be used to pass the quantization factor β to the decoder. Because factor β can represent (a) the respective contributions of the left and right channels to the main channel, and (b) an energy scaling factor that helps to more effectively allocate bits between the main channel Y and the auxiliary channel X, applies a mono signal version to the main channel to obtain sound, or provides correlation / energy information, the index passed to the decoder conveys two different information elements with the same number of bits.

[0096] In order to obtain the long-term correlation difference In this example embodiment, the mapping between the factor β and the converter and quantizer 455 first converts the long-term correlation difference... The long-term correlation difference is restricted to between -1.5 and 1.5, and then linearized between 0 and 2 to obtain the time-linearized long-term correlation difference G. L ′ R (t), as shown in relation (7):

[0097]

[0098] In an alternative implementation, it can be determined that only the long-term correlation difference G filled with linearization is used by further restricting its value to, for example, between 0.4 and 0.6. L ′ R This is a portion of the space of (t). This additional constraint will reduce stereo image localization and save some quantization bits. This option can be considered depending on the design choice.

[0099] After linearization, the converter and quantizer 455 use relation (8) to perform the long-term correlation difference G of the linearization. L ′ R (t) Mapping to the "cosine" domain:

[0100]

[0101] In order to perform the time-domain mixing sub-operation 406, the time-domain mixer 456 uses relations (9) and (10) to generate the main channel Y and the auxiliary channel X as a mix of the right R and left L channels:

[0102] Y(i)=R(i)·(1-β(t))+L(i)·β(t) (9)

[0103] X(i)=L(i)·(1-β(t))-R(i)·β(t) (10)

[0104] Where i = 0, ..., N-1 are the sample indices in the frame and t is the frame index.

[0105] Figure 13 It also demonstrates the use of pre-adjustment factors to enhance stereo image stability. Figure 2 and 3 The time-domain mixing operation of the stereo sound coding method, including sub-operations of 201 / 301, and... Figure 2 and 3 Block diagram of another embodiment of the channel mixer 251 / 351 module of the stereo sound encoding system. Figure 13 In the alternative implementation shown, the time-domain mixing operation 201 / 301 includes the following sub-operations: energy analysis sub-operation 1301, energy trend analysis sub-operation 1302, L and R channel normalized correlation analysis sub-operation 1303, pre-adjustment factor calculation sub-operation 1304, operation of applying the pre-adjustment factor to the normalized correlation 1305, long-term (LT) correlation difference calculation sub-operation 1306, gain to factor β conversion and quantization sub-operation 1307, and time-domain mixing sub-operation 1308.

[0106] Sub-operations 1301, 1302, and 1303 are basically in accordance with... Figure 4 The suboperations 401, 402 and 403, and the analyzers 451, 452 and 453, are performed in the same manner as explained above by the energy analyzer 1351, the energy trend analyzer 1352, and the L and R normalized correlation analyzer 1353, respectively.

[0107] To perform sub-operation 1305, the channel mixer 251 / 351 includes a calculator 1355 for feeding data to the correlation G according to relation (4). L|R (G L (t) and G R (t) Direct application of pre-regulation factor a r This makes the evolution of the correlation gain smoother, depending on the energy and characteristics of the two channels. If the signal energy is low or if it has some unvoiced characteristics, the evolution of the correlation gain can be slower.

[0108] To perform the pre-adjustment factor calculation sub-operation 1304, the channel mixers 251 / 351 include a pre-adjustment factor calculator 1354, which is supplied with (a) long-term left and right channel energy values ​​from relation (2) of the energy analyzer 1351, (b) the frame classification of the previous frame, and (c) speech activity information of the previous frame. The pre-adjustment factor calculator 1354 calculates the pre-adjustment factor a using relation (6a). r It can depend on the minimum long-term RMS values ​​of the left and right channels from the analyzer 1351. Linearized between 0.1 and 1:

[0109]

[0110] In the embodiment, the coefficient M a It can have a value of 0.0009, coefficient B a It can have a value of 0.16. In a variant, for example, if the previous classification of the two channels R and L indicates silent characteristics and active signals, then the pre-adjustment factor a... r It can be forced to 0.15. The Voice Activity Detection (VAD) hangover flag can also be used to determine that the first part of a frame is an active segment.

[0111] Pre-regulation factor a r Normalized correlation G applied to the left L and right R channels L|R (G from relation (4) L (t) and G R Operation 1305 of (t)) and Figure 4 The operation is different from 404. Instead, it is done by directing the normalized correlation G... L|R (G L (t) and G R (t)) Applying factor (1-α), where α is the convergence rate defined above (relationship (5)), to calculate the normalized correlation for long-term (LT) smoothing, calculator 1355 uses relation (11b) to calculate the normalized correlation G of the left L and right R channels. L|R (G L (t) and G R (t) Direct application of pre-regulation factor a r :

[0112]

[0113] The output of calculator 1355 provides a modulated correlation gain τ to the calculator with long-term (LT) correlation difference 1356. L|R .exist Figure 13 In the implementation, the operation of mixing 201 / 301 in the time domain ( Figure 2 and 3 ) including with Figure 4 Suboperations 404, 405, and 406 are similar to suboperation 1306 for calculating long-term (LT) correlation difference, suboperation 1307 for converting and quantizing long-term correlation difference to factor β, and suboperation 1358 for mixing in the time domain, respectively.

[0114] exist Figure 13 In the implementation, the operation of mixing 201 / 301 in the time domain ( Figure 2 and 3 ) including with Figure 4Suboperations 404, 405, and 406 are similar to suboperation 1306 for calculating long-term (LT) correlation difference, suboperation 1307 for converting long-term correlation difference to factor β and quantizing, and suboperation 1358 for mixing in the time domain, respectively.

[0115] Suboperations 1306, 1307, and 1308 are performed by calculator 1356, converter and quantizer 1357, and time-domain mixer 1358, respectively, in essentially the same manner as explained in the preceding descriptions of suboperations 404, 405, and 405, calculator 454, converter and quantizer 455, and time-domain mixer 456.

[0116] Figure 5 This demonstrates how to linearize the long-term correlation difference G′ LR (t) is mapped to the factor β and energy scaling. It can be observed that for a linearized long-term correlation difference G′ of 1.0... LR (t), which means that the right R and left L channels have almost the same energy / correlation, the factor β equals 0.5 and the energy normalization (rescaling) factor ε is 1.0. In this case, the content of the main channel Y is essentially a mono mix, and the auxiliary channel X forms a side channel. The calculation of the energy normalization (rescaling) factor ε is described below.

[0117] On the other hand, if the linearized long-term correlation difference G′ LR (t) equals 2, meaning most of the energy is in the left channel L, so factor β is 1, and the energy normalization (rescaling) factor is 0.5. This indicates that the main channel Y essentially comprises the left channel L in the integrated design implementation, or a downscaled representation of the left channel L in the embedded design implementation. In this case, the auxiliary channel X comprises the right channel R. In the example embodiment, the converter and quantizer 455 or 1357 use 31 possible quantization entries to quantize factor β. The quantized version of factor β is represented using a 5-bit index and, as described above, is supplied to the multiplexer for integration into the multiplexed bitstream 207 / 307 and transmitted to the decoder via a communication link.

[0118] In an embodiment, the factor β can also be used as an indicator for both the main channel encoder 252 / 352 and the auxiliary channel encoder 253 / 353 to determine bit rate allocation. For example, if the β factor is close to 0.5, which means that the energy / correlation of the two (2) input channels is close to each other, more bits will be allocated to the auxiliary channel X and fewer bits to the main channel Y, unless if the contents of the two channels are very close, the contents of the auxiliary channel will be actually low in energy and may be considered inactive, thus allowing very few bits to be encoded for it. On the other hand, if the factor β is close to 0 or 1, the bit rate allocation will favor the main channel Y.

[0119] Figure 6 The above-described PCA / KLT scheme is shown using the entire frame. Figure 6 The two curves above) and the "cosine" function developed in relation (8) for calculating factor β ( Figure 6 The difference between the curves below. Essentially, the PCA / KLT scheme tends to search for either a minimum or maximum value. This is in Figure 6 The intermediate curve shows that it works well for active speech, but it doesn't actually work very well for speech with background noise because it tends to switch continuously from 0 to 1, as... Figure 6 The intermediate curve is shown. Switching to endpoints 0 and 1 too frequently can cause a lot of artifacts when encoding at low bit rates. A potential solution would be to smooth out the decision of the PCA / KLT scheme, but this would negatively affect the detection of voice bursts and their correct positions, while the "cosine" function of relation (8) is more effective in this regard.

[0120] Figure 7 The diagram illustrates the generation of the main channel Y, the auxiliary channel X, and their spectra by applying temporal downmixing to recorded stereo samples in a small echo chamber using a binaural microphone setup with background office noise. After the temporal downmixing operation, it can be seen that both channels still have similar spectral shapes, and the auxiliary channel X still possesses speech with similar temporal content, thus allowing for the use of speech-based models to encode the auxiliary channel X.

[0121] The temporal mixing presented in the preceding description may exhibit some problems in certain cases where the right R and left L channels are out of phase. Adding the right R and left L channels to obtain a mono signal will result in the right R and left L channels canceling each other out. To address this potential problem, in the embodiment, the channel mixer 251 / 351 compares the energy of the mono signal with the energies of both the right R and left L channels. The energy of the mono signal should be at least greater than the energy of one of the right R and left L channels. Otherwise, in this embodiment, the temporal mixing model enters the special case of phase inversion. When this special case occurs, the factor β is forced to 1, and the consonant channel X is forced to use a generic or silent mode encoding, thereby preventing inactive encoding modes and ensuring the correct encoding of the consonant channel X. This special case (where no energy rescaling is applied) signaling is transmitted to the decoder by using the last bit combination (index value) available for the transmission factor β. (Basically, since β is quantized with 5 bits and 31 entries (quantization levels) are used for quantization as described above, the 32nd possible bit combination (entry or index value) is used for signaling this special case.)

[0122] In alternative implementations, more emphasis can be placed on detecting signals for which the downmixing and coding techniques described above are suboptimal, such as in the case of out-of-phase or near-out-of-phase signals. Once these signals are detected, the underlying coding techniques can be adjusted if necessary.

[0123] Typically, for temporal downmixing as described herein, some cancellation may occur during downmixing processing when the left (L) and right (R) channels of the input stereo signal are out of phase, which can result in suboptimal quality. In the example above, the detection of these signals is straightforward, and the encoding strategy involves encoding the two channels separately. However, sometimes, utilizing special signals (e.g., out-of-phase signals), it may be more efficient to still perform downmixing similar to mono / side channel (β = 0.5), where greater emphasis is placed on the side channels. Given that certain special processing of these signals may be beneficial, the detection of these signals needs to be carefully performed. Furthermore, the transition from the normal temporal downmixing model described above to the temporal downmixing model that processes these special signals can be triggered in very low-energy regions or in regions where the pitch of the two channels is unstable, minimizing the subjective effect of switching between the two models.

[0124] Time Delay Correction (TDC) between L and R channels (see...) Figure 17 and 18 The time delay corrector 1750 in the above-described reference [8] or a similar technique (the entire contents of which are incorporated herein by reference) may be executed before entering the downmixing modules 201 / 301, 251 / 351. In such an embodiment, the factor β may end-up with a meaning different from that described above. For this type of implementation, the factor β may become close to 0.5 if the time delay correction operates as expected, which means that the configuration of the downmixing in the time domain is close to a mono / side channel configuration. With proper operation of the time delay correction (TDC), the side channel may include a signal containing a smaller amount of important information. In this case, the bit rate of the side channel X may be minimal when the factor β is close to 0.5. On the other hand, if the factor β is close to 0 or 1, it means that the time delay correction (TDC) may not properly overcome the delay misalignment, and the content of the side channel X may be more complex, thus requiring a higher bit rate. For both types of implementations, a factor β and an associated energy normalization (rescaling) factor ε can be used to improve the bit allocation between the main channel Y and the secondary channel X.

[0125] Figure 14This is a block diagram illustrating the operation of out-of-phase signal detection and the out-of-phase signal detector 1450, which are part of the lower mixing operation 201 / 301 and the channel mixer 251 / 351. (See diagram for details.) Figure 14 As shown, the out-of-phase signal detection operation includes an out-of-phase signal detection operation 1401, a switch position detection operation 1402, and a channel mixer selection operation 1403, to select between a time-domain mixing operation 201 / 301 and an out-of-phase specific time-domain mixing operation 1404. These operations are performed by an out-of-phase signal detector 1451, a switch position detector 1452, a channel mixer selector 1453, the previously described time-domain channel mixers 251 / 351, and the out-of-phase specific time-domain channel mixer 1454, respectively.

[0126] The out-of-phase signal detection 1401 is based on the open-loop correlation between the main and auxiliary channels in the previous frame. To this end, the detector 1451 uses equations (12a) and (12b) to calculate the energy difference S between the side channel signal s(i) and the mono signal m(i) in the previous frame. m (t):

[0127]

[0128] and

[0129] Then, detector 1451 uses relation (12c) to calculate the long-term side channel to mono energy difference.

[0130]

[0131] Where t indicates the current frame, t -1 Indicates the previous frame, and the inactive content therein can be derived from the Voice Activity Detector (VAD) trailing flag or from the VAD trailing counter.

[0132] Besides the long-term energy difference between side channels and mono channels In addition, the maximum open-loop correlation C of the last pitch of each channel Y and X as defined in Clause 5.1.10 of reference [1] is also considered. F|L This is to determine when to consider the current model as suboptimal. This indicates the maximum open-loop correlation of the pitch of the main channel Y in the previous frame, and This indicates the maximum open-loop correlation of the pitch of the consonant channel X in the previous frame. Suboptimal label F sub The switching position detector 1452 calculates according to the following criteria:

[0133] If there is a long-term energy difference between the side channel and the mono channel Above a certain threshold, for example when At that time, if the open-loop maximum correlation of pitch and Both values ​​are between 0.85 and 0.92, meaning these signals have good correlation, but not as good as speech signals. Therefore, the suboptimal label F... sub Set to 1, this indicates the out-of-phase condition between the left L and right R channels.

[0134] Otherwise, the suboptimal label F sub Setting it to 0 indicates that there is no out-of-phase condition between the left L and right R channels.

[0135] To add stability to the suboptimal marker determination, the switching position detector 1452 implements a standard pitch contour for each channel Y and X. When the suboptimal marker F is set in the example embodiment... sub At least three (3) consecutive instances are set to 1 and the main channel p pc(t-1) or consonant channel p sc(t-1) When the pitch stability of the last frame of one of the signals is greater than 64, the switching position detector 1452 determines that the channel mixer 1454 will be used to encode the suboptimal signal. Pitch stability is determined by the three open-loop pitches p1, p2, p3, p4, p6, p7, p8, p9, p1, p1, p1, p2, p1, p2, p1, p2, p1, p2, p3, p4, p1, p2, p1, p2, p3, p1, p2 ... 0|1|2 The sum of absolute differences:

[0136] p pc =|p1-p0|+|p2-p1|and p sc =|p1-p0|+|p2-p1| (12d)

[0137] The switching position detector 1452 provides a decision to the channel mixer selector 1453, which then selects either channel mixer 251 / 351 or channel mixer 1454. The channel mixer selector 1453 implements a hysteresis mechanism, ensuring that when channel mixer 1454 is selected, the decision holds until the following condition is met: for example, multiple consecutive frames of 20 frames are considered optimal, and the main channel p... pc(t-1) or consonant channel p sc(t-1) One of the last frames has a pitch stability greater than, for example, a predetermined number of 64, and a long-term side channel to mono energy difference. Less than or equal to 0.

[0138] 2) Dynamic coding between the main and auxiliary channels

[0139] Figure 8This is a block diagram illustrating a stereo sound coding method and system, showing the possible implementation of optimized coding of both the main Y and auxiliary X channels of a stereo signal (such as speech or audio).

[0140] refer to Figure 8 The stereo sound coding method includes a low-complexity preprocessing operation 801 implemented by a low-complexity preprocessor 851, a signal classification operation 802 implemented by a signal classifier 852, a judgment operation 803 implemented by a judgment module 853, a four (4) subframe model universal unique coding operation 804 implemented by a four (4) subframe model universal unique coding module 854, a two (2) subframe model coding operation 805 implemented by a two (2) subframe model coding module 855, and an LP filter coherence analysis operation 806 implemented by an LP filter coherence analyzer 856.

[0141] After temporal mixing 301 has been performed by channel mixer 351, in the case of the embedded model, (a) the main channel Y is encoded using a conventional encoder such as a conventional EVS encoder or any other suitable conventional voice encoder as the main channel encoder 352 (main channel encoding operation 302) (it should be remembered that, as mentioned in the preceding description, any suitable type of encoder can be used as the main channel encoder 352). In the case of the integrated architecture, a dedicated voice codec is used as the main channel encoder 252. The dedicated voice encoder 252 can be a variable bit rate (VBR) based encoder, such as a modified version of the conventional EVS encoder, which has been modified to have greater bit rate scalability, allowing for variable bit rate handling at the frame level (again, it should be remembered that, as mentioned in the preceding description, any suitable type of encoder can be used as the main channel encoder 252). This allows the minimum number of bits used to encode the secondary channel X to vary in each frame and adapt to the characteristics of the audio signal to be encoded. Finally, the signature of the secondary channel X will be as uniform as possible.

[0142] The encoding of the consonant channel X (i.e., lower energy / correlation with the mono input) is optimized to use the minimum bit rate, particularly but not exclusively for speech-like content. For this purpose, consonant channel encoding can utilize parameters already encoded in the main channel Y, such as LP filter coefficients (LPC) and / or pitch hysteresis 807. Specifically, as described later, it is determined whether the parameters calculated during main channel encoding are sufficiently close to the corresponding parameters calculated during consonant channel encoding for reuse during consonant channel encoding.

[0143] First, a low-complexity preprocessing operation 801 is applied to the consonant channel X using a low-complexity preprocessor 851, where an LP filter, speech activity detection (VAD), and open-loop pitch are calculated in response to the consonant channel X. The subsequent calculations can be implemented, for example, by those performed in a conventional EVS encoder and described in clauses 5.1.9, 5.1.12, and 5.1.10 of reference [1], as stated above, the entire contents of which are incorporated herein by reference. As mentioned in the preceding description, since any suitable type of encoder can be used as the main channel encoder 252 / 352, the above calculations can be implemented by those performed in such a main channel encoder.

[0144] Then, signal classifier 852 analyzes the characteristics of the consonant channel X signal to classify the consonant channel X as silent, general, or inactive using a technique similar to that of the EVS signal classification function in clause 5.1.13 of the same reference [1]. These operations are known to those skilled in the art and can be extracted from standard 3GPP TS26.445v.12.0.0 for simplicity, but alternative implementations may also be used.

[0145] a. Reuse the main channel LP filter coefficients

[0146] A significant portion of bit rate consumption lies in the quantization of LP filter coefficients (LPC). At low bit rates, full quantization of LP filter coefficients can consume nearly 25% of the bit budget. Given that the frequency content of the auxiliary channel X is typically close to that of the main channel Y, but has the lowest energy level, it is necessary to examine whether it is possible to reuse the LP filter coefficients of the main channel Y. To do this, such as... Figure 8 As shown, an LP filter coherence analysis operation 806 has been developed, implemented by an LP filter coherence analyzer 856, in which several parameters are calculated and compared to verify the possibility of reusing the LP filter coefficients (LPC) 807 of the main channel Y.

[0147] Figure 9 It's a diagram. Figure 8 The block diagram of the stereo sound coding method and system LP filter coherence analysis operation 806 and the corresponding LP filter coherence analyzer 856 is shown.

[0148] like Figure 9 As shown, Figure 8The stereo sound coding method and system includes an LP filter coherence analysis operation 806 and a corresponding LP filter coherence analyzer 856, comprising a main channel LP (linear prediction) filter analysis sub-operation 903 implemented by an LP filter analyzer 953, a weighting sub-operation 904 implemented by a weighting filter 954, an auxiliary channel LP filter analysis sub-operation 912 implemented by an LP filter analyzer 962, a weighting sub-operation 901 implemented by a weighting filter 951, an Euclidean distance analysis sub-operation 902 implemented by an Euclidean distance analyzer 952, a residual filter sub-operation 913 implemented by a residual filter 963, a residual energy calculation sub-operation 914 implemented by a residual energy calculator 964, and a subtractor 965. The following are subtraction sub-operations: 915, 910 (sound energy calculation such as speech and / or audio) implemented by energy calculator 960, 906 (auxiliary channel residual filtering operation implemented by auxiliary channel residual filter 956), 907 (residual energy calculation sub-operation implemented by residual energy calculator 957), 908 (subtraction sub-operation implemented by subtractor 958), 911 (gain ratio calculation sub-operation implemented by gain ratio calculator), 916 (comparison sub-operation implemented by comparator 966), 917 (comparison sub-operation implemented by comparator 967), 918 (auxiliary channel LP filter usage judgment sub-operation implemented by judgment module 968), and 919 (main channel LP filter reuse judgment sub-operation implemented by judgment module 969).

[0149] refer to Figure 9 LP filter analyzer 953 performs LP filter analysis on the main channel Y, while LP filter analyzer 962 performs LP filter analysis on the auxiliary channel X. The LP filter analysis performed on each main Y and auxiliary X channel is similar to the analysis described in section 5.1.9 of reference [1].

[0150] Then, the LP filter coefficients A from the LP filter analyzer 953 y It is supplied to the residual filter 956 for the first residual filter r of the auxiliary channel X. Y In the same way, the optimal LP filter coefficients A from the LP filter analyzer 962... x It is supplied to residual filter 963 for the second residual filter r of auxiliary channel X. X Using relation (11), perform the operation with filter coefficient A. Y Or A X Residual filtering:

[0151]

[0152] In this example, s xThe auxiliary channel is indicated by LP filter order 16, and N is the number of samples in the frame (frame size), which is typically 256 corresponding to a 20ms frame duration at a 12.8kHz sampling rate.

[0153] Calculator 910 uses formula (14) to calculate the energy E of the sound signal in the consonant channel X. x :

[0154]

[0155] Furthermore, calculator 957 uses relation (15) to calculate the energy E of the residual from residual filter 956. ry :

[0156]

[0157] Subtractor 958 subtracts the residual energy from calculator 957 from the sound energy from calculator 960 to produce the prediction gain G. Y .

[0158] In the same manner, calculator 964 uses relation (16) to calculate the energy E of the residual from residual filter 963. rx :

[0159]

[0160] Furthermore, subtractor 965 subtracts the residual energy from the sound energy from calculator 960 to produce prediction gain G. X .

[0161] Calculator 961 calculates the gain ratio G Y / G X Comparator 966 compares the gain ratio G. Y / G X With a threshold τ, which is 0.92 in this example embodiment. If the ratio G Y / G X If the result is less than the threshold τ, the comparison result is sent to the judgment module 968. The judgment module 968 forces the use of the auxiliary channel LP filter coefficients for encoding the auxiliary channel X.

[0162] Euclidean distance analyzer 952 performs LP filter similarity measurements, such as line spectrum pairs (lsp) calculated by LP filter analyzer 953 in response to the main channel Y. Y And the line spectrum of lsp calculated by the LP filter analyzer 962 in response to the auxiliary channel X. X The Euclidean distance between them. As those skilled in the art know, the line spectrum corresponds to the lsp... Y and lsp XThis represents the LP filter coefficients in the quantization domain. Analyzer 952 uses relation (17) to determine the Euclidean distance dist:

[0163]

[0164] Where M represents the filter order, and lsp Y and lsp X These represent the line spectrum pairs calculated for the main Y and auxiliary X channels, respectively.

[0165] Before calculating the Euclidean distance in analyzer 952, the two sets of line spectrum pairs may be weighted by appropriate weighting factors. Y and lsp X This allows for varying degrees of focus on certain parts of the spectrum. Other LP filter representations can also be used to compute LP filter similarity metrics.

[0166] Once the Euclidean distance dist is known, it is compared with a threshold σ in comparator 967. In the example embodiment, the threshold σ has a value of 0.08. When comparator 966 determines the ratio G... Y / G X When the Euclidean distance dist is equal to or greater than the threshold τ, and comparator 967 determines that the Euclidean distance dist is equal to or greater than the threshold σ, the comparison result is transmitted to the judgment module 968, which forces the use of the auxiliary channel LP filter coefficients for encoding the auxiliary channel X. When comparator 966 determines the ratio G... Y / G X When the Euclidean distance dist is equal to or greater than the threshold τ, and comparator 967 determines that the Euclidean distance dist is less than the threshold σ, the results of these comparisons are transmitted to the decision module 969. The decision module 969 forces the reuse of the main channel LP filter coefficients for encoding the secondary channel X. In the latter case, the main channel LP filter coefficients are reused as part of the secondary channel encoding.

[0167] In specific cases where the signal is sufficiently easy to encode and a static bit rate exists suitable for encoding LP filter coefficients, such as in a silent coding mode, additional tests can be performed to limit the reuse of main channel LP filter coefficients for encoding consonant channel X. Reuse of main channel LP filter coefficients may also be forced when very low residual gain has been achieved using consonant channel LP filter coefficients, or when consonant channel X has a very low energy level. Finally, the variables τ, σ, residual gain level, or very low energy level that can force the reuse of LP filter coefficients can all be adjusted based on the available bit budget and / or based on the content type. For example, if the consonant channel content is considered inactive, reuse of main channel LP filter coefficients can be determined even if the energy is high.

[0168] b. Low bit rate encoding of the consonant channel

[0169] Since the primary Y and secondary X channels can be a mixture of the right R and left L input channels, this implies that even if the energy content of the secondary X channel is lower than that of the primary Y channel, coding artifacts can be perceived once channel upmixing is performed. To limit this potential artifact, the coding signature of the secondary X channel is kept as constant as possible to restrict any unexpected energy variations. Figure 7 As shown, the content of the secondary channel X has similar characteristics to the content of the primary channel Y, and for this reason, a coding model similar to that of very low bit rate speech has been developed.

[0170] Return to reference Figure 8 The LP filter coherence analyzer 856 sends a decision from the decision module 969 regarding the reuse of the main channel LP filter coefficients, or a decision from the decision module 968 regarding the use of the auxiliary channel LP filter coefficients, to the decision module 853. The decision module 803 then determines that when the main channel LP filter coefficients are reused, the auxiliary channel LP filter coefficients are not quantized, and when the decision indicates that auxiliary channel LP filter coefficients are being used, the auxiliary channel LP filter coefficients are quantized. In the latter case, the quantized auxiliary channel LP filter coefficients are sent to multiplexers 254 / 354 for inclusion in the multiplexed bitstreams 207 / 307.

[0171] In the four (4) subframe model universal unique coding operation 804 and the corresponding four (4) subframe model universal unique coding module 854, in order to keep the bit rate as low as possible, the ACELP search described in section 5.2.3.1 of reference [1] is used only when the LP filter coefficients from the main channel Y can be reused, when the signal classifier 852 classifies the auxiliary channel X as universal, and when the energy of the input right R and left L channels is close to the center (meaning that the energy of the right R and left L channels is close to each other). The coding parameters obtained during the ACELP search in the four (4) subframe model universal unique coding module 854 are then used to construct the auxiliary channel bitstream 206 / 306 and send it to the multiplexers 254 / 354 for inclusion in the multiplexed side bitstream 207 / 307.

[0172] Otherwise, in the two (2) subframe model encoding operation 805 and the corresponding two (2) subframe model encoding module 855, when the LP filter coefficients from the main channel Y cannot be reused, a half-band model is used to encode the consonant channel X with general content. For inactive and silent content, only the spectral shape is encoded.

[0173] In the encoding module 855, inactive content encoding includes (a) frequency domain spectral band gain encoding with noise padding and (b) encoding of auxiliary channel LP filter coefficients when needed, as described in (a) 5.2.3.5.7 and 5.2.3.5.11 and (b) 5.2.2.1 of reference [1], respectively. Inactive content can be encoded at bit rates as low as 1.5 kb / s.

[0174] In the encoding module 855, the consonant channel X silent encoding is similar to the consonant channel X inactive encoding, except that the silent encoding uses an additional number of bits to quantize the consonant channel LP filter coefficients for the silent consonant channel encoding.

[0175] The half-band universal coding model is constructed similarly to ACELP as described in section 5.2.3.1 of reference [1], but it is used only frame-by-frame with two (2) subframes. Therefore, to do so, the residuals described in section 5.2.3.1.1 of reference [1], the memory of the adaptive codebook described in section 5.2.3.1.4 of reference [1], and the input consonant channel are first downsampled by a factor of 2. Using the technique described in section 5.4.4.2 of reference [1], the LP filter coefficients are also modified to represent the downsampled domain, instead of the 12.8 kHz sampling frequency.

[0176] Following the ACELP search, bandwidth expansion is performed in the frequency domain of the excitation. Bandwidth expansion first copies the lower spectral band energies to the higher bands. To copy the spectral band energies, the energies G of the first 9 (9) spectral bands are... bd (i) As described in section 5.2.3.5.7 of reference [1], and the following band is filled as shown in relation (18):

[0177] G bd (i)=G bd (16-i-1), where i=8,…,15. (18)

[0178] Then, using relation (19), the high-frequency content f of the excitation vector represented in the frequency domain as described in section 5.2.3.5.9 of reference [1] is populated with the lower-band frequency content. d (k):

[0179] f d (k)=f d (kP b ), where k = 128, ..., 255, (19)

[0180] Pitch shift P bBased on the multiples of pitch information described in section 5.2.3.1.4.1 of reference [1], and converted into offsets of frequency bins as shown in relation (20):

[0181]

[0182] in F represents the average value of the decoded pitch information for each subframe. s This is the internal sampling frequency, which is 12.8 kHz in this example embodiment, and F r It refers to frequency resolution.

[0183] Then, the coding parameters obtained during the low-rate inactive coding, low-rate silent coding, or half-band general coding performed in the two (2) subframe model coding modules 855 are used to construct the auxiliary channel bitstream 206 / 306 sent to the multiplexer 254 / 354 to be included in the multiplexed bitstream 207 / 307.

[0184] c. Replacement implementation of low bit rate coding for consonant channels

[0185] The encoding of the consonant channel X can be implemented in different ways, with the same goal: to achieve the best possible quality while using the fewest possible bits and maintaining a constant signature. Independent of the potential reuse of LP filter coefficients and pitch information, the encoding of the consonant channel X can be partially driven by the available bit budget. Furthermore, the two (2) subframe model encoding (operation 805) can be half-band or full-band. In this alternative implementation of low bit-rate encoding of the consonant channel, the LP filter coefficients and / or pitch information of the main channel can be reused, and the two (2) subframe model encoding can be selected based on the available bit budget for encoding the consonant channel X. Furthermore, the 2-subframe model encoding presented below has been created by doubling the subframe length instead of downsampling / upsampling its input / output parameters.

[0186] Figure 15 The concurrent diagram illustrates the block diagram of the replacement stereo sound coding method and the replacement stereo sound coding system. Figure 15 Stereo sound coding methods and systems include Figure 8 The methods and several operations and modules of the system are identified using the same reference numerals, and for the sake of brevity, their descriptions are not repeated here. Additionally, Figure 15 The stereo sound coding method includes a preprocessing operation 1501 applied to the main channel Y before operation 202 / 302, a pitch coherence analysis operation 1502, a silent / inactive judgment operation 1504, a silent / inactive coding judgment operation 1505, and a 2 / 4 subframe model judgment operation 1506.

[0187] Suboperations 1501, 1502, 1503, 1504, 1505, and 1506 are executed by a preprocessor 1551 (similar to a low-complexity preprocessor 851), a pitch coherence analyzer 1552, a bit allocation estimator 1553, a silent / inactive judgment module 1554, a silent / inactive coding judgment module 1555, and a 2 / 4 subframe model judgment module 1556, respectively.

[0188] To perform pitch coherence analysis operation 1502, preprocessors 851 and 1551 provide the pitch coherence analyzer 1552 with the open-loop pitches of both the main Y and auxiliary X channels, OLpitch respectively. pri and OLpitch sec .exist Figure 16 It shows more details in the middle. Figure 15 The pitch coherence analyzer 1552, Figure 16 This is a block diagram illustrating the sub-operations of pitch coherence analysis operation 1502 and the modules of pitch coherence analyzer 1552.

[0189] Pitch coherence analysis operation 1502 evaluates the similarity of open-loop pitches between the main channel Y and the consonant channel X to determine under what circumstances the main open-loop pitch can be reused when encoding the consonant channel X. To this end, pitch coherence analysis operation 1502 includes a main channel open-loop pitch addition sub-operation 1601 performed by the main channel open-loop pitch adder 1651 and a consonant channel open-loop pitch addition sub-operation 1602 performed by the consonant channel open-loop pitch adder 1652. The sum from adder 1652 is subtracted from the sum from adder 1651 using sub-operation 1603. The subtraction result from sub-operation 1603 provides stereo pitch coherence. As a non-limiting example, the sums in sub-operations 1601 and 1602 are based on three (3) previous consecutive open-loop pitches available for each channel Y and X. The open-loop pitch can be calculated, for example, as defined in section 5.1.10 of reference [1]. The stereo pitch coherence S is calculated in suboperations 1601, 1602, and 1603 using relation (21). pc :

[0190]

[0191] in p|s(i) represents the open-loop pitch of the main Y and auxiliary X channels, and i represents the position of the open-loop pitch.

[0192] When the stereo coherence is below a predetermined threshold Δ, the pitch information from the main channel Y may be reused to encode the secondary channel X, depending on the available bit budget. Furthermore, the reuse of pitch information for signals with acoustic characteristics of both the main Y and secondary X channels may be limited, depending on the available bit budget.

[0193] To this end, the pitch coherence analysis operation 1502 includes a judgment sub-operation 1604 performed by the judgment module 1654, which considers the available bit budget and the characteristics of the audio signal (e.g., indicated by the encoding patterns of the main and auxiliary channels). When the judgment module 1654 detects that the available bit budget is sufficient, or that the audio signals of both the main Y and auxiliary X channels do not have acoustic characteristics, it determines that pitch information related to the auxiliary channel X is to be encoded (1605).

[0194] When the judgment module 1654 detects a low available bit budget for the purpose of encoding the pitch information of the auxiliary channel X, or when the sound signals used for both the main Y and auxiliary X channels have acoustic characteristics, the judgment module compares the stereo sound high coherence S. pc With respect to the threshold Δ. When the bit budget is low, the threshold Δ is set to a larger value compared to cases where the bit budget is more important (sufficient to encode the pitch information of the secondary channel X). When the stereo sound has high coherence S pc When the absolute value is less than or equal to the threshold Δ, module 1654 determines to reuse the pitch information from the main channel Y to encode the secondary channel X (1607). When the stereo high coherence S... pc When the value is higher than the threshold Δ, module 1654 determines the pitch information of the encoded consonant channel X (1605).

[0195] Ensuring that the channels possess acoustic characteristics increases the possibility of smooth pitch evolution, thereby reducing the risk of added artifacts by reusing the pitch of the main channel. As a non-limiting example, when the stereo bit budget is below 14 kb / s and the stereo correlation S is high... pc When the value is less than or equal to 6 (Δ = 6), the main pitch information can be reused when encoding the secondary channel X. According to another non-limiting example, if the stereo bit budget is greater than 14 kb / s and less than 26 kb / s, both the main Y and secondary X channels are considered audible, and the stereo high coherence S... pc This results in a lower reuse rate of pitch information for the main channel Y at a bit rate of 22kb / s compared to a lower threshold Δ=3.

[0196] Return to reference Figure 15The bit allocation estimator 1553 is supplied with the factor β from the channel mixers 251 / 351, a decision from the LP filter coherence analyzer 856 regarding the reuse of the main channel LP filter coefficients or the use and encoding of the auxiliary channel LP filter coefficients, and pitch information determined by the pitch coherence analyzer 1552. Depending on the encoding requirements of the main and auxiliary channels, the bit allocation estimator 1553 provides the main channel encoder 252 / 352 with a bit budget for encoding the main channel Y and the decision module 1556 with a bit budget for encoding the auxiliary channel X. In one possible implementation, for all inactive content, a portion of the total bit rate is allocated to the auxiliary channels. The auxiliary channel bit rate is then increased by an amount related to the energy normalization (rescaling) factor ε described earlier.

[0197] B x =B M +(0.25·ε-0.125)·(B t -2·B M (21a)

[0198] Among them B x B represents the bit rate allocated to the consonant channel X. t B represents the total available stereo bit rate. M This represents the minimum bit rate allocated to the secondary channel, and is typically approximately 20% of the total stereo bit rate. Finally, ε represents the energy normalization factor mentioned above. Therefore, the bit rate allocated to the primary channel corresponds to the difference between the total stereo bit rate and the secondary channel stereo bit rate. In an alternative implementation, the secondary channel bit rate allocation can be described as:

[0199]

[0200] Among them B x Again, B represents the bit rate allocated to the consonant channel X. t Indicates the total available stereo bit rate and B M This represents the minimum bit rate allocated to the consonant channel. Finally, ε idx This represents the index of the transmission of the aforementioned energy normalization factor. Therefore, the bit rate allocated to the main channel corresponds to the difference between the total stereo bit rate and the auxiliary channel bit rate. In all cases, for inactive content, the auxiliary channel bit rate is set to the minimum bit rate required to encode the spectral shape of the auxiliary channels for a given bit rate generally close to 2 kb / s.

[0201] During this process, signal classifier 852 provides the signal classification of consonant channel X to decision module 1554. If decision module 1554 determines that the sound signal is inactive or silent, silent / inactive encoding module 1555 provides the spectral shape of consonant channel X to multiplexers 254 / 354. Alternatively, decision module 1554 informs decision module 1556 when the sound signal is neither inactive nor silent. For such a sound signal, using the bit budget used to encode consonant channel X, decision module 1556 determines whether there are sufficient bits available for encoding consonant channel X using four (4) subframe model universal unique encoding module 854; otherwise, decision module 1556 selects to use two (2) subframe model encoding module 855 to encode consonant channel X. In order to select four subframe model universal unique encoding module, the bit budget available for consonant channel must be high enough once all other parts have been quantized or reused to allocate at least 40 bits to the algebraic book, including LP coefficients and pitch information and gain.

[0202] As will be understood from the above description, in the four (4) subframe model universal unique coding operation 804 and the corresponding four (4) subframe model universal unique coding module 854, the ACELP search described in section 5.2.3.1 of reference [1] is used to maintain the bit rate as low as possible. In the four (4) subframe model universal unique coding, the pitch information from the main channel may or may not be reused. The coding parameters obtained during the ACELP search in the four (4) subframe model universal unique coding module 854 are then used to construct the secondary channel bitstream 206 / 306, and the coding parameters are sent to multiplexers 254 / 354 to be included in the multiplexed bitstream 207 / 307.

[0203] In the alternative two (2) subframe model coding operation 805 and the corresponding alternative two (2) subframe model coding module 855, a general coding model is constructed similarly to that described in clause 5.2.3.1 of reference [1], but it is used only frame-by-frame with two (2) subframes. Therefore, in order to do so, the length of the subframe is increased from 64 samples to 128 samples, while still maintaining an internal sampling rate of 12.8 kHz. If the pitch coherence analyzer 1552 has determined that the pitch information from the main channel Y is reused for coding the consonant channel X, the average pitch of the first two subframes of the main channel Y is calculated and used as the pitch estimate for the first half of the consonant channel X. Similarly, the average pitch of the last two subframes of the main channel Y is calculated and used for the second half of the consonant channel X. When reused from the main channel Y, the LP filter coefficients are interpolated, and the interpolation of the LP filter coefficients as described in Clause 5.2.2.1 of Reference [1] is modified to accommodate the two (2) subframe scheme by replacing the first and third interpolation factors with the second and fourth interpolation factors.

[0204] exist Figure 15 In the embodiments, the processing for determining between four (4) subframe and two (2) subframe coding schemes is driven by the bit budget available for encoding the consonant channel X. As previously described, the bit budget for the consonant channel X is derived from various elements, such as the available total bit budget, the factor β or energy normalization factor ε, the presence of a time delay correction (TDC) module, the possibility of reusing LP filter coefficients and / or pitch information from the main channel Y.

[0205] When both the LP filter coefficients and pitch information are reused from the main channel Y, the absolute minimum bit rate used by the two (2) subframe coding model of the auxiliary channel X is approximately 2 kb / s for a general signal, and approximately 3.6 kb / s for a four (4) subframe coding scheme. For encoders like ACELP, using two (2) or four (4) subframe coding models, much of the quality comes from the number of bits that can be searched and allocated to the algebraic book (ACB), as defined in clause 5.2.3.1.5 of reference [1].

[0206] Then, to maximize quality, the idea is to compare the bit budget available for the four (4) subframe algebraic book (ACB) search and the two (2) subframe algebraic book (ACB) search, and then consider all that will be encoded. For example, for a given frame, if there exists a 4kb / s (80 bits / 20ms frame) available for encoding consonant channel X, and the LP filter coefficients can be reused while the pitch information needs to be transmitted, then the minimum number of bits needed to encode the consonant channel signaling, consonant channel pitch information, gain, and algebraic book for both two (2) and four (4) subframes is removed from the 80 bits to obtain the bit budget available for encoding the algebraic book. For example, if at least 40 bits are available for encoding the four (4) subframe algebraic book, the four (4) subframe coding model is chosen; otherwise, the two (2) subframe scheme is used.

[0207] 3) A mono signal approximately derived from a partial bitstream

[0208] As described in the preceding description, temporal mixing is mono-friendly, meaning that in an embedded structure where the main channel Y is encoded using a conventional codec (it should be remembered that, as mentioned in the preceding description, any suitable type of encoder can be used as the main channel encoder 252 / 352) and stereo bits are appended to the main channel bitstream, the stereo bits can be stripped, and the conventional decoder can create a synthesis that subjectively approximates the assumed mono synthesis. For this purpose, a simple energy normalization is required on the encoder side before encoding the main channel Y. By rescaling the energy of the main channel Y to a value sufficiently close to the energy of the mono signal version of the sound, decoding the main channel Y using a conventional decoder can be similar to decoding the mono signal version of the sound using a conventional decoder. The energy normalization function is directly linked to the linearized long-term correlation difference G calculated using relation (7). L ′ R (t), and calculate using relation (22):

[0209] ε=-0.485·G L ′ R (t) 2 +0.9765·G L ′ R (t)+0.5. (22)

[0210] Figure 5 The normalization levels are shown in the figure. In practice, instead of using relation (22), a lookup table is used to correlate the normalized value ε with each possible value of the factor β (31 values ​​in this example embodiment). This extra step may be helpful when only decoding mono signals without decoding stereo bits, even though it is not required when encoding stereo sound signals (e.g., speech and / or audio) using an integrated model.

[0211] 4) Stereo decoding and upmixing

[0212] Figure 10 This is a block diagram illustrating a stereo sound decoding method and a stereo sound decoding system. Figure 11 It's a diagram. Figure 10 A block diagram of the stereo sound decoding method and additional features of the stereo sound decoding system.

[0213] Figure 10 and 11 The stereo sound decoding method includes a demultiplexing operation 1007 implemented by a demultiplexer 1057, a main channel decoding operation 1004 implemented by a main channel decoder 1054, a secondary channel decoding operation 1005 implemented by a secondary channel decoder 1055, and a time-domain mixing operation 1006 implemented by a time-domain channel mixer 1056. The secondary channel decoding operation 1005 includes, for example... Figure 11 The judgment operation 1101 performed by the judgment module 1151, the four (4) subframe general decoding operation 1102 implemented by the four (4) subframe general decoder 1152, and the two (2) subframe general / silent / inactive decoding operation 1103 implemented by the two (2) subframe general / silent / inactive decoder 1153 are shown.

[0214] In a stereo audio decoding system, bitstream 1001 is received from an encoder. Demultiplexer 1057 receives bitstream 1001 and extracts from it the encoding parameters (bitstream 1002) for the main channel Y, the encoding parameters (bitstream 1003) for the secondary channel X, and a factor β, which are supplied to the main channel decoder 1054, the secondary channel decoder 1055, and the channel mixer 1056. As previously described, factor β is used as an indicator for both the main channel encoder 252 / 352 and the secondary channel encoder 253 / 353 to determine the bitrate allocation, thereby reusing factor β by both the main channel decoder 1054 and the secondary channel decoder 1055 to properly decode the bitstream.

[0215] The main channel coding parameters correspond to the ACELP coding model at the received bit rate and can be associated with a conventional or modified EVS encoder (it should be remembered here that, as mentioned in the preceding description, any suitable type of encoder can be used as the main channel encoder 252). The bitstream 1002 is supplied to the main channel decoder 1054 to decode the main channel coding parameters (coder mode 1, β, LPC 1, pitch 1, fixed codebook index 1, and gain 1, as in reference [1]) using a method similar to that in reference [1]. Figure 11 (as shown), to generate the decoded main channel Y'.

[0216] The auxiliary channel decoder 1055 uses auxiliary channel encoding parameters that correspond to the model used to encode the second channel X, and may include:

[0217] (a) A reused generic coding model with LP filter coefficients (LPC1) from the main channel Y and / or other coding parameters (e.g., pitch lag pitch 1). The four (4) subframe generic decoder 1152 of the auxiliary channel decoder 1055. Figure 11 The LP filter coefficients (LPC1) and / or other encoding parameters (e.g., pitch lag pitch 1) from the main channel Y of decoder 1054 are supplied and / or the bitstream 1003 is supplied. Figure 11 The β, pitch 2, fixed codebook index 2, and gain 2 shown are used in conjunction with the encoding module 854. Figure 8 The opposite method is used to generate the decoded consonant channel X'.

[0218] (b) Other coding models may or may not reuse the LP filter coefficients (LPC1) from the main channel Y and / or other coding parameters (e.g., pitch lag pitch 1), including half-band universal coding models, low-rate silent coding models, and low-rate inactive coding models. As an example, the inactive coding model may reuse the main channel LP filter coefficients LPC1. The two (2) subframe universal / silent / inactive decoder 1153 of the auxiliary channel decoder 1055 Figure 11 The LP filter coefficients (LPC1) and / or other coding parameters (e.g., pitch lag pitch 1) from the main channel Y and / or the auxiliary channel coding parameters from bitstream 1003 are supplied. Figure 11 The encoding modes shown are 2, β, LPC2, pitch 2, fixed codebook index 2, and gain 2, and are used with encoding module 855 ( Figure 8 The opposite method is used to generate the decoded consonant channel X'.

[0219] The received encoding parameters (bitstream 1003) corresponding to the consonant channel X contain information related to the encoding model being used (coder mode 2). The determination module 1151 uses this information (coder mode 2) to determine and indicate to the four (4) subframe general decoder 1152 and the two (2) subframe general / silent / inactive decoder 1153 which encoding model will be used.

[0220] In the case of the embedded structure, factor β is used to recover the energy scaling index stored in the lookup table (not shown) on the decoder side, and to rescale the main channel Y' before performing the temporal upmixing operation 1006. Finally, factor β is supplied to the channel upmixer 1056 and used to upmix the decoded main Y' and auxiliary X' channels. Using relations (23) and (24), the temporal upmixing operation 1006 is performed as the inverse of the downmixing relations (9) and (10) to obtain the decoded right R' and left L' channels:

[0221]

[0222]

[0223] Where n = 0, ..., N-1 are the indices of the samples in the frame, and t is the frame index.

[0224] 5) Integration of time-domain and frequency-domain coding

[0225] For applications of this technique that utilize frequency-domain coding modes, time-domain mixing is also envisioned to save some complexity or simplify the data flow. In this case, the same mixing factor is applied to all spectral coefficients to preserve the advantages of time-domain mixing. It can be observed that this differs from applying spectral coefficients to each frequency band, as is the case with most frequency-domain mixing applications. The downmixer 456 can be adapted to compute relations (25.1) and (25.2):

[0226] F Y (k)=F R (k)·(1-β(t))+F L (k)·β(t) (25.1)

[0227] F X (k)=F L (k)·(1-β(t))-F R (k)·β(t), (25.2)

[0228] Where F R (k) represents the frequency coefficient k of the right channel R, and similarly, F L (k) represents the frequency coefficient k of the left channel L. Then, the main Y and auxiliary X channels are calculated by applying inverse frequency transformation to obtain the time representation of the submixed signal.

[0229] Figure 17 and 18 A possible implementation of a time-domain stereo coding method and system using frequency-domain mixing is shown, which can switch between time-domain and frequency-domain coding of the main Y and auxiliary X channels.

[0230] Figure 17 This illustrates a first variation of the method and system. Figure 17 The concurrent diagram illustrates a block diagram of a stereo coding method and system that uses time-domain hybridization and has operational capabilities in both the time and frequency domains.

[0231] exist Figure 17 In this stereo coding method and system, numerous previously described operations and modules, referred to in the preceding figures and identified by the same reference numerals, are included. The decision module 1751 (decision operation 1701) determines whether the left L' and right R' channels from the time delay corrector 1750 should be encoded in the time domain or the frequency domain. If time-domain encoding is selected, then... Figure 17 The stereo coding method and system operate in essentially the same manner as the stereo coding method and system shown in the previous figures, for example, but not limited to, such as... Figure 15 As in the embodiments.

[0232] If the decision module 1751 selects frequency encoding, the time-to-frequency converter 1752 (time-to-frequency conversion operation 1702) converts the left L' and right R' channels to the frequency domain. The frequency domain mixer 1753 (frequency domain mixing operation 1703) outputs the main Y and auxiliary X frequency domain channels. The frequency domain main channel is converted back to the time domain by the frequency-to-time converter 1754 (frequency-to-time conversion operation 1704), and the resulting time domain main channel Y is applied to the main channel encoder 252 / 352. The frequency domain auxiliary channel X from the frequency domain mixer 1753 is processed by the conventional parametric and / or residual encoder 1755 (parametric and / or residual encoding operation 1705).

[0233] Figure 18 This is a block diagram illustrating other stereo coding methods and systems that utilize frequency domain mixing and have operational capabilities in both the time and frequency domains. Figure 18 In this context, the stereo coding method and system are similar to... Figure 17 The stereo coding method and system are similar, and only new operations and modules will be described.

[0234] A time-domain analyzer 1851 (time-domain analysis operation 1801) replaces the previously described time-domain channel mixers 251 / 351 (time-domain down-mixing operations 201 / 301). The time-domain analyzer 1851 includes... Figure 4Most of the modules are present, but not the time-domain mixer 456. Therefore, its main function is to provide the calculation of factor β. This factor β is supplied to the preprocessor 851 and frequency-domain to time-domain converters 1852 and 1853 (frequency-domain to time-domain conversion operations 1802 and 1803), which respectively convert the frequency-domain auxiliary X and main Y channels received from the frequency-domain mixer 1753 to the time domain for time-domain encoding. Therefore, the output of converter 1852 is the time-domain auxiliary channel X supplied to the preprocessor 851, while the output of converter 1852 is the time-domain main channel Y, which is supplied to both the preprocessor 1551 and the encoders 252 / 352.

[0235] 6) Example Hardware Configuration

[0236] Figure 12 This is a simplified block diagram of an example configuration of the hardware components that make up each of the stereo audio encoding system and stereo audio decoding system described above.

[0237] Each of the stereo sound encoding system and stereo sound decoding system can be implemented as part of a mobile terminal, a portable media player, or any similar device. Each of the stereo sound encoding system and stereo sound decoding system (in...) Figure 12 The part (marked as 1200) includes input 1202, output 1204, processor 1206, and memory 1208.

[0238] Input 1202 is configured to receive the left (L) and right (R) channels of an input stereo audio signal in digital or analog form in the case of a stereo audio encoding system, or to receive bitstream 1001 in the case of a stereo audio decoding system. Output 1204 is configured to supply multiplexed bitstream 207 / 307 in the case of a stereo audio encoding system, or to supply decoded left channel L' and right channel R' in the case of a stereo audio decoding system. Input 1202 and output 1204 can be implemented in a common module, such as a serial input / output device.

[0239] Processor 1206 is operatively connected to input 1202, output 1204, and memory 1208. Processor 1206 is implemented for performing operations supporting, such as... Figure 2 , 3 The stereo sound coding systems shown in 1, 4, 8, 9, 13, 14, 15, 16, 17 and 18, and as shown in 18, are stereo sound coding systems. Figure 10 and 11 The stereo sound decoding system shown represents the functions of each module of each system, and one or more processors with code instructions.

[0240] Memory 1208 may include non-transitory memory for storing code instructions executable by processor 1206, specifically, processor-readable memory including non-transitory instructions that, when executed, cause the processor to implement the operations and modules of the stereo sound encoding and decoding methods and systems described in this disclosure. Memory 1208 may also include random access memory or buffers(s) for storing intermediate processing data from various functions performed by processor 1206.

[0241] Those skilled in the art will recognize that the descriptions of stereo sound encoding and decoding methods and systems are merely illustrative and not intended to be limiting in any way. Other embodiments will readily conceive of those skilled in the art upon receiving this disclosure. Furthermore, the disclosed stereo sound encoding and decoding methods and systems can be customized to provide valuable solutions to existing needs and problems in encoding and decoding stereo sound.

[0242] For clarity, not all conventional features of the implementation of stereo sound coding methods and systems, and stereo sound decoding methods and systems, are shown or described. It will be understood, of course, that in the development of any such actual implementation of stereo sound coding methods and systems, and stereo sound decoding methods and systems, many implementation-specific judgments may need to be made to achieve the developer's specific goals, such as complying with constraints related to the application, system, network, and business, and these specific goals will vary depending on the implementation and the developer. Furthermore, it will be recognized that the development work can be complex and time-consuming, but remains a routine engineering task for those skilled in the art of sound processing who will benefit from this disclosure.

[0243] According to this disclosure, the modules, processing operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. Furthermore, those skilled in the art will recognize that devices with less general-purpose characteristics, such as hardwired devices, field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs), can also be used. Where a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and these operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, they can be stored on tangible and / or non-transitory media.

[0244] The stereo sound encoding methods and systems described herein, as well as the stereo sound decoding methods and decoders, may include software, firmware, hardware, or any combination of software, firmware, or hardware as suited to the purposes described herein.

[0245] In the stereo sound encoding and decoding methods described herein, various operations and sub-operations can be performed in various orders, and some operations and sub-operations can be optional.

[0246] Although the present disclosure has been described above by way of its non-limiting illustrative embodiments, these embodiments may be modified freely within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0247] References

[0248] The following references are cited in this application, and their entire contents are incorporated herein by reference.

[0249] [1] 3GPP TS 26.445, v.12.0.0, "Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description", Sep 2014.

[0250] [2]M.Neuendorf,M.Multrus,N.Rettelbach,G.Fuchs,J.Robillard,J.Lecompte,S.Wilde,S.Bayer,S.Disch,C.Helmrich,R.Lefevbre,P.Gournay,etal.,"The ISO / MPEGUnified Speech and Audio Coding Standard-Consistent High Quality for AllContent Types and at All Bit Rates", J.Audio Eng.Soc., vol.61, no.12, pp.956-977, Dec.2013.

[0251] [3]B.Bessette, R.Salami, R.Lefebvre, M.Jelinek, J.Rotola-Pukkila, J.Vainio, H.Mikkola, and K. "The Adaptive Multi-Rate Wideband SpeechCodec(AMR-WB)," Special Issue of IEEE Trans.Speech and Audio Proc., Vol.10, pp.620-636, November 2002.

[0252] [4]R.G.van der Waal&R.N.J.Veldhuis,”Subband coding of stereophonicdigital audio signals”,Proc.IEEE ICASSP,Vol.5,pp.3601-3604,April 1991

[0253] [5]Dai Yang,Hongmei Ai,Chris Kyriakakis and C.-C.Jay Kuo,“High-Fidelity Multichannel Audio Coding With Karhunen-Loève Transform”,IEEETrans.Speech and Audio Proc.,Vol.11,No.4,pp.365-379,July 2003.

[0254] [6]J.Breebaart,S.van de Par,A.Kohlrausch and E.Schuijers,“ParametricCoding of Stereo Audio”,EURASIP Journal on Applied Signal Processing,Issue 9,pp.1305-1322,2005

[0255] [7]3GPP TS 26.290 V9.0.0,“Extended Adaptive Multi-Rate–Wideband(AMR-WB+)codec;Transcoding functions(Release 9)”,September 2009.

[0256] [8]Jonathan A.Gibbs,“Apparatus and method for encoding a multi-channelaudio signal”,US 8577045 B2。

Claims

1. A stereo sound decoding method, comprising: The system receives encoding parameters including encoding parameters for the main channel and encoding parameters for the auxiliary channel, wherein the encoding parameters for the main channel include the LP filter coefficients for the main channel. Decode the main channel in response to the main channel encoding parameters; and The secondary channel is decoded using one of a plurality of encoding models, wherein (a) at least one of the encoding models uses the primary channel LP filter coefficients to decode the secondary channel, and (b) at least one of the encoding models uses primary channel encoding parameters other than the LP filter coefficients to decode the secondary channel.

2. The stereo sound decoding method according to claim 1, wherein the encoding model includes a general encoding model, a silent encoding model, and an inactive encoding model.

3. The stereo sound decoding method according to claim 1, wherein the auxiliary channel encoding parameters include information identifying one of the encoding models to be used when decoding the auxiliary channel.

4. The stereo sound decoding method according to any one of claims 1 to 3, wherein: - The received encoded parameters include the factor β; Decoding the main channel in response to the main channel encoding parameters includes: using the factor β as an indicator for the bit rate allocation of the main channel; and Decoding the consonant channel includes using the factor β as an indicator of the bit rate allocation for the consonant channel.

5. The stereo sound decoding method according to claim 4, wherein decoding the left and right channels of the stereo sound signal includes: Use factor β to recover the energy scaling factor and then use the energy scaling factor to rescale the decoded main channel; as well as The decoded left and right channels are generated by rescaling the decoded main channel and the decoded auxiliary channel.

6. The stereo sound decoding method according to claim 4, wherein the received encoding parameters include: Receive a bitstream from a stereo audio encoder and extract the encoding parameters from the bitstream.

7. A stereo sound decoding system, comprising: A component for receiving encoding parameters including encoding parameters for the main channel and encoding parameters for the auxiliary channel, wherein the encoding parameters for the main channel include the LP filter coefficients of the main channel. The decoder for the main channel in response to the main channel encoding parameters; and A decoder for the secondary channel using one of a plurality of encoding models, wherein (a) at least one of the encoding models uses the primary channel LP filter coefficients to decode the secondary channel, and (b) at least one of the encoding models uses primary channel encoding parameters other than the LP filter coefficients to decode the secondary channel.

8. The stereo sound decoding system of claim 7, wherein the auxiliary channel decoder comprises a first decoder using a universal coding model and a second decoder using one of a universal coding model, a silent coding model, and an inactive coding model.

9. The stereo sound decoding system of claim 7, wherein the auxiliary channel encoding parameters include information identifying one of the encoding models to be used when decoding the auxiliary channel, and wherein the stereo sound decoding system includes a determination module for instructing the first and second decoders on the encoding model to be used when decoding the auxiliary channel.

10. The stereo sound decoding system according to any one of claims 7 to 9, wherein: - The received encoded parameters include the factor β; - The decoder of the main channel uses the factor β as an indicator for the bit rate allocation of the main channel; and - The decoder for the consonant channel uses the factor β as an indicator of the bit rate allocation for the consonant channel.

11. The stereo sound decoding system according to claim 10, wherein decoding the left and right channels of the stereo sound signal includes: The component used to recover the energy scaling factor using factor β and to rescale the decoded main channel using the energy scaling factor; as well as A component used to generate the decoded left and right channels by using the rescaled decoded main channel and decoded auxiliary channel.

12. The stereo sound decoding system according to claim 10, wherein the encoding parameter receiving unit receives a bit stream from the stereo sound encoder and extracts the encoding parameters from the bit stream.

13. The stereo sound decoding system according to claim 12, wherein the encoding parameter receiving component includes a demultiplexer.