DEVICE, METHOD OF DECODING A SIGNAL, AND COMPUTER-READABLE MEMORY

The decoder addresses inefficiencies in audio encoding by determining ICBWE gain mapping parameters, enhancing encoding efficiency and audio quality through aligned and equalized audio channels, specifically in scenarios with temporal mismatches and spatial effects.

BR112019020643B1Active Publication Date: 2026-07-28QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112019020643
Authority / Receiving Office
BR · BR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-03-26
Filing Date
2018-03-27
Publication Date
2026-07-28
Estimated Expiration
2038-03-27

AI Technical Summary

Technical Problem

Existing audio encoding technologies face inefficiencies in handling temporal mismatches and spatial effects between multiple audio channels, leading to reduced coding gains and increased bit usage due to uncompensated temporal and phase shifts in stereo encoding.

Method used

The implementation of a decoder that determines inter-channel bandwidth extension (ICBWE) gain mapping parameters based on frequency domain gain parameters, allowing for efficient decoding of low-band and high-band audio signals by generating synthesized high-band signals using non-linear harmonic extensions and gain scaling operations, thereby reducing redundant parameter transmission.

Benefits of technology

This approach enhances encoding efficiency by aligning and equalizing audio channels, minimizing bit usage, and improving the quality of decoded audio signals, particularly in scenarios with temporal mismatches and spatial effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000081_0000
    Figure 00000081_0000
  • Figure 00000082_0000
    Figure 00000082_0000
  • Figure 00000083_0000
    Figure 00000083_0000
Patent Text Reader

Abstract

One method involves decoding a low-band center channel bitstream to generate a low-band center signal and a low-band center excitation signal. The method further involves decoding a high-band center channel bandwidth extension bitstream to generate a synthesized high-band center signal. The method also involves determining an inter-channel bandwidth extension (icbwe) gain mapping parameter corresponding to the synthesized high-band center signal. The icbwe gain mapping parameter is based on a selected frequency domain gain parameter that is extracted from a stereo downmix / upmix parameter bitstream. The method further involves performing a gain scaling operation on the synthesized high-band center signal based on the icbwe gain mapping parameter to generate a reference high-band channel and a target high-band channel.The method involves emitting a first audio channel and a second audio channel. The first audio channel is based on the reference high-band channel, and the second audio channel is based on the target high-band channel.
Need to check novelty before this filing date? Find Prior Art

Description

1 / 75 “DEVICE, METHOD OF DECODING A SIGNAL, AND COMPUTER-READABLE MEMORY I. Claim for Priority

[0001] This application claims the benefit of priority of the provisional patent application assigned to the same assignee US No. 62 / 482,150, filed on April 5, 2017, entitled “INTER-CHANNEL BANDWIDTH EXTENSION”, and non-provisional patent application US No. 15 / 935,952, filed on March 26, 2018, entitled “INTER-CHANNEL BANDWIDTH EXTENSION”, the contents of each of the aforementioned applications are expressly incorporated herein by reference in their entirety. II. Field

[0002] The present disclosure relates, in general, to the encoding of multiple audio signals. III. Description of the Related Technique

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there is now a variety of portable personal computing devices, including cordless phones such as cell phones and smartphones, tablets, and laptops that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Furthermore, many devices incorporate additional functionality such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Also, such devices can process executable instructions, including software applications, such as a [missing word - likely "application" or similar]. Petition 870250025693, dated 03 / 31 / 2025, page 6 / 182 2 / 75 web browser, which can be used to access the Internet. As such, these devices may include significant computing capabilities.

[0004] A computing device may include multiple microphones to receive audio channels. For example, a first microphone may receive a left audio channel, and a second microphone may receive a corresponding right audio channel. In stereo encoding, an encoder may transform the left audio channel and the corresponding right audio channel in a frequency domain to generate a left frequency domain channel and a right frequency domain channel, respectively. The encoder may downmix the frequency domain channels to generate a center channel. An inverse transform may be applied to the center channel to generate a time domain center channel, and a low-band encoder may encode a low-band portion of the center channel in the time domain to generate a low-band encoded center channel.A media channel bandwidth extension (BWE) encoder can generate center channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain formats, a gain frame, etc.) based on the time-domain center channel and an excitation of the encoded low-band center channel. The encoder can generate a bitstream that includes the encoded low-band center channel and the center channel BWE parameters.

[0005] The encoder can also extract stereo parameters (e.g., Discrete Fourier Transform (DFT) downmix parameters) from the channels in the domain. Petition 870250025693, dated 03 / 31 / 2025, page 7 / 182 3 / 75 of the frequency (e.g., the channel in the left frequency domain and the channel in the right frequency domain). Stereo parameters may include frequency domain gain parameters (e.g., side gains), interchannel phase difference (IPD) parameters, interchannel level differences (ILD), spread / diffusion gains, and interchannel BWE gain mapping (ICBWE) parameters. Stereo parameters may also include estimated interchannel time differences (ITD) based on time domain and / or frequency domain analysis of the left and right stereo channels. Stereo parameters may be embedded (e.g., included or encoded) in the bitstream, and the bitstream may be transmitted from the encoder to a decoder. IV. Summary

[0006] According to one implementation, a device includes a receiver configured to receive a bitstream from an encoder. The bitstream includes at least one low-band center channel bitstream, one high-band center channel bandwidth extension (BWE) bitstream, and one stereo downmix / upmix parameter bitstream. The device also includes a decoder configured to decode the low-band center channel bitstream to generate a low-band center signal and a low-band center excitation signal. The decoder is further configured to generate a non-linear harmonic extension of the low-band center excitation signal corresponding to a high-band BWE portion. The decoder is further configured to decode the center channel BWE bitstream of Petition 870250025693, dated 03 / 31 / 2025, page 8 / 182 4 / 75 high band to generate a synthesized high-band center signal based at least on the non-linear harmonic excitation signal and high-band center channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, and gain frame parameters). The decoder is also configured to determine an interchannel bandwidth extension (ICBWE) gain mapping parameter corresponding to the synthesized high-band center signal. The ICBWE gain mapping parameter is determined (e.g., predicted, derived, guided, or mapped) based on a frequency domain gain parameter (e.g., a group of sub-bands or frequency collectors corresponding to the high-band BWE portion) that is extracted from the stereo downmix / upmix parameter bitstream.For broadband content, the decoder is additionally configured to perform a gain scaling operation on the synthesized highband center signal based on the ICBWE gain mapping parameter to generate a reference highband channel and a target highband channel. The device also includes one or more speakers configured to output a first audio channel and a second audio channel. The first audio channel is based on the reference highband channel, and the second audio channel is based on the target highband channel.

[0007] According to another implementation, a method of decoding a signal involves receiving a bit stream from an encoder. The bit stream includes at least one low-bandwidth center-channel bit stream, one center-channel bandwidth extension bit stream Petition 870250025693, dated 03 / 31 / 2025, page 9 / 182 The method involves decoding the low-band center channel bitstream to generate a low-band center signal and a low-band center excitation signal. It also includes generating a non-linear harmonic extension of the low-band center excitation signal corresponding to a portion of the high-band BWE. The method further includes decoding the high-band center channel BWE bitstream to generate a synthesized high-band center signal based at least on the non-linear harmonic excitation signal and high-band center channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, and gain frame parameters). The method also includes determining an inter-channel bandwidth extension (ICBWE) gain mapping parameter corresponding to the synthesized high-band center signal.The ICBWE gain mapping parameter is determined (e.g., predicted, derived, guided, or mapped) based on a frequency domain gain parameter (e.g., a group of sub-bands or frequency collectors corresponding to the high-band BWE portion) that is extracted from the stereo downmix / upmix parameter bitstream. The method further includes performing a gain scaling operation on the synthesized high-band center signal based on the ICBWE gain mapping parameter to generate a reference high-band channel and a target high-band channel. The method also includes outputting a first audio channel and a second audio channel. The first audio channel is based on the reference high-band channel, and the second audio channel is based on the target channel. Petition 870250025693, dated 03 / 31 / 2025, page 10 / 182 6 / 75 high target band.

[0008] According to another implementation, a non-temporary computer-readable medium includes instructions for decoding a signal. The instructions, when executed by a processor within a decoder, cause the processor to perform operations including receiving a bitstream from an encoder. The bitstream includes at least one low-band center channel bitstream, one high-band center channel bandwidth extension (BWE) bitstream, and one stereo downmix / upmix parameter bitstream. The operations also include decoding the low-band center channel bitstream to generate a low-band center signal and a low-band center excitation signal. The operations also include generating a non-linear harmonic extension of the low-band center excitation signal corresponding to a high-band BWE portion.The operations also include decoding the high-band center-channel BWE bitstream to generate a synthesized high-band center signal based at least on the non-linear harmonic excitation signal and high-band center-channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, and gain frame parameters). The operations also include determining an inter-channel bandwidth extension (ICBWE) gain mapping parameter corresponding to the synthesized high-band center signal. The ICBWE gain mapping parameter is determined (e.g., predicted, derived, guided, or mapped) based on a frequency domain gain parameter (e.g., a group of sub-bands or frequency collectors corresponding to the BWE portion). Petition 870250025693, dated 03 / 31 / 2025, page 11 / 182 The operation involves a 7 / 75 high-band signal) extracted from the stereo downmix / upmix parameter bitstream. It further includes performing a gain scaling operation on the synthesized high-band center signal based on the ICBWE gain mapping parameter to generate a reference high-band channel and a target high-band channel. The operation also includes outputting a first audio channel and a second audio channel. The first audio channel is based on the reference high-band channel, and the second audio channel is based on the target high-band channel.

[0009] According to another implementation, an apparatus includes means for receiving a bitstream from an encoder. The bitstream includes at least one low-band center channel bitstream, one high-band center channel bandwidth extension (BWE) bitstream, and one stereo downmix / upmix parameter bitstream. The apparatus also includes means for decoding the low-band center channel bitstream to generate a low-band center signal and a low-band center excitation signal. The apparatus also includes means for generating a non-linear harmonic extension of the low-band center excitation signal corresponding to a high-band BWE portion.The apparatus also includes means for decoding the high-band center-channel BWE bitstream to generate a synthesized high-band center signal based at least on the non-linear harmonic excitation signal and high-band center-channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, and gain frame parameters). The apparatus also includes means for determining a mapping parameter. Petition 870250025693, dated 03 / 31 / 2025, page 12 / 182 8 / 75 Interchannel bandwidth extension (ICBWE) gain corresponding to the synthesized high-band center signal. The ICBWE gain mapping parameter is determined (e.g., predicted, derived, guided, or mapped) based on a frequency domain gain parameter (e.g., a group of sub-bands or frequency collectors corresponding to the high-band BWE portion) that is extracted from the stereo downmix / upmix parameter bitstream. The device also includes means for performing a gain scaling operation on the synthesized high-band center signal based on the gain mapping parameter. ICBWE is used to generate a reference high-band channel and a target high-band channel. The device also includes means for outputting a first audio channel and a second audio channel. The first audio channel is based on the reference high-band channel, and the second audio channel is based on the target high-band channel.

[0010] Other implementations, advantages and features of the present disclosure will become apparent after analysis of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description and the Claims. V. Brief Description of the Drawings

[0011] Figure 1 is a block diagram of a specific illustrative example of a system that includes an operable decoder to determine the interchannel bandwidth extension (ICBWE) mapping parameters based on a frequency domain gain parameter transmitted from an encoder;

[0012] Figure 2 is a diagram that illustrates the Petition 870250025693, dated 03 / 31 / 2025, page 13 / 182 9 / 75 encoder of Figure 1;

[0013] Figure 3 is a diagram illustrating the decoder in Figure 1;

[0014] Figure 4 is a flowchart illustrating a specific method for determining ICBWE mapping parameters based on a frequency domain gain parameter transmitted from an encoder;

[0015] Figure 5 is a block diagram of a specific illustrative example of a device that is operable to determine ICBWE mapping parameters based on a frequency domain gain parameter transmitted from an encoder; and

[0016] Figure 6 is a block diagram of a base station that is operable to determine ICBWE mapping parameters based on a frequency domain gain parameter transmitted from an encoder. VI. Detailed Description

[0017] Specific aspects of the present disclosure are described below with reference to the drawings. In the description, common features are referred to as common numerical references. As used in this document, various terminologies are used only for the purpose of describing specific implementations and are not intended to be limiting of implementations. For example, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, except where the context indicates otherwise. It may be further understood that the terms “comprises” and “comprising” may be interchangeable with “includes” or “including.” Additionally, it will be understood that the term “in which” may be interchangeable Petition 870250025693, dated 03 / 31 / 2025, p. 14 / 182 10 / 75 with “where”. As used in this document, an ordinal term (e.g., “first”, “second”, “third”, etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element in relation to another element, but instead merely distinguishes the element from another element that has the same name (but for use of the ordinal term). As used in this document, the term “set” refers to one or more of a specific element, and the term “plurality” refers to a multiple (e.g., two or more) of a specific element.

[0018] In the present disclosure, terms such as “determine”, “calculate”, “displace”, “adjust”, etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not interpreted as limiting and other techniques may be used to perform similar operations. Additionally, as referred to herein, “generate”, “calculate”, “use”, “select”, “access”, “identify”, and “determine” may be used interchangeably. For example, “generate”, “calculate”, or “determine” a parameter (or signal) may refer to actively generating, calculating, or determining the parameter (or signal) or may refer to using, selecting, or accessing the parameter (or signal) that has already been generated, such as by another component or device.

[0019] Operable systems and devices for encoding multiple audio signals are disclosed. A device may include an encoder configured to encode multiple audio signals. The multiple audio signals may be captured simultaneously using Petition 870250025693, dated 03 / 31 / 2025, page 15 / 182 11 / 75 multiple recording devices, for example, multiple microphones. In some examples, multiple audio signals (or multichannel audio) can be synthetically (e.g., artificially) generated by multiplexing several audio channels that are recorded simultaneously or at different times. As illustrative examples, simultaneous recording or multiplexing of audio channels can result in a 2-channel configuration (i.e., Stereo: Left and Right), a 5.1 channel configuration (Left, Right, Center, Left Surround, Right Surround, and low-frequency emphasis (LFE) channels), a 7.1 channel configuration, a 7.1+4 channel configuration, a 22.2 channel configuration, or an N channel configuration.

[0020] Audio capture devices in teleconferencing rooms (or telepresence rooms) may include multiple microphones that acquire spatial audio. Spatial audio may include speech as well as background audio that is encoded and transmitted. Speech / audio from a given source (e.g., a speaker) may reach the multiple microphones at different times depending on how the microphones are arranged as well as where the source (e.g., the speaker) is located relative to the microphones and room dimensions. For example, a sound source (e.g., a speaker) may be closer to a first microphone associated with the device than to a second microphone associated with the device. In this way, a sound emitted from the sound source may reach the first microphone before the second microphone. The device may receive a first audio signal through the first microphone and may receive a second audio signal through the second microphone. Petition 870250025693, dated 03 / 31 / 2025, p. 16 / 182 12 / 75

[0021] Mid-side (MS) coding and parametric stereo (PS) coding are stereo coding techniques that can provide improved efficiency over dual-mono coding techniques. In dual-mono coding, the Left (L) channel (or signal) and the Right (R) channel (or signal) are independently coded without making use of inter-channel correlation. MS coding reduces redundancy between a pair of correlated L / R channels by transforming the Left and Right channels into a sum channel and a difference channel (e.g., a side channel) before coding. The sum signal and the difference signal are waveform-coded or model-coded in MS coding. Relatively more bits are spent on the sum signal than on the side signal. PS coding reduces redundancy in each sub-band or frequency band by transforming the L / R signals into a sum signal and a set of side parameters.Side parameters can indicate an interchannel intensity difference (IID), an interchannel phase difference (IPD), an interchannel time difference (ITD), lateral or residual prediction gains, etc. The sum signal is waveform-encoded and transmitted along with the side parameters. In a hybrid system, the side channel can be waveform-encoded in the lower bands (e.g., less than 2 kilohertz (kHz)) and PS-encoded in the upper bands (e.g., greater than or equal to 2 kHz) where interchannel phase preservation is perceptually less critical. In some implementations, PS encoding can be used in the lower bands as well to reduce interchannel redundancy before waveform encoding. Petition 870250025693, dated 03 / 31 / 2025, page 17 / 182 13 / 75

[0022] MS coding and PS coding can be performed in both the frequency domain and the sub-band domain. In some examples, the Left and Right channels may be uncorrelated. For example, the Left and Right channels may include uncorrelated synthetic signals. When the Left and Right channels are uncorrelated, the coding efficiency of MS coding, PS coding, or both, can approach the coding efficiency of dual-mono coding.

[0023] Depending on a recording setup, there may be a temporal mismatch between a Left channel and a Right channel, as well as other spatial effects such as echo and room reverb. If the temporal and phase mismatch between the channels is not compensated for, the sum channel and the difference channel may contain comparable energies that reduce the coding gains associated with MS or PS techniques. The reduction in coding gains may be based on the amount of temporal (or phase) shift. The comparable energies of the sum signal and the difference signal may limit the use of MS coding in certain frames where the channels are temporally shifted, but are highly correlated. In stereo coding, a center channel (e.g., a sum channel) and a side channel (e.g., a difference channel) can be generated based on the following formula: M= (L+R) / 2, S= (LR) / 2, Formula 1

[0024] where M corresponds to the central channel, S corresponds to the lateral channel, L corresponds to the left channel Petition 870250025693, dated 03 / 31 / 2025, p. 18 / 182 14 / 75 and R corresponds to the Right channel.

[0025] In some cases, the center channel and the side channel can be generated based on the following formula: M=c (L+R), S= c (LR), Formula 2

[0026] where c corresponds to a complex value that is frequency-dependent. Generating the center and side channels based on Formula 1 or Formula 2 can be called implementing a “down-mixing” algorithm. A reverse process of generating the left and side channels Directing traffic from the center channel and side channel, based on Formula 1 or Formula 2, can be called the implementation of an "up-mixing" algorithm.

[0027] In some cases, the central channel may be based on other formulas such as: M = (L+gDR) / 2, or Formula 3 M = g1L + g2R Formula 4

[0028] where g1 + g2 = 1.0, and where gD is a gain parameter. In other examples, the down-mix can be performed in bands, where mid(b) = c1L(b) + C2R(b), where c1 and C2 are complex numbers, where side(b) = C3L(b) C4R(b), where C3 and C4 are complex numbers.

[0029] An ad-hoc approach used to choose between MS encoding or dual-mono encoding for a specific frame might involve generating a center channel and a side channel, calculating the energies of the center and side channels, and determining whether to perform MS encoding based on the energies. For example, MS encoding might be performed in response to the determination that the ratio of side channel to center channel energies is less than a threshold. To illustrate, if a Right channel is shifted Petition 870250025693, dated 03 / 31 / 2025, p. 19 / 182 15 / 75 at least once (e.g., about 0.001 seconds or 48 samples at 48 kHz), a first center channel energy (corresponding to the sum of the left and right signals) can be comparable to a second side channel energy (corresponding to the difference between the left and right signals) for voice-over frames. When the first energy is comparable to the second energy, a larger number of bits can be used to encode the side channel, thus reducing the encoding efficiency of MS encoding compared to dual-mono encoding. In this way, dual-mono encoding can be used when the first energy is comparable to the second energy (e.g., when the ratio of the first energy to the second energy is greater than or equal to a threshold).In an alternative approach, the decision between MS encoding and dual-mono encoding for a specific frame can be made based on a comparison of a threshold and normalized cross-correlation values ​​of the Left and Right channels.

[0030] In some examples, the encoder may determine a mismatch value indicating an amount of temporal mismatch between the first audio signal and the second audio signal. As used in this document, a “temporal offset value,” an “offset value,” and a “mismatch value” may be used interchangeably. For example, the encoder may determine a temporal offset value indicating an offset (e.g., temporal mismatch) of the first audio signal relative to the second audio signal. The offset value Petition 870250025693, dated 03 / 31 / 2025, page 20 / 182 16 / 75 may correspond to a time delay between the reception of the first audio signal at the first microphone and the reception of the second audio signal at the second microphone. Additionally, the encoder may determine the offset value on a frame-by-frame basis, for example, based on each 20-millisecond (ms) speech / audio frame. For instance, the offset value may correspond to the amount of time that a second frame of the second audio signal is delayed relative to a first frame of the first audio signal. Alternatively, the offset value may correspond to the amount of time that the first frame of the first audio signal is delayed relative to the second frame of the second audio signal.

[0031] When the sound source is closer to the first microphone than to the second microphone, the frames of the second audio signal may be delayed relative to frames of the first audio signal. In this case, the first audio signal may be called the “reference audio signal” or “reference channel” and the delayed second audio signal may be called the “target audio signal” or “target channel”. Alternatively, when the sound source is closer to the second microphone than to the first microphone, the frames of the first audio signal may be delayed relative to frames of the second audio signal. In this case, the second audio signal may be called the “reference audio signal” or “reference channel” and the delayed first audio signal may be called the “target audio signal” or “target channel”.

[0032] Depending on where the sound sources (e.g., speakers) are located in a room of Petition 870250025693, dated 03 / 31 / 2025, p. 21 / 182 17 / 75 conference or telepresence or how the position of the sound source (e.g., speaker) changes relative to the microphones, the reference channel and the target channel may change from one frame to another; similarly, the time mismatch value may also change from one frame to another. However, in some implementations, the offset value may always be positive to indicate an amount of delay of the “target” channel relative to the “reference” channel. Furthermore, the offset value may correspond to a “non-causal offset” value whereby the delayed target channel is “set back” in time so that the target channel is aligned (e.g., maximally aligned) with the “reference” channel in the encoder. The down-mix algorithm to determine the center channel and the side channel may be performed on the non-causal offset reference channel and target channel.

[0033] The encoder can determine the offset value based on the reference audio channel and a plurality of offset values ​​applied to the target audio channel. For example, a first frame from the reference audio channel, X, can be received at a first time (mi). A specific first frame from the target audio channel, Y, can be received in a second time (ni) corresponding to a first shift value, for example, shift1 = ni - mi. Furthermore, a second frame from the reference audio channel can be received in a third time (rm). A second specific frame from the target audio channel can be received in a fourth time (m) corresponding to a second shift value, for example, shift2 = - rri2.

[0034] The device can perform a Petition 870250025693, dated 03 / 31 / 2025, page 22 / 182 18 / 75 framing or a buffering algorithm to generate a frame (e.g., 20 ms samples) at a first sampling rate (e.g., 32 kHz sampling rate (i.e., 640 samples per frame)). The encoder can, in response to the determination that a first frame of the first audio signal and a second frame of the second audio signal arrive at the same time at the device, estimate a shift value (e.g., shift1) as zero samples. A Left channel (e.g., corresponding to the first audio signal) and a Right channel (e.g., corresponding to the second audio signal) can be temporarily aligned. In some cases, the Left and Right channels, even when aligned, may differ in power due to various reasons (e.g., microphone calibration).

[0035] In some instances, the Left and Right channels may be temporarily misaligned due to various reasons (for example, a sound source, such as a speaker, may be closer to one microphone than the other, and the two microphones may be larger than a threshold (e.g., 1 to 20 centimeters) apart). The location of the sound source relative to the microphones may introduce different delays in the first and second channels. In addition, there may be a difference in gain, a difference in power, or a difference in level between the first and second channels.

[0036] In some examples, when there are more than two channels, a reference channel is initially selected based on the channel levels or energies and subsequently refined based on the values ​​of Petition 870250025693, dated 03 / 31 / 2025, p. 23 / 182 19 / 75 time mismatch between different pairs of channels, for example, t1(ref, ch2), t2(ref, ch3), t3(ref, ch4),... t3(ref, chN), where ch1 is the initial ref channel and t1(.), t2(.), etc. are the functions to estimate the mismatch values. If all time mismatch values ​​are positive, then ch1 is treated as the reference channel. If any of the mismatch values ​​is a negative value, the reference channel is reconfigured to the channel that was associated with a mismatch value that resulted in a negative value, and the above process continues until the best selection (i.e., based maximally on the maximum number of side channel decorrelations) of the reference channel is obtained. A hysteresis can be used to overcome any sudden variations in the selection of the reference channel.

[0037] In some examples, the arrival time of audio signals at microphones from multiple sound sources (e.g., announcers) may vary when multiple announcers are speaking alternately (e.g., without overlapping). In such a case, the encoder can dynamically adjust a time shift value based on the announcer to identify the reference channel. In some other examples, multiple announcers may be speaking simultaneously, which may result in variable time shift values ​​depending on who is the loudest announcer, closest to the microphone, etc. In such a case, the identification of reference and target channels may be based on the variable time shift values ​​in the current frame, the estimated time mismatch values ​​in previous frames, and the energy (or temporal evolution) of the first Petition 870250025693, dated 03 / 31 / 2025, page 24 / 182 20 / 75 and the second audio signals.

[0038] In some examples, the first audio signal and the second audio signal may be synthesized or artificially generated when the two signals potentially show less (e.g., no) correlation. It should be understood that the examples described in this document are illustrative and may be instructive in determining a relationship between the first audio signal and the second audio signal in similar or different situations.

[0039] The encoder can generate comparison values ​​(e.g., difference values ​​or cross-correlation values) based on a comparison of a first frame of the first audio signal and a plurality of frames of the second audio signal. Each frame of the plurality of frames can correspond to a specific offset value. The encoder can generate an estimated first offset value based on the comparison values. For example, the estimated first offset value might correspond to a comparison value indicating greater temporal similarity (or less difference) between the first frame of the first audio signal and a corresponding first frame of the second audio signal.

[0040] The encoder can determine the final offset value by refining, in multiple stages, a series of estimated offset values. For example, the encoder can first estimate a “provisional” offset value based on comparison values ​​generated from pre-processed and resampled stereo versions of the first audio signal and the second audio signal. The encoder can generate interpolated comparison values. Petition 870250025693, dated 03 / 31 / 2025, page 25 / 182 21 / 75 associated with displacement values ​​close to the estimated “provisional” displacement value. The coder can determine a second estimated “interpolated” displacement value based on the interpolated comparison values. For example, the second estimated “interpolated” displacement value might correspond to a specific interpolated comparison value that indicates greater temporal similarity (or less difference) than the remaining interpolated comparison values ​​and the first estimated “provisional” displacement value.If the second estimated “interpolated” offset value of the current frame (e.g., the first frame of the first audio signal) differs from a final offset value of a previous frame (e.g., a frame of the first audio signal that precedes the first frame), then the “interpolated” offset value of the current frame is further “altered” to enhance the temporal similarity between the first audio signal and the second shifted audio signal. In particular, a third estimated “altered” offset value may correspond to a more accurate measurement of temporal similarity by searching over the second estimated “interpolated” offset value of the current frame and the final estimated offset value of the previous frame.The third estimated “changed” value is further conditioned to estimate the final displacement value by limiting any spurious changes in the displacement value between frames and further controlled to not change from a negative displacement value to a positive displacement value (or vice versa) in two successive (or consecutive) frames as described in this document. Petition 870250025693, dated 03 / 31 / 2025, page 26 / 182 22 / 75

[0041] In some instances, the encoder may refrain from switching between a positive offset value and a negative offset value or vice versa in consecutive frames or in adjacent frames. For example, the encoder may set the final offset value to a specific value (e.g., 0) not indicating a time shift based on the estimated interpolated or “altered” offset value of the first frame and a corresponding estimated interpolated or altered offset value or final offset value in a specific frame preceding the first frame.To illustrate, the encoder can adjust the final shift value of the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0, in response to the determination that one of the estimated provisional, interpolated, or altered values ​​of the current frame is positive and the other of the estimated provisional, interpolated, altered, or final values ​​of the previous frame (e.g., the frame preceding the first frame) is negative. Alternatively, the encoder can also adjust the final shift value of the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0, in response to the determination that one of the estimated provisional, interpolated, or altered values ​​of the current frame is negative and the other of the estimated provisional, interpolated, altered, or final values ​​of the previous frame (e.g., the frame preceding the first frame) is positive.

[0042] It should be noted that in some Petition 870250025693, dated 03 / 31 / 2025, page 27 / 182 In 23 / 75 implementations, the final shift value estimate can be performed in the transform domain, where interchannel cross-correlations can be estimated in the frequency domain. As an example, the final shift value estimate can be largely based on the Generalized Cross-Correlation - Phase Transform (GCC-PHAT) algorithm.

[0043] The encoder can select a frame from the first audio signal or the second audio signal as a “reference” or “target” based on the offset value. For example, in response to the determination that the final offset value is positive, the encoder can generate a reference channel or signal indicator that has a first value (e.g., 0) indicating that the first audio signal is a “reference” channel and that the second audio signal is the “target” channel. Alternatively, in response to the determination that the final offset value is negative, the encoder can generate the reference channel or signal indicator that has a second value (e.g., 1) indicating that the second audio signal is the “reference” channel and that the first audio signal is the “target” channel.

[0044] The encoder can estimate a relative gain (e.g., a relative gain parameter) associated with the non-causally offset reference channel and target channel. For example, in response to the determination that the final offset value is positive, the encoder can estimate a gain value to normalize or equalize the energy or power levels of the first audio signal relative to the second audio signal that is offset by the non-causally offset value (e.g., an absolute value). Petition 870250025693, dated 03 / 31 / 2025, p. 28 / 182 24 / 75 of the final offset value). Alternatively, in response to the determination that the final offset value is negative, the encoder may estimate a gain value to normalize or equalize the power or amplitude levels of the first audio signal relative to the second audio signal. In some examples, the encoder may estimate a gain value to normalize or equalize the amplitude or power levels of the “reference” channel relative to the non-causally offset “target” channel. In other examples, the encoder may estimate the gain value (e.g., a relative gain value) based on the reference channel relative to the target channel (e.g., the non-offset target channel).

[0045] The encoder can generate at least one encoded signal (e.g., a center channel, a side channel, or both) based on the reference channel, the target channel, the non-causal offset value, and the relative gain parameter. In other implementations, the encoder can generate at least one encoded signal (e.g., a center channel, a side channel, or both) based on the reference channel and the target channel adjusted for temporal mismatch. The side channel can correspond to a difference between the first samples of the first frame of the first audio signal and selected samples from a selected frame of the second audio signal. The encoder can select the selected frame based on the final offset value. Fewer bits can be used to encode the side channel signal due to the reduced difference between the first samples and the selected samples compared to other samples of the second. Petition 870250025693, dated 03 / 31 / 2025, page 29 / 182 25 / 75 audio signal corresponding to a frame of the second audio signal that is received by the device at the same time as the first frame. A device transmitter can transmit at least one encoded signal, the non-causal offset value, the relative gain parameter, the reference channel or signal indicator, or a combination thereof.

[0046] The encoder can generate at least one encoded signal (e.g., a center channel, a side channel, or both) based on the reference channel, the target channel, the non-causal offset value, the relative gain parameter, low-band parameters of a specific frame of the first audio signal, high-band parameters of the specific frame, or a combination thereof. The specific frame may precede the first frame. Certain low-band parameters, high-band parameters, or a combination thereof, from one or more previous frames may be used to encode a center channel, a side channel, or both, of the first frame. The encoding of the center channel, the side channel, or both, based on the low-band parameters, the high-band parameters, or a combination thereof, may include estimates of the non-causal offset value and the inter-channel relative gain parameter.Low-band parameters, high-band parameters, or a combination thereof, may include a pitch parameter, a voice parameter, an encoder type parameter, a low-band power parameter, a high-band power parameter, a slope parameter, a pitch gain parameter, an FCB gain parameter, and a mode parameter. Petition 870250025693, dated 03 / 31 / 2025, page 30 / 182 26 / 75 encoding, a voice activity parameter, a noise estimation parameter, a signal-to-noise ratio parameter, a formant shaping parameter, a speech / music decision parameter, the non-causal offset, the interchannel gain parameter, or a combination thereof. A device transmitter may transmit at least one encoded signal, the non-causal offset value, the relative gain parameter, the reference channel (or signal) indicator, or a combination thereof.

[0047] According to some implementations, the encoder can transform a left audio channel and a corresponding right audio channel in the frequency domain to generate a left frequency domain channel and a right frequency domain channel, respectively. The encoder can downmix the frequency domain channels to generate a center channel. An inverse transform can be applied to the center channel to generate a time domain center channel, and a low-band encoder can encode a low-band portion of the center channel in the time domain to generate a coded low-band center channel. A center channel bandwidth extension (BWE) encoder can generate center channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, a gain frame, etc.).In some implementations, the center-channel BWE encoder generates the center-channel BWE parameters based on the time-domain center channel and an excitation of the encoded low-band center channel. The encoder can generate a bitstream that includes the encoded low-band center channel and the... Petition 870250025693, dated 03 / 31 / 2025, page 31 / 182 27 / 75 center channel BWE parameters.

[0048] The encoder can also extract stereo parameters (e.g., Discrete Fourier Transform (DFT) downmix parameters) from the frequency domain channels (e.g., the left frequency domain channel and the right frequency domain channel). Stereo parameters may include frequency domain gain parameters (e.g., side gains or interchannel level differences (ILDs)), interchannel phase difference (IPD) parameters, stereo fill gains, etc. Stereo parameters may be inserted (e.g., included or encoded) into the bitstream, and the bitstream may be transmitted from the encoder to a decoder. Depending on one implementation, stereo parameters may include interchannel BWE gain mapping (ICBWE) parameters. However, ICBWE gain mapping parameters may be somewhat “redundant” with respect to other stereo parameters.Therefore, to reduce encoding complexity and redundant transmission, ICBWE gain mapping parameters may not be extracted from the frequency domain channels. For example, the encoder may ignore the determination of ICBWE gain parameters for the frequency domain channels.

[0049] Upon receiving the bit stream from the encoder, the decoder can decode the encoded low-band center channel to generate a low-band center signal and a low-band center excitation signal. The center channel BWE parameters (received from the encoder) can be decoded using the excitation of Petition 870250025693, dated 03 / 31 / 2025, page 32 / 182 A 28 / 75 low-band center channel is used to generate a synthesized high-band center signal. A left high-band channel and a right high-band channel can be generated by applying ICBWE gain mapping parameters to the synthesized high-band center signal. However, because ICBWE gain mapping parameters are not included as part of the bitstream, the decoder can generate an ICBWE gain mapping parameter based on the frequency domain gain parameters (e.g., side gains or ILDs). The decoder can also generate ICBWE gain mapping parameters based on the high-band center synthesis signal (or excitation), the low-band center synthesis signal (or excitation), and the low-band side synthesis signal (e.g., residual prediction).

[0050] For example, the decoder can extract the frequency domain gain parameters from the bitstream and select a frequency domain gain parameter that is associated with a frequency range of the synthesized high-band center signal. To illustrate, for wideband encoding, the synthesized high-band center signal might have a frequency range between 6.4 kilohertz (kHz) and 8 kHz. If a frequency domain gain parameter is associated with a frequency range between 5.2 kHz and 8.56 kHz, the specific frequency domain gain parameter can be selected to generate the ICBWE gain mapping parameter. In another example, if one or more groups of frequency domain gain parameters are associated with one or more sets of frequency ranges, Petition 870250025693, dated 03 / 31 / 2025, page 33 / 182 29 / 75 for example, 6.0 to 7.0 kHz, 7.0 to 8.0 kHz, then one or more groups of stereo downmix / upmix gain parameters are selected to generate the ICBWE gain mapping parameter. According to one implementation, the ICBWE gain mapping parameter (gsMapping) can be determined based on the selected frequency domain gain parameter (sidegain) using the following example: ICBWE gain Mapping parameter, gsMapping =(1— sidegairi)

[0051] Once the ICBWE gain mapping parameter is determined (e.g., extracted), the left highband channel and the right highband channel can be synthesized using a gain scaling operation. For example, the synthesized highband center signal can be scaled by the ICBWE gain mapping parameter to generate the target highband channel, and the synthesized highband center signal can be scaled by a modified ICBWE gain mapping parameter (by Ί2 — gsMapping or J2 — gsMapping2,Ί, example,vb &) to generate the reference highband channel.

[0052] A left lowband channel and a right lowband channel can be generated based on an upmix operation associated with a frequency domain version of the lowband center signal. For example, the lowband center signal can be converted to the frequency domain, stereo parameters can be used to perform the upmix of the frequency domain version of the lowband center signal to generate left and right frequency domain lowband channels, and inverse transform operations can be performed on the left and right frequency domain lowband channels to Petition 870250025693, dated 03 / 31 / 2025, page 34 / 182 30 / 75 generates the left low-band channel and the right low-band channel, respectively. The left low-band channel can be combined with the left high-band channel to generate a left channel that is substantially similar to the left audio channel, and the right low-band channel can be combined with the right high-band channel to generate a right channel (which is substantially similar to the right audio channel).

[0053] In this way, the coding complexity and transmission bandwidth can be reduced by omitting the extraction and transmission of ICBWE gain mapping parameters in the encoder depending on the input content bandwidth. For example, ICBWE gain mapping parameters may not be transmitted for WB multichannel coding, however, they are transmitted for super-wideband or full-band multichannel coding. In particular, ICBWE gain mapping parameters can be generated in the decoder for wideband signals based on other stereo parameters (e.g., frequency domain gain parameters) included in the bitstream.In other implementations, the ICBWE gain mapping parameters can also be generated based on the high-band center synthesis signal (i.e., BWE), the low-band center synthesis signal (or excitation), and the low-band side synthesis signal (e.g., residual prediction).

[0054] With reference to Figure 1, a specific illustrative example of a system is revealed and generally designated 100. System 100 includes a first device 104 communicatively coupled, through a network 120, to Petition 870250025693, dated 03 / 31 / 2025, page 35 / 182 31 / 75 a second device 106. The network 120 may include one or more wireless networks, one or more wired networks, or a combination thereof.

[0055] The first device 104 may include an encoder 114, a transmitter 110, one or more input interfaces 112, or a combination thereof. A first input interface of the input interfaces 112 may be coupled to a first microphone 146. A second input interface of the input interface(s) 112 may be coupled to a second microphone 148. The first device 104 may also include a memory 153 configured to store analysis data 191. The second device 106 may include a decoder 118. The decoder 118 may include an extension gain mapping parameter generator (ICBWE) 322. The second device 106 may be coupled to a first loudspeaker 142, a second loudspeaker 144, or both.

[0056] During operation, the first device 104 can receive a first audio channel 130 through the first input interface of the first microphone 146 and can receive a second audio channel 132 through the second input interface of the second microphone 148. The first audio channel 130 can correspond to either a right-channel signal or a left-channel signal. The second audio channel 132 can correspond to either the right-channel signal or the left-channel signal. For ease of description and illustration, except where otherwise specified, the first audio channel 130 corresponds to the left audio channel, and the second audio channel 132 corresponds to the right audio channel. A source Petition 870250025693, dated 03 / 31 / 2025, page 36 / 182 32 / 75 sound source 152 (e.g., a user, a loudspeaker, ambient noise, a musical instrument, etc.) may be closer to the first microphone 146 than to the second microphone 148. Consequently, an audio signal from the sound source 152 may be received at the input interface(s) 112 via the first microphone 146 before it is received via the second microphone 148. This natural delay in acquiring a multichannel signal through multiple microphones may introduce a time shift between the first audio channel 130 and the second audio channel 132.

[0057] Encoder 114 can be configured to determine an offset value (e.g., a final offset value 116) indicating a time offset between audio channels 130, 132. The final offset value 116 can be stored in memory 153 as analysis data 191 and encoded into a stereo downmix / upmix parameter bitstream 290 as a stereo parameter. Encoder 114 can also be configured to transform audio channels 130, 132 into the frequency domain to generate frequency domain audio channels. The frequency domain audio channels can be downmixed to generate a center channel, and a low-band portion of a time domain version of the center channel can be encoded into a low-band center channel bitstream 292.The 114 encoder can also generate center-channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain shapes, a gain frame, etc.) based on the time-domain center channel and an encoded low-band center channel excitation. The 114 encoder can encode the... Petition 870250025693, dated 03 / 31 / 2025, page 37 / 182 33 / 75 center channel BWE parameters as a bit stream of High-band 294 center channel BWE.

[0058] Encoder 114 can also extract stereo parameters (e.g., Discrete Fourier Transform (DFT) downmix parameters) from the audio channels in the frequency domain. Stereo parameters can include frequency domain gain parameters (e.g., side gains), interchannel phase difference (IPD) parameters, stereo fill gains, etc. Stereo parameters can be inserted into the stereo downmix / upmix parameter bitstream 290. Because of this, ICBWE gain mapping parameters can be determined or estimated using the other stereo parameters; ICBWE gain mapping parameters may not be extracted from the audio channels in the frequency domain to reduce encoding complexity and redundant transmission.The transmitter can transmit the stereo downmix / upmix parameter bitstream 290, the low-band center channel bitstream 292, and the high-band center channel BWE bitstream 294 to the second device 106 via the network. 120. The operations associated with encoder 114 are described in more detail in relation to Figure 2.

[0059] The 118 decoder can perform decoding operations based on the stereo downmix / upmix parameter bitstream 290, the low-band center channel bitstream 292, and the high-band center channel BWE bitstream 294. The 118 decoder can decode the low-band center channel bitstream. 292 to generate a low-band center signal and a low-band center excitation signal. The BWE bit stream of Petition 870250025693, dated 03 / 31 / 2025, page 38 / 182 34 / 75 high-band center channel 294 can be decoded using the low-band center excitation signal to generate a synthesized high-band center signal. A left high-band channel and a right high-band channel can be generated by applying the ICBWE gain mapping parameters to the synthesized high-band center signal. However, because the ICBWE gain mapping parameters are not included as part of the bitstream, the decoder 118 can generate an ICBWE gain mapping parameter based on the frequency domain gain parameters associated with the stereo downmix / upmix parameter bitstream 290.

[0060] For example, the decoder may include a generator of ICBWE 322 spatial gain mapping parameters configured to extract the frequency domain gain parameters from the stereo downmix / upmix parameter bitstream 290 and configured to select a frequency domain gain parameter that is associated with a frequency range of the synthesized high-band center signal. To illustrate, for wideband encoding, the synthesized high-band center signal may have a frequency range between 6.4 kilohertz (kHz) and 8 kHz. If a frequency domain gain parameter is associated with a frequency range between 5.2 kHz and 8.56 kHz, the specific frequency domain gain parameter can be selected to generate the ICBWE gain mapping parameter.According to one implementation, the ICBWE gain mapping parameter (gsMapping) can be determined based on the selected frequency domain gain parameter (sidegain). Petition 870250025693, dated 03 / 31 / 2025, page 39 / 182 35 / 75 using the following equation: gsMapplng =T1 - sidegain

[0061] Once the ICBWE gain mapping parameter is determined, the left high-band channel and the right high-band channel can be synthesized using a gain scaling operation. A left low-band channel and a right low-band channel can be generated based on an upmix operation associated with a frequency-domain version of the low-band center signal. The left low-band channel can be combined with the left high-band channel to generate a first output channel 126 (e.g., a left channel) that is substantially similar to the first audio channel 130, and the right low-band channel can be combined with the right high-band channel to generate a second output channel 128 (e.g., a right channel) that is substantially similar to the second audio channel 132. The first speaker 142 can output the first output channel 126, and the second speaker 144 can output the second output channel 128.The operations associated with decoder 118 are described in more detail in relation to Figure 3.

[0062] In this way, the coding complexity and transmission bandwidth can be reduced by omitting the extraction and transmission of ICBWE gain mapping parameters in the encoder. The ICBWE gain mapping parameters can be generated in the decoder based on other stereo parameters (e.g., frequency domain gain parameters). Petition 870250025693, dated 03 / 31 / 2025, page 40 / 182 36 / 75 included in the bitstream.

[0063] With reference to Figure 2, a specific implementation of encoder 114 is shown. Encoder 114 includes a transform unit 202, a transform unit 204, a stereo track estimator 206, a center channel generator 208, an inverse transform unit 210, a center channel encoder 212, and a center channel BWE encoder 214.

[0064] The first audio channel 130 (e.g., the left channel) can be supplied to transform unit 202, and the second audio channel 132 (e.g., the right channel) can be supplied to transform unit 204. Transform unit 202 can be configured to perform a windowing operation and a transform operation on the first audio channel 130 to generate a first frequency-domain audio channel Lfr(b) 252, and transform unit 204 can be configured to perform a windowing operation and a transform operation on the second audio channel 132 to generate a second frequency-domain audio channel Rfr(b) 254. For example, transform units 202, 204 can apply Discrete Fourier Transform (DFT) operations, Fast Fourier Transform (FFT) operations, MDCT operations, etc., on audio channels 130, 132, respectively.According to some implementations, Quadrature Mirrored Filter Bank (QMF) operations can be used to divide audio channel 130, 132 into multiple sub-bands. The first audio channel in the frequency domain 252 is provided to the stereo track estimator 206 and the center channel generator 208. The second channel of... Petition 870250025693, dated 03 / 31 / 2025, page 41 / 182 37 / 75 audio in the frequency domain 254 is also provided to the stereo track estimator 206 and the center channel generator 208.

[0065] The stereo track estimator 206 can be configured to extract (e.g., generate) stereo tracks from the frequency domain audio channels 252, 254 to generate the stereo downmix / upmix parameter bitstream 290. Non-limiting examples of stereo tracks (e.g., DFT downmix parameters) encoded in the stereo downmix / upmix parameter bitstream 290 may include frequency domain gain parameters (e.g., side gains), interchannel phase difference (IPD) parameters, stereo fill or residual prediction gains, etc. Depending on the implementation, the stereo tracks may include ICBWE gain mapping parameters. However, the ICBWE gain mapping parameters may be determined or estimated based on the other stereo tracks.Thus, to reduce encoding complexity and redundant transmission, the ICBWE gain mapping parameters may not be extracted (e.g., the ICBWE gain mapping parameters are not encoded in the stereo downmix / upmix parameter bitstream 290). Stereo tracks may be inserted (e.g., included or encoded) into the stereo downmix / upmix parameter bitstream 290, and the stereo downmix / upmix parameter bitstream 290 may be transmitted from encoder 114 to decoder 118. Stereo tracks may also be supplied to the center channel generator 208.

[0066] The 208 center channel generator can generate Petition 870250025693, dated 03 / 31 / 2025, page 42 / 182 38 / 75 a center channel in the frequency domain Mfr(b) 256 based on the first audio channel in the frequency domain 252 and the second audio channel in the frequency domain 254. According to some implementations, the center channel in the frequency domain Mfr(b) 256 can also be generated based on the stereo tracks. Some methods for generating the center channel in the frequency domain 256 based on the audio channels in the frequency domains 252, 254 and the stereo tracks are as follows: Mfr(b) = (Lfr(b) + R(b)) / 2 Mfr(b) = c1(b)*Lfr(b) + C2*Rfr(b), where c1(b) and c2(b) are downmix parameters per frequency band. In some implementations, the downmix parameters c1(b) and C2(b) are based on the stereo tracks. For example, in a mid-side downmix implementation when IPDs are estimated, c1(b) = (cos(-y)-i*sin(-y)) / 20.5 and c2(b) = (cos(IPD(b)-Y) + i *sin(IPD(b)-Y)) / 20.5 where i is the imaginary number meaning the square root of -1. In other examples, the center channel may also be based on an offset value (e.g., the final offset value 116). In such implementations, the left and right channels may be temporally aligned based on an estimate of the offset value before the frequency domain estimate of the center channel. In some implementations, this temporal alignment may be performed in the time domain on the first and second audio channels 130, 132 directly.In other implementations, time alignment can be performed in the transform domain in Lfr(b) and Rfr(b) by applying phase rotation to obtain the time-shift effect. In some. Petition 870250025693, dated 03 / 31 / 2025, page 43 / 182 In 39 / 75 implementations, the temporal alignment of the channels can be performed as a non-causal shift operation on the target channel. While in other implementations, the temporal alignment can be performed as a causal shift operation on the reference channel or a causal / non-causal shift operation on the reference / target channels, respectively. In some implementations, information about the reference and target channels can be captured as a reference channel indicator (which could be estimated based on the sign of the final shift value 116). In some implementations, information about the reference channel indicator and the shift value may be included as part of the encoder's bitstream output.

[0067] The frequency domain center channel 256 is provided to the inverse transform unit 210. The inverse transform unit 210 can perform an inverse transform operation on the frequency domain center channel 256 to generate a time domain center channel M(t) 258. In this way, the frequency domain center channel 256 can be inversely transformed in the time domain, or transformed in the MDCT domain for encoding. The time domain center channel 258 is provided to the center channel encoder 212 and the center channel BWE encoder 214.

[0068] The center channel encoder 212 can be configured to encode a low-band portion of the center channel in the 258 time domain to generate the low-band center channel bitstream 292. The low-band center channel bitstream 292 can be transmitted to Petition 870250025693, dated 03 / 31 / 2025, page 44 / 182 40 / 75 from encoder 114 to decoder 118. Center channel encoder 212 can be configured to generate a 260 low-band center channel excitation. The 260 low-band center channel excitation is provided to the 214 center channel BWE encoder.

[0069] The center-channel BWE encoder 214 can generate center-channel BWE parameters (e.g., linear prediction coefficients (LPCs), gain formats, a gain frame, etc.) based on the time-domain center channel 258 and the low-band center-channel excitation 260. The center-channel BWE encoder 214 can encode the center-channel BWE parameters into the high-band center-channel BWE bitstream 294. The high-band center-channel bitstream 294 can be transmitted from the encoder 114 to the decoder 116.

[0070] According to one implementation, the 214 center-channel BWE encoder can encode the high-band center channel using a high-band encoding algorithm based on a time-domain bandwidth extension (TBE) model. The TBE encoding of the high-band center channel can produce a set of LPC parameters, a high-band total gain parameter, and high-band temporal gain format parameters. The 214 center-channel BWE encoder can generate a set of high-band center gain parameters corresponding to the high-band center channel. For example, the 214 center-channel BWE encoder can generate a synthesized high-band center channel based on the Petition 870250025693, dated 03 / 31 / 2025, page 45 / 182 The 214-channel center-channel BWE encoder can generate the high-band center gain parameter based on a comparison between the high-band center signal and the synthesized high-band center signal. The 214-channel center-channel BWE encoder can also generate at least one tuning gain parameter, at least one tuning spectral shape parameter, or a combination thereof, as described herein. The 214-channel center-channel BWE encoder can transmit the LPC parameters (e.g., high-band center LPC parameters), the set of high-band center gain parameters, at least one tuning gain parameter, at least one spectral shape parameter, or a combination thereof. The LPC parameters, the high-band center gain parameter, or both, can correspond to an encoded version of the high-band center signal.

[0071] In this way, encoder 114 can generate the stereo downmix / upmix parameter bitstream 290, the low-band center channel bitstream 292, and the high-band center channel BWE bitstream 294. The 290, 292, 294 bitstreams can be multiplexed into a single bitstream, and the single bitstream can be transmitted to decoder 118. To reduce encoding complexity and redundant transmission, the ICBWE gain mapping parameters are not encoded in the stereo downmix / upmix parameter bitstream 290. As described in detail in relation to Figure 3, the ICBWE gain mapping parameters can be generated in decoder 118 based on other stereo tracks (e.g., stereo downmix DFT parameters).

[0072] With reference to Figure 3, it is shown Petition 870250025693, dated 03 / 31 / 2025, page 46 / 182 42 / 75 a specific implementation of the 118 decoder. The 118 decoder includes a low-band center channel decoder 302, a center channel BWE decoder 304, a transform unit 306, an ICBWE spatial balancer 308, a stereo upmixer 310, an inverse transform unit 312, an inverse transform unit 314, a combiner 316, and a shifter 320.

[0073] The lowband center channel bit stream 292 can be supplied from encoder 114 in Figure 2 to the lowband center channel decoder 302. The lowband center channel decoder 302 can be configured to decode the lowband center channel bit stream 292 to generate a lowband center signal 350. The lowband center channel decoder 302 can also be configured to generate an excitation of the lowband center signal 350. For example, the lowband center channel decoder 302 can generate a lowband center excitation signal 352. The lowband center signal 350 is supplied to the transformer unit 306, and the lowband center excitation signal 352 is supplied to the center channel BWE decoder 304.

[0074] The transform unit 306 can be configured to perform a transform operation on the low-band center signal 350 to generate a low-band center signal in the frequency domain 354. For example, the transform unit 306 can transform the low-band center signal 350 from the time domain to the frequency domain. The low-band center signal in the frequency domain 354 is provided to the stereo upmixer 310.

[0075] The stereo upmixer 310 can be Petition 870250025693, dated 03 / 31 / 2025, page 47 / 182 43 / 75 is configured to perform an upmix operation on the low-band center signal in frequency domain 354 using the stereo tracks extracted from the stereo downmix / upmix parameter bitstream 290. For example, the stereo downmix / upmix parameter bitstream 290 can be supplied (from encoder 114) to stereo upmixer 310. Stereo upmixer 310 can use the stereo tracks associated with the stereo downmix / upmix parameter bitstream 290 to perform the upmix of the low-band center signal in frequency domain 354 and to generate a first low-band channel in frequency domain 356 and a second low-band channel in frequency domain 358. The first low-band channel in frequency domain 356 is supplied to the inverse transform unit 312, and the second low-band channel in frequency domain 358 is supplied to the inverse transform unit 314.

[0076] The inverse transform unit 312 can be configured to perform an inverse transform operation on the first low-band channel in the frequency domain 356 to generate a first low-band channel 360 (e.g., a time-domain channel). The first low-band channel 360 (e.g., a left low-band channel) is provided to the combiner 316. The inverse transform unit 314 can be configured to perform an inverse transform operation on the second low-band channel in the frequency domain 358 to generate a second low-band channel 362 (e.g., a time-domain channel). The second low-band channel 362 (e.g., a right low-band channel) is also provided to the combiner 316. Petition 870250025693, dated 03 / 31 / 2025, page 48 / 182 44 / 75

[0077] The center-channel BWE decoder 304 can be configured to generate a synthesized high-band center signal 364 based on the low-band center excitation signal 352 and the center-channel BWE parameters encoded in the high-band center-channel BWE bitstream 294. For example, the high-band center-channel BWE bitstream 294 is supplied (from encoder 114) to the center-channel BWE decoder 304. A synthesis operation can be performed on the center-channel BWE decoder 304 by applying the center-channel BWE parameters to the low-band center excitation signal 352. Based on the synthesis operation, the center-channel BWE decoder 304 can generate the synthesized high-band center signal 364. The synthesized high-band center signal 364 is supplied to the ICBWE spatial balancer 308.In some implementations, the 304 center channel BWE decoder may be included in the 308 ICBWE spatial balancer. In other implementations, the 308 ICBWE spatial balancer may be included in the 304 center channel BWE decoder. In some specific implementations, the center channel BWE parameters may not be explicitly determined, but instead, the first and second high-band channels may be directly generated.

[0078] The stereo downmix / upmix parameter bitstream 290 is provided (from encoder 114) to decoder 118. As described in Figure 2, the ICBWE gain mapping parameters are not included in the bitstream (e.g., the stereo downmix / upmix parameter bitstream 290) provided to Petition 870250025693, dated 03 / 31 / 2025, page 49 / 182 45 / 75 decoder 118. Therefore, to generate a first high-band channel 366 and a second high-band channel using an ICBWE 308 spatial balancer, the ICBWE 308 spatial balancer (or another component of the 118 decoder) can generate an ICBWE 332 gain mapping parameter based on other stereo tracks (e.g., stereo DFT parameters) encoded in the stereo downmix / upmix parameter bitstream 290.

[0079] The ICBWE 308 spatial balancer includes the ICBWE 322 gain mapping parameter generator. Although the ICBWE 322 gain mapping parameter generator is included in the ICBWE 308 spatial balancer, in another implementation, the ICBWE 322 gain mapping parameter generator may be included in a different component of the decoder 118, may be outside the decoder 118, or may be a separate component of the decoder 118. The ICBWE 322 gain mapping parameter generator includes an extractor 324 and a selector 326. The extractor 324 can be configured to extract one or more gain parameters in the frequency domain 328 from the stereo downmix / upmix parameter bitstream 290.Selector 326 can be configured to select a group of frequency domain group parameters 330 (from one or more extracted frequency domain group parameters 328) for use in generating the ICBWE gain mapping parameter 332.

[0080] According to one implementation, the ICBWE 322 gain mapping parameter generator can generate the ICBWE 332 gain mapping parameter for broadband content using the following Petition 870250025693, dated 03 / 31 / 2025, page 50 / 182 46 / 75 pseudocode: if( st->bwidth == WB ) { / * copy to outputHB and reset hb_synth values ​​* / mvr2r( synthRef, synth, output_-Frame ); if( st->element_mode == IVAS_CPE_TD ) / * time-domain stereo * / { hStereoICBWE->prevSpecMapping = 0.0-F; hStereoICBWE->prevgsMapping = 1.0f; hStereoICBWE->icbweM2Ref_prev = 1.0f; } else if( st->element_mode == IVAS_CPE_DFT ) / * -Frequency-domain stereo * / ΐ hStereoICBWE->refchanlndx_bwe = L_CH_INDX; hStereoICBWE->prevSpecMapping = Θ.Ξ-F; prevgsMappíng = hStereoICBWE->prevgsHapping; tempi = hStereoD-Ft->side_gain[ 2*STEREO_DFT_BAND_MAX + 9 ]; temp2 = (l+templ+STEREO_DFT_FLT_MIN) / (l-templ+STEREO_DFT_FLT_MIN); hStereoICBWE->prevgsMapping = 2.0f / ( 1.0f + temp2 ); gsMapping = hStereoICBWE->prevgsMapping; winLen = (short)((SHE_OVERLAP_LEN * st->output_Fs) / 16000); winSlope = 1.0f / winLen; alpha = winSlope; icbweM2Ref = (float)sqrt(2.0f - gsMapping * gsMapping); for( i =θ; i < winLen; i++ ) í synthRef[i] *= (alpha * ( icbweM2Ref ) + (1.0f - alpha) * ( hStereoICBWE->icbweM2Ref_prev )); synth[i] *= (alpha * ( gsMapping ) + (1.0-F - alpha) * ( prevgsMapping )); alpha += winSlope; } for( ; i < NS2SA( st- / -output—Fs, FRAME_SIZE_NS); i++) í synthRef[i] *= ( icbweM2Ref ); synth[i] *= ( gsMapping ); } hStereoICBWE->icbweM2Ref_prev = icbweM2Ref;} return; }

[0081] The selected frequency domain gain parameter 330 can be selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter 330 Petition 870250025693, dated 03 / 31 / 2025, page 51 / 182 47 / 75 and a frequency range of the synthesized high-band center signal 364. For example, a first frequency range of a first specific frequency domain gain parameter can overlap the frequency range of the synthesized high-band center signal 364 by a first amount, and a second frequency range of a second specific frequency domain gain parameter can overlap the frequency range of the synthesized high-band center signal 364 by a second amount. For example, if the first amount is greater than the second amount, the first specific frequency domain gain parameter can be selected as the selected frequency domain gain parameter 330.In an implementation where no frequency domain gain parameter (from the extracted frequency domain gain parameters 328) has a frequency range that overlaps the frequency range of the synthesized high-band center signal 364, the frequency domain gain parameter that has a frequency range that is closest to the frequency range of the synthesized high-band center signal 364 can be selected as the selected frequency domain gain parameter 330.

[0082] As a non-limiting example of frequency domain gain parameter selection, for wideband encoding, the synthesized high-band center signal 364 may have a frequency range between 6.4 kilohertz (kHz) and 8 kHz. If the frequency domain gain parameter 330 is associated with a frequency range between 5.2 kHz and 8.56 kHz, the frequency domain gain parameter 330 may be selected for Petition 870250025693, dated 03 / 31 / 2025, p. 52 / 182 48 / 75 generate the ICBWE 332 gain mapping parameter. For example, in current implementations, the band number (b) = corresponds to the frequency range between 5.28 and 8.56 kHz. Since the band includes the frequency range (6.4 to 8 kHz), the sidegain of this band can be used directly to derive the ICBWE 322 gain mapping parameter. In the case where there are no bands covering the frequency range corresponding to the high band (6.4 to 8 kHz), the band closest to the high band frequency range can be used. In an example implementation where there are multiple frequency bands corresponding to the high band, then the sidegains of each of the frequency bands are weighted according to the frequency bandwidth to generate the final ICBWE gain mapping parameter, i.e., gsMapping = weight[b] * sidegain[b] + weight[b+1] * sidegain[b+1].

[0083] After selector 326 selects the frequency domain gain parameter 330, the ICBWE gain mapping parameter generator 322 can generate the ICBWE gain mapping parameter 332 using the frequency domain gain parameter 330. According to one implementation, the ICBWE gain mapping parameter (gsMapping) 332 can be determined based on the selected frequency domain gain parameter (sidegain) 330 using the following equation: gsMapping = (1 — sidegain)

[0084] For example, side gains can be alternative representations of ILDs. ILDs can be extracted (by the stereo track estimator 206) in frequency bands based on the audio channels in the domain of Petition 870250025693, dated 03 / 31 / 2025, page 53 / 182 49 / 75 frequency 252, 254. The relationship between ILDs and side gains can be approximately as follows: , ,=1+ sidegainíb) — sidegainfb') Therefore, the ICBWE 322 gain mapping parameter can also be expressed as: gsMapping = XT i LU

[0085] Since the ICBWE 322 gain mapping parameter generator generates the ICBWE gain mapping parameter (gsMapping) 322, the ICBWE 308 spatial balancer can generate the first high-band channel 366 and the second high-band channel 368. For example, the ICBWE 308 spatial balancer can be configured to perform a gain scaling operation on the synthesized high-band center signal 364 based on the ICBWE gain mapping parameter (gsMapping) 322 to generate the high-band channels 366.To illustrate, the ICBWE 308 spatial balancer can scale the synthesized high-band center signal 364 by the difference between two and the ICBWE 332 gain mapping parameter (e.g., 2, 72-gsMapping2gsMapping or ) to generate the first high-band channel 366 (e.g., the left high-band channel), and the ICBWE 308 spatial balancer can scale the synthesized high-band center signal 364 by the ICBWE 332 gain mapping parameter to generate the second high-band channel 368 (e.g., the right high-band channel). The high-band channels 366, 368 are provided to the combiner 316. To minimize gain variation artifacts. Petition 870250025693, dated 03 / 31 / 2025, page 54 / 182 50 / 75 interframe with ICBWE gain mapping, an overlap-add with a tapered window (e.g., a Sine(.) window or a triangular window) can be used at frame boundaries when transitioning from the i-th gsMapping parameter of the frame to the (i+1)-th gsMapping parameter of the frame.

[0086] The ICBWE reference channel can be used in combiner 316. For example, combiner 316 can determine that the high-band channel 366, 368 corresponds to the left channel and that the high-band channel 366, 368 corresponds to the right channel. In this way, a reference channel indicator can be provided to the ICBWE 308 spatial balancer to indicate whether the left high-band channel corresponds to the first high-band channel 366 or the second high-band channel 368. Combiner 316 can be configured to combine the first high-band channel 366 and the first low-band channel 360 to generate a first channel 370. For example, combiner 316 can combine the left high-band channel and the left low-band channel 360 to generate a left channel. The 316 combiner can also be configured to combine the second high-band channel 368 and the second low-band channel 362 to generate a second channel 372.For example, combiner 316 can combine the right high-band channel and the right low-band channel to generate a right channel. The first and second channels 370, 372 are provided to shifter 320.

[0087] As an example, the first channel can be designated as the reference channel, and the second channel can be designated as the non-reference channel or the “target” channel. In this way, the second channel 372 can be submitted Petition 870250025693, dated 03 / 31 / 2025, page 55 / 182 51 / 75 to a shift operation on shifter 320. Shifter 320 can extract a shift value (e.g., the final shift value 116) from the stereo downmix / upmix parameter bitstream 290 and can shift the second channel 372 by the shift value to generate the second output channel 128. Shifter 320 can pass the first high-band channel 366 as the first output channel 126. In some implementations, shifter 320 can be configured to perform a causal shift on the target channel. In other implementations, shifter 320 can be configured to perform a non-causal shift on the reference channel. While in other implementations, shifter 320 can be configured to perform a causal / non-causal shift on the target / reference channels, respectively. Information indicating which channel is the target channel and which channel is the reference channel may be included as part of the received bitstream.In some implementations, the 320 shifter can perform the shift operation in the time domain. In other implementations, the shift operation can be performed in the frequency domain. In some implementations, the 320 shifter may be included in the 310 stereo upmixer. In this way, the shift operation can be performed on low-band signals.

[0088] According to one implementation, the shift operation may be independent of ICBWE operations. For example, the high band reference channel indicator may not be the same as the shifter 320 reference channel indicator. To illustrate, the high band reference channel (e.g., the reference channel) Petition 870250025693, dated 03 / 31 / 2025, page 56 / 182 The 52 / 75 associated with ICBWE operations) may be different from the reference channel in shifter 320. According to some implementations, a reference channel may not be assigned in shifter 320, and shifter 320 may be configured to shift both channels 370 and 372.

[0089] In this way, the encoding complexity and transmission bandwidth can be reduced by omitting the extraction and transmission of the ICBWE gain mapping parameters in the encoder 114. The ICBWE gain mapping parameters can be generated in the decoder 118 based on other stereo parameters (e.g., frequency domain gain parameters 328) included in the bitstream 290.

[0090] With reference to Figure 4, a method 400 is shown for determining the ICBWE mapping parameters based on a frequency domain gain parameter transmitted from an encoder. The method 400 can be performed by the decoder 118 of Figures 1 and 3.

[0091] Method 400 involves receiving a bitstream from an encoder, at numerical reference 402. The bitstream may include at least one low-band center channel bitstream, one high-band center channel BWE bitstream, and one stereo downmix / upmix parameter bitstream. For example, with reference to Figure 3, decoder 118 may receive stereo downmix / upmix parameter bitstream 290, low-band center channel bitstream 292, and high-band center channel BWE bitstream 294.

[0092] Method 400 also includes decoding the low-bandwidth center-channel bitstream to generate a Petition 870250025693, dated 03 / 31 / 2025, page 57 / 182 53 / 75 low-band center signal and a low-band center excitation signal, in numerical reference 404. For example, with reference to Figure 3, the low-band center channel decoder 302 can decode the low-band center channel bit stream 292 to generate the low-band center signal 350. The low-band center channel decoder 302 can also generate the low-band center excitation signal 352.

[0093] Method 400 additionally includes decoding the high-band center-channel BWE bitstream to generate a synthesized high-band center signal based on a non-linear harmonic extension of the low-band center excitation signal and based on the high-band channel BWE parameters, in numerical reference 406. For example, center-channel BWE decoder 304 can generate synthesized high-band center signal 364 based on low-band center excitation signal 352 and center-channel BWE parameters encoded in the high-band center-channel BWE bitstream 294. To illustrate, a synthesis operation can be performed on center-channel BWE decoder 304 by applying the center-channel BWE parameters to the low-band center excitation signal 352. Based on the synthesis operation, center-channel BWE decoder 304 can generate synthesized high-band center signal 364.

[0094] Method 400 also includes determining an ICBWE gain mapping parameter for the synthesized high-band center signal based on a selected frequency-domain gain parameter that is extracted from the stereo downmix / upmix parameter bitstream, in Petition 870250025693, dated 03 / 31 / 2025, page 58 / 182 54 / 75 numerical reference 408. The selected frequency domain gain parameter can be selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter and a frequency range of the synthesized high-band center signal. For example, with reference to Figure 3, the extractor can extract the frequency domain gain parameters 328 from the stereo downmix / upmix parameter bitstream 290, and the selector 326 can select the frequency domain gain parameter 330 (from one or more extracted frequency domain gain parameters 328) for use in generating the ICBWE gain mapping parameter 332. Thus, according to one implementation, method 400 can also include extracting one or more frequency domain gain parameters from the stereo parameter bitstream.The selected frequency domain gain parameter can be chosen from one or more frequency domain gain parameters.

[0095] The selected frequency domain gain parameter 330 can be selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter 330 and a frequency range of the synthesized high-band center signal 364. To illustrate, for wideband encoding, the synthesized high-band center signal 364 may have a frequency range between 6.4 kilohertz (kHz) and 8 kHz. If the frequency domain gain parameter 330 is associated with a frequency range between 5.2 kHz and 8.56 kHz, the frequency domain gain parameter 330 can be selected to generate the Petition 870250025693, dated 03 / 31 / 2025, page 59 / 182 55 / 75 gain mapping parameter of ICBWE 332.

[0096] After selector 326 selects the frequency domain gain parameter 330, the ICBWE gain mapping parameter generator 322 can generate the ICBWE gain mapping parameter 332 using the frequency domain gain parameter 330. According to one implementation, the ICBWE gain mapping parameter (gsMapping) 332 can be determined based on the selected frequency domain gain parameter (sidegain) 330 using the following equation: gsMapping =T1 - sidegain

[0097] Method 400 additionally includes performing a gain scaling operation on the synthesized highband center signal based on the ICBWE gain mapping parameter to generate a reference highband channel and a target highband channel, in numerical reference 410. Performing the gain scaling operation may include scaling the synthesized highband center signal by the ICBWE gain mapping parameter to generate the right highband channel. For example, with reference to Figure 3, the ICBWE spatial balancer 308 may scale the synthesized highband center signal 364 by the ICBWE gain mapping parameter 332 to generate the second highband channel 368 (e.g., the right highband channel). Performing the gain scaling operation may also include scaling the synthesized highband center signal by a difference between two and the gain mapping parameter of Petition 870250025693, dated 03 / 31 / 2025, page 60 / 182 56 / 75 ICBWE to generate the left high-band channel. For example, with reference to Figure 3, the ICBWE 308 spatial balancer can scale the synthesized high-band center signal 364 by the difference between two and the ICBWE 332 gain mapping parameter (e.g., 2-gsMapping) to generate the first high-band channel 366 (e.g., the left high-band channel).

[0098] Method 400 also includes outputting a first audio channel and a second audio channel, at numerical reference 412. The first audio channel can be based on the reference high-band channel, and the second audio channel can be based on the target high-band channel. For example, with reference to Figure 1, the second device 106 can output the first output channel 126 (e.g., the first audio channel based on the left channel 370) and the second output channel 128 (e.g., the second audio channel based on the right channel 372).

[0099] Thus, according to method 400, the encoding complexity and transmission bandwidth can be reduced by omitting the extraction and transmission of ICBWE gain mapping parameters in encoder 114. The ICBWE gain mapping parameters can be generated in decoder 118 based on other stereo parameters (e.g., frequency domain gain parameters 328) included in bitstream 290.

[00100] With reference to Figure 5, a block diagram of a specific illustrative example of a device (e.g., a wireless communication device) is shown and generally designated 500. In several Petition 870250025693, dated 03 / 31 / 2025, page 61 / 182 In 57 / 75 implementations, device 500 may have fewer or more components than illustrated in Figure 5. In an illustrative implementation, device 500 may correspond to the second device 106 in Figure 1. In an illustrative implementation, device 500 may perform one or more operations described with reference to systems and methods in Figures 1 to 4.

[00101] In a specific implementation, device 500 includes a processor 506 (e.g., a central processing unit (CPU)). Device 500 may include one or more additional processors 510 (e.g., one or more digital signal processors (DSPs)). Processors 510 may include a media encoder / decoder (e.g., speech and music) (CODEC) 508, and an echo canceller 512. The media CODEC 508 may include the decoder 118, the encoder 114, or both, of Figure 1. The decoder 118 may include the ICBWE gain mapping parameter generator 322.

[00102] Device 500 may include a memory 153 and a CODEC 534. Although media CODEC 508 is illustrated as a component of processors 510 (e.g., dedicated circuitry and / or executable programming code), in other implementations, one or more components of media CODEC 508, such as decoder 118, encoder 114, or both, may be included in processor 506, CODEC 534, another processing component, or a combination thereof.

[00103] Device 500 may include a transceiver 590 coupled to an antenna 542. Device 500 may include a screen 528 coupled to a controller of Petition 870250025693, dated 03 / 31 / 2025, p. 62 / 182 58 / 75 display 526. One or more speakers 548 can be connected to the CODEC 534. One or more microphones 546 can be connected, via input interface(s) 592, to the CODEC 534. In a specific implementation, the loudspeakers 548 may include the first loudspeaker 142, the second loudspeaker 144 of Figure 1, or a combination thereof. The CODEC 534 may include a digital-to-analog converter (DAC) 502 and an analog-to-digital converter (ADC) 504.

[00104] Memory 153 may contain instructions 560 executable by the decoder 118, the processor 506, the processors 510, the CODEC 534, another processing unit of the device 500, or a combination thereof, to perform one or more operations described with reference to Figures 1 to 4.

[00105] For example, instructions 560 can be executable to cause processor 510 to decode the low-band center channel bitstream 292 to generate the low-band center signal 350 and the low-band center excitation signal 352. Instructions 560 can additionally be executable to cause processor 510 to decode the high-band center channel BWE bitstream 294 based on the low-band center excitation signal 352 to generate the synthesized high-band center signal 364. Instructions 560 can also be executable to cause processor 510 to determine the ICBWE gain mapping parameter 332 for the synthesized high-band center signal 364 based on the selected frequency domain gain parameter 330 that is extracted from the stereo downmix / upmix parameter bitstream 290. The frequency domain gain parameter Petition 870250025693, dated 03 / 31 / 2025, page 63 / 182 The selected frequency 330 can be selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter 330 and a frequency range of the synthesized high-band center signal 364. Instructions 560 can additionally be executed to cause the processor 510 to perform a gain scaling operation on the synthesized high-band center signal 364 based on the ICBWE gain mapping parameter 332 to generate the first high-band channel 366 (e.g., the left high-band channel) and the second high-band channel 368 (e.g., the right high-band channel). Instructions 560 can also be executed to cause the processor 510 to generate the first output channel 326 and the second output channel 328.

[00106] One or more components of the 500 device may be implemented via dedicated hardware (e.g., circuit assembly), by a processor that executes instructions to perform one or more tasks, or a combination thereof. As an example, the 153 memory or one or more components of the 506 processor, the 510 processors, and / or the 534 CODEC may be a memory device, such as random access memory (RAM), magnetoresistive random access memory (MRAM), STT-MRAM, flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, or a compact disc read-only memory (CDROM). The memory device may include instructions (e.g., the 560 instructions) that, when executed by a Petition 870250025693, dated 03 / 31 / 2025, page 64 / 182 60 / 75 computer (for example, a processor in CODEC 534, decoder 118, processor 506, and / or processors 510), can cause the computer to perform one or more operations described with reference to Figures 1 to 4. As an example, memory 153 or one or more components of processor 506, processors 510, and / or CODEC 534 can be a non-temporary, computer-readable medium that includes instructions (for example, instructions 560) which, when executed by a computer (for example, a processor in CODEC 534, decoder 118, processor 506, and / or processors 510), cause the computer to perform one or more operations described with reference to Figures 1 to 4.

[00107] In a specific implementation, device 500 may be included in a packaged system or system-on-a-chip device (e.g., a mobile station modem (MSM)) 522. In a specific implementation, the processor 506, processors 510, display controller 526, memory 153, CODEC 534, and transceiver 590 are included in a packaged system or system-on-a-chip device 522. In a specific implementation, an input device 530, such as a touch screen and / or numeric keypad, and a power supply 544 are coupled to the system-on-a-chip device 522. Furthermore, in a specific implementation, as illustrated in Figure 5, the screen 528, the input device 530, the speakers 548, the microphones 546, the antenna 542, and the power supply 544 are outside the system-on-a-chip device. 522.However, each of the following is a list of components: screen 528, input device 530, speakers 548, microphones 546, antenna 542, and power supply. Petition 870250025693, dated 03 / 31 / 2025, page 65 / 182 61 / 75 power supply 544 can be coupled to a system device component on a chip 522, such as an interface or a controller.

[00108] Device 500 may include a cordless phone, a mobile communication device, a mobile phone, a smartphone, a cellular phone, a laptop, a desktop computer, a computer, a tablet, a signal decoder, a personal digital assistant (PDA), a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a communication device, a fixed location data unit, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a decoder system, an encoder system, or any combination thereof.

[00109] In a specific implementation, one or more components of the systems and devices disclosed herein may be integrated into a decoding system or apparatus (for example, an electronic device, a CODEC, or a processor therein), into an encoding system or apparatus, or both. In other implementations, one or more components of the systems and devices disclosed herein may be integrated into a cordless phone, a tablet, a desktop computer, a laptop, a signal decoder, a music player, a video player, an entertainment unit, a television, a game console, a navigation device, a communication device, a personal digital assistant (PDA), a fixed location data unit, a Petition 870250025693, dated 03 / 31 / 2025, page 66 / 182 62 / 75 personal media player, or other type of device.

[00110] It should be noted that several functions performed by one or more components of the systems and devices disclosed in this document are described as being performed by specific components or modules. This division of components and modules is for illustrative purposes only. In an alternative implementation, a function performed by a specific component or module may be divided among multiple components or modules. Furthermore, in an alternative implementation, two or more components or modules may be integrated into a single component or module. Each component or module may be implemented using hardware (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a DSP, a controller, etc.), software (e.g., instructions executable by a processor), or any combination thereof.

[00111] In conjunction with the implementations described, an apparatus includes means for receiving a bit stream from an encoder. The bit stream may include a low-band center-channel bit stream, a center-channel BWE bit stream, and a stereo parameter bit stream. For example, the receiving means may include the second device of Figure 1, the 542 antenna of Figure 5, the 590 transceiver of Figure 5, one or more other devices, modules, circuits, components, or a combination thereof.

[00112] The device may also include means for decoding the band-center channel bit stream. Petition 870250025693, dated 03 / 31 / 2025, page 67 / 182 63 / 75 low to generate a low-band center signal and a low-band center channel excitation of the low-band center signal. For example, the means of decoding the low-band center channel bit stream may include the decoder 118 of Figures 1, 3 and 5, the low-band center channel decoder 302 of Figure 3, the CODEC 508 of Figure 5, the processors 510, the processor 506 of Figure 5, the device 500, the instructions 560 executable by a processor, one or more other devices, modules, circuits, components, or a combination thereof.

[00113] The apparatus may also include means for decoding the center-channel BWE bitstream based on low-band center-channel excitation to generate a synthesized high-band center signal. For example, means for decoding the center-channel BWE bitstream may include the decoder 118 of Figures 1, 3 and 5, the center-channel BWE decoder 304 of Figure 3, the CODEC 508 of Figure 5, the processors 510, the processor 506 of Figure 5, the device 500, the instructions 560 executable by a processor, one or more other devices, modules, circuits, components, or a combination thereof.

[00114] The apparatus may also include means for determining an ICBWE gain mapping parameter for the synthesized high-band center signal based on a selected frequency-domain gain parameter that is extracted from the stereo parameter bitstream. The selected frequency-domain gain parameter may be selected based on a spectral proximity of a frequency band to the frequency-domain gain parameter. Petition 870250025693, dated 03 / 31 / 2025, page 68 / 182 64 / 75 selected frequency and a frequency range of the synthesized high-band center signal. For example, the means for determining the ICBWE gain mapping parameter may include the decoder 118 of Figures 1, 3 and 5, the ICBWE spatial balancer 308 of Figure 3, the ICBWE gain mapping parameter generator 322 of Figure 3, the extractor 324 of Figure 3, the selector 326 of Figure 3, the CODEC 508 of Figure 5, the processors 510, the processor 506 of Figure 5, the device 500, the instructions 560 executable by a processor, one or more other devices, modules, circuits, components, or a combination thereof.

[00115] The apparatus may also include means for performing a gain scaling operation on the synthesized high-band center signal based on the ICBWE gain mapping parameter to generate a left high-band channel and a right high-band channel. For example, the means for performing the gain scaling operation may include the decoder 118 of Figures 1, 3 and 5, the ICBWE spatial balancer 308 of Figure 3, the CODEC 508 of Figure 5, the processors 510, the processor 506 of Figure 5, the device 500, the instructions 560 executable by a processor, one or more other devices, modules, circuits, components, or a combination thereof.

[00116] The device may also include means for outputting a first audio channel and a second audio channel. The first audio channel may be based on the left high-band channel, and the second audio channel may be based on the right high-band channel. For example, the output means may include the first speaker 142 of Figure Petition 870250025693, dated 03 / 31 / 2025, page 69 / 182 65 / 75 1, the second loudspeaker 144 in Figure 1, the loudspeakers 548 in Figure 5, one or more other devices, modules, circuits, components, or a combination thereof.

[00117] With reference to Figure 6, a block diagram of a specific illustrative example of a 600 base station is shown. In various implementations, the 600 base station may have more or fewer components than illustrated in Figure 6. In one illustrative example, the 600 base station may include the second device in Figure 1. In another illustrative example, the 600 base station may operate according to one or more of the methods or systems described with reference to Figures 1 to 5.

[00118] Base station 600 can be part of a wireless communication system. The wireless communication system can include multiple base stations and multiple wireless devices. The wireless communication system can be a Long-Term Evolution (LTE) system, a Code Division Multiple Access (CDMA) system, a Global System for Mobile Communications (GSM) system, a wireless local area network (WLAN) system, or some other wireless system. A CDMA system can implement Wideband CDMA (WCDMA), CDMA IX, Evolution Optimized Data (EVDO), Time Division Synchronous CDMA (TD-SCDMA), or some other version of CDMA.

[00119] Wireless devices may also be called user equipment (UE), a mobile station, a terminal, an access terminal, a subscriber unit, a station, etc. Wireless devices may include a mobile phone, a smartphone, a tablet, a wireless modem, a personal digital assistant (PDA), a device Petition 870250025693, dated 03 / 31 / 2025, page 70 / 182 66 / 75 portable, a laptop, a smartbook, a netbook, a tablet, a cordless phone, a wireless local area network (WLL) station, a Bluetooth device, etc. Wireless devices may include or correspond to device 500 in Figure 5.

[00120] Several functions can be performed by one or more components of the base station 600 (and / or other components not shown), such as sending and receiving messages and data (e.g., audio data). In one specific example, the base station 600 includes a processor. 606 (for example, a CPU). Base station 600 may include a transcoder 610. The transcoder 610 may include an audio CODEC 608. For example, the transcoder 610 may include one or more components (e.g., circuit assembly) configured to perform audio CODEC 608 operations. As another example, the transcoder 610 may be configured to execute one or more computer-readable instructions to perform audio CODEC 608 operations. Although audio CODEC 608 is illustrated as a component of transcoder 610, in other examples, one or more components of audio CODEC 608 may be included in processor 606, another processing component, or a combination thereof. For example, a decoder 638 (e.g., a vocoder decoder) may be included in a receiver data processor 664.As another example, a 636 encoder (for example, a vocoder encoder) may be included in a 682 transmission data processor. The 636 encoder may include the 114 encoder of Figure 1. The 638 decoder may include the 118 decoder of Figure 1. Petition 870250025693, dated 03 / 31 / 2025, page 71 / 182 67 / 75

[00121] The 610 transcoder can operate to transcode messages and data between two or more networks. The 610 transcoder can be configured to convert audio message and data from a first format (e.g., a digital format) to a second format. For example, the 638 decoder can decode encoded signals that have a first format, and the 636 encoder can encode the decoded signals into encoded signals that have a second format. Additionally or alternatively, the 610 transcoder can be configured to perform data rate adaptation. For example, the 610 transcoder can down-convert a data rate or up-convert a data rate without changing the format of the audio data. For example, the 610 transcoder can down-convert 64 kbit / s signals to 16 kbit / s signals.

[00122] Base station 600 may include memory 632. Memory 632, as a computer-readable storage device, may include instructions. Instructions may include one or more instructions that are executable by processor 606, transcoder 610, or a combination thereof, to perform one or more operations described with reference to the methods and systems of Figures 1 to 5.

[00123] Base station 600 may include multiple transmitters and receivers (e.g., transceivers), such as a first transceiver 652 and a second transceiver 654, coupled to an antenna array. The antenna array may include a first antenna 642 and a second antenna 644. The antenna array may be configured to be Petition 870250025693, dated 03 / 31 / 2025, page 72 / 182 68 / 75 communicate wirelessly with one or more wireless devices, such as device 500 in Figure 5. For example, the second antenna 644 can receive a data stream 614 (e.g., a bit stream) from a wireless device. The data stream 614 can include messages, data (e.g., encoded speech data), or a combination thereof.

[00124] Base station 600 may include a network connection 660, as a backhaul connection. The network connection 660 may be configured to communicate with a core network or one or more base stations of the wireless communication network. For example, base station 600 may receive a second data stream (e.g., messages or audio data) from a core network via the network connection 660. Base station 600 may process the second data stream to generate messages or audio data and provide the messages or audio data to one or more wireless devices via one or more antennas or to another base station via the network connection 660. In a specific implementation, the network connection 660 may be a wide area network (WAN) connection, as an illustrative, non-limiting example. In some implementations, the core network may include or correspond to a Public Switched Telephone Network (PSTN), a packet backbone network, or both.

[00125] Base station 600 may include a media gateway 670 that is coupled to network connection 660 and processor 606. The media gateway 670 may be configured for conversion between media streams of different telecommunications technologies. For example, the media gateway 670 may convert between different transmission protocols, different encoding schemes, Petition 870250025693, dated 03 / 31 / 2025, page 73 / 182 69 / 75 or both. To illustrate, the 670 media gateway can convert PCM signals into Real-Time Transport Protocol (RTP) signals, as a non-limiting illustrative example. The 670 media gateway can convert data between packet-switched networks (e.g., a Voice over Internet Protocol (VoIP) network, an IP Multimedia Subsystem (IMS), a fourth-generation (4G) wireless network such as LTE, WiMax, and UMB, etc.), circuit-switched networks (e.g., a PSTN), and hybrid networks (e.g., a second-generation (2G) wireless network such as GSM, GPRS, and EDGE, a third-generation (3G) wireless network such as WCDMA, EVDO, and HSPA, etc.).

[00126] Additionally, the 670 media gateway may include a transcoder, such as the 610 transcoder, and may be configured to transcode data when codecs are incompatible. For example, the 670 media gateway may transcode between an Adaptive Multi-Rate (AMR) codec and a G.711 codec, as a non-limiting illustrative example. The 670 media gateway may include a router and a plurality of physical interfaces. In some implementations, the 670 media gateway may also include a controller (not shown). In a specific implementation, the media gateway controller may be outside the 670 media gateway, outside the 600 base station, or both. The media gateway controller may control and coordinate the operations of multiple media gateways.The 670 media gateway can receive control signals from the media gateway controller and can act as a bridge between different transmission technologies, adding services to end-user resources and connections. Petition 870250025693, dated 03 / 31 / 2025, page 74 / 182 70 / 75

[00127] Base station 600 may include a demodulator 662 that is coupled to transceivers 652, 654, receiver data processor 664, and processor 606, and receiver data processor 664 may be coupled to processor 606. Demodulator 662 may be configured to demodulate modulated signals received from transceivers 652, 654 and provide demodulated data to receiver data processor 664. Receiver data processor 664 may be configured to extract a message or audio data from the demodulated data and send the message or audio data to processor 606.

[00128] Base station 600 may include a 682 transmission data processor and a 684 transmission multiple-input multiple-output (MIMO) processor. The 682 transmission data processor may be coupled to the 606 processor and the 684 transmission MIMO processor. The 684 transmission MIMO processor may be coupled to the 652, 654 transceivers and the processor. 606. In some implementations, the 684 transmission MIMO processor may be coupled to the 670 media gateway. The 682 transmission data processor may be configured to receive messages or audio data from the 606 processor and encode the messages or audio data based on an encoding scheme such as CDMA or orthogonal frequency division multiplexing (OFDM), as a non-limiting illustrative example. The 682 transmission data processor may provide the encoded data to the 684 transmission MIMO processor.

[00129] Encoded data can be multiplexed with other data, such as pilot data, using Petition 870250025693, dated 03 / 31 / 2025, page 75 / 182 71 / 75 CDMA or OFDM techniques are used to generate multiplexed data. The multiplexed data can then be modulated (i.e., symbol-mapped) by the 682 transmission data processor based on a specific modulation scheme (e.g., Binary Phase Shift Keying (BPSK), Quadrature Phase Shift Keying (QSPK), M-ary Phase Shift Keying (M-PSK), M-ary Quadrature Amplitude Modulation (M-QAM), etc.) to generate modulation symbols. In a specific implementation, the encoded data and other data can be modulated using different modulation schemes. The data rate, encoding, and modulation for each data stream can be determined by instructions executed by the 606 processor.

[00130] The 684 transmission MIMO processor can be configured to receive modulation symbols from the 682 transmission data processor and can additionally process the modulation symbols and perform beamforming on the data. For example, the 684 transmission MIMO processor can apply beamforming weights to the modulation symbols.

[00131] During operation, the second antenna 644 of base station 600 can receive a data stream 614. The second transceiver 654 can receive the data stream 614 from the second antenna 644 and can provide the data stream 614 to the demodulator 662. The demodulator 662 can demodulate modulated signals from the data stream 614 and provide demodulated data to the receiver data processor 664. The receiver data processor 664 can extract audio data from the demodulated data and provide the extracted audio data to Petition 870250025693, dated 03 / 31 / 2025, page 76 / 182 72 / 75 processor 606.

[00132] Processor 606 can provide audio data to transcoder 610 for transcoding. Transcoder 610's decoder 638 can decode audio data from a first format into decoded audio data, and encoder 636 can encode the decoded audio data into a second format. In some implementations, encoder 636 can encode audio data using a higher data rate (e.g., upconvert) or a lower data rate (e.g., downconvert) than that received from the wireless device. In other implementations, audio data cannot be transcoded. Although transcoding (e.g., decoding and encoding) is illustrated as performed by a transcoder 610, transcoding operations (e.g., decoding and encoding) can be performed by multiple base station 600 components.For example, decoding can be performed by receiver data processor 664 and encoding can be performed by transmission data processor 682. In other implementations, processor 606 can provide the audio data to media gateway 670 for conversion to another transmission protocol, encoding scheme, or both. Media gateway 670 can provide the converted data to another base station or core network via network connection 660.

[00133] The encoded audio data generated in encoder 636 can be supplied to transmission data processor 682 or to network connection 660 via processor 606. The transcoded audio data from Petition 870250025693, dated 03 / 31 / 2025, page 77 / 182 Transcoder 610 (73 / 75) can be supplied to the transmission data processor 682 for encoding according to a modulation scheme, such as OFDM, to generate the modulation symbols. The transmission data processor 682 can supply the modulation symbols to the transmission MIMO processor 684 for further processing and beamforming. The transmission MIMO processor 684 can apply beamforming weights and can supply the modulation symbols to one or more antennas in the antenna array, such as the first antenna 642 via the first transceiver 652. In this way, base station 600 can supply a transcoded data stream 616, which corresponds to the data stream 614 received from the wireless device, to another wireless device. The transcoded data stream 616 can have a different encoding format, data rate, or both, than the data stream 614.In other implementations, the transcoded data stream 616 can be provided to network connection 660 for transmission to another base station or a core network.

[00134] Those skilled in the art might further appreciate that the various illustrative logic block steps, configurations, modules, circuits, and algorithms described in conjunction with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above in general terms of their functionality. The possibility of such functionality being implemented as executable hardware or software depends on Petition 870250025693, dated 03 / 31 / 2025, page 78 / 182 74 / 75 specific application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each specific application, but such implementation decisions should not be interpreted as causing a deviation from the scope of the present disclosure.

[00135] The steps of a method or algorithm described together with the implementations disclosed herein may be incorporated directly into hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in a memory device such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, or a compact disc read-only memory (CD-ROM). An exemplary memory device is coupled to the processor so that the processor can read information from and write information to the memory device.Alternatively, the memory device may be integral to the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as separate components in a computing device or a user terminal. Petition 870250025693, dated 03 / 31 / 2025, page 79 / 182 75 / 75

[00136] The preceding description of the disclosed implementations is provided to enable one skilled in the art to make or use the disclosed implementations. Various modifications to these implementations will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other implementations without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the implementations shown herein, but should conform to the broadest possible scope compatible with the innovative principles and features as defined by the following claims. Petition 870250025693, dated 03 / 31 / 2025, page 80 / 182

Claims

1 / 5 CLAIMS 1. Device characterized in that it comprises: a receiver configured to receive a bit stream from an encoder, the bit stream comprising at least one low-band center channel bit stream (292), one high-band center channel bandwidth extension, BWE, bit stream (294), and one stereo downmix / upmix parameter bit stream (290); a decoder configured to: decode the low-band center channel bit stream to generate a low-band center signal (350) and a low-band center excitation signal (352); generate a non-linear harmonic extension of the low-band center excitation signal corresponding to a high-band BWE portion;decode the high-band center channel BWE bitstream to generate a synthesized high-band center signal (364) based on the non-linear harmonic extension of the low-band center excitation signal and based on high-band center channel BWE parameters; determine an inter-channel bandwidth extension gain mapping parameter, ICBWE, (332) corresponding to the synthesized high-band center signal, the ICBWE gain mapping parameter based on a selected frequency domain gain parameter (330) that is extracted from the stereo downmix / upmix parameter bitstream, wherein the selected frequency domain gain parameter is selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter and a frequency range of the synthesized high-band center signal;and perform a gain scaling operation on the synthesized high-band center signal based on the ICBWE gain mapping parameter to generate a left high-band channel (366) and a right high-band channel (368); and one or more loudspeakers configured to output a first audio channel and a second audio channel, the first audio channel (126) based on the left high-band channel, and the second audio channel (128) based on the right high-band channel.; 2. Device according to claim 1, characterized in that the selected frequency domain gain parameter corresponds to a side gain of the stereo downmix / upmix parameter bitstream or interchannel level difference, ILD, of the stereo downmix / upmix parameter bitstream.

3. Device according to claim 1, characterized in that the left high-band channel corresponds to a reference high-band channel or a target high-band channel, and in that the right high-band channel corresponds to the other between the reference high-band channel or the target high-band channel.

4. Device according to claim 3, characterized in that the decoder is additionally configured to generate, based on the low-band center signal, a left low-band channel and a right low-band channel.

5. Device according to claim 4, Petition 870250025693, dated 31 / 03 / 2025, pp. 162 / 182 3 / 5 characterized in that the decoder is additionally configured to: combine the left low-band channel and the left high-band channel to generate the first audio channel; and combine the right low-band channel and the right high-band channel to generate the second audio channel.

6. Device according to claim 1, characterized in that the decoder is configured to extract one or more frequency domain gain parameters from the stereo downmix / upmix parameter bitstream and select a set of frequency domain gain parameters from the one or more frequency domain gain parameters, the set of frequency domain gain parameters including the selected frequency domain gain parameter.

7. Device according to claim 1, characterized in that the decoder is configured to scale the synthesized high-band center signal using the ICBWE gain mapping parameter to generate a target high-band channel.

8. Device, according to claim 1, characterized in that the side gains of multiple frequency bands of a high band are weighted based on the frequency bandwidths of each frequency band of the multiple frequency bands to generate the ICBWE gain mapping parameter.

9. Device according to claim 1, characterized in that the decoder is integrated into a base station. Petition 870250025693, dated 31 / 03 / 2025, pp. 163 / 182 4 / 5 10. Device according to claim 1, characterized in that the decoder is integrated into a mobile device.

11. A method for decoding a signal, the method characterized in that it comprises: receiving a bit stream from an encoder, the bit stream comprising at least one low-band center channel bit stream (292), a high-band center channel bandwidth extension, BWE, bit stream (294), and a stereo downmix / upmix parameter bit stream (290); decoding, in a decoder, the low-band center channel bit stream to generate a low-band center signal (350) and a low-band center excitation signal (352); generating a non-linear harmonic extension of the low-band center excitation signal corresponding to a high-band BWE portion; decoding the high-band center channel BWE bit stream to generate a synthesized high-band center signal (364) based on the non-linear harmonic extension of the low-band center excitation signal and based on high-band center channel BWE parameters;determine an interchannel bandwidth extension gain mapping parameter, ICBWE, (332) corresponding to the synthesized highband center signal, the ICBWE gain mapping parameter based on a selected frequency domain gain parameter (330) that is extracted from the stereo downmix / upmix parameter bitstream, wherein the selected frequency domain gain parameter is selected based on a spectral proximity of a frequency range of the selected frequency domain gain parameter and a frequency range of the synthesized highband center signal; perform a gain scaling operation on the synthesized highband center signal based on the ICBWE gain mapping parameter to generate a left highband channel (366) and a right highband channel (368);and output a first audio channel and a second audio channel, the first audio channel (126) based on the left high-band channel, and the second audio channel (128) based on the right high-band channel.; 12. Method, according to claim 11, characterized in that the left high-band channel corresponds to a reference high-band channel or a target high-band channel, and in that the right high-band channel corresponds to the other between the reference high-band channel or the target high-band channel.

13. Computer-readable memory characterized in that it comprises instructions stored therein, the instructions being executed by a processor within a decoder to perform the method as defined in any one of claims 11 to 12. Petition 870250025693, dated 31 / 03 / 2025, pp. 165 / 182