Related methods using multi-signal encoders, multi-signal decoders, and signal whitening or signal post-processing.

Adaptive joint signal processing on pre-processed audio signals with perceptually whitened spectra and broadband energy normalization addresses inefficiencies in multi-channel audio coding, enhancing efficiency and quality in immersive 3D audio formats.

JP2026076313APending Publication Date: 2026-05-11FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2026-02-13
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing multi-channel audio coding technologies are limited in flexibility and efficiency, particularly in immersive 3D audio formats, due to insufficient joint stereo coding and inefficient use of channel dependencies, leading to increased data transmission requirements and reduced perceptual audio quality.

Method used

Adaptive joint signal processing is applied on pre-processed audio signals to improve coding efficiency, using perceptually whitened spectra and broadband energy normalization, with post-processing on the decoder side to maintain audio quality while reducing data transmission.

Benefits of technology

This approach enhances multi-channel coding efficiency by reducing data transmission while maintaining perceptual audio quality, allowing for flexible joint coding of arbitrary channel setups and improving stereo coding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026076313000001_ABST
    Figure 2026076313000001_ABST
Patent Text Reader

Abstract

It provides an improved and more flexible concept for multi-signal coding or decoding. [Solution] The multi-signal encoder includes a signal preprocessor (100) for individually preprocessing each audio signal to obtain at least three preprocessed audio signals. The preprocessing is performed so that the preprocessed audio signals are whited out compared to the unprocessed signals. The encoder also includes an adaptive joint signal processor (200) for obtaining at least three jointly processed signals or at least two jointly processed signals and an unprocessed signal, a signal encoder (300) for encoding each signal to obtain an encoded signal, and an output interface (400) for transmitting or storing an encoded multi-signal audio signal including an encoded signal, side information regarding preprocessing, and side information regarding processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments relate to MDCT-based multi-signal encoding and decoding systems having signal adaptive joint channel processing, where the signal is a channel and the multi-signal is a multi-channel signal or, alternatively, an audio signal that is a component of a sound field representation, such as the W, X, Y, Z of first-order ambisonics or any other component of a higher-order ambisonics representation. The signal can also be a signal representing the A format or B format or any other format of the sound field.

Background Art

[0002] ·In MPEG USAC [1], joint stereo encoding of two channels is performed using complex prediction (Complex Prediction), MPS2-1-2, or Unified Stereo with band-limited or full-band residual signals.

[0003] [[ID=,16]]·MPEG Surround [2] hierarchically combines OTT and TTT boxes for joint encoding of multi-channel audio, with or without transmission of residual signals.

[0004] ·MPEG-H Quad Channel Elements [3] hierarchically applies MPS2-1-2 stereo boxes following complex prediction / MS stereo boxes that construct a "fixed" 4x4 remix tree. [[ID=,21]]

[0005] ·AC4 [4] introduces new 3-channel, 4-channel, and 5-channel elements that enable remixing of transmitted channels via a transmitted mix matrix and subsequent joint stereo encoding information.

[0006] Previous publications have suggested using orthogonal transforms such as the Karhunen-Loeve Transform (KLT) for Enhanced Multichannel Audio Coding[5].

[0007] A Multichannel Coding Tool (MCT) [6] that supports joint coding of three or more channels enables flexible, signal-adaptive joint channel coding in the MDCT domain. This is achieved by the iterative combination and concatenation of complex stereo predictions of real values ​​of two specified channels, as well as stereo coding techniques such as rotational stereo coding (KLT).

[0008] In the context of 3D audio, loudspeaker channels are distributed across several height layers, resulting in horizontal and vertical channel pairs. The two-channel joint coding defined in USAC is insufficient to account for the spatial and perceptual relationships between channels. MPEG surround is applied with additional pre- / post-processing steps, and residual signals are transmitted separately, without the possibility of joint stereo coding that utilizes dependencies between, for example, left and right vertical residual signals. AC-4 introduces a dedicated N-channel element that allows for sufficient coding of joint coding parameters, but fails in common speaker setups with more channels, as proposed in new immersive playback scenarios (7.1+4, 22.2). MPEG-H is also limited to only four channels and cannot be dynamically applied to arbitrary channels, but only to a pre-configured fixed number of channels. MCT introduces the flexibility of signal-adaptive joint channel coding for arbitrary channels, but stereo processing is performed on windowed and transformed denormalized (non-whitened) signals. Furthermore, encoding the predictive counts or angles in each band of each stereo box requires a large number of bits. [Overview of the project] [Problems that the invention aims to solve]

[0009] The object of the present invention is to provide an improved and more flexible concept for multi-signal coding or decoding. [Means for solving the problem]

[0010] This objective is achieved by the multi-signal encoder of claim 1, the multi-signal decoder of claim 32, the method for performing multi-signal coding of claim 44, the method for performing multi-signal decoding of claim 45, the computer program of claim 46, or the coded signal of claim 47.

[0011] This invention is based on the discovery that multi-signal coding efficiency can be substantially improved by performing adaptive joint signal processing on a pre-processed audio signal rather than the original signal, such that the pre-processed audio signal is whiter than the unprocessed signal. On the decoder side, this means that post-processing is performed following the joint signal processing to obtain at least three processed decoded signals. These at least three processed decoded signals are post-processed according to the side information contained in the coded signal, such that the post-processed signals are no longer whiter than the unprocessed signal. The post-processed signals ultimately represent a decoded audio signal, i.e., a decoded multi-signal, either directly or following further signal processing operations.

[0012] Particularly in immersive 3D audio formats, efficient multi-channel coding is obtained that leverages the characteristics of multiple signals to reduce the amount of data transmitted while maintaining overall perceptual audio quality. In a preferred implementation, signal-adaptive joint coding within a multi-channel system is performed using a perceptually whitened spectrum, in addition to correcting for inter-channel level difference (ILD). The joint coding is preferably performed using a simple per-band M / S transform decision driven on the estimated number of bits of the entropy coder.

[0013] A multi-signal encoder for encoding at least three audio signals includes a signal preprocessor for individually preprocessing each audio signal to obtain at least three preprocessed audio signals, the preprocessing being performed so that the preprocessed audio signals are whitened relative to the unprocessed signals. Adaptive joint signal processing of at least three preprocessed audio signals is performed to obtain at least three jointly processed signals. This processing acts on the whitened signals. Preprocessing will result in the extraction of certain signal characteristics, such as spectral envelopes, or, if not extracted, will reduce the efficiency of joint signal processing, such as joint stereo or joint multichannel processing. In addition, to improve joint signal processing efficiency, broadband energy normalization is performed on at least three preprocessed audio signals so that each preprocessed audio signal has normalized energy. This broadband energy normalization is signaled to the encoded audio signals as side information so that this broadband energy normalization can be inverted on the decoder side following inverse joint stereo or joint multichannel signal processing. This preferred additional broadband energy normalization procedure improves adaptive joint signal processing efficiency to such an extent that the number of bandwidths, or even full frames, that can undergo mid / side processing, as opposed to left / right processing (dual-mono processing), is substantially increased. The overall efficiency of the stereo coding process improves even more as the number of bandwidths, or even full frames, that undergo common stereo or multi-channel processing, such as mid / side processing, increases.

[0014] The lowest efficiency is obtained when, from a stereo processing perspective, the adaptive joint signal processor needs to adaptively determine that a band or frame should be processed in "dual mono" or left / right processing. Here, the left and right channels are processed as is, but naturally within the whitened and energy-normalized region. However, when the adaptive joint signal processor determines that mid / side processing should be performed for a particular band or frame, the mid signal is calculated by adding the first and second channels, and the side signal is calculated by calculating the difference between the first and second channels of the channel pair. Typically, the mid signal is comparable to one of the first and second channels in terms of its range of values, while the side signal is typically a low-energy signal that can be encoded with high efficiency, or, in the most favorable circumstances, the side signal is zero, or so close to zero that the spectral domain of the side signal is quantized to zero, and therefore can be entropy encoded very efficiently. This entropy coding is performed by a signal encoder on each signal to obtain one or more coded signals, and the output interface of the multi-signal encoder transmits or stores coded multi-signal audio signals containing one or more coded signals, side information regarding preprocessing, and side information regarding adaptive joint signal processing.

[0015] On the decoder side, a signal decoder, typically including an entropy decoder, decodes at least three encoded signals, which typically depend on the bit distribution information preferably contained within. This bit distribution information is included as side information in the encoded multi-signal audio signal and can be derived on the encoder side, for example, by examining the energy of the signal at the input to the signal (entropy) encoder. The output of the signal decoder in the multi-signal decoder is input to a joint signal processor to perform joint signal processing according to the side information contained in the encoded signal to obtain at least three processed decoded signals. This joint signal processor preferably reverses the joint signal processing performed on the encoder side, typically performing inverse stereo or inverse multi-channel processing. In a preferred implementation, the joint signal processor applies processing operations to compute the left / right signals from the mid / side signals. However, if the joint signal processor determines from the side information that dual-mono processing already exists for a particular channel pair, this situation is recorded and used by the decoder for further processing.

[0016] The decoder-side joint signal processor may be a processor that operates in cascaded channel pair tree or simplified tree mode, similar to the encoder-side adaptive joint signal processor. A simplified tree also represents a kind of cascading process, but differs from a cascaded channel pair tree in that the output of a processed pair cannot become the input to another pair that will be processed.

[0017] With respect to the first channel pair used by the joint signal processor on the multi-signal decoder side to initiate joint signal processing, this first channel pair, which was the last channel pair processed on the encoder side, may have side information indicating dual mono in a certain bandwidth, but these dual mono signals may be used later in channel pair processing as mid or side signals. This is signaled by the corresponding side information regarding pairwise processing performed to obtain at least three individually encoded channels to be decoded on the decoder side.

[0018] The embodiment relates to an MDCT-based multi-signal coding and decoding system having signal adaptive joint channel processing, where the signal is a channel, and the multi-signal may be a multi-channel signal or, instead, an audio signal which is a component of a sound field representation, such as ambisonic components, i.e., W, X, Y, Z of first-order ambisonics, or any other arbitrary component of a higher-order ambisonic representation. The signal may also be a signal of a sound field representation in A-format or B-format or any other arbitrary format.

[0019] Next, further advantages of preferred embodiments are shown. The codec uses new concepts to combine the flexibility of signal-adaptive joint coding of any channel, as described in [6], by introducing the concepts described in [7] for joint stereo coding. These are, a) Use of perceptually whitened signals for further encoding (similar to the method used in speech coders). This has several advantages.

[0020] • Simplification of codec architecture • Compact representation of noise shaping characteristics / masking threshold (e.g., as LPC coefficients) • Integrates conversion and audio codec architectures, thus enabling audio / speech coding combinations. b) Use of ILD parameters of any channel for efficiently encoding the panned source c) Flexible bit distribution between energy-based processed channels.

[0021] The codec further uses frequency domain noise shaping (FDNS) to perceptually whiten the signal in a rate loop as described in [8] in combination with spectral envelope warping as described in [9]. The codec further normalizes the spectrum whitened by FDNS towards the average energy level using the ILD parameters. The channel pairs for joint encoding are adaptively selected as described in [6], and the stereo encoding consists of a per-band determination of M / S vs. L / R. The per-band determination of M / S is based on the estimated bit rate of each band when encoded in L / R and M / S modes as described in [7]. The bit rate distribution between the per-band M / S processed channels is energy-based.

[0022] Preferred embodiments of the present invention will be further described below with reference to the following attached drawings.

Brief Description of the Drawings

[0023] [Figure 1] Shows a block diagram of single-channel preprocessing in a preferred implementation. [Figure 2] Shows a preferred implementation of a block diagram of a multi-signal encoder. [Figure 3] Shows a preferred implementation of the cross-correlation vector and channel pair selection procedure of FIG. 2. [Figure 4] Shows an indexing scheme for channel pairs in a preferred implementation. [Figure 5a] Shows a preferred implementation of a multi-signal encoder according to the present invention. [Figure 5b] Shows a schematic diagram of an encoded multi-channel audio signal frame. [Figure 6]The procedure performed by the adaptive joint signal processor shown in Figure 5a is illustrated. [Figure 7] Figure 8 shows a preferred implementation configuration performed by the adaptive joint signal processor. [Figure 8] Figure 5 shows another preferred implementation configuration performed by the adaptive joint signal processor. [Figure 9] Figure 5 shows another procedure for performing the bit allocation used by the quantization coding processor. [Figure 10] A block diagram of a suitable implementation configuration for a multi-signal decoder is shown. [Figure 11] Figure 10 shows a preferred implementation configuration performed by the joint signal processor. [Figure 12] Figure 10 shows a preferred implementation configuration for the signal decoder. [Figure 13] This presents another preferred implementation of a joint signal processor in the context of bandwidth expansion or intelligent gap filling (IGF). [Figure 14] Figure 10 shows a further preferred implementation of the joint signal processor. [Figure 15a] Figure 10 shows a suitable processing block performed by the signal decoder and joint signal processor. [Figure 15b] This document describes implementations of post-processors for performing dewhitening operations and other optional procedures. [Modes for carrying out the invention]

[0024] Figure 5 shows a preferred implementation of a multi-signal encoder for encoding at least three audio signals. The at least three audio signals are input to a signal processor 100 for individually pre-processing each audio signal to obtain at least three pre-processed audio signals 180, the pre-processing being performed so that the pre-processed audio signals are whitened relative to the corresponding signals before pre-processing. The at least three pre-processed audio signals 180 are input to an adaptive joint signal processor 200 configured to perform processing on the at least three pre-processed audio signals to obtain at least three jointly processed signals, and in one embodiment, as will be described later, at least two jointly processed signals and an unprocessed signal. The multi-signal encoder includes a signal encoder 300 connected to the output of the adaptive joint signal processor 200 and configured to encode each signal output by the adaptive joint signal processor 200 to obtain one or more encoded signals. These encoded signals at the output of the signal encoder 300 are transferred to an output interface 400. The output interface 400 is configured to transmit or store encoded multi-signal audio signals 500, the encoded multi-signal audio signals 500 at the output of the output interface 400, include one or more encoded signals as generated by the signal encoder 300, side information 520 relating to pre-processing performed by the signal preprocessor 200, i.e., whitening information, and in addition, the encoded multi-signal audio signals also include side information 530 relating to processing performed by the adaptive joint signal processor 200, i.e., side information relating to adaptive joint signal processing.

[0025] In a preferred implementation, the signal encoder 300 includes a rate loop processor controlled by bit distribution information 536, which is generated by the adaptive joint signal processor 200 and transferred from block 200 to block 300, as well as to the output interface 400 within side information 530, and therefore also within the encoded multi-signal audio signal. The encoded multi-signal audio signal 500 is typically generated in a frame-by-frame manner, and framing, and typically corresponding windowing and time-frequency conversion, are performed within the signal preprocessor 100.

[0026] An illustrative diagram of a frame of the encoded multi-signal audio signal 500 is shown in Figure 5b. Figure 5b shows the bitstream portion 510 of the individually encoded signals as generated by block 300. Block 520 is for pre-processing side information generated by block 100 and transferred to the output interface 400. In addition, joint processing side information 530 is generated by the adaptive joint signal processor 200 in Figure 5a and introduced into the encoded multi-signal audio signal frame shown in Figure 5b. On the right side of Figure 5b, the next frame of the encoded multi-signal audio signal will be written to the serial bitstream, and on the left side of Figure 5b, the previous frame of the encoded multi-signal audio signal will be written.

[0027] As will be shown later, preprocessing includes time noise shaping and / or frequency domain noise shaping or LTP (Long-Term Prediction) or windowing operations. The corresponding preprocessing side information 550 may include at least one of time noise shaping (TNS) information, frequency domain noise shaping (FDNS) information, long-term prediction (LTP) information, or windowing or windowing information.

[0028] Time noise shaping involves predicting spectral frames against frequency. Spectral values ​​with higher frequencies are predicted using weighted combinations of spectral values ​​with lower frequencies. TNS side information includes the weights of the weighted combinations, also known as LPC coefficients, derived from the frequency predictions. The whitened spectral values ​​are the predicted residuals, or differences, for each spectral value between the original spectral values ​​and the predicted spectral values. On the decoder side, inverse prediction of LPC synthesis filtering is performed to reverse the TNS processing on the encoder side.

[0029] FDNS processing involves weighting the spectral values ​​of a frame using weighting coefficients for the corresponding spectral values, which are derived from LPC coefficients calculated from the windowed time-domain signal blocks / frames. FDNS side information includes a representation of the LPC coefficients derived from the time-domain signal.

[0030] Another whitening procedure useful in this invention is spectral equalization using a scale factor so that the equalized spectrum represents a whiter version than the unequalized version. The side information is the scale factor used for weighting, and the reverse procedure involves reversing the equalization on the decoder side using the transmitted scale factor.

[0031] Another whitening procedure involves performing inverse filtering of the spectrum using an inverse filter controlled by LPC coefficients derived from time-domain frames, as is known in the field of speech coding. The side information is the inverse filter information, and this inverse filtering is reversed in the decoder using the transmitted side information.

[0032] Another whitening procedure involves performing an LPC analysis in the time domain and calculating time-domain residual values, which are later converted to spectral bands. Typically, the spectral values ​​thus obtained are similar to those obtained by FDNS. On the decoder side, post-processing involves performing LPC synthesis using the transmitted LPC coefficient representation.

[0033] In a preferred implementation, the joint processing side information 530 includes pairwise processing side information 532, energy scaling information 534, and bit distribution information 536. The pairwise processing side information may include at least one of the following: channel pair-side information bits, full mid / side or dual-mono or per-band mid / side information, and, in the case of per-band mid / side display, a mid / side mask indicating on a per-band basis whether the bands in the frame are processed by mid / side or L / R processing. The pairwise processing side information may additionally include other bandwidth expansion information, such as intelligent gap filling (IGF) or SBR (spectral band replication) information.

[0034] Energy scaling information 534 may include, for each whitened, i.e., pre-processed signal 180, an energy scaling value and a flag indicating whether the energy scaling is upscaling or downscaling. For example, in the case of eight channels, block 534 includes eight scaling values, such as eight quantized ILD values, and for each of the eight channels, eight flags indicating whether the upscaling or downscaling was performed in the encoder or decoder. Encoder upscaling is required when the actual energy of a particular pre-processed channel in a frame is below the average energy of the frame across all channels, and downscaling is required when the actual energy of a particular channel in a frame is above the average energy across all channels in the frame. Joint processing side information may include bit distribution information for each jointly processed signal, or for each jointly processed signal, and the unprocessed signal if available, which is used by the signal encoder 300 as shown in Figure 5a and accordingly used by the signal decoder used as shown in Figure 10, which receives this bitstream information from the encoded signal via an input interface.

[0035] Figure 6 shows a preferred implementation of the adaptive joint signal processor. The adaptive joint signal processor 200 is configured to perform broadband energy normalization of at least three preprocessed audio signals so that each preprocessed audio signal has normalized energy. The output interface 400 is configured to include, as further side information, the broadband energy normalized value of each preprocessed audio signal, which corresponds to the energy scaling information 534 in Figure 5b. Figure 6 shows a preferred implementation of broadband energy normalization. In step 211, the broadband energy of each channel is calculated. The input to block 211 consists of the preprocessed (whitened) channels. As a result, C totalThe broadband energy value for each channel is obtained. In block 212, the average broadband energy is typically calculated by summing the individual values ​​and dividing the individual values ​​by the number of channels. However, other averaging procedures, such as geometric mean, can also be performed.

[0036] In step 213, each channel is normalized. For this purpose, a scaling factor or value and upscaling or downscaling information are determined. Thus, block 213 is configured to output a scaling flag for each channel, as shown in 534a. In block 214, the actual quantization of the scaling ratio determined in block 212 is performed, and this quantized scaling ratio is output for each channel in 534b. This quantized scaling ratio is the inter-channel level difference This is also shown as TIFF2026076313000002.tif515, i.e., for a specific channel k relative to a reference channel having an average energy. In block 215, the spectrum of each channel is scaled using a quantization scaling ratio. The scaling operation in block 215 is controlled by block 213, i.e., by information on whether upscaling or downscaling should be performed. The output of block 215 represents the scaled spectrum of each channel.

[0037] Figure 7 shows a preferred implementation of the adaptive joint signal processor 200 for cascaded pair processing. The adaptive joint signal processor 200 is configured to calculate the cross-correlation value of each possible channel pair, as shown in block 221. Block 229 shows the selection of the pair with the highest cross-correlation value, and in block 232a, the joint stereo processing mode is determined for this pair. The joint stereo processing mode may consist of mid / side coding for the full frame, or mid / side coding per band, i.e., for each of the multiple bands, it is determined whether this band should be processed in mid / side mode or L / R mode, or whether full-band dual-mono processing should be performed for this particular pair under consideration in the actual frame. In block 232b, the joint stereo processing of the selected pair is actually performed using the mode determined in block 232a.

[0038] In blocks 235 and 238, cascaded or non-cascaded processing using full-tree or simplified tree processing continues until a specific termination criterion is reached. At this criterion, for example, the pair display output by block 229 and the stereo mode processing information output by block 232a are generated and input to the bitstream of pairwise processing side information 532, as described with respect to Figure 5b.

[0039] Figure 8 shows a preferred implementation of an adaptive joint signal processor intended to prepare the signal encoding to be performed by the signal encoder 300 in Figure 5a. For this purpose, the adaptive joint signal processor 200 calculates the signal energy of each stereo-processed signal in block 282. Block 282 receives a joint stereo-processed signal as input, and if a channel has not been stereo-processed since it was found that this channel does not have sufficient cross-correlation with any other channel to form a useful channel pair, this channel is input to block 282 with inverted or modified or unnormalized energy. This is generally referred to as the “energy-recovered signal,” although the energy normalization performed in block 215 in Figure 6 does not necessarily have to be fully recovered. There are specific alternatives for processing channel signals that are not found to be useful for channel pairing with other channels. One procedure is to invert the scaling, which is initially performed in block 215 in Figure 6. Another procedure is to invert the scaling only partially, or another procedure is to weight the scaled channel in a particular different way, depending on the case.

[0040] In block 284, the total energy of all signals output by the adaptive joint signal processor 200 is calculated. Based on the signal energy of each stereo processed signal, or, if available, an energy recovery or energy weighted signal, and based on the total energy output by block 284, bit distribution information for each signal is calculated in block 286. The side information 536 generated by block 286 is transferred, on the one hand, to the signal encoder 300 in Figure 5a, and in addition, to the output interface 400 via the logic connection 530, so that this bit distribution information is included in the encoded multi-signal audio signal 500 in Figure 5a or Figure 5b.

[0041] The actual bit allocation is performed in a preferred embodiment based on the procedure shown in Figure 9. In the first step, the minimum number of bits for the non-LFE (low-frequency emphasis) channel is allocated, and if available, the low-frequency emphasis channel bits are allocated. These minimum number of bits are required by the signal encoder 300 regardless of the specific signal content. The remaining bits are allocated according to the bit distribution information 536 generated by block 286 in Figure 8 and input to block 291. The allocation is based on the quantized energy ratio, and it is preferable to use the quantized energy ratio rather than the unquantized energy ratio.

[0042] In step 292, the refinement is performed. If the remaining bits are allocated and the quantization results in a number of bits higher than the number of available bits, the bits allocated in block 291 must be subtracted. However, if the quantization of the energy ratio still has bits that need to be allocated in the allocation procedure in block 291, these bits may be additionally allocated or distributed in refinement step 292. Following the refinement step, if there are still bits available for use in the signal encoder, the final donation step 293 is performed, and the final donation is made for the channel with the highest energy. At the output of step 293, the bit allocations allocated to each signal are available.

[0043] In step 300, quantization and entropy coding are performed on each channel using the assigned bit allocation generated by the processes in steps 290, 291, 292, and 293. Essentially, the bit allocation is performed so that higher energy channels / signals are quantized more accurately than lower energy channels / signals. Importantly, the bit allocation is not performed using the original signal or the whitened signal, but using the signal at the output of the adaptive joint signal processor 200, which has a different energy than the signal input to the adaptive joint signal processor for joint channel processing. In this regard, while channel pair processing is a preferred implementation, it should also be noted that other groups of channels may be selected and processed by cross-correlation. For example, groups of three or even four channels may be formed by the adaptive joint signal processor and processed accordingly in a cascaded full procedure, a cascaded procedure using a simplified tree, or a non-cascaded procedure.

[0044] The bit assignments shown in blocks 290, 291, 292, and 293 are performed in the same manner on the decoder side by the signal decoder 700 in Figure 10, using distribution information 536 extracted from the encoded multi-signal audio signal 500.

[0045] Preferred Embodiment In this implementation, the codec uses new concepts to combine the flexibility of arbitrary channel signal-adaptive joint coding as described in [6] by introducing the concepts described in [7] for joint stereo coding. These are, a) Use of perceptually whitened signals for further encoding (similar to the method used in speech coders). This has several advantages.

[0046] • Simplification of codec architecture • Compact representation of noise shaping characteristics / masking threshold (e.g., as LPC coefficients) • Integrates conversion and audio codec architectures, thus enabling audio / speech coding combinations. b) Using ILD parameters on any channel to efficiently encode the panned source c) Flexible bit distribution between processed channels based on energy.

[0047] The codec uses frequency-domain noise shaping (FDNS) to perceptually whiten the signal in a rate loop as described in [8], in combination with spectral envelope warping as described in [9]. The codec further normalizes the FDNS-whitened spectrum toward the average energy level using ILD parameters. Channel pairs for joint coding are adaptively selected as described in [6], and stereo coding consists of band-by-band M / S vs L / R determinations. Band-by-band M / S determinations are based on the estimated bitrate of each band when coded in L / R and M / S modes as described in [7]. The band-by-band M / S processed bitrate distribution between channels is energy-based.

[0048] The embodiment relates to an MDCT-based multi-signal coding and decoding system having signal-adaptive joint channel processing, where a signal is a channel, and a multi-signal can be a multi-channel signal, or instead, an audio signal that is a component of a sound field representation, such as ambisonic components, i.e., W, X, Y, Z of first-order ambisonics, or any other arbitrary component of a higher-order ambisonic representation. A signal can also be a signal of a sound field representation in A-format or B-format or any other arbitrary format. Thus, the same disclosure given to “channels” is also valid for “components” or other “signals” of a multi-signal audio signal.

[0049] Single-channel encoder processing up to the whitening spectrum Following the processing steps shown in the block diagram of Figure 1, each single channel TIFF2026076313000003.tif43 is analyzed and converted into a whitened MDCT region spectrum.

[0050] The processing blocks for the time-domain transient detector, windowing, MDCT, MDST, and OLA are described in [8]. The MDCT and MDST form a Modulated Complex Lapped Transform (MCLT), and performing the MDCT and MDST separately is equivalent to performing the MCLT, while "MCLT to MDCT" means taking only the MDCT portion of the MCLT and discarding the MDST.

[0051] Time-domain noise shaping (TNS) is performed as described in [8], with the addition that the order of TNS and frequency-domain noise shaping (FDNS) is adaptive. The presence of the two TNS boxes in the figure should be understood as the possibility of changing the order of FDNS and TNS. The determination of the order of FDNS and TNS may be as described, for example, in [9].

[0052] Frequency-domain noise shaping (FDNS) and the calculation of FDNS parameters are similar to the procedure described in [9]. One difference is that for frames where TNS is inactive, the FDNS parameters are calculated from the MCLT spectrum. For frames where TNS is active, the MDST spectrum is estimated from the MDCT spectrum.

[0053] Figure 1 shows a preferred implementation of a signal processor 100 that performs whitening of at least three audio signals to obtain individually preprocessed whitened signals 180. The signal preprocessor 100 includes an input for a time-domain input signal for channel k. This signal is input to a windower 102, a transient detector 104, and an LTP parameter calculator 106. The transient detector 104 detects whether the current portion of the input signal is transient, and if so, controls the windower 102 to set a shorter window length. The window indication, i.e., which window length was selected, is also included in the side information, particularly the preprocessing side information 520 in Figure 5b. In addition, LTP parameters calculated by block 106 are also introduced into the side information block, and these LTP parameters may be used, for example, to perform certain post-processing of the decoded signal or other procedures known in the art. The windower 140 generates windowed time-domain frames which are introduced into the time-spectral converter 108. The time-spectrum converter 108 preferably performs a complex wrap transform. From this complex wrap transform, the real part can be derived to obtain the result of the MDCT transform, as shown in block 112. The result of block 112, i.e., the MDCT spectrum, is input to the TNS block 114a and subsequently the combined FDNS block 116. Alternatively, only FDNS may be performed without the TNS block 114a, or vice versa, or the TNS processing may be performed after the FDNS processing, as shown by block 114b. Typically, either block 114a or block 114b is present. At the output of block 114b, when block 114a is not present, or at the output of block 116 when block 114b is not present, whitened and individually processed signals, i.e., preprocessed signals, are obtained for each channel k. The TNS block 114a or 114b and the FDNS block 116 generate preprocessing information and transfer it to the side information 520.

[0054] In no case is it necessary to perform a complex transformation within block 108. In addition, a time-spectrum transformer that performs only MDCT is also sufficient for certain applications, and if the imaginary part of the transformation is required, this imaginary part can, in some cases, also be estimated from the real part. A feature of the TNS / FDNS process is that when TNS is inactive, the FDNS parameters are calculated from the complex spectrum, i.e., from the MCLT spectrum, and in frames where TNS is active, the MDST spectrum is estimated from the MDCT spectrum, so that the full complex spectrum is always available in the frequency-domain noise shaping operation.

[0055] Description of the Joint Channel Coding System In the described system, after each channel is converted to a whitened MDCT region, a signal-adaptive utilization of various similarities between any two channels for joint coding is applied based on the algorithm described in [6]. From this procedure, each channel pair is detected and selected to be jointly coded using band-by-band M / S conversion.

[0056] An overview of the encoding system is shown in Figure 2. For simplicity, block arrows represent single-channel processing (i.e., processing blocks are applied to each channel), and the "MDCT region analysis" block is shown in detail in Figure 1.

[0057] The following paragraphs detail the individual steps of the algorithm applied to each frame. The dataflow graph of the algorithm described is shown in Figure 3.

[0058] It should be noted that the initial system configuration includes a channel mask that indicates which channels the multi-channel joint coding tool will be active on. Therefore, for inputs where LFE (Low-Frequency Effect / Enhancement) channels exist, these are not considered in the tool's processing steps.

[0059] Energy normalization of all channels toward average energy M / S conversion is inefficient when ILD is present, i.e., when the channels are panned. The amplitude of the perceptually whitened spectrum of all channels is averaged to the energy level. This problem can be avoided by normalizing the file to TIFF2026076313000004.tif43.

[0060] Each channel Regarding TIFF2026076313000005.tif529, energy Calculate TIFF2026076313000006.tif45. TIFF2026076313000007.tif2040 Here, TIFF2026076313000008.tif44 is the total number of spectral coefficients.

[0061] • Calculate the average energy. TIFF2026076313000009.tif1228 · Normalize the spectrum of each channel toward the average energy. In the case of TIFF2026076313000010.tif514 (downscaling), TIFF2026076313000011.tif1116 Here, TIFF2026076313000012.tif33 is the scaling ratio. The scaling ratio is uniformly quantized and sent to the decoder as side information bits. TIFF2026076313000013.tif5126 Here, TIFF2026076313000014.tif447 Next, the quantization scaling ratio at which the spectrum is ultimately scaled is given by: In the case of TIFF2026076313000015.tif1135TIFF2026076313000016.tif514 (upscaling), TIFF2026076313000017.tif1015 and TIFF2026076313000018.tif1135 Here, TIFF2026076313000019.tif514 is calculated in the same way as before.

[0062] To distinguish whether to perform downscaling / upscaling in the decoder, and to restore normalization, each channel In addition to the TIFF2026076313000020.tif47 value, a 1-bit flag (0 = downscaling / 1 = upscaling) is sent. TIFF2026076313000021.tif418 is a transmitted and quantized scaling value. This indicates the number of bits used in TIFF2026076313000022.tif47. This value is known to the encoder and decoder and does not need to be transmitted in the encoded audio signal.

[0063] Calculation of normalized inter-channel cross-correlation values ​​for all possible channel pairs In this step, the normalized cross-correlation values ​​between the channels of each possible channel pair are calculated in order to determine and select which channel pair has the highest similarity and is therefore suitable to be selected as a pair for stereo joint coding. The normalized cross-correlation values ​​for each channel pair are given by the cross-spectrum as follows: TIFF2026076313000023.tif1029 Here, TIFF2026076313000024.tif1557TIFF2026076313000025.tif44 are the total number of spectral counts per frame. TIFF2026076313000026.tif412 and TIFF2026076313000027.tif411 contains the spectra of each channel pair under consideration.

[0064] The normalized cross-correlation values ​​for each paired channel are stored in the cross-correlation vector. TIFF2026076313000028.tif535 Here, TIFF2026076313000029.tif554 has the maximum number of possible pairs.

[0065] As shown in Figure 1, the transient detector can have different block sizes (e.g., a window block size of 10 or 20 ms). Therefore, the inter-channel cross-correlation is calculated assuming that the spectral resolution of both channels is the same. Otherwise, the value is set to 0, so such a channel pair is not reliably selected for joint coding.

[0066] An indexing scheme is used to uniquely represent each channel pair. An example of such a scheme for indexing six input channels is shown in Figure 4.

[0067] The same indexing scheme used to signal channel pairs to the decoder is maintained throughout the algorithm. The amount of bits required to signal one channel is: TIFF2026076313000030.tif552 Channel pair selection and co-encoded stereo processing After calculating the cross-correlation vectors, the first channel pair to consider for joint coding is the one with the highest cross-correlation value and preferably a minimum threshold of 0.3.

[0068] The selected pair of channels serves as input to the stereo coding procedure, i.e., band-by-band M / S conversion. For each spectral band, the decision of whether the channels are coded using M / S or discrete L / R coding depends on the estimated bitrate in each case. The coding method that is less demanding in terms of bits is selected. This procedure is described in detail in [7].

[0069] The output of this process yields updated spectra for each channel of the selected channel pair. Additionally, information that needs to be shared with the decoder regarding this channel pair (side information) is created, namely which stereo mode is selected (full M / S, dual mono, or per-band M / S), and, if per-band M / S is the selected mode, masks are created indicating whether M / S coding is selected (1) or L / R coding is selected (0).

[0070] In the next step, there are two variations of the algorithm.

[0071] Cascade Channel Pair Tree In this variation, the cross-correlation vector is updated for the channel pair affected by the modified spectrum (if M / S transformation is involved) of the selected channel pair. For example, in the case of six channels, if the selected and processed channel pair is indexed at 0 in Figure 4, i.e., channel 0 is encoded by channel 1, then after stereo processing, the cross-correlation of the affected channel pair needs to be recalculated at indices 0, 1, 2, 3, 4, 5, 6, 7, and 8.

[0072] Next, the procedure continues as described above. The channel pair with the greatest cross-correlation is selected, ensuring it exceeds the minimum threshold, and the stereo operation is applied. This means that channels that were part of the previous channel pair may be re-selected to function as inputs to the new channel pair, a phenomenon called "cascading." This can occur because there is still correlation between the output of the channel pair and any other arbitrary channel representing a different direction in the spatial domain. Naturally, the same channel pair should not be selected twice.

[0073] The maximum number of allowed iterations (absolute maximum is The procedure continues when the value (TIFF2026076313000031.tif43) is reached, or when no channel pair values ​​exceed the threshold of 0.3 after updating the cross-correlation vector (no correlation between any given channels).

[0074] • Simplified tree The cascaded channel pair tree process is theoretically optimal because it attempts to remove correlations from all arbitrary channels and provide maximum energy compression. On the other hand, the number of channel pairs selected It can become quite complex as it may exceed TIFF2026076313000032.tif76, resulting in even more complex calculations (due to the M / S determination process for stereo operation) and additional metadata that needs to be sent to the decoder for each channel pair.

[0075] In the simplified tree variation, "cascading" is not permitted. This is ensured from the process described above, where the values ​​of channel pairs affected by the previous channel pair stereo operation are not recalculated and are set to 0 while the cross-correlation vector is being updated. Therefore, it is not possible to select a channel pair in which one of the channels was already part of an existing channel pair.

[0076] This is a variation illustrating the "adaptive joint channel processing" shown in Figure 2.

[0077] In this case, the maximum number of channel pairs that can be selected is Since it is TIFF2026076313000033.tif77, similar complexity arises in systems with specific channel pairs (for example, L and R, rear L and rear R).

[0078] It should be noted that the stereo operation of a selected channel pair may not change the channel spectrum. This occurs when the M / S determination algorithm decides to set the encoding mode to "dual mono". In this case, any channels involved are encoded separately and are no longer considered a channel pair. Updating the cross-correlation vector also has no effect. To continue the process, the next highest channel pair is considered. The steps in this case are continued as described above.

[0079] Maintain the channel pair selection (stereo tree) from the previous frame. In many cases, the normalized cross-correlation values ​​of any channel pair per frame may be close, and therefore the selection may frequently switch between these close values. This can lead to frequent channel pair tree switching, resulting in unstable audibility of the output system. Therefore, it is chosen to use a stabilization mechanism in which a new set of channel pairs is selected only when there is a significant change in the signal and the similarity between any two channels changes. To detect this, the cross-correlation vector of the current frame is compared to the vector of the previous frame, and if the difference is greater than a certain threshold, the selection of a new channel pair is permitted.

[0080] The time variation of the cross-correlation vector is calculated as follows: In the case of TIFF2026076313000034.tif1560TIFF2026076313000035.tif517, the selection of a new channel pair to be co-encoded is permitted, as described in the previous step. The selected threshold is: TIFF2026076313000036.tif548 On the other hand, if the difference is small, the same channel pair tree as the previous frame is used. For each given channel pair, the bandwidth-based M / S operation is applied as described above. However, if the normalized cross-correlation value of a given channel pair does not exceed the threshold of 0.3, the selection of a new channel pair to create a new tree is initiated.

[0081] Restore single-channel energy After the iterative process for channel pair selection is complete, there may be channels that are not part of any channel pair and are therefore encoded separately. For these channels, the initial normalization of the energy level toward the average energy level is reversed to the original energy level. Depending on the flag that signals upscaling or downscaling, the energy of these channels is the reciprocal of the quantization scaling ratio. It will be restored using TIFF2026076313000037.tif99.

[0082] IGF for multi-channel processing For IGF analysis, in the case of stereo channel pairs, additional joint stereo processing is applied, as fully described in

[10] . This is necessary because, in a particular target range of the IGF spectrum, the signals may be highly correlated panned sources. If the source regions selected for this particular region are not well correlated, the spatial image may be impaired due to the uncorrelated source regions, even if the energies match in the target region.

[0083] Therefore, if the stereo mode of the core region differs from the stereo mode of the IGF region, or if the core stereo mode is flagged as M / S per band, stereo IGF is applied to each channel pair. If these conditions do not apply, single-channel IGF analysis is performed. If there are single channels within a channel pair that are not co-encoded, these also undergo single-channel IGF analysis.

[0084] The distribution of bits available to encode the spectrum of each channel After the joint channel pair stereo processing process, each channel is quantized and encoded separately by an entropy coder. Therefore, each channel should be assigned a number of available bits. In this step, the energy of the processed channels is used to distribute the total number of available bits to each channel.

[0085] The energy of each channel, although its calculation is described above in the normalization step, is recalculated as a spectrum because each channel may have changed due to the jointing process. The new energy is: It is represented as TIFF2026076313000038.tif536. As a first step, the energy-based ratios for distributing the bits are calculated. TIFF2026076313000039.tif1327 Note that if the input also consists of an LFE channel, this is not considered in the ratio calculation. In an LFE channel, the minimum number of bits is considered only if the channel has non-zero content. The file TIFF2026076313000040.tif514 is assigned. The ratios are quantized uniformly. TIFF2026076313000041.tif5101TIFF2026076313000042.tif440 Quantized ratio TIFF2026076313000043.tif44 is stored within a bitstream used by the decoder to allocate the same amount of bits to each channel in order to read the transmitted channel spectral coefficients.

[0086] The bit distribution scheme is described below.

[0087] • Entropy coder for each channel Allocate the minimum number of bits required by TIFF2026076313000044.tif514. • The remaining bits, i.e. TIFF2026076313000045.tif697 is a quantized ratio The file will be split using TIFF2026076313000046.tif56. TIFF2026076313000047.tif1159 · Due to the quantized ratio, the bits are almost dispersed, therefore It could be TIFF2026076313000048.tif664. Therefore, in the second improvement step, the difference TIFF2026076313000049.tif558 is channel bit It is subtracted proportionally from TIFF2026076313000050.tif510. TIFF2026076313000051.tif1165 · After the improvement step, Compared to TIFF2026076313000052.tif516, still If there is a mismatch in TIFF2026076313000053.tif515, the difference (usually a very small number of bits) is donated to the channel with the highest energy.

[0088] The exact same procedure is repeated from the decoder to determine the amount of bits to be read in order to decode the spectral coefficients of each channel. TIFF2026076313000054.tif415 contains bit distribution information. This indicates the number of bits used in TIFF2026076313000055.tif510. This value is known to the encoder and decoder and does not need to be transmitted in the encoded audio signal.

[0089] Quantization and encoding of each channel The entropy coding, including quantization, noise filling, and rate loop, is described in [8]. The rate loop is estimated Optimization is possible using TIFF2026076313000056.tif48. The power spectrum P (magnitude of MCLT) is used for quantization and intelligent gap-filling (IGF) tone / noise measurements, as described in [8]. Since the whitened and band-by-band M / S processed MDCT spectrum is used for the power spectrum, the same FDNS and M / S processing must be applied to the MDST spectrum. The same ILD-based normalization scaling applied to the MDCT must also be applied to the MDST spectrum. In frames where TNS is active, the MDST spectrum used for power spectrum calculations is estimated from the whitened and M / S processed MDCT spectrum.

[0090] Figure 2 shows a block diagram of a preferred implementation of an encoder, in particular the adaptive joint signal processor 200 shown in Figure 2. At least three preprocessed audio signals 180 are all input to an energy normalization block 210, which generates channel energy ratio side bits 534 at its output, consisting of a quantized ratio on the one hand and a flag for each channel indicating upscaling or downscaling on the other. However, other procedures without explicit upscaling or downscaling flags may also be performed.

[0091] The normalized channels are input to block 220 to perform cross-correlation vector calculation and channel pair selection. Based on the procedure in block 220, which is preferably an iterative procedure using cascaded full tree or cascaded and simplified tree processing, or a non-iterative non-cascaded process, the corresponding stereo operation is performed in block 240, which may perform mid / side processing for the entire band or per band, or any other corresponding stereo processing operation such as rotation, scaling, or any weighted or unweighted linear or non-linear combination.

[0092] At the output of block 240, stereo intelligent gap filling (IGF) processing, or any other arbitrary bandwidth expansion processing such as spectral band duplication or harmonic band processing, may be performed. Processing of individual channel pairs is transmitted via channel pair side information bits, and although not shown in Figure 2, IGF or general bandwidth expansion parameters generated by block 260 are also written to a bitstream for joint processing side information 530, and specifically for pairwise processing side information 532 in Figure 5b.

[0093] The final stage of Figure 2 is a channel bit distribution processor 280 that calculates bit allocations, as described with respect to Figure 9, for example. Figure 2 shows a schematic diagram of a signal encoder 300 as a quantizer and encoder controlled by channel bitrate side information 530, and further, an output interface 400 or bitstream writer 400 that combines the results of the signal encoder 300 with all the necessary side information bits 520, 530 from Figure 5b.

[0094] Figure 3 shows a preferred implementation of the substantial procedure performed by blocks 210, 220, and 240. Following the start of the procedure, ILD normalization is performed as shown in 210 in Figure 2 or Figure 3. In step 221, the cross-correlation vector is calculated. The cross-correlation vector consists of the normalized cross-correlation values ​​for each possible channel pair of channels from 0 to N output by block 210. For example, in Figure 4, where there are six channels, 15 different possibilities from 0 to 14 can be examined. The first element of the cross-correlation vector has the cross-correlation value between channel 0 and channel 1, and for example, the element of the cross-correlation vector with index 11 has the cross-correlation between channel 2 and channel 5.

[0095] In step 222, a calculation is performed to determine whether the tree determined in the previous frame should be maintained. For this purpose, the time variation of the cross-correlation vector is calculated, preferably the sum of the individual differences of the cross-correlation vector, in particular the magnitude of the differences. In step 223, it is determined whether the sum of the differences is greater than a threshold. If so, in step 224, the flag keepTree is set to 0, which means that the tree is not maintained, but a new tree is calculated. However, if it is determined that the sum is less than the threshold, block 225 sets the flag keepTree=1 so that the tree determined from the previous frame is also applied to the current frame.

[0096] In step 226, the iteration termination criterion is checked. If it is determined that the maximum number of channel pairs (CPs) has not been reached, which is naturally the first time block 226 has been accessed, and furthermore, when the flag keepTree is set to 0 as determined by block 228, the procedure proceeds to block 229 for selecting the channel pair with the greatest cross-correlation from the cross-correlation vector. However, when the tree from previous frames is maintained, i.e., when keepTree is equal to 1 as checked in block 225, block 230 determines whether the cross-correlation of the "forced" channel pair is greater than a threshold. If this is not the case, the procedure proceeds to step 227, which means that a new tree should still be determined, although the procedure in block 223 determined it in reverse. The evaluation in block 230, and the corresponding result in block 227, may overturn the decisions in blocks 223 and 225.

[0097] In block 231, it is determined whether the channel pair with the highest cross-correlation has a value greater than 0.3. If so, the stereo operation in block 232 is performed, which is also shown as 240 in Figure 2. In block 233, if it is determined that the stereo operation was dual mono, a value equal to 0, keepTree, is set in block 234. However, if it is determined that the stereo mode was different from dual mono, a mid / side operation is performed, and the cross-correlation vector 235 needs to be recalculated because the output of the stereo operation block 240 (or 232) is different for processing. Updating the CC vector 235 is only necessary when there was actually a mid / side stereo operation, or a stereo operation that was different from dual mono in general.

[0098] However, if the check in block 226 or block 231 results in a "no" answer, control proceeds to block 236 to check whether a single channel exists. If this is the case, i.e., if a single channel is found that has not been processed with other channels in the channel pair processing, the ILD normalization is inverted in block 237. Alternatively, the inversion in block 237 may be merely a partial inversion, or it may be some kind of weighting.

[0099] If the iteration is complete, and if blocks 236 and 237 are also complete, the procedure is finished, all channel pairs have been processed, and at the output of the adaptive joint signal processor, there are at least three jointly processed signals if block 236 gives a "no" answer, and at least two jointly processed signals and an unprocessed signal corresponding to a "single channel" if block 236 gives a "yes" answer.

[0100] Description of the decoding system The decoding process begins with decoding and dequantizing the spectra of the co-encoded channels, followed by noise filling, as described in 6.2.2. “MDCT-based TCX” in

[11] or

[12] . The number of bits assigned to each channel is encoded in the bitstream, along with the window length, stereo mode, and bitrate ratio. This is determined based on TIFF2026076313000057.tif56. The number of bits assigned to each channel must be known before the bitstream is fully decoded.

[0101] In the Intelligent Gap Filling (IGF) block, lines quantized to zero within a specific range of the spectrum, called target tiles, are filled with processed content from different ranges of the spectrum, called source tiles. Due to per-band stereo processing, the stereo representation (i.e., L / R or M / S) may differ between the source and target tiles. To ensure good quality, if the representation of the source tiles differs from that of the target tiles, the source tiles are processed to be converted to the representation of the target file before gap filling in the decoder. This procedure has already been described in

[10] . In contrast to

[11] and

[12] , the IGF itself is applied in a whitened spectral region rather than the original spectral region. In contrast to known stereo codecs (e.g.,

[10] ), the IGF is applied in a whitened and ILD-corrected spectral region.

[0102] Bitstream signaling also reveals whether there are co-encoded channel pairs. The inverse process begins with the last channel pair formed by the encoder, especially in a cascaded channel pair tree, to convert each channel back to its original whitened spectrum. For each channel pair, inverse stereo processing is applied based on the determination of the stereo mode and M / S per band.

[0103] For all channels involved in a channel pair and co-encoded, the spectrum is sent from the encoder. Based on the TIFF2026076313000058.tif514 value, it is denormalized to the original energy level.

[0104] Figure 10 shows a preferred implementation of a multi-signal decoder for decoding an encoded signal 500. The multi-signal decoder includes an input interface 600 and a signal decoder 700 for decoding at least three encoded signals output by the input interface 600. The multi-signal decoder includes a joint signal processor 800 for performing joint signal processing according to side information contained in the encoded signal to obtain at least three processed decoded signals. The multi-signal decoder includes a post-processor 900 for post-processing at least three processed decoded signals according to side information contained in the encoded signal. In particular, the post-processing is performed such that the post-processed signal is no whiter than the signal before post-processing. The post-processed signal directly or indirectly represents the decoded audio signal 1000.

[0105] The side information extracted by the input interface 600 and transferred to the joint signal processor 800 is the side information 530 shown in Figure 5b, and the side information extracted by the input interface 600 from the encoded multi-signal audio signal transferred to the post-processor 900 to perform a dewhitening operation is the side information 520 illustrated and described with respect to Figure 5b.

[0106] The joint signal processor 800 is configured to extract and receive the energy normalized value of each joint stereo decoded signal from the input interface 600. This energy normalized value of each joint stereo decoded signal corresponds to the energy scaling information 530 in Figure 5b. The adaptive joint signal processor 200 is configured to perform pairwise processing 820 of the decoded signal using the joint stereo side information or joint stereo mode indicated by the joint stereo side information 532 contained in the encoded audio signal 500, in order to obtain the joint stereo decoded signal at the output of block 820. In block 830, a rescaling operation, specifically energy rescaling of the joint stereo decoded signal, is performed using the energy normalized value to obtain the decoded signal processed in block 800 in Figure 10.

[0107] As described in Block 237 with respect to Figure 3, in order to ensure that channels receive inverse ILD normalization, the joint signal processor 800 is configured to check whether the energy normalized value extracted from the encoded signal of a particular signal has a predetermined value. If this is the case, either no energy rescaling is performed, reduced energy rescaling is performed on the particular signal, or any other weighting operation is performed on this individual channel when the energy normalized value has this predetermined value.

[0108] In one embodiment, the signal decoder 700 is configured to receive bit distribution values ​​for each encoded signal from the input interface 600, as shown in block 620. These bit distribution values, shown in 536 of Figure 12, are transferred to block 720 so that the signal decoder 700 can determine the bit distribution to be used. Preferably, for determining the bit distribution to be used in block 720 of Figure 12, the same steps described for the encoders of Figures 6 and 9, namely steps 290, 291, 292, and 293, are performed by the signal decoder 700. In blocks 710 / 730, individual decodings are performed to obtain input to the joint signal processor 800 of Figure 10.

[0109] The joint signal processor 800 has bandwidth duplication, bandwidth expansion, or intelligent gap-filling processing functions that use specific side information contained in side information block 532. This side information is transferred to block 810, and block 820 performs joint stereo (decoder) processing using the results of the bandwidth expansion procedure applied by block 810. In block 810, the intelligent gap-filling procedure is configured to convert the source range from one stereo representation to another when the target range for bandwidth expansion or IGF processing is indicated to have a different stereo representation. When the target range is indicated to have a mid / side stereo mode and the source range is indicated to have an L / R stereo mode, the stereo mode of the L / R source range is converted to the stereo mode of the mid / side source range, and then IGF processing is performed using the mid / side stereo mode representation of the source range.

[0110] Figure 14 shows a preferred implementation of the joint signal processor 800. The joint signal processor is configured to extract ordered signal pair information, as shown in block 630. This extraction can be performed by the input interface 600, or the joint signal processor can extract this information from the output of the input interface, or the information can be extracted directly without a specific input interface, as in the case of other extraction procedures described with respect to the joint signal processor or signal decoder.

[0111] In block 820, the joint signal processor performs a preferred cascaded inverse process, starting with the last signal pair, where the term “last” refers to the processing order determined and performed by the encoder. In the decoder, the “last” signal pair is the one processed first. Block 820 receives side information 532 indicating, for each signal pair, as indicated by the signal pair information shown in block 630 and implemented, for example, in the manner described with respect to Figure 4, that a particular pair is a dual-mono, full-MS, or per-band MS procedure with an associated MS mask.

[0112] Following the inverse processing in block 820, the denormalization of the signals contained in the channel pair is performed again in block 830, depending on the side information 534 which indicates the normalization information for each channel. The denormalization shown with respect to block 830 in Figure 14 is preferably a rescaling using the energy normalized value as a downscaling when flag 534a has a first value, and a rescaling as an upscaling when flag 534a has a second value different from the first value.

[0113] Figure 15a shows a preferred implementation configuration as a block diagram of the signal decoder and joint signal processor of Figure 10, and Figure 15b shows a preferred implementation configuration block diagram representation of the post-processor 900 of Figure 10.

[0114] The signal decoder 700 includes a decoder and inverse quantizer stage 710 for the spectrum contained in the encoded signal 500. The signal decoder 700 includes a bit assigner 720 that receives, preferably, a window length, a specific stereo mode, and bit assignment information for each encoded signal as side information. In a preferred implementation, the bit assigner 720 performs bit assignment, particularly using steps 290, 291, 292, and 293, where the bit assignment information for each encoded signal is used in step 291, and the window length and stereo mode information is used in block 290 or 291.

[0115] In block 730, noise filling, which also preferably uses noise-filling side information, is performed on the spectral range that is quantized to zero and is not within the IGF range. Noise filling is preferably limited to the low-bandwidth portion of the signal output by block 710. In block 810, intelligent gap filling, or generally bandwidth grading, is performed using specific side information, which importantly acts on the whitened spectrum.

[0116] In block 820, using the side information, the inverse stereo processor performs steps to reverse the processing performed in item 240 of Figure 2. Final descaling is performed using the per-channel transmit and quantized ILD parameters contained in the side information. The output of block 830 is input to block 910, a post-processor that performs inverse TNS processing and / or inverse frequency-domain noise shaping processing or any other dewhitening operation. The output of block 910 is a simple spectrum that is converted to the time domain by the frequency-time converter 920. The outputs of block 920 of adjacent frames are finally superimposed in the superimposed-adding processor 930 according to a specific encoding or decoding rule to obtain a number of decoded audio signals, or generally a decoded audio signal 1000, from the superimposed operation. This signal 1000 may consist of individual channels, or components of a sound field representation such as ambisonic components, or any other components of a higher-order ambisonic representation. The signal can also be a representation of the sound field in A-format, B-format, or any other format. All of these alternatives are shown together as the decoded audio signal 1000 in Figure 15b.

[0117] Next, further advantages and specific features of preferred embodiments are shown.

[0118] The scope of the present invention is to provide a solution to the principle from [6] when processing perceptually whitened and ILD parameter-corrected signals.

[0119] The combination of FDNS using rate loops as described in [8] and spectral envelope warping as described in [9] provides a simple but highly effective method for separating quantization noise and perceptual shaping of rate loops.

[0120] Using the average energy level for all channels of the spectrum whitened with FDNS provides a simple yet effective way to determine whether the M / S processing described in [7] is beneficial for each channel pair selected for joint coding.

[0121] It is sufficient to encode a single broadband ILD for each channel of the described system, thus achieving bit savings in contrast to known approaches.

[0122] By selecting channel pairs for joint coding using highly cross-correlated signals, a full-spectrum M / S conversion is typically achieved, and therefore, transmitting M / S or L / R in each band is almost always replaced by a single bit transmitting a complete M / S conversion, resulting in further average bit savings.

[0123] • Flexible and simple bit distribution based on the energy of the processed channels.

[0124] Features of a Preferred Embodiment As described in the previous paragraph, in this implementation, the codec uses novel means to combine the flexibility of signal-adaptive joint coding of any channel, as described in [6], by introducing the concept of joint stereo coding as described in [7]. The novelty of the proposed invention is summarized in the following differences:

[0125] • The joining process for each channel pair differs from the multi-channel processing described in [6] with respect to global ILD correction. Global ILD equalizes the levels of the channels before selecting channel pairs for M / S determination and processing, thus enabling more efficient stereo coding, especially for panned sources.

[0126] • The jointing process for each channel pair differs from the stereo processing described in [7] with respect to global ILD correction. The proposed system does not have global ILD correction for each channel pair. There is a normalization that brings all channels to a single energy level, i.e., an average energy level, so that the M / S determination mechanism described in [7] can be used for any channel. This normalization is performed before selecting channel pairs for jointing.

[0127] After the adaptive channel pair selection process, if there are channels that are not part of the channel pair for joint processing, their energy levels are returned to their initial energy levels.

[0128] As described in [7], the bit distribution of entropy coding is not implemented for each channel pair. Instead, all channel energies are taken into account and the bits are distributed as described in each paragraph of this document.

[0129] • There is an explicit “low complexity” mode for adaptive channel pair selection as described in [6], in which a single channel that is part of a channel pair during an iterative channel pair selection process cannot be part of another channel pair during the next iteration of the channel pair selection process.

[0130] The advantages of using a simple per-band M / S for each channel pair, and thus reducing the amount of information that needs to be transmitted within the bitstream, are enhanced by the fact that signal-adaptive channel pair selection is used in [6]. By selecting highly correlated channels to co-encode, wideband M / S conversion is optimal in most cases, i.e., M / S coding is used across all bands. This is because the signal can be transmitted in single bits, and therefore requires significantly less signaling information compared to per-band M / S determination. This significantly reduces the total amount of information bits that need to be transmitted for all channel pairs.

[0131] Embodiments of the present invention relate to signal-adaptive joint coding for a multichannel system having a perceptually whitened and ILD-corrected spectrum, wherein the joint coding consists of a simple band-by-band M / S conversion decision based on the estimated number of bits of the entropy coder.

[0132] While some embodiments have been described in the context of apparatus, it is clear that these embodiments also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such devices.

[0133] The encoded audio signal of the present invention can be stored on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium such as the Internet or a wired transmission medium.

[0134] Depending on the specific implementation requirements, the embodiments of the present invention may be implemented in hardware or software. These embodiments may be implemented using a digital storage medium such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which stores electronically readable control signals and cooperates (or is capable of cooperating) with a computer system programmable to perform each method. Thus, the digital storage medium may be computer-readable.

[0135] Some embodiments of the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system so that one of the methods described herein can be performed.

[0136] Generally, embodiments of the present invention can be implemented as a computer program product having program code, which operates to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0137] Another embodiment includes a computer program stored on a machine-readable carrier for performing one of the methods described herein.

[0138] Therefore, in other words, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.

[0139] Accordingly, further embodiments of the methods of the present invention include a computer program for performing one of the methods described herein, on which it is recorded, a data carrier (or digital storage medium or computer-readable medium). The data carrier, digital storage medium, or recording medium is typically tangible and / or non-temporary.

[0140] Accordingly, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.

[0141] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0142] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.

[0143] Further embodiments of the present invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0144] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0145] The apparatus described herein may be implemented using hardware devices, or using a computer, or using a combination of hardware devices and a computer.

[0146] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.

[0147] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the arrangements and details described herein will be obvious to those skilled in the art. Therefore, it is intended that the invention is limited only by the scope of the immediate claims and not by the specific details presented herein.

[0148] References (all are incorporated herein by reference in their entirety) [1] “Information technology - MPEG audio technologies Part 3: Unified speech and audio coding,” ISO / IEC 23003-3, 2012

[0149] [2] “Information technology - MPEG audio technologies Part 1: MPEG Surround,” ISO / IEC 23003-1, 2007

[0150] [3] J. Herre, J. Hilpert, K. Achim and J. Plogsties, “MPEG-H 3D Audio-The New Standard for Coding of Immersive Spatial Audio,” Journal of Selected Topics in Signal Processing, vol. 5, no. 9, pp. 770-779, August 2015.

[0151] [4] “Digital Audio Compression (AC-4) Standard,” ETSI TS 103 190 V1.1.1, 2014-04

[0152] [5] D. Yang, H. Ai, C. Kyriakakis and C. Kuo, “High-fidelity multichannel audio coding with Karhunen-Loeve transform,” Transactions on Speech and Audio Processing, vol. 11, no. 4, pp. 365-380, July 2003.

[0153] [6] F. Schuh, S. Dick, R. Fueg, C. R. Helmrich, N. Rettelbach and T. Schwegler, “Efficient Multichannel Audio Transform Coding with Low Delay and Complexity,” in AES Convention, Los Angeles, September 20, 2016.

[0154] [7] G. Markovic, E. Fotopoulou, M. Multrus, S. Bayer, G. Fuchs, J. Herre, E. Ravelli, M. Schnell, S. Doehla, W. Jaegers, M. Dietz and C. Helmrich, “Apparatus and method for mdct m / s stereo with global ild with improved mid / side decision”. International Patent WO2017125544A1, 27 July 2017

[0155] [8] 3GPP TS 26.445, Codec for Enhanced Voice Services (EVS); Detailed algorithmic description.

[0156] [9] G. Markovic, F. Guillaume, N. Rettelbach, C. Helmrich and B. Schubert, “Linear prediction based coding scheme using spectral domain noise shaping”. EU Patent 2676266 B1, 14 February 2011

[0157]

[10] S. Disch, F. Nagel, R. Geiger, B. N. Thoshkahna, K. Schmidt, S. Bayer, C. Neukam, B. Edler and C. Helmrich, “Audio Encoder, Audio Decoder and Related Methods Using Two-Channel Processing Within an Intelligent Gap Filling Framework”. International Patent PCT / EP2014 / 065106, 15 07 2014

[0158]

[11] “Codec for Encanced Voice Services (EVS); Detailed algorithmic description,” 3GPP TS 26.445 V 12.5.0, December 2015

[0159]

[12] “Codec for Encanced Voice Services (EVS); Detailed algorithmic description,” 3GPP TS 26.445 V 13.3.0, September 2016

[0160]

[13] Sascha Dick, F. Schuh, N. Rettelbach, T. Schwegler, R. Fueg, J. Hilpert and M. Neusinger, “APPARATUS AND METHOD FOR ENCODING OR DECODING A MULTI-CHANNEL SIGNAL”. International Patent PCT / EP2016 / 054900, 08 March 2016.

Claims

1. A multi-signal encoder for encoding at least three audio signals, A signal preprocessor (100) for individually preprocessing each audio signal to obtain at least three preprocessed audio signals, wherein the preprocessing is performed such that the preprocessed audio signals are whitened relative to the original signals, An adaptive joint signal processor (200) for performing processing of the at least three pre-processed audio signals in order to obtain at least three jointly processed signals or at least two jointly processed signals and an unprocessed signal, A signal encoder (300) for encoding each signal in order to obtain one or more encoded signals, An output interface (400) for transmitting or storing an encoded multi-signal audio signal including one or more encoded signals, side information relating to the preprocessing, and side information relating to the processing. A multi-signal encoder that includes [specific component].

2. The adaptive joint signal processor (200) is configured to perform broadband energy normalization (210) of the at least three preprocessed audio signals such that each preprocessed audio signal has normalized energy. The multi-signal encoder according to claim 1, wherein the output interface (400) is configured to include, as further side information, a broadband energy normalized value (534) of each preprocessed audio signal.

3. The adaptive joint signal processor (200) is, The information regarding the average energy of the pre-processed audio signal is calculated (212), Calculate the energy information for each pre-processed audio signal (211), The energy normalized value is calculated based on the information relating to the average energy and the information relating to the energy of a specific preprocessed audio signal (213, 214). A multi-signal encoder according to claim 2, configured as described above.

4. The adaptive joint signal processor (200) is configured to calculate a scaling ratio (534b) between a specific preprocessed audio signal from the average energy and the energy of the preprocessed audio signal (213, 214), The adaptive joint signal processor (200) is configured to determine a flag (534a) indicating whether the scaling ratio is upscaling or downscaling, and the flag for each signal is included in the encoded signal. A multi-signal encoder according to any one of claims 1 to 3.

5. The adaptive joint signal processor (200) is configured to quantize the scaling ratio to the same quantization range (214) regardless of whether the scaling is upscaling or downscaling. The multi-signal encoder according to claim 4.

6. The adaptive joint signal processor (200) is, To obtain at least three normalized signals, each preprocessed audio signal is normalized relative to a reference energy (210), The cross-correlation value of the normalized signals of each possible pair of the at least three normalized signals is calculated (220), Select the signal pair with the highest cross-correlation value (229), The joint stereo processing mode for the selected signal pair is determined (232a), To obtain a processed signal pair, the selected signal pair is subjected to joint stereo processing according to the determined joint stereo processing mode (232b). A multi-signal encoder according to any one of claims 1 to 5, configured as described above.

7. The adaptive joint signal processor (200) is configured to apply cascaded signal pair preprocessing, or the adaptive joint signal processor (200) is configured to apply non-cascaded signal pair processing. In the cascaded signal pair preprocessing, the signals of the processed signal pair are selectable in a further iterative step comprising: calculating updated cross-correlation values; selecting the signal pair having the highest cross-correlation value; determining the joint stereo processing mode for the selected signal pair; and performing the joint stereo processing on the selected signal pair according to the determined joint stereo processing mode, or In the non-cascaded signal pair processing, the signals of the processed signal pair are not selectable in further selecting the signal pair having the highest cross-correlation value, determining the joint stereo processing mode for the selected signal pair, and performing the joint stereo processing on the selected signal pair according to the determined joint stereo processing mode. The multi-signal encoder according to claim 6.

8. The adaptive joint signal processor (200) is configured to determine the signals to be individually encoded as signals remaining after the pairwise processing procedure. The adaptive joint signal processor (200) is configured to correct the energy normalization applied to the signal before performing the pairwise processing procedure, such as a return (237), or to at least partially return the energy normalization applied to the signal before performing the pairwise processing procedure. A multi-signal encoder according to any one of claims 1 to 7.

9. The adaptive joint signal processor (200) is configured to determine bit distribution information (536) for each signal processed by the signal encoder (300), and the output interface (400) is configured to introduce the bit distribution information (536) into the encoded signal for each signal. A multi-signal encoder according to any one of claims 1 to 8.

10. The adaptive joint signal processor (200) calculates the signal energy information of each signal processed by the signal encoder (300) (282), The total energy of the plurality of signals encoded by the signal encoder (300) is calculated (284), Based on the signal energy information and the total energy information, the system is configured to calculate bit distribution information (536) for each signal (286). The output interface (400) is configured to introduce the bit distribution information into the encoded signal for each signal. A multi-signal encoder according to any one of claims 1 to 9.

11. The adaptive joint signal processor (200) is configured to optionally assign an initial number of bits to each signal (290), assign a number of bits based on the bit distribution information (291), optionally perform a further improvement step (292), or optionally perform a final donation step (292), The signal encoder (300) is configured to perform the signal coding using the assigned bits for each signal. The multi-signal encoder according to claim 10.

12. The signal preprocessor (100) performs the following for each audio signal: Time-spectral transformation operations (108, 110, 112) to obtain the spectrum of each audio signal, Time noise shaping operations (114a, 114b) and / or frequency domain noise shaping operations (116) for each signal spectrum It is configured to perform the following actions: The signal preprocessor (100) is configured to supply the signal spectrum to the adaptive joint signal processor (200) following the time noise shaping operation and / or the frequency domain noise shaping operation. The adaptive joint signal processor (200) is configured to perform the joint signal processing on the received signal spectrum. A multi-signal encoder according to any one of claims 1 to 11.

13. The adaptive joint signal processor (200) is, For each signal in the selected signal pair, determine the required bitrate for a full-band decoupled coding mode such as L / R, the required bitrate for a full-band joint coding mode such as M / S, or the bitrate for a band-by-band joint coding mode such as M / S plus the required bits for band-by-band signal transmission, such as the M / S mask. When the majority of the bandwidth is determined for a particular mode, and a small portion of the bandwidth, less than 10% of the total bandwidth, is determined for other encoding modes, determine the decoupled encoding mode or the joint encoding mode as the particular mode for all bandwidths of the signal pair, or determine the encoding mode that requires the fewest number of bits. It is configured in such a way, The output interface (400) is configured to include a display in the encoded signal, the display indicating the specific mode of all bandwidths of the frame instead of the encoded mode mask of the frame. A multi-signal encoder according to any one of claims 1 to 12.

15. The signal encoder (300) includes a rate loop processor for each individual signal or for two or more signals, and the rate loop processor is configured to receive and use bit distribution information (536) for the particular signal or two or more signals. A multi-signal encoder according to any one of claims 1 to 14.

16. The adaptive joint signal processor (200) is configured to adaptively select signal pairs for joint coding, or to determine, for each selected signal pair, a band-by-band mid / side coding mode, a full-band mid / side coding mode, or a full-band left / right coding mode, and the output interface (400) is configured to display the selected coding mode in the coded multi-signal audio signal as side information (532). A multi-signal encoder according to any one of claims 1 to 15.

17. The adaptive joint signal processor (200) is configured to form a band-by-band mid / side decision versus left / right decision based on the estimated bit rate in each band when encoded in mid / side mode or left / right mode, and the final joint coding mode is determined based on the results of the band-by-band mid / side decision versus left / right decision. A multi-signal encoder according to any one of claims 1 to 16.

18. The adaptive joint signal processor (200) is configured to perform the spectral band replication process or the intelligent gap-filling process (260) in order to determine parameter side information for the spectral band replication process or the intelligent gap-filling process, and the output interface (400) is configured to include the spectral band replication or intelligent gap-filling side information (532) as additional side information in the encoded signal, the multi-signal encoder according to any one of claims 1 to 17.

19. The adaptive joint signal processor (200) is configured to perform stereo intelligent gap-filling processing on an encoded signal pair and to perform single-signal intelligent gap-filling processing on at least one of the individually encoded signals. The multi-signal encoder according to claim 18.

20. The at least three audio signals include a low-frequency boosted signal, the adaptive joint signal processor (200) is configured to apply a signal mask, the signal mask indicates for which the adaptive joint signal processor (200) is activated, and the signal mask indicates that the low-frequency boosted signal should not be used in the pairwise processing of the at least three preprocessed audio signals. A multi-signal encoder according to any one of claims 1 to 19.

21. The adaptive joint signal processor (200) calculates the energy of the MDCT spectrum of the signal as the information relating to the energy of the signal, or The information relating to the average energy of the at least three preprocessed audio signals is configured to calculate the average energy of the MDCT spectra of the at least three preprocessed audio signals. A multi-signal encoder according to any one of claims 1 to 5.

22. The adaptive joint signal processor (200) is configured to calculate a scaling factor for each signal (213) based on the energy information of a specific signal and energy information relating to the average energy of the at least three audio signals. The adaptive joint signal processor (200) is configured to quantize the scaling ratio (214) in order to obtain a quantization scaling ratio value, the quantization scaling ratio value is used to induce side information of the scaling ratio of each included signal into the encoded signal. The adaptive joint signal processor (200) is configured to derive a quantization scaling ratio from the quantization scaling ratio value, and the preprocessed audio signal is scaled using the quantization scaling ratio before being used for the pairwise processing of the scaled signal together with other appropriately scaled signals. A multi-signal encoder according to any one of claims 1 to 5.

23. The adaptive joint signal processor (200) is configured to calculate normalized inter-signal cross-correlation values ​​(221) of possible signal pairs in order to determine and select which signal pairs have the highest similarity and are therefore suitable to be selected as pairs for pairwise processing of the at least three pre-processed audio signals. The normalized cross-correlation values ​​for each signal pair are stored in the cross-correlation vector. The adaptive joint signal processor (200) is configured to determine whether one or more signal pair selections from the previous frame should be maintained by comparing the cross-correlation vector of the previous frame with the cross-correlation vector of the current frame (222, 223), wherein the signal pair selections from the previous frame are maintained when the difference between the cross-correlation vector of the current frame and the cross-correlation vector of the previous frame falls below a predetermined threshold (225). A multi-signal encoder according to any one of claims 1 to 22.

24. The signal preprocessor (100) is configured to perform time-frequency conversion using a specific window length selected from a plurality of different window lengths. The adaptive joint signal processor (200) is configured to determine whether the pairs of signals have the same associated window length when comparing the preprocessed audio signals to determine pairs of signals to be processed pairwise, The adaptive joint signal processor (200) is configured to enable pairwise processing of the two signals only when the two signals are associated with the same window length applied by the signal preprocessor (100). A multi-signal encoder according to any one of claims 1 to 23.

25. The adaptive joint signal processor (200) is configured to apply non-cascaded signal pair processing such that the signals of the processed signal pair are not selectable in further signal pair processing, the adaptive joint signal processor (200) is configured to select the signal pairs based on the cross-correlation between the signal pairs for pairwise processing, and the pairwise processing of several selected signal pairs is performed in parallel. A multi-signal encoder according to any one of claims 1 to 24.

26. The adaptive joint signal processor (200) is configured to determine a stereo encoding mode for a selected signal pair, and when the stereo encoding mode is determined to be dual mono mode, the signals included in the signal pair are at least partially rescaled and displayed as individually encoded signals. The multi-signal encoder according to claim 25.

27. The adaptive joint signal processor (200) is configured to perform a stereo intelligent gap-filling (IGF) operation on a pairwise processed signal pair if the stereo mode of the core region is different from the stereo mode of the IGF region, or if the stereo mode of the core has a band-by-band mid / side coding flag set, or The adaptive joint signal processor (200) is configured to apply single-signal IGF analysis to the signals of a pairwise processed signal pair if the stereo mode of the core region is no different from the stereo mode of the IGF region, or if the stereo mode of the core is not flagged as a band-by-band mid / side coding mode. The multi-signal encoder according to claim 18 or 19.

28. The adaptive joint signal processor (200) is configured to perform an intelligent gap-filling operation before the results of the IGF operation are individually encoded by the signal encoder (300). The power spectrum is used for quantization and intelligent gap-filling (IGF) tone / noise determination, and the signal preprocessor (100) is configured to perform the same frequency-domain noise shaping on the MDST spectrum that was used on the MDCT spectrum. The adaptive joint signal processor (200) is configured to perform the same mid / side processing on a preprocessed MDST spectrum so that the results of the processed MDST spectrum are used in the quantization performed by the signal encoder (300) or in the intelligent gap-filling processing performed by the adaptive joint signal processor (200), or The adaptive joint signal processor (200) is configured to apply the same normalization scaling that was performed on the MDCT spectrum using the same quantized scaling vector, based on the full-band scaling vector of the MDST spectrum. A multi-signal encoder according to any one of claims 1 to 27.

29. The multi-signal encoder according to any one of claims 1 to 28, wherein the adaptive joint signal processor (200) is configured to perform pairwise processing of the at least three pre-processed audio signals to obtain the at least three jointly processed signals or at least two jointly processed signals and individually encoded signals.

30. The audio signals of the at least three audio signals are either audio channels or The audio signals of the at least three audio signals are audio component signals of an audio field representation such as an ambisonic sound field representation, a B-format representation, an A-format representation, or any other arbitrary sound field representation such as a sound field representation that represents a sound field relative to a reference position. A multi-signal encoder according to any one of claims 1 to 29.

31. The signal encoder (300) is configured to encode each signal individually to obtain at least three individually encoded signals, or to perform (entropy) encoding together with two or more signals. A multi-signal encoder according to any one of claims 1 to 30.

32. A multi-signal decoder for decoding encoded signals, A signal decoder (700) for decoding at least three encoded signals, A joint signal processor (800) for performing joint signal processing according to side information contained in the encoded signal in order to obtain at least three processed decoded signals, A post-processor (900) for post-processing the at least three processed decoded signals according to the side information contained in the encoded signal, wherein the post-processing is performed such that the post-processed signal is no whiter than the signal before post-processing, and the post-processed signal represents the decoded audio signal. A multi-signal decoder, including...

33. The aforementioned joint signal processor (800) The system is configured to extract the energy normalized value of each joint stereo decoded signal from the encoded signal (610), The system is configured to process the decoded signal pairwise (820) using the joint stereo mode indicated by the side information in the encoded signal in order to obtain the joint stereo decoded signal. To obtain the processed decoded signal, the system is configured to energy rescale the joint stereo decoded signal using the energy normalization value (830), The multi-signal decoder according to claim 32.

34. The joint signal processor (800) is configured to check whether the energy normalized value extracted from the encoded signal of a particular signal has a predetermined value. The joint signal processor (800) is configured to either not perform energy rescaling on the specific signal, or to perform only reduced energy rescaling, when the energy normalized value has the predetermined value. The multi-signal decoder according to claim 32.

35. The aforementioned signal decoder (700) From the above encoded signals, the bit distribution values ​​of each encoded signal are extracted (620), The bit distribution of the signal, the number of remaining bits of all signals, and optionally a further refinement step, or optionally a final contribution step, are used to determine the bit distribution of the signal to be used (720). Based on the bit distribution used for each signal, the individual decodings are performed (710, 730). A multi-signal decoder according to any one of claims 32 to 34, configured as described above.

36. The aforementioned joint signal processor (800) To obtain spectrally enhanced individual signals, use the side information of the encoded signal to perform band duplication or band duplication on the individually decoded signals (820), Using the individual signals whose spectra have been enhanced, perform joint processing (820) according to the joint processing mode. A multi-signal decoder according to any one of claims 32 to 35, configured as described above.

37. The joint signal processor (800) is configured to convert a source range from one stereo representation to the other stereo representation when the target range is indicated to have a different stereo representation. The multi-signal decoder according to claim 36.

38. The aforementioned joint signal processor (800) From the encoded signal, the energy normalized value (534b) of each joint stereo decoded signal is extracted, and in addition, a flag (534a) indicating whether the energy normalized value is an upscaled value or a downscaled value is extracted. When the flag has a first value, rescaling is performed as downscaling, and when the flag has a second value different from the first value, rescaling is performed as upscaling, using the energy normalized value (830). A multi-signal decoder according to any one of claims 32 to 37, configured as described above.

39. The aforementioned joint signal processor (800) From the encoded signal, side information indicating the signal pair obtained from the co-encoding operation is extracted (630), To return each signal to its original preprocessed spectrum, inverse stereo or multichannel processing is performed starting from the last signal pair to obtain the encoded signal (820), and the inverse stereo processing is performed based on the stereo mode and / or per-band mid / side determination shown in the side information (532) of the encoded signal. A multi-signal decoder according to any one of claims 32 to 38, configured as described above.

40. The joint signal processor (800) is configured to denormalize all signals in a signal pair to their corresponding original energy levels based on quantized energy scaling information contained for each individual signal (830), and other signals that were not involved in the signal pair processing are not denormalized in the same way as the signals that were involved in the signal pair processing. A multi-signal decoder according to any one of claims 32 to 39.

41. The post-processor (900) is configured to perform, for each individual processed decoded signal, a time noise shaping operation (910) or a frequency domain noise shaping operation (910), and a conversion from the spectral domain to the time domain (920), as well as a subsequent superposition addition operation (930) between subsequent time frames of the post-processed signal. A multi-signal decoder according to any one of claims 32 to 40.

42. The joint signal processor (800) is configured to extract a flag from the encoded signal indicating whether certain bands of the time frame of the signal pair are to be inverted using mid / side or left / right encoding, and the joint signal processor (800) is configured to use the flag to cause the corresponding bands of the signal pair to undergo either mid / side or left / right processing, depending on the value of the flag. For different time frames of the same signal pair, or for different signal pairs of the same time frame, an encoding mode mask is extracted from the side information of the encoded signal, indicating a separate encoding mode for each individual bandwidth, and the joint signal processor (800) is configured to determine whether to apply inverse mid / side processing or mid / side processing to the corresponding bandwidth indicated for the bits associated with that bandwidth. A multi-signal decoder according to any one of claims 32 to 41.

43. The encoded signal is an encoded multichannel signal, the multi-signal decoder is a multichannel decoder, the encoded signal is an encoded multichannel signal, the signal decoder (700) is a channel decoder, the encoded signal is an encoded channel, the joint signal processing is joint channel processing, the at least three processed decoded signals are at least three processed decoded signals, the post-processed signal is a channel, or The encoded signal is an encoded multi-component signal representing an audio component signal of a sound field representation such as an ambisonic sound field representation, a B-format representation, an A-format representation, or any other arbitrary sound field representation such as a sound field representation representing a sound field relative to a reference position; the multi-signal decoder is a multi-component decoder; the encoded signal is an encoded multi-component signal; the signal decoder (700) is a component decoder; the encoded signal is an encoded component; the joint signal processing is joint component processing; the at least three processed decoded signals are at least three processed decoded components; and the post-processed signal is a component audio signal. A multi-signal decoder according to any one of claims 32 to 42.

44. A method for performing multi-signal coding of at least three audio signals, A step of pre-processing each audio signal individually in order to obtain at least three pre-processed audio signals, wherein the pre-processing is performed such that the pre-processed audio signals are whited out relative to the original signals. The steps include performing processing on the at least three preprocessed audio signals in order to obtain at least three jointly processed signals or at least two jointly processed signals and individually encoded signals, The steps include: encoding each signal in order to obtain one or more encoded signals, A step of transmitting or storing an encoded multi-signal audio signal including one or more encoded signals, side information relating to the preprocessing, and side information relating to the processing. A method that includes this.

45. A method for multi-signal decoding of an encoded signal, The steps include decoding at least three encoded signals individually, The steps include performing joint signal processing according to side information contained in the encoded signal in order to obtain at least three processed decoded signals, A step of post-processing the at least three processed decoded signals according to the side information contained in the encoded signal, wherein the post-processing is performed such that the post-processed signal is no whiter than the signal before post-processing, and the post-processed signal represents the decoded audio signal. A method that includes this.

46. A computer program for performing the method of claim 44 or the method of claim 45 when executed on a computer or processor.

47. It is an encoded signal, At least three individually encoded signals (510), Side information (520) relating to preprocessing performed to obtain the at least three individually encoded signals, Side information (532) relating to pairwise processing performed to obtain the at least three individually encoded signals, The encoded signal includes, for each of the at least three individually encoded signals obtained by multi-signal coding, an energy scaling value (534), or for each of the individually encoded signals, a bit distribution value (536).