Apparatus, method or computer program for processing an encoded audio scene using bandwidth extension
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2026-08-11
Smart Images

Figure CN116457878B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to audio processing, and more particularly to processing coded audio scenes to generate processed audio scenes for rendering, transmission or storage. Background Technology
[0002] Traditionally, audio applications providing means of user communication (such as telephone or teleconferencing) have been primarily limited to mono recording and playback. However, in recent years, the emergence of new immersive VR / AR technologies has sparked interest in spatial rendering of communication scenarios. To meet this interest, a new 3GPP audio standard called Immersive Voice and Audio Services (IVAS) is currently under development. Based on the recently released Enhanced Voice Services (EVS) standard, IVAS provides multi-channel and VR extensions capable of rendering immersive audio scenarios for purposes such as spatial teleconferencing, while still meeting the low-latency requirements for smooth audio communication. This ongoing need to keep the total latency of the codec to a minimum without sacrificing playback quality has driven the work described below.
[0003] Encoding scene-based audio (SBA) material (such as third-order surround sound content) using systems employing parametric audio coding (such as DirAC)[1][2] at low bit rates (e.g., 32 kbps and below) only allows direct encoding of a single (transmission) channel, while recovering spatial information via side parameters at the decoder in the filter bank domain. A complete reconstruction of the 3D audio scene is not required where the speaker setup at the decoder is only capable of stereo playback. Since higher bit rate encoding of two or more transmission channels is possible, in these cases, a stereo reproduction of the scene can be directly extracted and played back without any parametric spatial upmixing (completely skipping the spatial renderer) and the additional delays that come with it (e.g., due to additional filter bank analysis / synthesis, such as complex-valued low-delay filter banks (CLDFB)). However, this is not possible at low bit rates with only one transmission channel. Therefore, in the case of DirAC, until now, stereo output has required FOA (first-order surround sound) upmixing with subsequent L / R conversion. This is problematic because this configuration now has a higher total latency than other possible stereo output configurations in the system, and will require alignment of all stereo output configurations.
[0004] Example of DirAC stereo rendering with high latency
[0005] Figure 12 A block diagram example of conventional decoder processing for DirAC stereo upmixing with high latency is shown.
[0006] For example, at the encoder not depicted, a single downmix channel is derived via spatial downmixing in the DirAC encoder processing and then encoded using a core encoder such as Enhanced Voice Service (EVS) [3].
[0007] At the decoder, for example, using Figure 12 The conventional DirAC upmixing process described in the figure first decodes a usable transmission channel from the bitstream 1212 using a mono or IVAS mono decoder 1210, thereby generating a time-domain signal that can be regarded as the original audio scene in the decoded mono downmix 1214.
[0008] The decoded mono signal 1214 is input to the CLDFB 1220 for analysis of the delayed signal 1214 (converting the signal to the frequency domain). The significantly delayed output signal 1222 enters the DirAC renderer 1230. The DirAC renderer 1230 processes the delayed output signal 1222, and the side information sent (i.e., DirAC side parameters 1213) is used to transform the signal 1222 into a FOA representation (i.e., FOA upmix 1232 of the original scene, which has spatial information recovered from the DirAC side parameters 1213).
[0009] The transmitted parameter 1213 may include azimuth angles (e.g., an azimuth value for the horizontal plane and an elevation angle for the vertical plane) and a spread value for each frequency band to perceptually describe the entire 3D audio scene. Due to the frequency band-based processing of DirAC stereo upmixing, parameter 1213 is transmitted multiple times per frame, i.e., in a group for each frequency band. Furthermore, each group includes (e.g., 20ms in length) multiple directional parameters for each subframe throughout the entire frame to improve temporal resolution.
[0010] The result of the DirAC renderer 1230 can be, for example, a full 3D scene in FOA format (i.e., FOA upmixed 1232), which can now be converted using matrix transformation 1240 into an L / R signal 1242 suitable for playback on a stereo speaker setup. In other words, the L / R signal 1242 can be input to a stereo speaker or it can be input to a CLDFB synthesizer 1250 using predefined channel weights. The CLDFB synthesizer 1250 converts the two output channels (L / R signal 1242) input in the frequency domain to the time domain, thereby generating an output signal 1252 ready for stereo playback.
[0011] Alternatively, the same DirAC stereo upmix can be used to directly generate a render for the stereo output configuration, avoiding the intermediate step of generating the FOA signal. This reduces the algorithmic complexity that could potentially complicate the framework. However, both methods require an additional filter bank after the core encoding, which results in an additional 5ms delay. Other examples of DirAC rendering can be found in [2].
[0012] The DirAC stereo upmixing method is suboptimal in terms of both latency and complexity. Due to the use of the CLDFB filter bank, the output is significantly delayed (an additional 5ms delay in the DirAC example), resulting in the same total latency as full SBA upmixing (compared to the latency of a stereo output configuration where no additional rendering step is required). This also reasonably assumes that performing full SBA upmixing to generate a stereo signal is not ideal in terms of system complexity.
[0013] The purpose of this invention is to provide an improved concept for processing encoded audio scenarios.
[0014] This objective is achieved by the apparatus for processing encoded audio scenes according to claim 1, the method for processing encoded audio scenes according to claim 32, or the computer program according to claim 33.
[0015] This invention is based on the following discovery: According to a first aspect related to parameter transformation, an improved concept for processing coded audio scenes is obtained by converting given parameters in a coded audio scene related to the virtual listener's position into transformation parameters related to the channel representation of a given output format. This process provides a high degree of flexibility in processing and ultimately rendering the processed audio scene in a channel-based environment.
[0016] An embodiment of the first aspect of the invention includes an apparatus for processing an coded audio scene representing a sound field associated with a virtual listener's position, the coded audio scene including information about a transmitted signal (e.g., a core coded audio signal) and a first set of parameters associated with the virtual listener's position. The apparatus includes: a parameter converter for converting the first set of parameters (e.g., DirAC side parameters in B-format or first-order surround sound (FOA) format) into a second set of parameters (e.g., stereo parameters) associated with a channel representation including two or more channels, for reproducing the two or more channels at a predefined spatial location; and an output interface for generating the processed audio scene using the second set of parameters and the information about the transmitted signal.
[0017] In this embodiment, a short-time Fourier transform (STFT) filter bank is used for upmixing, rather than a directional audio coding (DirAC) renderer. Therefore, a downmix channel (included in the bitstream) can be upmixed to a stereo output without any additional total latency. This upmixing allows for analysis at the decoder using a window with very short overlap, keeping it within the total latency required by the communication codec or the upcoming Immersive Speech and Audio Services (IVAS). This value could be, for example, 32 milliseconds. In this embodiment, any post-processing for bandwidth expansion purposes can be avoided, as such processing can be performed in parallel with parameter transformation or parameter mapping.
[0018] By mapping listener-specific parameters for low-frequency (LB) signals to a channel-specific stereo parameter set for LB signals, low-delay upmixing for LB signals can be achieved in the DFT domain. For high-frequency signals, a single stereo parameter set allows upmixing of the high-frequency band to be performed in the time domain, preferably in parallel with spectral analysis, spectral upmixing, and spectral synthesis for LB signals.
[0019] For example, the parameter converter is configured to use a single-side gain parameter for translation, and a residual prediction parameter that is closely related to the stereo width and also closely related to the diffusion parameter used in DirAC (Directional Audio Coding).
[0020] In this embodiment, when processing encoded audio scenes (scene-based audio) to obtain stereo output, the "DFT-stereo" method allows the IVAS codec to remain within the same total latency (specifically, 32 milliseconds) as in EVS. Lower complexity of parametric stereo upmixing is achieved by implementing direct processing via DFT stereo instead of spatial DirAC rendering.
[0021] This invention is based on the following discovery: according to a second aspect related to bandwidth expansion, an improved concept for processing coded audio scenarios has been obtained.
[0022] An embodiment of the second aspect of the invention includes an apparatus for processing an audio scene representing a sound field, the audio scene including information about a transmitted signal and a parameter set. The apparatus further includes: an output interface for generating a processed audio scene using the parameter set and the information about the transmitted signal, wherein the output interface is configured to generate an original representation of two or more channels using the parameter set and the transmitted signal; a multichannel enhancer for generating an enhanced representation of two or more channels using the transmitted signal; and a signal combiner for combining the original representation of two or more channels and the enhanced representation of two or more channels to obtain the processed audio scene.
[0023] The generation of the original representations for two or more channels on the one hand, and the separate generation of the enhanced representations for two or more channels on the other hand, allows for great flexibility in selecting algorithms for the original and enhanced representations. For each of one or more output channels—that is, in the multi-channel output domain rather than in the lower channel input or encoded scene domain—the final combination has already occurred. Therefore, after combination, two or more channels are synthesized and can be used in other processes, such as rendering, transmission, or storage.
[0024] In this embodiment, a portion of the core processing (e.g., bandwidth extension (BWE) of the algebraic text-excited linear prediction (ACELP) speech encoder for the enhanced representation) can be performed in parallel with DFT stereo processing for the original representation. Therefore, any delays caused by either algorithm do not accumulate, and a given delay caused by only one algorithm will be the final delay. In this embodiment, only the transmission signal (e.g., the low-frequency band (LB) signal (channels)) is input to the output interface, such as for DFT stereo processing, while the high-frequency band (HB) is upmixed separately in the time domain, for example, by using a multi-channel enhancer, allowing stereo decoding to be processed within a target time window of 32 milliseconds. By using wideband shifting, for example, based on mapped side gain from, for example, a parametric converter, direct time-domain upmixing for the entire high-frequency band is achieved without any significant delay.
[0025] In this embodiment, the delay reduction in DFT stereo may not be entirely due to the difference in overlap between the two transforms, for example, a 5ms transform delay caused by CLDFB and a 3.125ms transform delay caused by STFT. Instead, DFT stereo utilizes the fact that the last 3.25ms of the 32ms target EVS encoder delay essentially comes from ACELP BWE. Everything else (the remaining milliseconds before reaching the target EVS encoder delay) is merely artificially delayed to re-align the two transformed signals (HB stereo upmixed signal and HB fill signal with LB stereo core signal) at the very end. Thus, to avoid additional delay in DFT stereo, for example within a very short DFT window overlap, only all other components of the encoder are transformed, while ACELP BWE, for example using a multi-channel enhancer, is upmixed in the time domain with almost no delay.
[0026] This invention is based on the following discovery: According to a third aspect related to parametric smoothing, an improved concept for processing coded audio scenes is obtained by performing parametric smoothing on time according to smoothing rules. Therefore, the processed audio scene obtained by applying smoothing parameters instead of the original parameters to the transport channels will have improved audio quality. This is especially true when the smoothing parameter is an upmixing parameter, but for any other parameter (e.g., envelope parameter, LPC parameter, noise parameter, or scaling factor parameter), the use of the smoothing parameter obtained through smoothing rules will lead to a subjective improvement in the audio quality of the obtained processed audio scene.
[0027] An embodiment of the invention according to a third aspect includes an apparatus for processing an audio scene representing a sound field, the audio scene including information about a transmitted signal and a first parameter set. The apparatus further includes: a parameter processor for processing the first parameter set to obtain a second parameter set, wherein the parameter processor is configured to compute at least one original parameter for each input time frame using at least one parameter from the first parameter set for input time frames, compute smoothing information (e.g., a factor for each original parameter) according to a smoothing rule, and apply the corresponding smoothing information to the corresponding original parameter to derive parameters for the second parameter set for output time frames; and an output interface for generating the processed audio scene using the second parameter set and the information about the transmitted signal.
[0028] By smoothing the original parameters over time, strong fluctuations in gain or parameters from one frame to the next are avoided. The smoothing factor determines the strength of the smoothing, which is adaptively calculated in a preferred embodiment by a parameter processor that also functions as a parameter converter to transform listener position-related parameters into channel-related parameters. Adaptive calculation allows for a faster response whenever the audio scene changes abruptly. The adaptive smoothing factor is calculated band-wise based on changes in energy within the current frequency band. Band-wise energy is calculated across all subframes included in the frame. Furthermore, energy changes over time are characterized by two averages (a short-term average and a long-term average), so extreme cases have no effect on smoothing, and a less rapid increase in energy does not so drastically degrade smoothing. Therefore, the smoothing factor is calculated for each DTF stereo subframe in the current frame based on the quotient of these averages.
[0029] It should be noted that all the alternatives or aspects discussed previously and subsequently can be used individually, i.e., without any aspect. However, in other embodiments, two or more aspects are combined with each other, and in other embodiments, all aspects are combined with each other to achieve an improved trade-off between total latency, achievable audio quality, and required implementation effort. Attached Figure Description
[0030] Preferred embodiments of the invention will now be discussed with reference to the accompanying drawings, in which:
[0031] Figure 1 This is a block diagram of an apparatus for processing encoded audio scenes using a parameter converter according to an embodiment;
[0032] Figure 2a A schematic diagram of a first parameter set and a second parameter set according to an embodiment is shown;
[0033] Figure 2b This is an embodiment of a parameter converter or parameter processor used to calculate raw parameters;
[0034] Figure 2c This is an embodiment of a parameter converter or parameter processor used to combine raw parameters;
[0035] Figure 3 This is an embodiment of a parameter converter or parameter processor used to perform a weighted combination of raw parameters;
[0036] Figure 4 This is an example of a parameter converter used to generate side gain parameters and residual prediction parameters;
[0037] Figure 5a This is an embodiment of a parameter converter or parameter processor for calculating a smoothing factor for the original parameters;
[0038] Figure 5b This is an embodiment of a parameter converter or parameter processor used to calculate a smoothing factor for a frequency band;
[0039] Figure 6 A schematic diagram is shown showing the averaging of the transmitted signal with respect to a smoothing factor according to an embodiment;
[0040] Figure 7 This is an embodiment of a parameter converter or parameter processor used for calculating recursive smoothing;
[0041] Figure 8 This is an embodiment of a device for decoding transmitted signals;
[0042] Figure 9 This is an embodiment of a device that uses bandwidth extension to process encoded audio scenarios;
[0043] Figure 10 This is an embodiment of a device for obtaining a processed audio scene;
[0044] Figure 11 This is a block diagram of an embodiment of a multi-channel enhancer;
[0045] Figure 12 This is a block diagram of the standard DirAC stereo upmixing process;
[0046] Figure 13 This is an embodiment of a device that uses parameter mapping to obtain a processed audio scene; and
[0047] Figure 14 This is an embodiment of a device that uses bandwidth expansion to obtain a processed audio scene. Detailed Implementation
[0048] Figure 1 An apparatus is shown for processing an encoded audio scene 130, for example, representing a sound field associated with a virtual listener's position. The encoded audio scene 130 includes information about a transmitted signal 122 (e.g., a bitstream) and a first parameter set 112 associated with the virtual listener's position (e.g., also including multiple DirAC parameters in the bitstream). The first parameter set 112 is input to a parameter converter 110 or parameter processor, which converts the first parameter set 112 into a second parameter set 114 associated with a channel representation comprising at least two or more channels. The apparatus is capable of supporting different audio formats. The audio signals can be acoustic in nature, picked up by a microphone, or electrical in nature, which should be sent to a loudspeaker. Supported audio formats can be mono signals, low-frequency band signals, high-frequency band signals, multi-channel signals, first-order and higher-order surround sound components, and audio objects. Audio scenes can also be described by combining different input formats.
[0049] Parameter converter 110 is configured to compute a second parameter set 114 as parametric stereo or multi-channel (e.g., two or more channels) parameters input to output interface 120. Output interface 120 is configured to generate a processed audio scene 124 by combining the transmitted signal 122 or information about the transmitted signal with the second parameter set 114 to obtain a transcoded audio scene as the processed audio scene 124. Another embodiment includes upmixing the transmitted signal 122 using the second parameter set 114 to an upmixed signal including two or more channels. In other words, parameter converter 120 maps a first parameter set 112, for example, used for DirAC rendering, to the second parameter set 114. The second parameter set may include side gain parameters for translation and residual prediction parameters that result in spatial image improvement of the audio scene when applied to upmixing. For example, the parameters of the first parameter set 112 may include at least one of an arrival direction parameter, a diffusion parameter, a direction information parameter related to a sphere with the virtual listening position as the origin of the sphere, and a distance parameter. For example, the parameters of the second parameter set 114 may include at least one of the following: side gain parameter, residual prediction gain parameter, inter-channel level difference parameter, inter-channel time difference parameter, inter-channel phase difference parameter, and inter-channel coherence parameter.
[0050] Figure 2a A schematic diagram of a first parameter set 112 and a second parameter set 114 according to an embodiment is shown. Specifically, the parameter resolution for the two parameters (the first parameter and the second parameter) is depicted. Figure 2a Each horizontal axis represents time, and Figure 2a Each vertical axis represents frequency. For example... Figure 2a As shown, the input time frame 210 associated with the first parameter set 112 includes two or more input time subframes 212 and 213. Directly below, the output time frame 220 associated with the second parameter set 114 is shown in the corresponding diagram associated with the top diagram. This indicates that the output time frame 220 is smaller than the input time frame 210, and longer than the input time subframes 212 or 213. Note that the input time subframes 212 or 213 and the output time frame 220 may include multiple frequencies as frequency bands. The input frequency band 230 may include the same frequencies as the output frequency band 240. According to an embodiment, the frequency bands of the input frequency band 230 and the output frequency band 240 may not be connected to or related to each other.
[0051] It should be noted that Figure 4 The side gains and residual gains described herein are typically calculated frame-by-frame, such that a single side gain and a single residual gain are calculated for each input frame 210. However, in other embodiments, not only are single side gains and single residual gains calculated for each frame, but a set of side gains and a set of residual gains are calculated for the input time frame 210, wherein each side gain and each residual gain is associated with, for example, a certain input time subframe 212 or 213 of a frequency band. Therefore, in embodiments, the parameter converter 110 calculates a set of side gains and a set of residual gains for each frame of the first parameter set 112 and the second parameter set 114, wherein the number of side gains and residual gains for the input time frame 210 is typically equal to the number of input frequency bands 230.
[0052] Figure 2b An embodiment of a parameter converter 110 for calculating the raw parameters 252 of the second parameter set 114 is shown. The parameter converter 110 calculates the raw parameters 252 in a time-sequential manner for each of two or more input time subframes 212 and 213. For example, for each input frequency band 230 and time instance (input time subframes 212, 213), the calculation 250 derives the principal direction of arrival (DOA) of the azimuth angle θ and the principal direction of arrival of the elevation angle φ, as well as the spread parameter ψ.
[0053] For the directional components (such as X, Y, and Z), the first-order spherical harmonic at the center location can be derived using the omnidirectional components w(b, n) and the DirAC parameters using the following equation:
[0054]
[0055]
[0056]
[0057]
[0058] The W channel represents the non-directional mono component of the signal, corresponding to the output of an omnidirectional microphone. The X, Y, and Z channels are the three-dimensional directional components. From these four FOA channels, a stereo signal (stereo version, stereo output) can be obtained using a parametric converter 110 through decoding involving the W and Y channels, resulting in two cardioid azimuth angles of +90 degrees and -90 degrees. Due to this fact, the following equation shows the relationship between the left and right stereo signals, where the left channel L is represented by adding the Y channel to the W channel, and the right channel R is represented by subtracting the Y channel from the W channel.
[0059]
[0060] In other words, this decoding corresponds to first-order beamforming pointing in two directions, which can be expressed by the following equation:
[0061]
[0062] Therefore, there is a direct correlation between the stereo output (left and right channels) and the first parameter set 112 (i.e., the DirAC parameters).
[0063] However, on the other hand, the second parameter set 114 (i.e., the DFT parameters) depends on the left channel L and right channel R models based on the intermediate signal M and the side signals, which can be expressed by the following equation:
[0064]
[0065] Here, M is transmitted as a mono signal (channel), which corresponds to the omnidirectional channel W in the case of Scene-Based Audio (SBA) mode. Furthermore, in DFT stereo, S is predicted from M using the side gain parameter, which will be explained below.
[0066] Figure 4 An embodiment of a parameter converter 110 is shown for generating, for example, a side gain parameter 455 and a residual prediction parameter 456 using calculation process 450. The parameter converter 110 preferably processes calculations 250 and 450 to calculate the original parameter 252 (e.g., side parameter 455) for the output band 241 using the following equation:
[0067]
[0068] According to this equation, b is the output frequency band, sidegain is the side gain parameter 455, azimuth is the azimuth component of the direction of arrival parameter, and elevation is the elevation component of the direction of arrival parameter. For example... Figure 4 As shown, the first parameter set 112 includes a direction of arrival (DOA) parameter 456 for the input band 231 as described above, and the second parameter set 114 includes a side gain parameter 455 for each input band 230. However, if the first parameter set 112 additionally includes a spread parameter ψ453 for the input band 231, the parameter converter 110 is configured to calculate the side gain parameter 455 for the output band 241 using the following equation:
[0069]
[0070] According to this equation, diff(b) is the diffusion parameter ψ453 for the input frequency band b 230. It should be noted that the direction parameter 456 of the first parameter set 112 can include different value ranges, for example, the azimuth parameter 451 is [0; 360], the elevation parameter 452 is [0; 180], and the resulting side gain parameter 455 is [-1; 1]. Figure 2c As shown, parameter converter 110 uses combiner 260 to combine at least two original parameters 252 to derive parameters in second parameter set 114 related to output time frame 220.
[0071] According to an embodiment, the second parameter set 114 also includes residual prediction parameters 456 for output band 241 in output band 240, which in Figure 4 As shown in the diagram, parameter converter 110 can use the diffusion parameter ψ453 from input band 231 as the residual prediction parameter 456 for output band 241, as shown in residual selector 410. If input band 231 and output band 241 are equal to each other, parameter converter 110 uses the diffusion parameter ψ453 from input band 231. The diffusion parameter ψ453 for output band 241 is derived from the diffusion parameter ψ453 for input band 231, and the diffusion parameter ψ453 for output band 241 is used as the residual prediction parameter 456 for output band 241. Parameter converter 110 can then use the diffusion parameter ψ453 from input band 231.
[0072] In DFT stereo processing, the residuals predicted using residual selector 410 are assumed to be incoherent and modeled using their energy and the decorrelated residual signals going to the left channel L and right channel R. The residuals predicted together with the side signal S and the center signal M, which is a mono signal (channel), can be expressed as:
[0073] R(b) = S(b) - sidegain[b]M(b)
[0074] Its energy is modeled using the following equation in DFT stereo processing with residual prediction gain:
[0075] ||R(b)|| 2 =residual prediction[b]||M(b)|| 2
[0076] Since residual gain represents the inter-channel incoherent components and spatial width of a stereo signal, it is directly related to the diffusion component modeled by DirAC. Therefore, the residual energy can be rewritten as a function of the DirAC diffusion parameters:
[0077] ||R(b)|| 2 =ψ(b)||M(b)|| 2
[0078] Figure 3 A parameter converter 110 is shown according to an embodiment for performing a weighted combination 310 of the original parameters 252. At least two original parameters 252 are input to the weighted combination 310, wherein a weighting factor 324 for the weighted combination 310 is derived based on an amplitude correlation measurement 320 of the transmitted signal 122 in a corresponding input time subframe 212. Furthermore, the parameter converter 110 is configured to use the energy or power value of the transmitted signal 112 in a corresponding input time subframe 212 or 213 as the amplitude correlation measurement 320. The amplitude correlation measurement 320, for example, measures the energy or power of the transmitted signal 122 in a corresponding input time subframe 212 such that, compared to a weighting factor 324 for an input subframe 212 with lower energy or power of the transmitted signal 122, the weighting factor 324 for a corresponding input time subframe 212 with higher energy or power of the transmitted signal 122 is larger.
[0079] As previously mentioned, the direction parameter, azimuth parameter, and elevation parameter have corresponding value ranges. However, the direction parameters of the first parameter set 112 typically have a higher temporal resolution than the second parameter set 114, meaning that two or more azimuth and elevation values must be used to calculate a lateral gain value. According to an embodiment, this calculation is based on energy-related weights, which can be obtained as the output of the amplitude-related measurement 320. For example, for all input time subframes 212 and 213, the energy nrg of the subframe is calculated using the following equation:
[0080]
[0081] Where is the time-domain input signal, is the number of samples in each subframe, and is the sample index. Furthermore, for each output time frame 230, the weight 324 of the contribution of each input time subframe 212, 213 within each output time frame can be calculated as follows:
[0082]
[0083] Then, the side gain parameter 455 is finally calculated using the following equation:
[0084]
[0085] Due to the similarity between parameters, the spread parameter 453 of each frequency band is directly mapped to the residual prediction parameter 456 of all subframes in the same frequency band. The similarity can be expressed by the following equation:
[0086] residual prediction[i][b]=diffusion[b]
[0087] Figure 5a An embodiment of a parameter converter 110 or parameter processor is shown for calculating a smoothing factor 512 for each original parameter 252 according to a smoothing rule 514. Furthermore, the parameter converter 110 is configured to apply the smoothing factor 512 (a corresponding smoothing factor for an original parameter) to the original parameter 252 (an original parameter corresponding to the smoothing factor) to derive parameters for a second set 114 of parameters for the output time frame 220, i.e., the parameters of the output time frame.
[0088] Figure 5b An embodiment of a parameter converter 110 or parameter processor is shown for calculating a smoothing factor 522 for a frequency band using a compression function 540. The compression function 540 can be different for different frequency bands, such that the compression strength of the compression function 540 is stronger for lower frequency bands than for higher frequency bands. The parameter converter 110 is also configured to use a maximum limit selection 550 to calculate the smoothing factors 512 and 522. In other words, the parameter converter 110 can obtain the smoothing factors 512 and 522 by using different maximum limits for different frequency bands, such that the maximum limit for lower frequency bands is higher than the maximum limit for higher frequency bands.
[0089] Both compression function 540 and maximum limit selection 550 are input to calculation 520 to obtain a smoothing factor 522 for frequency band 522. For example, the parameter converter 110 is not limited to using two calculations 510 and 520 to calculate smoothing factors 512 and 522, but is configured to use only one calculation block to calculate smoothing factors 512, 522, which can output smoothing factors 512 and 522. In other words, the smoothing factors are calculated band-wise based on the energy changes in the current frequency band (for each original parameter 252). For example, by using parametric smoothing, the side gain parameter 455 and residual prediction parameter 456 are smoothed over time to avoid strong fluctuations in gain. Since relatively strong smoothing is required most of the time, but a faster response is needed whenever the audio scene 130 changes abruptly, smoothing factors 512, 522 are adaptively calculated to determine the smoothing intensity.
[0090] Therefore, the bandwise energy nrg is calculated in all subframes k using the following equation:
[0091]
[0092] Where x is the frequency range of the signal (real and imaginary parts) after DFT transformation, and is the interval index of all intervals in the current frequency band.
[0093] To capture the energy variation over time, amplitude correlation measurement 320 of the transmitted signal 122 is used to calculate two averages (a short-term average 331 and a long-term average 332), such as... Figure 3 As shown in the image.
[0094] Figure 6 A schematic diagram is shown of an amplitude correlation measurement 320 averaged for a smoothing factor 512 on a transmitted signal 122 according to an embodiment. The x-axis represents time, and the y-axis represents the energy (of the transmitted signal 122). The transmitted signal 122 is shown as a schematic portion of a sine function 122. Figure 6 As shown, the second time portion 631 is shorter than the first time portion 632. The changes in energy on average values 331 and 332 are calculated for each frequency band according to the following equation:
[0095]
[0096] as well as
[0097]
[0098] Where, N short and N long This is the number of previous time subframes k, for which individual averages are calculated. For example, in this particular embodiment, Nshort The value of is set to 3, and N long The value is set to 10.
[0099] Furthermore, the parameter converter or parameter processor 110 is configured to calculate smoothing factors 512 and 522 using calculation 510 based on the ratio between the long-term average 332 and the short-term average 331. In other words, the quotient of the two averages 331 and 332 is calculated such that a higher short-term average (which indicates a recent increase in energy) results in a decrease in smoothing. The following equation shows the correlation between the smoothing factor 512 and the two averages 331 and 312.
[0100]
[0101] Since the higher long-term average of 332 indicating a decrease in energy does not lead to a smoothing decrease, the smoothing factor 512 is set to a maximum of 1 (for now). Therefore, the above formula will fac smooth The minimum value of [b] is limited to (In this example, it is 0.3). However, in extreme cases, this factor must be close to 0, which is achieved by using the following equation to reduce the value from the range... The reason for converting to the range [0; 1]:
[0102]
[0103] In this embodiment, the smoothing is excessively reduced compared to the smoothing previously shown, causing the factor to be compressed towards the value 1 through the root function. Since stability is particularly important in the lowest frequency band, the fourth root is used in the bands b=0 and b=1. The equation for the lowest frequency band is:
[0104]
[0105] For all other frequency bands where b>1, compression is performed using the square root function using the following equation.
[0106]
[0107] By applying the square root function for all other frequency bands b>1, the energy can become smaller in the extreme case of exponentially increasing energy, while a less rapid increase in energy will not reduce smoothness so drastically.
[0108] Furthermore, the maximum smoothing is set for the following equation depending on the frequency band. Note that factor 1 will simply repeat the previous value without contributing to the current gain.
[0109] fac smooth [b] = min(fac) smooth [b],bounds[b])
[0110] Here, boubds[b] represents a given implementation with 5 frequency bands, which are set according to the following table:
[0111]
[0112] Calculate the smoothing factor for each DFT stereo subframe in the current frame.
[0113] Figure 7 A parameter converter 110 using recursive smoothing 710 according to an embodiment is shown, in which the side gain parameter g is adjusted according to the following equation. side [k][b]455 and residual prediction gain parameter g pred [k][b]456 is recursively smoothed:
[0114] g side [k][b]=fac smooth [k][b]g side [k-1][b]
[0115] +(1-fac smooth [k][b])g side [k][b]
[0116] as well as
[0117] g pred [k][b]=fac smooth [k][b]g pred [k-1][b]
[0118] +(1-fac smooth [k][b])g pred [k][b]
[0119] By combining the parameters of the previous output time frame 532 weighted by the first weighting value and the original parameters 252 of the current output time frame 220 weighted by the second weighting value, a recursive smoothing 710 is calculated for the current output time frame across temporally successive output time frames. In other words, the smoothing parameters of the current output time frame are calculated, thereby deriving the first and second weighting values from the smoothing factors of the current time frame.
[0120] These mapped and smoothed parameters (g) side g pred The signal is input to the DFT stereo processing, i.e., output interface 120, where the stereo signal (L / R) is based on the downmixed DMX, the residual prediction signal PRED, and the mapping parameter g. side and g predTo generate it. For example, downmix DMX is obtained from downmixing by using an all-pass filter to enhance stereo fill or by using delay to enhance stereo fill.
[0121] The superposition is described by the following equation:
[0122] L[k][b][i]=(1+g side [k][b])DMX[k][b][i]
[0123] +g pred [k][b]g norm PRED[k][b][i]
[0124] as well as
[0125] R[k][b][i]=(1-g side [k][b])DMX[k][b][i]
[0126] -g pred [k][b]g norm PRED[k][b][i]
[0127] Upmixing is performed for each subframe k in all intervals i within frequency band b, as described in the previously shown table. Furthermore, each side gain g... side By energy normalization factor g norm To weight the energy, the energy normalization factor is based on the energy of the downmixed DMX and the residual prediction gain parameter PRED or g. pred [k][b] is calculated as described above.
[0128] The mapped and smoothed side gain 755 and the mapped and smoothed residual gain 756 are input to the output interface 120 to obtain a smoothed audio scene. Therefore, using smoothing parameters to process the encoded audio scene, as described previously, results in an improvement in the trade-off between achievable audio quality and implementation effort.
[0129] Figure 8An apparatus for decoding a transmission signal 122 according to an embodiment is shown. An (encoded) audio signal 816 is input to a transmission signal core decoder 810 to perform core decoding on the (core-encoded) audio signal 816 to obtain a (decoded raw) transmission signal 812, which is input to an output interface 120. For example, the transmission signal 122 may be an encoded transmission signal 812 output from a transmission signal core encoder 810. The (decoded) transmission signal 812 is input to the output interface 120, which is configured to generate a raw representation 818 of two or more channels (e.g., left and right channels) using a parameter set 814 including a second parameter set 114. For example, the transmission signal core decoder 810 for decoding the core-encoded audio signal to obtain the transmission signal 122 is an ACELP decoder. Furthermore, the core decoder 810 is configured to feed decoded raw transmission signals 812 in two parallel branches. The first branch of the two parallel branches includes an output interface 120, and the second branch of the two parallel branches includes a transmission signal enhancer 820 or a multi-channel enhancer 990, or both. A signal combiner 940 is configured to receive a first input to be combined from the first branch and a second input to be combined from the second branch.
[0130] like Figure 9 As shown, the apparatus for processing the encoded audio scene 130 may use a bandwidth extension processor 910. A low-frequency transmission signal 901 is input to an output interface 120 to obtain a two-channel low-frequency representation 972 of the transmission signal. It should be noted that the output interface 120 processes the transmission signal 901 in the frequency domain 955, for example, during the upmixing process 960, and converts the two-channel transmission signal 901 in the time domain 966. This is done by a converter 970, which converts the upmixed spectral representation 962 appearing in the frequency domain 955 to the time domain to obtain the two-channel low-frequency representation 972 of the transmission signal.
[0131] like Figure 8 As shown, a mono low-frequency transmission signal 901 is input to a converter 950, which performs, for example, a conversion of the time portion of the transmission signal 901 corresponding to the output time frame 220 to a spectral representation 952 of the transmission signal 901, i.e., a conversion from the time domain 966 to the frequency domain 955. For example, as described in FIG2, the portion (of the output time frame) is shorter than the input time frame 210, in which parameters 252 of the first parameter set 112 are organized.
[0132] The spectral representation 952 is input to an upmixer 960 to upmix the spectral representation 952 using, for example, a second parameter set 114, thereby obtaining an upmixed spectral representation 962, which is (still) processed in the frequency domain 955. As previously described, the upmixed spectral representation 962 is input to a converter 970 for converting the upmixed spectral representation 962 (i.e., each of two or more channels) from the frequency domain 955 to the time domain 966 (time representation) to obtain a low-frequency band representation 972. Thus, two or more channels in the upmixed spectral representation 962 are computed. Preferably, the output interface 120 is configured to operate in the complex discrete Fourier transform domain, wherein the upmixing operation is performed in the complex discrete Fourier transform domain. A conversion from the complex discrete Fourier transform domain back to the real-valued time domain representation is performed using the converter 970. In other words, the output interface 120 is configured to use the upmixer 960 in the second domain (i.e., the frequency domain 955) to generate a raw representation of two or more channels, where the first domain represents the time domain 966.
[0133] In this embodiment, the upmixing operation of the upmixer 960 is based on the following equation:
[0134]
[0135] as well as
[0136]
[0137] in, It is a transmission signal 901 for frame t and frequency range k, where, This refers to the side gain parameter 455 for frame t and subband b, where... The residual prediction gain parameter is 456 for frame t and subband b, where g norm It is a dispensable energy adjustment factor, and among them, It refers to the original residual signal for frame t and frequency interval k.
[0138] Compared to the low-frequency transmission signal 901, transmission signals 902 and 122 are processed in the time domain 966. Transmission signal 902 is input to a bandwidth extension processor (BWE processor) 910 to generate a high-frequency band signal 912, and is input to a multi-channel filter 930 to apply a multi-channel fill operation. The high-frequency band signal 912 is input to an upmixer 920 to upmix the high-frequency band signal 912 into an upmixed high-frequency band signal 922 using a second parameter set 144 (i.e., the parameters of output time frames 262 and 532). For example, the upmixer 920 can apply broadband shift processing to the high-frequency band signal 912 in the time domain 966 using at least one parameter from the second parameter set 114.
[0139] Low-frequency band representation 972, upmixed high-frequency band signal 922, and multi-channel fill transmission signal 932 are input to signal combiner 940 to combine the result of broadband shift 922, the result of stereo fill 932, and the low-frequency band representation 972 of two or more channels in time domain 966. This combination results in a full-band multi-channel signal 942 in time domain 966 as a channel representation. As previously outlined, converter 970 converts each of the two or more channels in the spectral representation 962 to a time representation to obtain the original time representation 972 of the two or more channels. Therefore, signal combiner 940 combines the original time representation of two or more channels with the enhanced time representation of two or more channels.
[0140] In this embodiment, only the low-band (LB) transmission signal 901 is input to the output interface 120 (DFT stereo) processing, while the high-band (HB) transmission signal 912 is separately upmixed in the time domain (using upmixer 920). An ambient contribution is generated using the BWE processor 910 along with time-domain stereo fill (using multichannel filler 930), a process implemented via a translation operation. The translation process involves a wideband translation based on the mapped side gain per frame (e.g., a mapped and smoothed side gain 755). Here, only a single gain exists per frame covering the entire high-band frequency range, which simplifies the calculation of the left and right high-band channels in the downmixer based on the following equation:
[0141] H left [k][i]=HB dmx [k][i]+g side,hb [k]*HB dmx [k][i]
[0142] as well as
[0143] HB right [k][i]=HB dmx [k][i]-sidegain hb [k]*HB dmx [k][i]
[0144] For each sample in each subframe.
[0145] High-frequency stereo fill signal PRED hb (That is, the multi-channel fill transmission signal 932) is transmitted via delay HB. dmx and through g side,hb Weighted and additionally using the energy normalization factor g norm To obtain, as described in the following equation:
[0146] PRED hb,left [i] = gpred,hb *g norm *HB dmx [id]
[0147] as well as
[0148] PRE hb,right [i] = -g pred,hb *g norm *HB dmx [id]
[0149] For each sample i in the current time frame (performed for the full time frame 210, but not for time subframes 213 and 214), d is the number of samples for which the high-frequency downmixing is delayed to generate the fill signal 932 obtained by the multi-channel filler 930. Other methods besides delay can be performed to generate the fill signal, such as more advanced decorrelation processing or using a noise signal or any other signal derived from the transmitted signal in a different manner compared to the delay.
[0150] After DFT synthesis, signal combiner 940 is used to combine (mix back) the translated stereo signals 972 and 922 and the generated stereo fill signal 932 into the core signal.
[0151] This description process for the ACELP high-frequency band contrasts with the higher-latency DirAC processing, in which the ACELP core and TCX frames are artificially delayed to align with the ACELP high-frequency band. There, CLDFB (analysis) is performed on the complete signal, meaning that upmixing of the ACELP high-frequency band is also done in the CLDFB domain (frequency domain).
[0152] Figure 10 An embodiment of an apparatus for obtaining a processed audio scene 124 is shown. A transmission signal 122 is input to an output interface 120 to generate a raw representation 972 of two or more channels, and an enhanced representation 992 of two or more channels is generated using a second parameter set 114 and a multichannel enhancer 990. For example, the multichannel enhancer 990 is configured to perform at least one of a set of operations including bandwidth expansion, gap filling, quality enhancement, or interpolation. Both the raw representation 972 of the two or more channels and the enhanced representation 992 of the two or more channels are input to a signal combiner 940 to obtain the processed audio scene 124.
[0153] Figure 11A block diagram of an embodiment of a multichannel enhancer 990 for generating an enhanced representation 992 of two or more channels is shown. The multichannel enhancer 990 includes a transmit signal enhancer 820, an upmixer 830, and a multichannel filler 930. A transmit signal 122 and / or a decoded original transmit signal 812 are input to the transmit signal enhancer 820, which generates an enhanced transmit signal 822, which is then input to the upmixer 830 and the multichannel filler 930. For example, the transmit signal enhancer 820 is configured to perform at least one of a set of operations including bandwidth expansion, gap filling, quality enhancement, or interpolation.
[0154] like Figure 9 As seen in the diagram, the multichannel filler 930 uses the transmission signal 902 and at least one parameter 532 to generate a multichannel fill transmission signal 932. In other words, the multichannel enhancer 990 is configured to generate an enhanced representation of two or more channels 992 using either the enhanced transmission signal 822 and the second parameter set 114, or using the enhanced transmission signal 822 and the upmixed enhanced transmission signal 832. For example, the multichannel enhancer 990 includes an upmixer 830 or a multichannel filler 930, or both, for generating an enhanced representation of two or more channels 992 using either the transmission signal 122 or the enhanced transmission signal 933 and at least one parameter from the second parameter set 532. In embodiments, the transmission signal enhancer 820 or the multichannel enhancer 990 is configured to operate in parallel with the output interface 120 when generating the original representation 972, or the parameter converter 110 is configured to operate in parallel with the transmission signal enhancer 820.
[0155] exist Figure 13 In the process, the bitstream 1312 sent from the encoder to the decoder can be... Figure 12 The DirAC-based upmixing scheme shown is the same. The single transmission channel 1312 derived from the DirAC-based spatial downmixing process is input to the core decoder 1310, decoded using the core decoder (e.g., an EVS or IVAS mono decoder), and sent along with the corresponding DirAC side parameters 1313.
[0156] In this DFT stereo method used to handle audio scenarios without additional latency, the initial decoding in the mono core decoder (IVAS mono decoder) of the transmission channel also remains unchanged. Not through... Figure 12In the CLDFB filter bank 1220, the decoded downmixed signal 1314 is input to the DFT analysis 1320 to transform the decoded mono signal 1314 to the STFT domain (frequency domain), for example, by using windows with very short overlap. Therefore, using only the remaining margin between the total delay and the delay already caused by the MDCT analysis / synthesis of the core decoder, the DFT analysis 1320 does not introduce any additional delay relative to the target system delay of 32ms.
[0157] DirAC side parameters 1313 or the first parameter set 112 are input to parameter map 1360, which may include, for example, a parameter converter 110 or parameter processor for obtaining DFT stereo side parameters (i.e., the second parameter set 114). Frequency domain signal 1322 and DFT side parameters 1362 are input to DFT stereo decoder 1330 for example, by using… Figure 9 The upmixer 960 described herein generates a stereo upmix signal 1332. The two channels of the stereo upmix 1332 are input to a DFT synthesizer for, for example, using... Figure 9 The converter 970 described herein converts the stereo upmixer 1332 from the frequency domain to the time domain, thereby generating an output signal 1342 that can represent the processed audio scene 124.
[0158] Figure 14 An embodiment using bandwidth extension 1470 to process encoded audio scenarios is shown. The bitstream 1412 is input to the ACELP core or low-band decoder 1410 instead of... Figure 13 In the IVAS mono decoder described herein, a decoded low-frequency band signal 1414 is generated. The decoded low-frequency band signal 1414 is input to a DFT analyzer 1420 to convert the signal 1414 into a frequency domain signal 1422, for example, from... Figure 9 The spectrum representation 952 of the transmitted signal 901. The DFT stereo decoder 1430 can represent an upmixer 960, which uses the decoded low-frequency band signal 1442 in the frequency domain and the DFT stereo side parameters 1462 from the parameter map 1460 to generate an LB stereo upmix 1432. The generated LB stereo upmix 1432 is input to the DFT synthesizer 1440 for use, for example... Figure 9 A converter 970 is used to perform the time-domain conversion. The low-frequency band representation 972 of the transmission signal 122 (i.e., the output signal 1442 of the DFT synthesis stage 1440) is input to a signal combiner 940, which combines the upmixed high-frequency stereo signal 922, the multi-channel fill high-frequency transmission signal 932, and the low-frequency band representation 972 of the transmission signal to generate a full-band multi-channel signal 942.
[0159] Parameter 1415 of the decoded LB signal 1414 and BWE 1470 is input to the ACELP BWE decoder 910 to generate the decoded high-frequency band signal 912. Mapped side gain 1462 (e.g., mapped and smoothed side gain 755 for the low-frequency band spectrum region) is input to the DFT stereo block 1430, and mapped and smoothed single side gain for the entire high-frequency band is forwarded to the high-frequency band upmixing block 920 and the stereo fill block 930. The HB upmixing block 920, used to upmix the decoded HB signal 912 using the high-frequency band side gain 1472 (e.g., parameter 532 from the second parameter set 114 of the output time frame 262), generates the upmixed high-frequency band signal 922. The stereo fill block 930, used to fill the decoded high-frequency band transmission signals 912, 902, uses parameters 532, 456 from the second parameter set 114 of the output time frame 262 and generates the high-frequency band filled transmission signal 932.
[0160] In summary, embodiments of the invention have created a concept for processing encoded audio scenarios using parameter transformation and / or bandwidth expansion and / or parameter smoothing (which results in a trade-off improvement between total latency, achievable audio quality, and implementation effort).
[0161] Subsequently, other embodiments of various aspects of the invention, and specifically combinations thereof, are shown. The proposed solution for achieving low-latency upmixing is to use a parametric stereo approach, such as the one described in [4], which employs a short-time Fourier transform (STFT) filter bank instead of a DirAC renderer. In this “DFT stereo” approach, upmixing from a downmix channel to a stereo output is described. The advantage of this approach is that it has a very short overlapping window for DFT analysis at the decoder, which allows it to remain within the much lower total latency (32ms) required by communication codecs (such as EVS [3]) or the upcoming IVAS codec. Furthermore, unlike DirAC CLDFB, DFT stereo processing is not a post-processing step of the core encoder, but operates in parallel with a portion of the core processing (i.e., the bandwidth extension (BWE) of the algebraic text-excited linear prediction (ACELP) speech encoder) without exceeding that already given latency. DFT stereo processing can therefore be called latency-free relative to the 32ms latency of EVS, as it operates at the same total encoder latency. On the other hand, DirAC can be considered a post-processor that results in an additional 5ms latency due to CLDFB extending the total latency to 37ms.
[0162] Typically, latency gains are achieved. The low latency comes from processing steps that occur in parallel with core processing, while the exemplary CLDFB version is a post-processing step used for the required rendering after core encoding.
[0163] Unlike DirAC, all components except ACELP BWE are transformed to the DFT domain using only a very short overlapping window of 3.125 ms (within acceptable margin), and DFT stereo uses an artificial delay of 3.25 ms for these components without introducing further delay. Thus, only TCX and ACELP (without BWE) are upmixed in the frequency domain, while ACELPBWE is upmixed in the time domain via a separate delay-free processing step called Inter-Channel Bandwidth Extension (ICBWE) [5]. This time-domain BWE processing varies slightly in the specific stereo output case of a given embodiment, which will be described at the end of the embodiment.
[0164] The transmitted DirAC parameters cannot be directly used for DFT stereo upmixing. Therefore, a mapping from given DirAC parameters to corresponding DFT stereo parameters becomes necessary. While DirAC uses azimuth and elevation angles along with the spread parameters for spatial placement, DFT stereo has single-side gain parameters for translation and residual prediction parameters that are closely related to the stereo width and therefore closely related to the spread parameters of DirAC. In terms of parameter resolution, each frame is divided into two subframes and each subframe is divided into several frequency bands. [6] describes the side gain and residual gain used in DFT stereo.
[0165] The DirAC parameters are derived from a band-wise analysis of the audio scene initially in B format or FOA. Then, it derives the azimuth θ(bn) and elevation angle for each band k and time instance n. The main arrival direction and diffusion factor ψ(b,n) are given. For the directional components, the first-order spherical harmonic at the center position is derived from the omnidirectional component w(b,n) and the DirAC parameter:
[0166]
[0167]
[0168]
[0169]
[0170] Furthermore, a stereo version can be obtained from the FOA channel through decoding involving W and Y, resulting in two cardioid pointing azimuth angles of +90 and -90 degrees.
[0171]
[0172] This decoding corresponds to first-order beamforming pointing in two directions.
[0173]
[0174] Therefore, there is a direct correlation between stereo output and DirAC parameters. On the other hand, DFT parameters depend on models for the L and R channels based on the intermediate signal M and the side signals S.
[0175]
[0176] M is transmitted as a mono channel and corresponds to the omnidirectional channel W in SBA mode. In DFT stereo, S is predicted from M using the side gain, which can then be expressed using DirAC parameters as follows:
[0177]
[0178] In DFT stereo, the predicted residuals are assumed and expected to be incoherent, and are modeled using their energy and the decorrelated residual signals going to the left and right channels. The predicted residuals of S (using M) can be expressed as:
[0179] R(b) = S(b) - sidegain[b]M(b)
[0180] Furthermore, its energy is modeled in DFT stereo using the predicted gain as follows:
[0181] ||R(b)|| 2 =respred[b]||M(b)|| 2
[0182] Since residual gain represents the inter-channel incoherent components and spatial width of a stereo signal, it is directly related to the diffusion component modeled by DirAC. Therefore, the residual energy can be rewritten as a function of the DirAC diffusion parameters:
[0183] ||R(b)|| 2 =ψ(b)||M(b)|| 2
[0184] Because the frequency band configuration of DFT stereo is typically different from that of DirAC, it must be adjusted to cover the same frequency range as DirAC. For these bands, the DirAC's direction angle can then be mapped to the DFT stereo's side gain parameter via the following formula.
[0185]
[0186] Where is the current frequency band, and the parameter range is [0; 360] for azimuth, [0; 180] for elevation, and [-1; 1] for the resulting side gain value. However, DirAC's directional parameters typically have a higher temporal resolution than DFTSereo, meaning that two or more azimuth and elevation values must be used to calculate a side gain value. One approach is to average across subframes, but in this implementation, the calculation is based on energy-related weights. For all K DirAC subframes, the subframe energy is calculated as follows:
[0187]
[0188] Where x is the time-domain input signal, is the number of samples in each subframe, and is the sample index. For each DFT stereo subframe, the weight of the contribution included in each DirAC subframe can then be calculated as:
[0189]
[0190] The side gain is then finally calculated as:
[0191]
[0192] Due to the similarity between parameters, a spread value for each frequency band is directly mapped to the residual prediction parameters of all subframes in the same frequency band.
[0193] resprea[l][b]=diffuseness[b]
[0194] Furthermore, the parameters are smoothed over time to avoid sharp fluctuations in gain. Since relatively strong smoothing is needed most of the time, but a faster response is required whenever the scene changes abruptly, a smoothing factor is adaptively calculated to determine the smoothing strength. This adaptive smoothing factor is calculated band-wise based on the energy changes in the current frequency band. Therefore, the band-wise energy in all subframes must first be calculated:
[0195]
[0196] Where x is the frequency range (real and imaginary parts) of the signal after DFT transformation, and is the interval index of all intervals in the current frequency band.
[0197] To capture the energy variation over time, two average values (a short-term average and a long-term average) are then calculated for each frequency band b according to the following formula:
[0198]
[0199] as well as
[0200]
[0201] Where, N short and N long N is the number of previous subframes k, for which individual averages are calculated. In this particular implementation, N short Set to 3, and N long Set it to 10. Then calculate the smoothing factor based on the quotient of the averages, such that a higher short-term average (indicating a recent increase in energy) leads to a decrease in smoothing:
[0202]
[0203] The higher long-term average indicating a decrease in energy does not lead to a smoothing reduction, so the smoothing factor is now set to the maximum value of 1.
[0204] The above formula limits the minimum value to fac smooth [b]to (In this implementation, it is 0.3). However, in extreme cases, this factor must be close to 0, which is achieved by reducing the value from the range via the following formula. The reason for converting to the range [0; 1]:
[0205]
[0206] For less extreme cases, smoothing is now excessively reduced, so the factor is compressed towards a value of 1 via the root function. Since stability is particularly important in the lowest frequency bands, the fourth root is used in the bands b=0 and b=1:
[0207]
[0208] All other frequency bands with b>1 are compressed by the square root.
[0209]
[0210] In this way, extreme cases remain close to 0, and less rapid increases in energy do not so drastically reduce smoothness.
[0211] Finally, depending on the frequency band, the maximum smoothness is set (a factor of 1 will simply repeat the previous value without any contribution from the current gain):
[0212] fac smooth [b] = min(fac) smooth [b],bounds[b]
[0213] Wherein, bounds[b] is set according to the following table in a given implementation with 5 frequency bands.
[0214] b Boundary [b] 0 0.98 1 0.97 2 0.95 3 0.9 4 0.9
[0215] Calculate the smoothing factor for each DFT stereo subframe k in the current frame.
[0216] In the final step, both the side gain and the residual prediction gain are recursively smoothed according to the following formula.
[0217] g side [k][b]=fac smooth [k][b]g side [k-1][b]
[0218] +(1-fac smooth [k][b])g side [k][b]
[0219] as well as
[0220] g red [k][k]=fac smooth [k][b]g pred [k-1][b]
[0221] +(1-fac smooth [k][b])g pred [k][b]
[0222] These mapped and smoothed parameters are now fed into the DFT stereo processing, where the stereo signal L / R is obtained from the downmixer DMX, the residual prediction signal PRED (obtained from the downmixer via “enhanced stereo fill” using an all-pass filter [7] or via regular stereo fill using a delay), and the mapped parameter g. side and g pred Generation. The supermixing is usually described by the following formula [6]:
[0223] L[k][b][i]=(1+g side [k][b])DMX[k][b][i]
[0224] +g pred [k][b]g norm PRED[k][b][i]
[0225] as well as
[0226] R[k][b][i]=(1-g side [k][b])DMX[k][b][i]
[0227] -g pred [k][b]g norm PRED[k][b][i]
[0228] For each subframe k, all intervals i in frequency band b. Furthermore, each side gain g... side By energy normalization factor g norm The energy normalization factor is calculated based on the energies of DMX and PRED to weight the energy.
[0229] Finally, the upmixed signal is transformed back to the time domain via IDFT for playback on a given stereo setting.
[0230] Because the "Time-Domain Bandwidth Extension" (TBE)[8] used in ACELP generates its own delay (in the implementation, this example is based on exactly 2.3125ms), it cannot be transformed to the DFT domain while remaining within a total delay of 32ms (of which 3.25ms is left for the STFT, which already uses a 3.125ms stereo decoder). Therefore, only the low-frequency band (LB) is input to the DFT domain. Figure 14 The 1450 indicates DFT stereo processing, while the high-frequency band (HB) must be mixed separately in the time domain, such as... Figure 14 As shown in block 920. In conventional DFT stereo, this is done via translation via interchannel bandwidth extension (ICBWE) [5] plus temporal stereo fill environment. In the given case, the stereo fill in block 930 is calculated in the same way as in conventional DFT stereo. However, due to the lack of parameters, the ICBWE processing is completely skipped, and the mapping-based side gain 1472 is replaced by a low-resource approach that requires broadband translation in block 920. In the given embodiment, there is only a single gain covering the entire HB region, which simplifies the calculation of the left and right HB channels from the lower mix channel to the lower equation in block 920.
[0231] HB left [k][i]=HB dmx [k][i]+g side,hb [k]*HB dmx [k][i]
[0232] as well as
[0233] HB right [k][i]=HB dmx [k][i]-sidegain hb [k]*BH dmx [k][i]
[0234] For each sample i in each subframe k.
[0235] In block 930, HB is delayed. dmx and through g side,hb and energy normalization factor g normWeighting is performed to obtain the HB stereo fill signal PRED hb
[0236] PRED hb,left [i] = g pred,bb *g norm *HB dmx [id]
[0237] as well as
[0238] PRED hb,right [i] = -g pred,hb *g norm *HB dmx [id]
[0239] For each sample i in the current frame (performed on the full frame rather than a subframe), and where d is the number of samples undermixed with respect to the padding signal delay HB.
[0240] After DFT synthesis in combiner 940, the translated stereo signal and the generated stereo fill signal are finally mixed back into the core signal.
[0241] This special processing of ACELP HB contrasts with the higher-latency DirAC processing, in which the ACELP core and TCX frames are artificially delayed to align with ACELP HB. There, CLDFB is performed on the complete signal; that is, the upmixing of ACELP HB is also done in the CLDFB domain.
[0242] Advantages of the proposed method
[0243] In this special case where the SBA is input to the stereo output, no additional delay allows the IVAS codec to maintain the same total delay (32ms) as the EVS.
[0244] Due to its overall simpler and more straightforward processing, the complexity of parametric stereo upmixing space DirAC rendering via DFT is much lower.
[0245] Other preferred implementation schemes
[0246] 1. An apparatus, method, or computer program for encoding or decoding as described above.
[0247] 2. An apparatus or method for encoding or decoding, or a related computer program, comprising:
[0248] • A system in which the input is encoded using a model based on a spatial audio representation of a sound scene using a first set of parameters, and decoded at the output using a stereo model for two output channels or a multi-channel model for more than two output channels using a second set of parameters; and / or
[0249] • Mapping of spatial parameters to stereo parameters; and / or
[0250] • Conversion from an input representation / parameter based on one frequency domain to an output representation / parameter based on another frequency domain; and / or
[0251] • Conversion from parameters with higher time resolution to parameters with lower time resolution; and / or
[0252] • Lower output delay due to the shorter overlapping windows of the second frequency conversion; and / or
[0253] • Map DirAC parameters (direction angle, diffusion) to DFT stereo parameters (side gain, residual prediction gain) to output SBA DirAC encoded content as stereo; and / or
[0254] • Transformation from CLDFB-based input representations / parameters to DFT-based output representations / parameters; and / or
[0255] • Conversion of parameters with 5ms resolution to parameters with 10ms resolution; and / or
[0256] • Benefits: Compared to CLDFB, DFT has lower output latency due to shorter window overlap.
[0257] It should be noted here that all the alternatives or aspects discussed above, as well as all aspects defined by the independent claims in the appended claims, can be used individually; that is, there are no other alternatives or objectives different from the contemplated alternatives, objectives, or independent claims. However, in other embodiments, two or more alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.
[0258] As will be outlined, different aspects of the invention relate to parameter transformation, smoothing, and bandwidth expansion. In the embodiments described above, these aspects can be implemented individually or independently, or any two of the three aspects can be combined, or all three aspects can be combined.
[0259] The encoded signal of the present invention can be stored on a digital storage medium or a non-transitory storage medium, or it can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0260] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent descriptions of the corresponding methods, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent descriptions of features of the corresponding block or item or the corresponding apparatus.
[0261] Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or software. Implementations may be carried out using a digital storage medium (e.g., floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory) on which electronically readable control signals are stored, in cooperation with (or capable of cooperating with) a programmable computer system, such that the corresponding methods are executed.
[0262] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0263] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.
[0264] Other embodiments include a computer program stored on a machine-readable carrier or non-transitory storage medium for performing one of the methods described herein.
[0265] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.
[0266] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, the computer program being used to perform one of the methods described herein.
[0267] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0268] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.
[0269] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0270] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0271] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.
[0272] Bibliography or references
[0273] [1]V.Pulkki,M.-VVJLaitinen,J.Ahonen,T.Lokki and T. "Directional audio coding-perception-based reproduction of spatial sound," inINTERNATIONAL WORKSHOP ON THE PRINCIPLES AND APPLICATION ON SPATIAL HEARING, 2009.
[0274] [2]G.Fuchs,O.Thiergart,S.Korse,S. M. Multrus, F. Küch, Bouthéon, A. Eichenseer and S. Bayer, "Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding using low-order, mid-order and high-order components generators". WO Patent 2020115311A1,11 06 2020.
[0275] [3]3GPP TS 26.445,Codec for Enhanced Voice Services(EVS);Detailedalgorithmic description.
[0276] [4]S.Bayer,M.Dietz,S. E.Fotopoulou,G.Fuchs,W.Jaegers,G.Markovic,M.Multrus,E.Ravelli and M.Schnell,"APPARATUSANDMETHODFORESTIMATINGANINTER-CHANNELTIME DIFFERENCE".Patent WO17125563,27 07 2017.
[0277] [5]V.S.C.S.Chebiyyam and V.Atti,"Inter-channel bandwidth extension".WO Patent 2018187082A1,11 10 2018.
[0278] [6]J.Büthe,G.Fuchs,W. F.Reutelhuber,J.Herre,E.Fotopoulou,M.Multrus and S.Korse,"Apparatus and method for encoding or decoding amultichannel signal using a side gain and a residual gain".WO PatentWO2018086947A1,17 05 2018.
[0279] [7]J.Büthe,F.Reutelhuber,S.Disch,G.Fuchs,M.Multrus and R.Geiger,"Apparatus for Encoding or Decoding an Encoded Multichannel Signal Using aFilling Signal Generated by a Broad Band Filter".WO Patent WO2019020757A2,3101 2019.
[0280] [8]V.A.e.al.,"Super-wideband bandwidth extension for speech in the3GPP EVS codec,"in IEEE International Conference on Acoustics,Speech andSignal Processing(ICASSP),Brisbane,2015。
Claims
1. An apparatus for processing an audio scene representing a sound field, the audio scene including information and a set of parameters about transmitted signals, the apparatus comprising: An output interface is provided for generating a raw representation of two or more channels using the parameter set and the transmitted signal, wherein the output interface is configured as follows: The transmitted signal, which is a low-frequency band signal, is converted into a spectral representation; Upmix the spectral representation; and The upmixed spectral representation is transformed in the time domain to obtain the low-frequency band representation of the two or more channels, which serves as the original representation of the two or more channels. Multi-channel enhancers include: A transmission signal enhancer for generating an enhanced transmission signal, wherein the transmission signal enhancer includes a bandwidth extension processor for generating a high-frequency band signal from the transmission signal in the time domain as the enhanced transmission signal; An upmixer is configured to upmix the enhanced transmission signal to obtain an upmixed enhanced transmission signal, wherein the upmixer is configured to apply a broadband shift in the time domain to a high-frequency band signal of the transmission signal using at least one parameter from the parameter set to obtain the upmixed enhanced transmission signal; and A multi-channel filler for generating a multi-channel filled transmission signal using the enhanced transmission signal and at least one parameter from the parameter set, wherein the multi-channel filler is configured to apply a stereo fill operation to a high-frequency band signal of the transmission signal in the time domain to obtain the multi-channel filled transmission signal; and A signal combiner is used to combine the upmixed enhancement transmission signal, the multichannel fill transmission signal, and the low-frequency band representation of the two or more channels as the original representation of the two or more channels in the time domain to obtain a full-band multichannel signal in the time domain as a processed audio scene.
2. The apparatus of claim 1, wherein, The transmission signal enhancer and the upmixer are configured to operate in parallel with the output interface when generating the original representation of the two or more channels.
3. The apparatus according to claim 1, further comprising: A parameter converter is used to convert a received set of parameters into a parameter set associated with a channel representation comprising two or more channels, for reproducing the two or more channels at a predefined spatial location, the channel representation corresponding to a processed audio scene. The parameter converter is configured to operate in parallel with the transmission signal enhancer.
4. The apparatus according to claim 1, wherein, The device is configured to receive a set of parameters, and The device further includes: a parameter converter for converting a received parameter set into the parameter set, wherein the parameter set is associated with a channel representation including the two or more channels, for reproducing the two or more channels at a predefined spatial location, the channel representation corresponding to the processed audio scene; and The output interface is configured to generate the processed audio scene using the parameter set and information about the transmitted signal.
5. The apparatus according to claim 4, wherein, For each of the multiple input time frames and for each of the multiple input frequency bands, the received parameter set includes at least one DirAC parameter. The parameter converter is configured to calculate the parameter set into parametric stereo or multi-channel parameters.
6. The apparatus according to claim 5, wherein, The at least one parameter includes at least one of the following: arrival direction parameter, diffusion parameter, direction information parameter related to the sphere with the virtual listening position as the origin of the sphere, and distance parameter. The stereo or multi-channel parameters include at least one of the following: side gain parameter, residual prediction gain parameter, inter-channel level difference parameter, inter-channel time difference parameter, inter-channel phase difference parameter, and inter-channel coherence parameter.
7. The apparatus according to claim 4, wherein, The input time frame associated with the received parameter set comprises two or more input time subframes, wherein the output time frame associated with the parameter set is shorter than the input time frame but longer than one of the two or more input time subframes. The parameter converter is configured to: calculate original parameters in the parameter set for each of the two or more temporally successive input time subframes, and combine at least two original parameters to derive parameters in the parameter set related to the output subframe.
8. The apparatus according to claim 7, wherein, The parameter converter is configured to perform a weighted combination of the at least two original parameters in the combination of at least two original parameters, wherein the weighting factor for the weighted combination is derived based on amplitude correlation measurements of the transmitted signal in the corresponding input time subframe.
9. The apparatus according to claim 8, wherein, The parameter converter is configured to use energy or power as the amplitude-related measurement, and wherein, when the energy or power of the transmitted signal is high in the input time frame, the weighting factor for the corresponding input subframe is larger than the weighting factor for the corresponding input subframe where the energy or power of the transmitted signal is low in the input time frame.
10. The apparatus according to claim 5, wherein, The parameter converter is configured to compute at least one raw parameter for each output time frame using at least one parameter from the received parameter set for the input time frame. The parameter converter is configured to calculate a smoothing factor for each original parameter according to a smoothing rule, and The parameter converter is configured to apply a corresponding smoothing factor to the corresponding original parameter to derive parameters in the parameter set for the output time frame.
11. The apparatus according to claim 10, wherein, The parameter converter is configured as follows: The long-term average value is calculated from the amplitude correlation measurement of the first time portion of the transmitted signal, and A short-term average value is calculated for the amplitude correlation measurement of the second time portion of the transmitted signal, wherein the second time portion is shorter than the first time portion. The smoothing factor is calculated based on the ratio between the long-term average and the short-term average.
12. The apparatus according to claim 10, wherein, The parameter converter is configured to use a compression function to calculate the smoothing factor of a frequency band, the compression function being different for different frequency bands, and wherein the compression function has a stronger compression strength for lower frequency bands than the compression function has for higher frequency bands.
13. The apparatus according to claim 10, wherein, The parameter converter is configured to calculate the smoothing factor using different maximum limits for different frequency bands, wherein the maximum limit for lower frequency bands is higher than the maximum limit for higher frequency bands.
14. The apparatus according to claim 10, wherein, The parameter converter is configured to apply a recursive smoothing rule as the smoothing rule to time-series output time frames, such that the smoothing parameter of the current output time frame is calculated by combining the parameters of the previous output time frame weighted by a first weight and the original parameters of the current output time frame weighted by a second weight, wherein the first weight and the second weight are derived from the smoothing factor of the current output time frame.
15. The apparatus according to claim 1, wherein, The output interface is configured as follows: A conversion is performed from a time portion of the transmitted signal corresponding to an output time frame to the spectral representation, wherein the portion is shorter than the input time frame, and parameters from the parameter set received in the input time frame are organized.
16. The apparatus according to claim 15, wherein, The output interface is configured as follows: Convert to the complex discrete Fourier transform domain. The upmixing is performed in the complex discrete Fourier transform domain, and Perform the conversion from the complex discrete Fourier transform domain to the real-valued time-domain representation.
17. The apparatus according to claim 1, wherein, The output interface is configured to perform the upmixing based on the following equation: = as well as = , in, It refers to the transmitted signal for frame t and frequency range k, where It is the first channel in the spectral representation of the two or more channels in the frame t and frequency interval k, where It is the second channel in the spectral representation of the two or more channels in the frame t and frequency interval k, wherein, These are the side gain parameters for frame t and subband b, where, It is the residual prediction gain parameter for the frame t and the subband b, where g norm It is a dispensable energy adjustment factor, and among them, It refers to the original residual signal for the frame t and the frequency interval k.
18. The apparatus according to claim 4, in, The received parameter set is the direction-of-arrival parameter for the input frequency band, and wherein the parameter set includes a side gain parameter for each input frequency band, and The parameter converter is configured to calculate the side gain parameter for the output frequency band using the following equation: , Where b is the output frequency band, sidegain is the side gain parameter, azimuth is the azimuth component of the direction of arrival parameter, and elevation is the elevation component of the direction of arrival parameter.
19. The apparatus according to claim 18, in, The received parameter set additionally includes a spread parameter for the input frequency band, and wherein the parameter converter is configured to calculate a side gain parameter for the output frequency band using the following equation. Wherein, diff(b) is the diffusion parameter for the input frequency band b.
20. The apparatus according to claim 4, in, The received parameter set includes spread parameters for each input frequency band, and The parameter set includes residual prediction gain parameters for the output frequency band, and Specifically, when the input frequency band and the output frequency band are equal to each other, the parameter converter will use the diffusion parameter from the input frequency band as the residual prediction gain parameter for the output frequency band, or derive the diffusion parameter for the output frequency band from the diffusion parameter for the input frequency band, and then use the diffusion parameter for the output frequency band as the residual prediction gain parameter for the output frequency band.
21. The apparatus according to claim 1, in, Information regarding the transmitted signal includes a core encoded audio signal, and wherein the device further includes: A transmission signal core decoder is used to perform core decoding on the core encoded audio signal to obtain the transmission signal.
22. The apparatus according to claim 21, wherein, The core decoder for the transmitted signal is in the ACELP decoder.
23. The apparatus according to claim 1, wherein, The multichannel filler is configured to generate the left channel of the multichannel filler transmission signal using the following formula: ,as well as The multichannel filler is configured to generate the right channel of the multichannel filler transmission signal using the following formula: in, It is the left channel of the multi-channel fill transmission signal of sample number i in the current frame, where, It is the right channel of the multi-channel fill transmission signal of sample number i in the current frame, where, It is the high-frequency band prediction gain as at least one parameter, where, It is the energy normalization factor, where i is the sample number in the current frame, and d is the high-frequency band signal in the enhanced transmission signal. The number of samples that were delayed.
24. A method for processing an audio scene representing a sound field related to a virtual listener's position, the audio scene including information about transmitted signals and a set of parameters, the method comprising: Using the parameter set and the transmitted signal to generate a raw representation of two or more channels includes: The transmitted signal, which is a low-frequency band signal, is converted into a spectral representation; Upmix the spectral representation; and The upmixed spectral representation is transformed in the time domain to obtain the low-frequency band representation of the two or more channels, which serves as the original representation of the two or more channels. Multichannel generation, including: The enhanced transmission signal is generated by using a bandwidth expansion operation to obtain a high-frequency band signal from the transmitted signal in the time domain as an enhanced transmission signal. The enhanced transmission signal is upmixed to obtain an upmixed enhanced transmission signal, wherein the upmixing includes: applying a broadband shift in the time domain to a high-frequency band signal of the transmission signal using at least one parameter from the parameter set to obtain the upmixed enhanced transmission signal; and A multi-channel fill transmission signal is generated using the enhanced transmission signal and at least one parameter from the parameter set for multi-channel fill, wherein the multi-channel fill comprises: applying a stereo fill operation to the high-frequency band signal of the transmission signal in the time domain to obtain the multi-channel fill transmission signal, and The upmixing enhancement transmission signal, the multichannel fill transmission signal, and the low-frequency band representation of the two or more channels as the original representation of the two or more channels are combined in the time domain to obtain a full-band multichannel signal in the time domain as a processed audio scene.
25. A storage medium having a computer program stored thereon, which, when run on a computer or processor, performs the method according to claim 24.
Citation Information
Patent Citations
Apparatus for Encoding or Decoding an Encoded Multichannel Signal Using a Filling Signal Generated by a Broad Band Filter
US20200152209A1