Packet loss concealment for DirAC-based spatial audio coding

The method addresses packet loss in DirAC-encoded audio by replacing lost spatial parameters with previously received data, maintaining audio quality and stability through hold strategy and dithering, effectively overcoming transmission issues.

JP7828378B2Active Publication Date: 2026-03-11FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2026-03-11

AI Technical Summary

Technical Problem

Existing solutions fail to effectively protect Directional Audio Coding (DirAC) spatial parameters from packet loss during transmission, leading to quality degradation in spatial audio signals.

Method used

A method for loss concealment in DirAC-encoded audio scenes, where lost spatial audio parameters are replaced by previously received information, using strategies like hold strategy, directional dithering, and extrapolation based on previously received directional information.

Benefits of technology

Enhances the robustness of DirAC-encoded audio by maintaining spatial audio quality despite packet loss, ensuring a stable and natural sound field reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007828378000031
    Figure 0007828378000031
  • Figure 0007828378000032
    Figure 0007828378000032
  • Figure 0007828378000033
    Figure 0007828378000033
Patent Text Reader

Abstract

To provide a concept and a method of loss concealment in directional audio coding (DirAC).SOLUTION: A method for loss concealment of a spatial audio parameter includes the steps of: receiving a first set of spatial audio parameters including at least first arriving direction information; receiving a second set of spatial audio parameters including at least second arriving direction information; and replacing second arriving direction information of the second set with replacement arriving direction information derived from first arriving direction information when at least second arriving direction information or a part of the second arriving direction information is lost or damaged.SELECTED DRAWING: Figure 3a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0003] Embodiments of the present invention relate to a method for loss concealment of spatial audio parameters, a method for decoding DirAC-encoded audio scenes, and corresponding computer programs. Further embodiments relate to a loss concealment device for loss concealment of spatial audio parameters, and a decoder comprising a packet loss concealment device. Preferred embodiments describe concepts / methods for compensating for quality degradation due to frame or packet losses and corruptions that occur during the transmission of audio scenes whose spatial images are parametrically coded according to the Directional Audio Coding (DirAC) paradigm. Introduction

[0002] Speech and audio communications can suffer from different quality issues due to packet loss during transmission. Indeed, adverse network conditions, such as bit errors and jitter, can lead to the loss of several packets. These losses result in severe artifacts such as clicks, plops, or unwanted muffling, which significantly degrade the perceived quality of the reconstructed speech or audio signal at the receiver. To combat the adverse effects of packet loss, packet loss concealment (PLC) algorithms have been proposed in traditional speech and audio coding methods. Such algorithms typically operate at the receiver by generating a synthetic audio signal to conceal the missing data in the received bitstream.

[0003] DirAC is a perceptually motivated spatial audio processing technique that compactly and efficiently represents a sound field using a set of spatial parameters and a downmix signal. The downmix signal can be a mono, stereo, or multichannel signal in an audio format such as A-format or B-format, also known as first-order Ambisonics (FAO). The downmix signal is complemented by spatial DirAC parameters that describe the audio scene in terms of direction of arrival (DOA) and diffuseness per time / frequency unit. For storage, streaming, or communication applications, the downmix signal is coded by a conventional core coder (e.g., EVS, a stereo / multichannel extension of EVS, or any other mono / stereo / multichannel codec) with the goal of preserving the audio waveforms of each channel. The core coder can be built around a transform-based coding scheme or speech coding scheme operating in the time domain, such as CELP. The core coder can then integrate existing error recovery tools, such as the packet loss concealment (PLC) algorithm. However, there are no existing solutions to protect DirAC spatial parameters, and therefore, improved methods are needed. Summary of the Invention [Problem to be solved by the invention]

[0004] The aim of the present invention is to provide a concept of loss concealment in the context of DirAC. [Means for solving the problem]

[0005] This object is solved by the subject matter of the independent claims.

[0006] An embodiment of the present invention provides a method for loss concealment of spatial audio parameters, where the spatial audio parameters include at least direction of arrival information. The method includes the following steps: receiving a first set of spatial audio parameters including first direction of arrival information and first diffuseness information; receiving a second set of spatial audio parameters, the second set including second direction of arrival information and second diffuseness information; and

[0007] Replacing the second set of second direction of arrival information with replacement direction of arrival information derived from the first direction of arrival information when at least the second direction of arrival information or a portion of the second direction of arrival information is lost.

[0008] The present invention is based on the finding that, in the event of loss or damage to incoming information, the lost / damaged information can be replaced by information derived from other available incoming information. For example, if a second incoming information is lost, it can be replaced by the first incoming information. In other words, the present invention provides packet loss concealment for spatial parametric audio, where the directional information is recovered in the event of transmission loss by using previously successfully received directional information and dithering. Thus, the present invention makes it possible to counter packet loss in the transmission of spatial audio sound directly encoded by parameters.

[0009] A further embodiment provides a method in which the first and second sets of spatial audio parameters include first and second diffusion information, respectively. In such a case, the approach can be as follows: According to an embodiment, the first or second diffusion information is derived from at least one energy ratio associated with at least one direction of arrival information. According to an embodiment, the method further includes replacing the second diffusion information of the second set by replacement diffusion information derived from the first diffusion information. This is part of a so-called hold strategy, which is based on the assumption that diffusion does not change much between frames. For this reason, a simple but effective approach is to retain the parameters of the last successfully received frame of a frame lost during transmission. Another part of this overall approach is replacing the second arrival information with the first arrival information, which was described in the context of the basic embodiment. It is generally safe to assume that the spatial image should be relatively stable over time, which can be translated into DirAC parameters, i.e., directions of arrival that likely do not change much between frames.

[0010] According to a further embodiment, the replacement direction of arrival information depends on the first direction of arrival information. In such a case, a strategy called directional dithering can be used. Here, the replacing step can include, according to an embodiment, dithering the replacement direction of arrival information. Alternatively or additionally, the replacing step can include injecting noise into the first direction of arrival information to obtain the replacement direction of arrival information. Dithering can then help make the rendered sound field more natural and comfortable by injecting random noise in a previous direction before using it in the same frame. According to an embodiment, the injecting step is preferably performed if the first or second diffuseness information exhibits high diffuseness. Alternatively, it may be performed if the first or second diffuseness information exceeds a predetermined threshold for diffuseness information exhibiting high diffuseness. According to a further embodiment, the diffuseness information contains more space relative to the ratio between directional and non-directional components of the audio scene described by the first and / or second set of spatial audio parameters. According to an embodiment, the injected random noise depends on the first and second diffuseness information. Alternatively, the injected random noise is scaled by a factor that depends on the first and / or second diffuseness information. Thus, according to an embodiment, the method may further comprise analyzing the tonality of the audio scene described by the first and / or second set of spatial audio parameters, analyzing the tonality of the transmitted downmix belonging to the first and / or second spatial audio parameters, to obtain a tonality value that describes the tonality. The injected random noise then depends on the tonality value. According to an embodiment, the scaling down is performed by a factor that decreases with the inverse of the tonality value or if the tonality increases.

[0011] According to a further approach, a method can be used that includes a step of estimating first direction of arrival information to obtain replacement direction of arrival information. According to this approach, it can be envisaged to estimate a directory of sound events in an audio scene and extrapolate the estimated directory. This is particularly relevant when the sound events are well localized in space and as point sources (direct model with low diffuseness). According to an embodiment, the extrapolation is based on one or more additional direction of arrival information belonging to one or more sets of spatial audio parameters. According to an embodiment, the extrapolation is performed if the first and / or second diffuseness information exhibits low diffuseness or if the first and / or second diffuseness information is below a predetermined threshold value of the diffuseness information.

[0012] According to an embodiment, the first set of spatial audio parameters belongs to a first time point and / or a first frame, and the second set of spatial audio parameters both belong to a second time point or a second frame. Alternatively, the second time point is after the first time point, or the second frame is after the first frame. Returning to the embodiment in which most sets of spatial audio parameters are used for the extrapolation, it is clear that preferably more sets of spatial audio parameters are used, e.g. belonging to multiple time points / frames that follow each other.

[0013] According to a further embodiment, the first set of spatial audio parameters comprises a first subset of spatial audio parameters for a first frequency band and a second subset of spatial audio parameters for a second frequency band, and the second set of spatial audio parameters comprises another first subset of spatial audio parameters for the first frequency band and another second subset of spatial audio parameters for the second frequency band.

[0014] Another embodiment provides a method for decoding DirAC-encoded audio scenes, comprising the step of decoding a DirAC-encoded audio scene comprising a downmix, a first set of spatial audio parameters and a second set of spatial audio parameters, the method further comprising the steps of the method for lossy concealment described above.

[0015] According to embodiments, the above-mentioned methods may be computer-implemented. The embodiments therefore refer to a computer-readable storage medium storing a computer program having a program code for executing, when executed on a computer, the method according to any one of the preceding claims.

[0016] Another embodiment relates to a loss concealment device for concealing losses of spatial audio parameters (including at least direction of arrival information). The device comprises a receiver and a processor. The receiver is configured to receive a first set of spatial audio parameters and a second set of spatial audio parameters (see above). The processor is configured to replace second direction of arrival information of the second set by replacement direction of arrival information derived from the first direction of arrival information if the second direction of arrival information is lost or damaged. Another embodiment relates to a decoder for a DirAC coded audio system comprising the loss concealment device. Embodiments of the present invention are described below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0017] [Figure 1a] 1 shows a schematic block diagram illustrating DirAC analysis and synthesis. [Figure 1b] 1 shows a schematic block diagram illustrating DirAC analysis and synthesis. [Figure 2] 1 shows a schematic detailed block diagram of DirAC analysis and synthesis in a low bitrate 3D audio coder. [Figure 3a]1 shows a schematic flow chart of a method for loss concealment according to a basic embodiment; [Figure 3b] 1 shows a schematic loss concealment device according to a basic embodiment; [Figure 4a] To illustrate the embodiment, a schematic diagram of the measured spread function for DDR (window size W=16 in FIG. 4a) is shown. [Figure 4b] To illustrate the embodiment, a schematic diagram of the measured spread function for DDR (window size W=512 in FIG. 4b) is shown. [Figure 5] To illustrate the embodiment, a schematic diagram of the direction (azimuth and elevation) measured as a function of divergence is shown. [Figure 6a] 1 shows a schematic flowchart of a method for decoding DirAC encoded audio scenes according to an embodiment; [Figure 6b] 1 shows a schematic block diagram of a decoder for DirAC encoded audio scenes according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0018]

[0023] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings, in which objects / elements having the same or similar functions are given the same reference numerals, so that the descriptions are mutually applicable and interchangeable. Before describing the embodiments of the present invention in detail, an introduction to DirAC will be given.

[0019] Introduction to DirAC: DirAC is a perceptually motivated spatial audio reproduction. It assumes that at a given time, for one important band, the spatial resolution of the auditory system is limited to decoding one cue for direction and another cue for interaural coherence.

[0020] Based on these assumptions, DirAC represents spatial sound in a frequency band by crossfading two streams: an omnidirectional diffuse stream and a directional non-diffuse stream. DirAC processing is performed in two stages: The first step is analysis, illustrated by FIG. 1a, and the second step is synthesis, illustrated by FIG. 1b.

[0021] FIG. 1a shows an analysis stage 10 comprising one or more bandpass filters 12a-n receiving microphone signals W, X, Y, and Z, an energy analysis stage 14e, and an intensity analysis stage 14i. By arranging them in time, a diffuseness Ψ (see reference numeral 16d) can be determined. The diffuseness Ψ is determined based on an analysis of the energy 14c and the intensity 14i. Based on the intensity and the analysis 14i, a direction 16e can be determined. The results of the direction determination are the azimuth and elevation angles. Ψ, azimuth, and ele are output as metadata. These metadata are used by a synthesis entity 20 shown in FIG. 1b.

[0022] The synthesis entity 20 shown in Fig. 1b comprises a first stream 22a and a second stream 22b. The first stream comprises a number of bandpass filters 12a-n and a calculation entity 24 for a virtual microphone. The second stream 22b comprises means for processing metadata, namely 26 for diffuseness parameters and 27 for directional parameters. Furthermore, the synthesis stage 20 uses a decorrelation entity 28, which receives the data of the two streams 22a, 22b. The output of the decorrelation entity 28 can be supplied to a speaker 29. In the DirAC analysis stage, a B-format first-order coincident microphone is considered as input and the sound diffuseness and direction of arrival are analyzed in the frequency domain.

[0023] In the DirAC synthesis stage, the sound is split into two streams: a non-diffuse stream and a diffuse stream. The non-diffuse stream is played as a point source using amplitude panning, which can be done using Vector-Based Amplitude Panning (VBAP) [2]. The diffuse stream, responsible for the sensation of envelopment, is generated by transmitting uncorrelated signals to the speakers.

[0024] DirAC parameters, hereafter also called spatial metadata or DirAC metadata, consist of a tuple of diffusivity and direction: direction can be expressed in spherical coordinates by two angles, azimuth and elevation, and diffusivity is a scalar coefficient between 0 and 1.

[0025] In the following, a system for DirAC spatial audio coding will be described with reference to Fig. 2. Fig. 2 shows a two-stage DirAC analysis 10' and DirAC synthesis 20', where the DirAC analysis comprises a filter bank analysis 12, a direction estimator 16i and a diffuseness estimator 16d. Both 16i ​​and 16d output diffuseness / direction data as spatial metadata. This data can be coded using an encoder 17. The direct analysis 20' comprises a spatial metadata decoder 21, an output synthesis 23 and a filter bank synthesis 12 that allows outputting signals to the loudspeakers FOA / HOA.

[0026] In parallel with the above-mentioned direct analysis stage 10' and direct synthesis stage 20', which process the spatial metadata, an EVS encoder / decoder is used. On the analysis side, beamforming / signal selection is performed based on the input signal B format (see beamforming / signal selection entity 15). The signal is then EVS encoded (see reference numeral 17). On the synthesis side (see reference numeral 20'), an EVS decoder 25 is used. This EVS decoder outputs a signal to the filter bank analysis 12, which in turn outputs the signal to the output synthesis 23. Now that the structure of the Direct Analysis / Direct Synthesis 10' / 20' has been described, the functionality will be described in detail.

[0027] The encoder analysis 10' typically analyzes a spatial audio scene in B format. Alternatively, the DirAC analysis can be tailored to analyze different audio formats, such as audio objects or multi-channel signals or any combination of spatial audio formats. The DirAC analysis extracts a parametric representation from the input audio scene. The direction of arrival (DOA) and the measured diffuseness per time-frequency unit form the parameters. The DirAC analysis is followed by a spatial metadata encoder, which quantizes and encodes the DirAC parameters to obtain a low-bitrate parametric representation.

[0028] The downmix signal, derived from different sources or audio input signals along with the parameters, is encoded for transmission by a conventional audio core coder. In the preferred embodiment, an EVS audio coder is used to encode the downmix signal, but the invention is not limited to this core coder and can be applied to any audio core coder. The downmix signal consists of different channels, called transport channels: depending on the target bit rate, the signals can be, for example, a B-format signal, a stereo pair, or four coefficient signals constituting a mono downmix. The encoded spatial parameters and the encoded audio bitstream are multiplexed before being transmitted over a communication channel.

[0029] In the decoder, the transport channels are decoded by the core decoder and the DirAC metadata is first decoded before being carried by the decoded transport channels to the DirAC synthesis. The DirAC synthesis uses the decoded metadata to control the reproduction of the direct sound stream and its mixing with the diffuse sound stream. The reproduced sound field can be reproduced with any speaker layout or generated in any order in Ambisonics format (HOA / FOA).

[0030] DirAC parameter estimation: For each frequency band, the sound arrival direction is estimated along with the sound diffusion. From the time-frequency analysis of TIFF0007828378000001.tif752, the pressure and velocity vectors can be determined as follows: TIFF0007828378000002.tif738TIFF0007828378000003.tif790

[0031] where i is the index of the input, TIFF0007828378000004.tif73 and TIFF0007828378000005.tif73 is the time and frequency index of the time-frequency tile, TIFF0007828378000006.tif716 represents a Cartesian unit vector. TIFF0007828378000007.tif714 and TIFF0007828378000008.tif714 is used to calculate DirAC parameters, i.e. DOA and diffusivity, by calculating the intensity vector: TIFF0007828378000009.tif1357, where: TIFF0007828378000010.tif76 shows the complex conjugate. The diffuseness of the composite sound field is given by: TIFF0007828378000011.tif1351 where, TIFF0007828378000012.tif78 shows the time averaging operator, TIFF0007828378000013.tif73 shows the speed of sound, TIFF0007828378000014.tif714 shows the sound field energy given by: TIFF0007828378000015.tif1378The diffuseness of a sound field is defined as the ratio of sound intensity to energy density, which has a value between 0 and 1. The direction of arrival (DOA) is a unit vector defined as Represented by TIFF0007828378000016.tif731. TIFF0007828378000017.tif1358

[0032] The direction of arrival can be determined by energy analysis of the B-format input and defined as the opposite direction of the intensity vector. The direction is defined in Cartesian coordinates, but can be easily converted to spherical coordinates defined by unit radius, azimuth, and elevation.

[0033] For transmission, the parameters need to be transmitted to the receiver side via a bitstream. For robust transmission over networks with limited capacity, a low-bitrate bitstream is preferred, which can be achieved by designing an efficient coding scheme for the DirAC parameters. This can use techniques such as frequency band grouping by averaging parameters over different frequency bands and / or time units, prediction, quantization, and entropy coding. At the decoder, if no errors occur in the network, the transmitted parameters can be decoded for each time / frequency unit (k, n). However, if network conditions are not sufficient to ensure proper packet transmission, packets may be lost during transmission. The present invention aims to provide a solution for the latter case.

[0034] Originally, DirAC was designed to process B-format recorded signals, also known as first-order Ambisonics signals. However, the analysis can be easily extended to any microphone array combining omnidirectional or directional microphones. In this case, the essence of the DirAC parameters remains unchanged, so the present invention remains relevant.

[0035] Furthermore, DirAC parameters, also known as metadata, can be calculated directly during microphone signal processing before being conveyed to the spatial audio coder. A spatial coding system based on DirAC is then directly fed with metadata and spatial audio parameters equivalent or similar to the DirAC parameters in the form of an audio waveform of the downmix signal. The DoA and diffuseness can be easily derived for each parameter band from the input metadata. Such an input format is sometimes called the MASA (Metadata-Assisted Spatial Audio) format. MASA allows the system to ignore the specificity of the microphone arrays and their shape factors, which are necessary to calculate the spatial parameters. These are derived outside the spatial audio coding system using processing specific to the device incorporating the microphones.

[0036] Embodiments of the present invention may use spatial coding systems such as those shown in Figure 2, where a DirAC-based spatial audio encoder and decoder are shown. The embodiment is described with respect to Figures 3a and 3b, and extensions to the DirAC model are described earlier.

[0037] The DirAC model can also be extended, according to an embodiment, by allowing different directional components to have the same time / frequency tile. It can be extended in two main ways:

[0038] The first extension consists of transmitting two or more DoAs per T / F tile, and each DoA must be associated with an energy or energy ratio. For example, the first DoA may be the energy ratio between the energy of the directional component and the energy of the entire audio scene. TIFF0007828378000018.tif138 can be associated with: TIFF0007828378000019.tif1349

[0039] where: TIFF0007828378000020.tif715 is the intensity vector associated with the lth direction. If L DoAs are transmitted along with their L energy ratios, the diffuseness can be estimated from the L energy ratios as follows: TIFF0007828378000021.tif1953

[0040] The spatial parameters transmitted in the bitstream may be L directions along with L energy ratios, or these latest parameters can also be converted into L-1 energy ratios+spreadiness parameters. TIFF0007828378000022.tif1931

[0041] The second extension consists of dividing the 2D or 3D space into non-overlapping sectors and transmitting a set of DirAC parameters (DoA + spreading factor per sector) for each sector. We next describe higher-order DirAC, which was introduced in [5]. Both extensions can in fact be combined and the present invention relates to both extensions.

[0042] 3a and 3b show an embodiment of the present invention, with FIG. 3a showing an approach focusing on the basic concept / method 100 used and the apparatus 50 used being shown by FIG. 3b. FIG. 3 a shows a method 100 including basic steps 110 , 120 and 130 .

[0043] The first steps 110 and 120 are equivalent to each other, i.e., they refer to the reception of sets of spatial audio parameters. In the first step 110, the first set is received, and in the second step 120, the second set is received. Furthermore, there may be additional reception steps (not shown). Note that the first set may refer to a first time point / first frame, and the second set may refer to a second (subsequent) time point / second (subsequent) frame, etc. As mentioned above, the first and second sets may include diffusion information (Ψ) and / or directional information (azimuth and elevation angles). This information can be encoded using a spatial metadata encoder. Now, assume that the second information set is lost or damaged during transmission. In this case, the second set is replaced by the first set. This enables packet loss concealment of spatial audio parameters, such as DirAC parameters.

[0044] In the case of packet loss, the erased DirAC parameters of the lost frames need to be restored to limit the impact on quality. This can be achieved by synthetically generating the missing parameters by taking into account previously received parameters. An unstable spatial image can be perceived as unpleasant and artifactual, while a strictly constant spatial image can be perceived as unnatural.

[0045] The method 100 described by Fig. 3a can be performed by an entity 50 as shown by Fig. 3b. The apparatus 50 for loss concealment comprises an interface 52 and a processor 54. Via the interface, a set of spatial audio parameters Ψ1, azi1, ele1, Ψ2, azi2, ele2, Ψn, azin, ele can be received. The processor 54 analyzes the received set and, in case of a lost or damaged set, replaces the lost or damaged set, for example, by a previously received set or an equivalent set. Different strategies can be used, which will be described later.

[0046] Hold Strategy: It is generally safe to assume that the spatial image should be relatively stable over time, which can be translated into DirAC parameters, i.e., direction of arrival and spread, which do not change much between frames. For this reason, a simple but effective approach is to hold the parameters of the last well-received frame of a frame that was lost during transmission.

[0047] Direction estimation: Alternatively, it can be envisaged to estimate the trajectory of a sound event in an audio scene and then attempt to extrapolate the estimated trajectory. This is particularly relevant when the sound event is well localized in space as a point source, which is reflected in the DirAC model by low diffusivity. The estimated trajectory can be calculated from past direction observations, and a curve can be fitted between these points, and either interpolation or smoothing can be developed. Regression analysis can also be used. The extrapolation is then performed by evaluating the fitted curve beyond the range of the observed data.

[0048] In DirAC, directions are often represented, quantized, and encoded in polar coordinates. However, it is usually more convenient to process directions and then trajectories in Cartesian coordinates to avoid processing modulo 2π arithmetic.

[0049] Directional dithering: As sound events become more diffuse, direction becomes less meaningful and can be thought of as a realization of a stochastic process. Dithering can then help make the rendered sound field more natural and more pleasant by injecting random noise in the previous direction before using it for lost frames. The injected noise and its variance can be a function of the diffuseness.

[0050] Standard DirAC audio scene analysis can be used to investigate the effect of diffuseness on the accuracy and significance of the model's direction. Using an artificial B-format signal, where a direct diffuse energy ratio (DDR) is imposed between the plane wave and diffuse field components, the resulting DirAC parameters and their accuracy can be analyzed. Theoretical diffusion TIFF0007828378000023.tif74 is the direct diffusion energy ratio (DDR) It is a function of TIFF0007828378000024.tif75 and is expressed as: TIFF0007828378000025.tif1981 where, TIFF0007828378000026.tif78 and TIFF0007828378000027.tif711 are plane wave and diffusivity, respectively. TIFF0007828378000028.tif73 is the DDR expressed in dB scale.

[0051] Of course, one or a combination of the three strategies discussed can be used. The strategy to be used is selected by the processor 54 depending on the received spatial audio parameter set. To this end, according to an embodiment, the audio parameters can be analyzed to allow the application of different strategies according to the characteristics of the audio scene, more particularly according to the diffuseness.

[0052] This means that, according to an embodiment, the processor 54 is configured to provide packet loss concealment for spatial parametric audio by using previously successfully received directivity information and dithering. According to a further embodiment, the dithering is a function of an estimated diffuseness or energy ratio between directional and omnidirectional components of the audio scene. According to an embodiment, the dithering is a function of the measured tonality of the transmitted downmix signal. Thus, the analyzer performs the analysis based on the estimated diffuseness, energy ratio and / or tonality.

[0053] In Figures 3a and 3b, the measured diffusivity is given as a function of DDR by simulating a diffuse field with N=466 uncorrelated pink noises evenly spaced on a sphere and a plane wave, with independent pink noises placed at 0° azimuth and 0° elevation. The measured diffusivity in the DirAC analysis was confirmed to be a good estimate of the theoretical diffusivity when the observation window length W is sufficiently large. This means that the diffusivity has long-term characteristics, which confirms that the parameter in the case of packet loss can be well predicted by simply retaining previously successfully received values.

[0054] Meanwhile, the estimation of directional parameters can also be evaluated as a function of the true diffuseness, as reported in Figure 4. It can be shown that the elevation and azimuth angles of the estimated plane wave position deviate from the ground truth position (0° azimuth and 0° elevation) with a standard deviation that increases with diffuseness. When the diffuseness is 1, the standard deviation is approximately 90° for an azimuth angle defined between 0° and 360°, corresponding to a completely random angle with a uniform distribution. In other words, the azimuth angle is meaningless. A similar observation can be made for the elevation angle. In general, the accuracy of the estimated direction and its significance decrease with diffuseness. Furthermore, the direction in DirAC is expected to fluctuate over time and deviate from its expected value using the variance function of diffuseness. This natural variance is part of the DirAC model and is essential for faithful reproduction of audio scenes. Indeed, rendering the directional component of DirAC in a constant direction, even with high diffuseness, actually produces a point source that is perceived as wider.

[0055] For the reasons made clear above, we propose to apply dithering in the direction of the top of the hold strategy. The amplitude of the dithering is made a function of the spread and can for example follow the model depicted in Figure 4. Two models can be derived for the elevation angle and the elevation measurement angle, where the standard deviation is expressed as: TIFF0007828378000029.tif738TIFF0007828378000030.tif741The pseudocode for DirAC parameter hiding can be done as follows: for k in frame_start:frame_end { if(bad_frame_indicator[k]) { for band in band_start:band_end { diff_index = diffuseness_index[k-1][band]; diffuseness[k][band] = unquantize_diffuseness(diff_index); azimuth_index[k][b] = azimuth_index[k-1][b]; azimuth[k][b] = unquantize_azimuth(azimuth_index[k][b]) azimuth[k][b] = azimuth[k][b] + random() * dithering_azi_scale[diff_index] elevation_index[k][b] = elevation_index[k-1][b]; elevation[k][b] = unquantize_elevation(elevation_index[k][b]) elevation[k][b] = elevation[k][b] + random() * dithering_ele_scale[diff_index] } else { for band in band_start:band_end { diffuseness_index[k][b] = read_diffusess_index() azimuth_index[k][b] = read_azimuth _index() elevation_index[k][b] = read_elevation_index() diffuseness[k][b] = unquantize_diffuseness(diffuseness_index[k][b]) azimuth[k][b] = unquantize_azimuth(azimuth_index[k][b]) elevation[k][b] = unquantize_elevation(elevation_index[k][b]) } output_frame[k] = Dirac_synthesis(diffuseness[k][b], azimuth[k][b], elevation[k][b]) }

[0056] Here, bad_frame_indicator[k] is a flag indicating whether the frame with index k was received well. For a good frame, the DirAC parameters are read, decoded, and not quantized for each parameter band corresponding to a given frequency range. For a bad frame, the spread is retained directly from the last successfully received frame in the same parameter band, while the azimuth and elevation angles are derived from dequantizing the last successfully received index by injecting random values ​​scaled by a coefficient function of the spread index. The function random() outputs random values ​​according to a given distribution. The random process can, for example, follow a standard normal distribution with mean and unit variance 0. Alternatively, it can follow a uniform distribution between -1 and 1 or a triangular probability density, for example, using the following pseudocode: random() { rand_val = uniform_random(); if( rand_val <= 0.0f ) { return 0.5f * sqrt(rand_val + 1.0f) - 0.5f; } else { return 0.5f - 0.5f * sqrt(1.0f - rand_val); } }

[0057] The dithering scale is a function of the diffuseness index inherited from the last successfully received frame in the same parameter band and can be derived from the model estimated from Figure 4. For example, if the diffuseness is coded with 8 indices, they can correspond to the following table: dithering_azi_scale[8] = { 6.716062e-01f, 1.011837e+00f, 1.799065e+00f, 2.824915e+00f, 4.800879e+00f, 9.206031e+00f, 1.469832e+01f, 2.566224e+01f }; dithering_ele_scale[8] = { 6.716062e-01f, 1.011804e+00f, 1.796875e+00f, 2.804382e+00f, 4.623130e+00f, 7.802667e+00f, 1.045446e+01f, 1.379538e+01f };

[0058] Furthermore, the dithering strength can also be manipulated depending on the nature of the downmix signal. Indeed, highly tonal signals tend to be perceived as more localized sources than non-tonal signals. Therefore, dithering can then be adjusted in function of the tonality of the conveyed downmix by reducing the dithering effect of tonal items. Tonality can be measured, for example, in the time domain by calculating the long-term prediction gain, or in the frequency domain by measuring the spectral flatness.

[0059] With respect to Figures 6a and 6b, further embodiments are described which refer to a method for decoding DirAC-encoded audio scenes (see Figure 6a, method 200) and a decoder 17 for DirAC-encoded audio scenes (see Figure 6b).

[0060] Fig. 6a shows a new method 200 which comprises steps 110, 120 and 130 of method 100 and an additional step 210 of decoding. The decoding step allows decoding of a DirAC-encoded audio scene comprising a downmix (not shown) by use of a first set of spatial audio parameters and a second set of spatial audio parameters, where the replaced second set is used and output by step 130. This concept is used by an apparatus 17 shown in Fig. 6b. Fig. 6b shows a decoder 70 which comprises a processor for loss concealment of the spatial audio parameters 15 and a DirAC decoder 72. The DirAC decoder 72, or more precisely the processor of the DirAC decoder 72, receives the downmix signal and the set of spatial audio parameters, for example directly from the interface 52 and / or processed by the processor 52 according to the above-described techniques.

[0061] While some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with a block or apparatus corresponding to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or function of a corresponding apparatus. Some or all of the method steps can be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, some one or more of the most important method steps can be performed by such an apparatus.

[0062] The encoded audio signal of the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0063] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, flash memories, etc., on which electronically readable control signals are stored and which cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage media can be computer-readable.

[0064] Some embodiments of the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0065] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code may for example be stored on a machine-readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0066] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0067] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) comprising recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.

[0068] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transferred via a data communication connection, such as, for example, the Internet.

[0069] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0070] Further embodiments according to the present invention comprise an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0071] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0072] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0073] References [1] V. Pulkki, MV. Laitinen, J. Vilkamo, J. Ahonen, T. Lokki, and T. Pihlajamaeki, “Directional audio coding - perception-based reproduction of spatial sound”, International Workshop on the Principles and Application on Spatial Hearing, Nov. 2009, Zao; Miyagi, Japan.

[0074] [2] V. Pulkki, “Virtual source positioning using vector base amplitude panning”, J. Audio Eng. Soc., 45(6):456-466, June 1997.

[0075] [3] J. Ahonen and V. Pulkki, “Diffuseness estimation using temporal variation of intensity vectors”, in Workshop on Applications of Signal Processing to Audio and Acoustics WASPAA, Mohonk Mountain House, New Paltz, 2009.

[0076] [4] T. Hirvonen, J. Ahonen, and V. Pulkki, “Perceptual compression methods for metadata in Directional Audio Coding applied to audiovisual teleconference”, AES 126th Convention 2009, May 7-10, Munich, Germany.

[0077] [5] A. Politis, J. Vilkamo and V. Pulkki, “Sector-Based Parametric Sound Field Reproduction in the Spherical Harmonic Domain,“ in IEEE Journal of Selected Topics in Signal Processing, vol. 9, no. 5, pp. 852-866, Aug. 2015.

Claims

1. A method (100) for loss concealment of spatial audio parameters, said spatial audio parameters including at least direction of arrival information, said method comprising the computer-implemented steps of: receiving (110) a first set of spatial audio parameters including at least first direction of arrival information (azi1, ele1); receiving (120) a second set of spatial audio parameters including at least second direction of arrival information (azi2, ele2); replacing the second set of direction of arrival information (azi2, ele2) with replacement direction of arrival information derived from the first direction of arrival information (azi1, ele1) if at least the second direction of arrival information (azi2, ele2) or part of the second direction of arrival information (azi2, ele2) is lost or damaged; The method (100), wherein the replacing step is performed based on a hold strategy, a directional extrapolation strategy, or a directional dithering strategy, and the strategy used is selected depending on the received first set of spatial audio parameters.

2. 2. The method (100) of claim 1, wherein the first set (1st set) and second set (2nd set) of spatial audio parameters comprise first and second diffuseness information (Ψ1, Ψ2), respectively.

3. 3. The method (100) of claim 2, wherein the first or second diffusion information (Ψ1, Ψ2) is derived from at least one energy ratio relating to at least one direction of arrival information.

4. 4. The method (100) according to claim 2 or 3, further comprising replacing the second diffusion information (Ψ2) of a second set (second set) by replacement diffusion information derived from the first diffusion information (Ψ1).

5. The method (100) according to any one of claims 1 to 4, wherein the replacement direction of arrival information is in accordance with the first direction of arrival information (azi1, ele1).

6. the step of substituting comprises the step of dithering the replacement direction of arrival information; and / or 5. The method (100) of claim 2, wherein the replacing step comprises injecting random noise into the first direction of arrival information (azi1, ele1) to obtain the replaced direction of arrival information.

7. The method (100) of claim 6, wherein the injecting step is performed when the first or second diffusion information (Ψ1, Ψ2) indicates a high degree of diffusion and / or when the first or second diffusion information (Ψ1, Ψ2) exceeds a predetermined threshold value of the diffusion information.

8. 8. The method (100) of claim 7, wherein the diffuseness information comprises or is based on a ratio between directional and non-directional components of an audio scene described by the first set (first set) and / or the second set (second set) of spatial audio parameters.

9. the injected random noise depends on the first and / or second diffusion information (Ψ1, Ψ2); and / or 9. The method (100) according to any one of claims 6 to 8, wherein the injected random noise is scaled by a factor that depends on the first and / or second diffusion information (Ψ1, Ψ2).

10. analysing the tonality of an audio scene described by said first set (first set) and / or second set (second set) of spatial audio parameters or analysing the tonality of a transmitted downmix belonging to said first set (first set) and / or second set (second set) of spatial audio parameters to obtain a tonality value describing said tonality, The method (100) of any one of claims 6 to 9, wherein the injected random noise depends on the tonality value.

11. 11. The method (100) of claim 10, wherein the random noise is scaled down by a factor that decreases with the inverse of the tonality value or if the tonality increases.

12. 12. The method (100) of any one of claims 2 to 4 and 6 to 11, wherein the method (100) comprises a step of extrapolating the first direction of arrival information (azi1, ele1) to obtain the replacement direction of arrival information.

13. The method (100) of claim 12, wherein the extrapolating is based on one or more additional direction of arrival information belonging to one or more sets of spatial audio parameters.

14. 14. The method (100) according to claim 12 or 13, wherein the extrapolation is performed if the first and / or second diffusion information (Ψ1, Ψ2) indicates a low degree of diffusion or if the first and / or second diffusion information (Ψ1, Ψ2) is below a predetermined threshold value of diffusion information.

15. the first set of spatial audio parameters (first set) belongs to a first time point and / or a first frame, and the second set of spatial audio parameters (second set) belongs to a second time point and / or a second frame, or 15. The method (100) according to any one of claims 1 to 14, wherein the first set of spatial audio parameters belongs to a first time point, and the second time point is after the first time point or the second frame is after the first frame.

16. the first set of spatial audio parameters comprises a first subset of spatial audio parameters for a first frequency band and a second subset of spatial audio parameters for a second frequency band; and / or 16. The method (100) of claim 1, wherein the second set of spatial audio parameters comprises a different first subset of spatial audio parameters for the first frequency band and a different second subset of spatial audio parameters for the second frequency band.

17. A method (200) for decoding DirAC encoded audio scenes, comprising: The computer-implemented steps include: decoding said DirAC encoded audio scenes comprising a downmix, a first set of spatial audio parameters and a second set of spatial audio parameters; Executing the method (100) according to one of the steps of the method (100) of any one of claims 1 to 16.

18. When executed on a computer, claims 1 to 17 A computer-readable digital storage medium having stored thereon a computer program having program code for performing the method (100, 200) according to any one of claims 1 to 4.

19. A loss concealment device (50) for loss concealment of spatial audio parameters, said spatial audio parameters including at least direction of arrival information, said device comprising: a receiver (52) for receiving (110) a first set of spatial audio parameters comprising first direction of arrival information (azi1, ele1) and for receiving (120) a second set of spatial audio parameters comprising second direction of arrival information (azi2, ele2); a processor (54) for replacing the second set of second direction of arrival information (azi2, ele2) by replacement direction of arrival information derived from the first direction of arrival information (azi1, ele1) if at least the second direction of arrival information (azi2, ele2) or a part of the second direction of arrival information (azi2, ele2) is lost or damaged; The replacement is performed based on a hold strategy, a directional extrapolation strategy, or a directional dithering strategy, the strategy to be used being selected by the processor (54) in response to the received first set of spatial audio parameters.

20. A decoder (70) for DirAC encoded audio scenes, comprising a loss concealment device according to claim 19.

Citation Information

Patent Citations

  • robust decoder

    JP2008542838A

  • Apparatus and method for resolving ambiguity from estimated directions of arrival

    JP2013536477A

  • Apparatus and method for providing enhanced and guided downmixing capabilities for 3D audio

    JP2015532062A

  • Packet loss compensation device, packet loss compensation method, and audio processing system

    JP2016528535A