Packet loss concealment for dirac-based spatial audio coding
By replacing the lost spatial audio parameters and utilizing the directional jitter and diffusion strategies in DirAC coding technology, the audio quality degradation caused by packet loss in DirAC coding is solved, achieving a more natural and pleasant audio restoration effect.
Patent Information
- Application Number
- CN202080043012.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-12
- Filing Date
- 2020-06-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-06-05
AI Technical Summary
The existing DirAC coding technology lacks an effective method for hiding packet loss, which leads to a serious deterioration in audio quality when packets are lost during network transmission, producing pseudo-sounds such as clicking and thumping sounds, and affecting the perceived quality.
By receiving and replacing lost spatial audio parameters, the lost direction of arrival information is recovered using previously well-received directional information and a jitter strategy. In combination with diffusion information, the stability of the audio scene is maintained, and diffusion and directional jitter techniques are used to improve audio quality.
It effectively restored the lost spatial audio parameters, reduced artifacts during transmission, improved the perceived quality of audio signals, and ensured the naturalness and enjoyment of the audio scene.
Smart Images

Figure CN114097029B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a method for loss concealment of spatial audio parameters, a method for decoding DirAC-coded audio scenes, and a corresponding computer program. Other embodiments relate to a loss concealment apparatus for spatial audio parameter loss concealment and a decoder including a packet loss concealment apparatus. Preferred embodiments describe a concept / method for compensating for quality degradation caused by lost and corrupted frames or packets during the transmission of an audio scene, wherein, for said audio scene, a spatial image is parametrically encoded using the Directional Audio Coding (DirAC) paradigm. Background Technology
[0002] Voice and audio communications can suffer from various quality issues due to packet loss during transmission. In fact, adverse conditions in the network, such as bit errors and jitter, can cause some packets to be lost. These losses result in severe artifacts such as clicks, thuds, or unwanted silence, which significantly degrade the perceived quality of the reconstructed speech or audio signal at the receiver. To combat the adverse effects of packet loss, packet loss concealment (PLC) algorithms have been proposed in traditional speech and audio coding schemes. These algorithms typically operate at the receiver by generating a synthesized audio signal to hide missing data in the received bitstream.
[0003] DirAC is a perceptually excited spatial audio processing technique that compactly and efficiently represents a sound field using a set of spatial parameters and a downmixed signal. The downmixed signal can be a mono, stereo, or multichannel signal in an audio format such as A or B, also known as first-order stereo reverberation (FAO). The downmixed signal is supplemented by spatial DirAC parameters, which describe the audio scene in terms of direction of arrival (DOA) and diffusion per time / frequency unit. In storage, streaming, or communication applications, the downmixed signal is encoded using a conventional core encoder (e.g., EVS or its stereo / multichannel extension or any other mono / stereo / multichannel codec) designed to preserve the audio waveform of each channel. The core encoder can be built around a transform-based coding scheme or speech coding scheme that operates in the time domain, such as CELP. The core encoder can then integrate existing error recovery tools such as the Packet Loss Concealment (PLC) algorithm.
[0004] On the other hand, there is no existing solution for protecting DirAC space parameters. Therefore, an improved method is needed. Summary of the Invention
[0005] The purpose of this invention is to provide a concept of loss hiding in the context of DirAC.
[0006] This objective is addressed by the subject matter of the independent claims.
[0007] Embodiments of the present invention provide a method for concealing the loss of spatial audio parameters, wherein the spatial audio parameters include at least direction of arrival information. The method includes the following steps:
[0008] • Receive a first set of spatial audio parameters, including first direction of arrival information and first spread information;
[0009] • Receive a second set of spatial audio parameters, including second direction of arrival information and second spread information; and
[0010] If at least the second arrival direction information or a portion of the second arrival direction information is lost, the second arrival direction information of the second group is replaced by replacement arrival direction information derived from the first arrival direction information.
[0011] The embodiments of the present invention are based on the discovery that, in the event of lost or corrupted arrival information, the lost / corrupted arrival information can be replaced by arrival information derived from another available arrival information. For example, if second arrival information is lost, it can be replaced by first arrival information. In other words, this means that the embodiments provide a packet loss concealment scheme for spatial parameter audio, where, in the event of transmission loss, directional information is recovered using previously well-received directional information and jitter. Therefore, the embodiments enable resistance to packet loss during the transmission of spatial audio sound encoded with direct parameters.
[0012] A further embodiment provides a method wherein the first set of spatial audio parameters and the second set of spatial audio parameters respectively include first diffusion information and second diffusion information. In this case, the strategy may be as follows: According to an embodiment, the first diffusion information or the second diffusion information is derived from at least one energy ratio associated with at least one direction of arrival information. According to an embodiment, the method further includes replacing the second diffusion information of the second set with alternative diffusion information derived from the first diffusion information. This is part of a so-called preservation strategy based on the assumption that diffusion does not change much between frames. For this reason, a simple but effective approach is to preserve the parameters of the last well-received frame of a frame lost during transmission. Another part of this overall strategy is replacing the second arrival information with the first arrival information, however, this has already been discussed in the context of the basic embodiment. It is generally safe to assume that the spatial image must be relatively stable over time, and this can be translated for DirAC parameters (i.e., the direction of arrival, which may also not change much between frames).
[0013] According to a further embodiment, the replacement direction of arrival information conforms to the first direction of arrival information. In this case, a strategy called directional jitter can be used. Here, according to an embodiment, the replacement step may include the step of jittering the replacement direction of arrival information. Alternatively or additionally, the replacement step may include injection when noise is the first direction of arrival information to obtain the replacement direction of arrival information. Jitter can then help make the presented sound field more natural and pleasant by injecting random noise into the previous direction when using the previous direction for the frame. According to an embodiment, the injection step is preferably performed if the first diffusion information or the second diffusion information indicates high diffusion. Alternatively, the injection step may be performed if the first diffusion information or the second diffusion information is higher than a predetermined threshold indicating high diffusion. According to other embodiments, the diffusion information includes more space regarding the ratio between directional and non-directional components of the audio scene described by the first set of spatial audio parameters and / or the second set of spatial audio parameters. According to an embodiment, the random noise to be injected depends on the first diffusion information and the second diffusion information. Alternatively, the random noise to be injected may be scaled according to a factor depending on the first diffusion information and / or the second diffusion information. Therefore, according to an embodiment, the method may further include the steps of: analyzing the tonality of an audio scene described by the first set of spatial audio parameters and / or the second set of spatial audio parameters, or analyzing the tonality of the transmitted downmixed audio belonging to the first spatial audio parameters and / or the second spatial audio parameters to obtain a tonality value describing the tonality. The random noise to be injected thus depends on the tonality value. According to an embodiment, a proportional reduction may be performed according to a factor that decreases along with the reciprocal of the tonality value, or a proportional reduction may be performed if the tonality increases.
[0014] According to a further strategy, a method including the following steps can be used: extrapolating the first direction of arrival information to obtain the alternative direction of arrival information. According to this method, it is conceivable to estimate the directionality of sound events in the audio scene by extrapolation. This is particularly useful if the sound events are well localized in space and act as point sources (with a direct model of low diffusion). According to an embodiment, the extrapolation is based on one or more additional direction of arrival information belonging to one or more sets of spatial audio parameters. According to an embodiment, the extrapolation is performed if the first diffusion information and / or the second diffusion information indicates low diffusion or if the first diffusion information and / or the second diffusion information is below a predetermined threshold used for diffusion information.
[0015] According to an embodiment, the first set of spatial audio parameters belongs to a first time point and / or a first frame, and the second set of spatial audio parameters both belong to a second time point or a second frame. Alternatively, the second time point is after the first time point, or the second frame is after the first frame. Returning to the embodiment where most sets of spatial audio parameters are used for extrapolation, it is obvious that it is preferable to use more sets of spatial audio parameters belonging to, for example, multiple time points / frames that are, after each other.
[0016] According to a further embodiment, the first set of spatial audio parameters includes a first subset of spatial audio parameters for a first frequency band and a second subset of spatial audio parameters for a second frequency band. The second set of spatial audio parameters includes another first subset of spatial audio parameters for the first frequency band and another second subset of spatial audio parameters for the second frequency band.
[0017] Another embodiment provides a method for decoding a DirAC-coded audio scene, comprising the steps of: decoding a DirAC-coded audio scene including downmixing, a first set of spatial audio parameters, and a second set of spatial audio parameters. This method further includes the loss-of-concealment method steps as discussed above.
[0018] According to embodiments, the methods described above can be implemented by a computer. Therefore, embodiments relate to a computer-readable storage medium storing a computer program having program code for performing a method according to one of the foregoing technical solutions when executed on a computer.
[0019] Another embodiment relates to a loss-concealing device for concealing the loss of spatial audio parameters (including at least direction-of-arrival information). The device includes a receiver and a processor. The receiver is configured to receive a first set of spatial audio parameters and a second set of spatial audio parameters (see above). The processor is configured to replace the second set of second direction-of-arrival information with replacement direction-of-arrival information derived from the first direction-of-arrival information in the event that the second direction-of-arrival information is lost or corrupted. Another embodiment relates to a decoder for DirAC-encoded audio scenes, the decoder including the loss-concealing device. Attached Figure Description
[0020] Embodiments of the invention will then be described with reference to the accompanying drawings, in which:
[0021] Figure 1a , Figure 1b A schematic block diagram illustrating DirAC analysis and synthesis is shown;
[0022] Figure 2 A schematic, detailed block diagram illustrating DirAC analysis and synthesis in a low-bitrate 3D audio encoder is shown.
[0023] Figure 3a A schematic flowchart of a method for loss concealment according to a basic embodiment is shown;
[0024] Figure 3b An illustrative loss-hiding device according to a basic embodiment is shown;
[0025] Figure 4a , Figure 4b Showing DDR ( Figure 4a Window size W = 16, Figure 4b A schematic diagram of the diffusion measurement function with a window size of W = 512 is provided to illustrate the embodiment;
[0026] Figure 5 A schematic diagram showing the measurement directions (azimuth and elevation) in the diffusion function is provided to illustrate the embodiment;
[0027] Figure 6a A schematic flowchart illustrating a method for decoding a DirAC-encoded audio scene according to an embodiment is shown; and
[0028] Figure 6b A schematic block diagram of a decoder for a DirAC encoded audio scene is shown according to an embodiment. Detailed Implementation
[0029] Hereinafter, embodiments of the invention will be discussed with reference to the accompanying drawings, wherein the same reference numerals are provided for objects / elements having the same or similar functions so that their descriptions are mutually applicable and interchangeable. Before discussing the embodiments of the invention in detail, an introduction to DirAC is given.
[0030] Detailed description of the embodiments
[0031] Introduction to DirAC: DirAC is a perceptually excited spatial sound reproduction. It assumes that at any given moment and for a critical frequency band, the spatial resolution of the auditory system is limited to decoding a cue for direction and another cue for interaural coherence. Based on these assumptions, DirAC represents spatial sound within a frequency band by cross-attenuating two streams: a non-directional diffuse stream and a directional non-diffuse stream. DirAC processing is performed in two stages:
[0032] The first stage is as follows Figure 1a The analysis described herein, and the second stage is as follows: Figure 1b The synthesis described.
[0033] Figure 1aThe diagram illustrates an analysis phase 10 including one or more bandpass filters 12a-n receiving microphone signals W, X, Y, and Z; an energy analysis phase 14e; and an intensity analysis phase 14i. Dispersion Ψ (see reference numeral 16d) can be determined using time configuration. Dispersion Ψ is determined based on energy analysis 14c and intensity analysis 14i. Direction 16e can be determined based on intensity and analysis 14i. The direction determination results in azimuth and elevation angles. Ψ, azi, and ele are output as metadata. This metadata is generated by... Figure 1b The synthetic entity 20 shown is used.
[0034] Figure 1b The synthesizing entity 20 shown includes a first stream 22a and a second stream 22b. The first stream includes multiple bandpass filters 12a-n and a computational entity for a virtual microphone 24. The second stream 22b includes components for processing metadata, namely 26 for the diffusion parameter and 27 for the direction parameter. Furthermore, a decorrelation unit 28 is used in the synthesizing stage 20, wherein this decorrelation entity 28 receives data from both streams 22a and 22b. The output of the decorrelation unit 28 can be fed to a loudspeaker 29.
[0035] During the DirAC analysis phase, a first-order coincident microphone in B format is considered as input, and the spread and direction of sound arrival are analyzed in the frequency domain.
[0036] In the DirAC synthesis stage, the sound is split into two streams, a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude translation, which can be achieved by using vector-based amplitude translation (VBAP)[2]. The diffuse stream is responsible for creating a sense of envelopment and is generated by sending mutually decorrelated signals to the amplifier.
[0037] DirAC parameters (hereinafter also referred to as spatial metadata or DirAC metadata) consist of tuples of diffusivity and orientation. Orientation can be represented in spherical coordinates by two angles (azimuth and elevation), while diffusivity is a scalar factor between 0 and 1.
[0038] In the following text, relative to Figure 2 Discuss the system of DirAC spatial audio coding. Figure 2 The two-stage DirAC analysis 10' and DirAC synthesis 20' are illustrated. Here, the DirAC analysis includes filter bank analysis 12, direction estimator 16i, and spread estimator 16d. Both 16i and 16d output spread / direction data as spatial metadata. This data can be encoded using encoder 17. Direct analysis 20' includes spatial metadata decoder 21, output synthesis 23, and filter combination 12, enabling the signal to be output to loudspeaker FOA / HOA.
[0039] Alongside the discussed direct analysis stage 10' and direct synthesis stage 20' for processing spatial metadata, an EVS encoder / decoder is used. On the analysis side, beamforming / signal selection is performed based on the input signal in B format (see beamforming / signal selection entity 15). The signal is then EVS encoded (see reference numeral 17). On the synthesis side (see reference numeral 20'), an EVS decoder 25 is used. This EVS decoder outputs the signal to a filter bank analysis 12, which outputs its signal to an output synthesizer 23.
[0040] Since the structures of 10' / 20' directly analyzed / synthesized have now been discussed, functionality will be discussed in detail.
[0041] Encoder analysis 10' typically analyzes spatial audio scenes in B format. Alternatively, DirAC analysis can be adapted to analyze different audio formats, such as audio objects or multichannel signals, or any combination of spatial audio formats. DirAC analysis extracts a parametric representation from the input audio scene. The direction of arrival (DOA) and spread measured at each time-frequency unit form these parameters. Following the DirAC analysis is a spatial metadata encoder, which quantizes and encodes the DirAC parameters to obtain a low-bit-rate parametric representation.
[0042] Along with the aforementioned parameters, downmixed signals derived from different sources or audio input signals are also encoded for transmission via a conventional audio core encoder. In a preferred embodiment, the EVS audio encoder is preferably used to encode the downmixed signal, but the invention is not limited to this core encoder and can be applied to any audio core encoder. The downmixed signal consists of different channels referred to as delivery channels: the signals can be, for example, four coefficient signals constituting a B-format signal, stereo pairs depending on the target bit rate, or mono downmixing. The encoded spatial parameters and the encoded audio bitstream are multiplexed before transmission via the communication channel.
[0043] In the decoder, the delivery channels are decoded via the core decoder, while DirAC metadata is first decoded before being sent to the DirAC synthesizer using the decoded delivery channels. The DirAC synthesizer uses the decoded metadata to control the reproduction of the direct sound stream and its mixing with the diffuse sound stream. The reproduced sound field can be reproduced on any amplifier layout or generated in any order in a stereo reverb format (HOA / FOA).
[0044] DirAC parameter estimation: Estimates the direction of arrival and dispersion of sound within each frequency band. This is derived from time-frequency analysis of the input B-format components. i (n),x i (n),yi (n),z i (n), the pressure and velocity vectors can be determined as follows:
[0045] P i (n,k)=W i (n,k)
[0046] U i (n,k)=X i (n,k)e x +Y i (n,k)e y +Z i (n,k)e z ,
[0047] Where i is the input index, and k and n are the time and frequency indices of the time-frequency data block, and e x e y e z Represents the Cartesian unit vector. P(n,k) and U(n,k) are used to calculate the DirAC parameters, namely DOA and diffusivity, by calculating the intensity vector:
[0048]
[0049] in This represents complex conjugation. The diffusion of the combined sound field is given by the following equation:
[0050]
[0051] Where Ε{.} denotes the time averaging operator, c is the speed of sound, and E(k,n) is the sound field energy, given by the following formula:
[0052]
[0053] The diffusion of a sound field is defined as the ratio between sound intensity and energy density, with a value between 0 and 1.
[0054] The direction of arrival (DOA) is expressed using the direction of the unit vector (n,k), and is limited to...
[0055]
[0056] The direction of arrival is determined through energy analysis of the B-format input and can be constrained as the relative direction of the intensity vector. The direction is constrained in Cartesian coordinates but can be easily transformed into spherical coordinates defined by unit radius, azimuth, and elevation.
[0057] In transmission scenarios, parameters need to be transmitted to the receiver via a bitstream. For robust transmission over capacity-constrained networks, low bit-rate bitstreams are preferred, which can be achieved by designing an efficient coding scheme for the DirAC parameters. This can be achieved, for example, by using techniques such as band grouping, to average, predict, quantize, and entropy encode the parameters over different frequency bands and / or time units. At the decoder, assuming no errors occur in the network, the transmitted parameters can be decoded for each time / frequency unit (k, n). However, if network conditions are not good enough to ensure proper packet transmission, packets may be lost during transmission. This invention aims to provide a solution to the latter situation.
[0058] Originally, DirAC was intended for processing B-format recorded signals, also known as first-order stereo reverberation signals. However, the analysis can be easily extended to any microphone array combining omnidirectional or directional microphones. In this case, the present invention remains appropriate because the nature of the DirAC parameters remains unchanged.
[0059] Furthermore, DirAC parameters, also known as metadata, can be directly calculated during microphone signal processing before the microphone signal is fed to the spatial audio encoder. The spatial audio parameters, or similar to the DirAC parameters in metadata form, and the audio waveform of the downmixed signal are then directly fed to the DirAC-based spatial coding system. DoA and spread can be easily derived from the input metadata for each parameter range. This type of input format is sometimes called MASA (Metadata-Aided Spatial Audio) format. MASA allows the system to ignore the specificity of the microphone array and its physical dimensions required to calculate the spatial parameters. These will be derived externally to the spatial audio coding system using device-specific processing with microphones.
[0060] Embodiments of the present invention may use, as follows Figure 2 The spatial coding system described herein depicts a DirAC-based spatial audio encoder and decoder. (The remaining text appears to be unrelated and possibly machine-generated.) Figure 3a and Figure 3b Examples are described below, with an extension of the DirAC model discussed beforehand.
[0061] According to embodiments, the DirAC model can also be extended by allowing different directional components to have the same time / frequency data blocks. This can be extended in two main ways:
[0062] The first extension consists of two or more DoA's sent per T / F data block. Each DoA must subsequently be associated with an energy or energy ratio. For example, the l-th DoA could be associated with the energy ratio Γ between the energy of the directional component and the overall audio scene energy. l Related:
[0063]
[0064] Where I l (k,n) is the intensity vector associated with the l-th direction. If L DoA's are transmitted along with their L energy ratios, the diffusion can then be inferred from the L energy ratios as follows:
[0065]
[0066] The spatial parameters transmitted in the potential stream can be L directions along with L energy ratios, or these latest parameters can be converted into L-1 energy ratios plus diffusivity parameters.
[0067]
[0068] The second extension consists of splitting the 2D or 3D space into non-overlapping sectors and transmitting a set of DirAC parameters (DoA + sector-by-sector spread) for each sector. We will then discuss the higher-order DirAC as described in [5].
[0069] Both extensions can actually be combined, and this invention is related to both extensions.
[0070] Figure 3a and Figure 3b The embodiments of the present invention are described, wherein Figure 3a This illustrates a method focused on the basic concept / method 100, wherein the apparatus 50 used is comprised of... Figure 3b As shown.
[0071] Figure 3a The instructions include method 100, which includes basic steps 110, 120, and 130.
[0072] Steps 110 and 120 are equivalent to each other, involving the reception of several sets of spatial audio parameters. In step 110, a first set is received, and in step 120, a second set is received. Further reception steps (not shown) may also exist. It should be noted that the first set may relate to a first time point / first frame, the second set may relate to a second (subsequent) time point / second (subsequent) frame, and so on. As discussed above, the first and second sets may include spread information (Ψ) and / or direction information (azimuth and elevation). This information can be encoded using a spatial metadata encoder. Now, suppose the second set of information is lost or corrupted during transmission. In this case, the second set is replaced by the first set. This implementation is used for packet loss concealment of spatial audio parameters such as DirAC parameters.
[0073] In the event of packet loss, the erased DirAC parameters of the lost frames need to be replaced to limit the impact on quality. This can be achieved by synthesizing the missing parameters incorporating previously received parameters. Unstable spatial images may be perceived as unpleasant and as artifacts, but strictly constant spatial images may be perceived as unnatural.
[0074] like Figure 3a The method 100 discussed can be derived from, for example Figure 3b The entity 50 shown performs the operation. The device 50 for loss concealment includes an interface 52 and a processor 54. Through the interface, sets of spatial audio parameters Ψ1, azi1, ele1, Ψ2, azi2, ele2, Ψn, azin, and ele can be received. The processor 54 analyzes the received sets and, in the event of a lost or corrupted set, replaces the lost or corrupted set, for example, with a previously received set or an equivalent set. These different strategies can be used, which will be discussed below.
[0075] Preservation Strategy: It is generally safe to assume that the spatial image must be relatively stable over time, and that it can be translated with respect to DirAC parameters (i.e., direction of arrival and spread, which do not change much between frames). For this reason, a simple but effective approach is to preserve the parameters of the last well-received frame from which frames were lost during transmission.
[0076] Orientational Extrapolation: Alternatively, one can envision estimating the trajectories of sound events in the audio scene and then attempt to extrapolate these estimated trajectories. This is particularly appropriate if the sound events are well localized as point sources in space, as reflected in the DirAC model by low diffusion. The estimated trajectory can be computed from observations of past directions and by fitting curves between these points, which may evolve with interpolation or smoothing. Regression analysis can also be used. Extrapolation is then performed by evaluating the fitted curves beyond the range of observed data.
[0077] In DirAC, direction is often expressed, quantized, and encoded in polar coordinates. However, it is generally more convenient to process direction in Cartesian coordinates and then process the trajectory to avoid dealing with modulo 2π operations.
[0078] Directional jitter: When sound events are widely diffused, direction is less significant and can be viewed as the realization of a random process. By injecting random noise into the previous direction before using it for a lost frame, jitter can help make the presented sound field more natural and pleasant. The injected noise and its variance can be a function of diffuseness.
[0079] Using standard DirAC audio scene analysis, we can study the impact of diffusion on the accuracy and significance of model orientation. Using an artificial B-format signal (which gives the direct to diffuse energy ratio (DDR) between the plane wave component and the diffuse field component), we can analyze the obtained DirAC parameters and their accuracy.
[0080] The theoretical diffusivity Ψ varies with the direct to diffusion energy ratio (DDR)Γ, and is expressed as:
[0081]
[0082] Where P pw With P diff Let Γ be the plane wave power and the spread power, respectively, and let Γ be the DDR expressed in dB.
[0083] Of course, it is possible to use one or a combination of the three strategies discussed. The processor 54 selects the strategy to be used based on the set of received spatial audio parameters. To this end, according to an embodiment, the audio parameters can be analyzed to enable the application of different strategies based on the characteristics of the audio scene, and more specifically, based on the diffusion.
[0084] This means that, according to one embodiment, processor 54 is configured to provide packet loss concealment of spatial parameter audio by using previously well-received directional information and jitter. According to another embodiment, the jitter varies with the estimated spread or the energy ratio between the directional and non-directional components of the audio scene. According to another embodiment, the jitter varies with the measured tonality of the transmitted downmixed signal. Therefore, the analyzer performs its analysis based on the estimated spread, energy ratio, and / or tonality.
[0085] exist Figure 3a and Figure 3b In the study, the diffusivity was measured using a simulated diffusion field of N = 466 uncorrelated pink noise particles uniformly positioned on a sphere, given by DDR, and by plane wave measurements using independent pink noise particles placed at 0 degrees azimuth and 0 degrees elevation. This confirmed that the diffusivity measured in the DirAC analysis is a good estimate of the theoretical diffusivity if the observation window length W is sufficiently large. This implies that the diffusivity has long-term characteristics, confirming that parameters can be well predicted even in the event of packet loss, provided that previously well-received values are maintained.
[0086] On the other hand, the direction parameter estimation can also be evaluated based on the true diffusion, as reported in Figure 4. It can be seen that the estimated elevation and azimuth angles of the plane wave position deviate from the true positions (0-degree azimuth and 0-degree elevation), with the standard deviation increasing with diffusion. For diffusion 1, the standard deviation is approximately 90 degrees for the azimuth angle, confined to between 0 and 360 degrees, corresponding to a uniformly distributed, completely random angle. In other words, the azimuth angle is thus meaningless. The same observation can be made for the elevation angle. Generally, the accuracy and significance of the estimated direction decrease with diffusion. Therefore, it is expected that the direction in DirAC will fluctuate over time and deviate from its expected value as diffusion changes. This natural dispersion is part of the DirAC model and is crucial for realistically reproducing audio scenes. In fact, presenting the directional components of DirAC in a constant direction (even with high diffusion) will produce a point source that should actually be perceived as wider.
[0087] For the reasons stated above, we propose applying jitter to the direction in addition to the hold strategy. The amplitude of the jitter is determined by the spread and can, for example, follow the model plotted in Figure 4. Two models for elevation angle and elevation angle measurement can be derived, with their standard deviations expressed as:
[0088] σ azi =65Ψ 3.5 +σ ele
[0089] σ ele =33.25Ψ+1.25
[0090] The pseudocode hidden by DirAC parameters can therefore be:
[0091]
[0092] Where `bad_frame_indicator[k]` is a flag indicating whether the frame at index `k` was well received. In the case of a good frame, the DirAC parameters are read, decoded, and dequantized for each parameter range corresponding to a given frequency range. In the case of a bad frame, the spread from the last well received frame in the same parameter range is directly preserved, while the azimuth and elevation angles are derived by dequantizing the last well received index using injected random values scaled according to a factor of the spread index. The function `random()` outputs random values based on a given distribution. The random process can, for example, follow a standard normal distribution with a mean of zero and unit variation. Alternatively, it can follow a uniform distribution between -1 and 1, or a triangular probability density using, for example, the following pseudocode:
[0093]
[0094] The jitter scale varies with the spread index inherited from the last good received frame within the same parameter range, and can be derived from the model inferred from Figure 4. For example, when the spread is encoded on 8 indices, they can correspond to the following table:
[0095]
[0096] Alternatively, the jitter intensity can be manipulated depending on the nature of the downmixed signal. In fact, tonally variable signals tend to be perceived as more localized sources, just like non-tonal signals. Therefore, the jitter can be adjusted according to the tonality of the transmitted downmixed signal by reducing its impact on the tonal term. Tonality can be measured, for example, by calculating the long-term prediction gain in the time domain or by measuring the spectral flatness in the frequency domain.
[0097] about Figure 6a and Figure 6b The discussion will cover methods for decoding DirAC-encoded audio scenes (see [link]). Figure 6a Method 200) and decoder 17 for DirAC encoded audio scenarios (see Method 200) Figure 6b Other embodiments of ).
[0098] Figure 6a The description includes a new method 200 comprising steps 110, 120, and 130 of method 100 and an additional decoding step 210. The decoding steps enable the decoding of a DirAC-coded audio scene, including downmixing (not shown), using a first set of spatial audio parameters and a second set of spatial audio parameters, wherein here, the replaced second set, output from step 130, is used. This concept is derived from... Figure 6b The device 17 shown is in use. Figure 6b Decoder 70 is shown, including a lost-hidden processor for spatial audio parameters 15 and a DirAC decoder 72. DirAC decoder 72, or more specifically, the processor of DirAC decoder 72, receives, for example, downmixed signals directly from interface 52 and / or processed by processor 52 according to the methods described above, and the various sets of spatial audio parameters.
[0099] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or a feature of a corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps can be performed by this apparatus.
[0100] The encoded audio signal of this invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0101] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementations can be executed using digital storage media, such as floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, storing electronically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to perform corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0102] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0103] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when executed on a computer, is operatively used to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0104] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0105] In other words, embodiments of the methods of the present invention are therefore computer programs having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0106] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0107] Therefore, a further embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.
[0108] Further embodiments include processing components, such as a computer or programmable logic device configured or adapted to perform one of the methods described herein.
[0109] Further embodiments include a computer having a computer program installed thereon for performing one of the methods described herein.
[0110] Further embodiments of the invention include means or systems configured to (e.g., electronically or optically) transmit a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, mobile device, memory device, etc. The means or system may, for example, include a file server for transmitting the computer program to the receiver.
[0111] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0112] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations to the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the following claims, and not by the specific details presented as interpreted through the description of the embodiments herein.
[0113] References
[0114] ·[1]V.Pulkki,MV.Laitinen,J.Vilkamo,J.Ahonen,T.Lokki,and T. "Directional audio coding-perception-based reproduction of spatialsound", International Workshop on the Principles and Application on SpatialHearing, Nov. 2009, Zao; Miyagi, Japan.
[0115] ·[2]V.Pulkki, "Virtual source positioning using vector base amplitudepanning", J.AudioEng.Soc., 45(6):456-466, June 1997.
[0116] ·[3]J.Ahonen and V.Pulkki,“Diffuseness estimation using temporalvariation of intensity vectors”,in Workshop on Applications of SignalProcessing to Audio and Acoustics WASPAA,Mohonk Mountain House,New Paltz,2009.
[0117] ·[4]T.Hirvonen,J.Ahonen,and V.Pulkki,“Perceptual compression methodsfor metadata in Directional Audio Coding applied to audiovisualteleconference”,AES 126th Convention 2009,May 7–10,Munich,Germany.
[0118] ·[5]A.Politis,J.Vilkamo and V.Pulkki,"Sector-Based Parametric SoundField Reproduction in the Spherical Harmonic Domain,"in IEEE Journal ofSelected Topics in Signal Processing,vol.9,no.5,pp.852-866,Aug.2015.
Claims
1. A method (100) for concealing the loss of spatial audio parameters, said spatial audio parameters including at least direction of arrival information, said method comprising the following steps: Receive (110) a first set of spatial audio parameters including at least a first direction of arrival (azi1, ele1); Receive (120) a second set of spatial audio parameters including at least a second direction of arrival (azi2, ele2); and If at least the second direction of arrival (azi2, ele2) information or a portion thereof is lost or corrupted, the second group of second direction of arrival (azi2, ele2) information is replaced with replacement direction of arrival information derived from the first direction of arrival (azi1, ele1) information. The replacement step includes a step of causing the replacement to jitter the direction information; and / or The replacement step includes injecting random noise into the first arrival direction (azi1, ele1) information to obtain the replacement arrival direction information.
2. The method (100) according to claim 1, wherein the first set of spatial audio parameters and the second set of spatial audio parameters respectively include first diffusion information and second diffusion information (Ψ1, Ψ2).
3. The method (100) according to claim 2, wherein the first diffusion information or the second diffusion information (Ψ1, Ψ2) is derived from at least one energy ratio associated with at least one direction of arrival information.
4. The method (100) according to claim 2, wherein the method further comprises replacing the second diffusion information (Ψ2) of the second group with replacement diffusion information derived from the first diffusion information (Ψ1).
5. The method (100) according to claim 1, wherein the replacement arrival direction information conforms to the first arrival direction (azi1, ele1) information.
6. The method (100) according to claim 2, wherein the injection step is performed if the first diffusion information or the second diffusion information (Ψ1, Ψ2) indicates high diffusion and / or if the first diffusion information or the second diffusion information (Ψ1, Ψ2) is higher than a predetermined threshold for diffusion information.
7. The method (100) of claim 6, wherein the diffusion information includes or is based on the ratio between the directional and non-directional components of an audio scene described by the first set of spatial audio parameters and / or the second set of spatial audio parameters.
8. The method (100) according to claim 2, wherein the random noise to be injected depends on the first diffusion information and / or the second diffusion information (Ψ1, Ψ2); and / or The random noise to be injected is scaled according to a factor that depends on the first diffusion information and / or the second diffusion information (Ψ1, Ψ2).
9. The method (100) according to claim 1, further comprising the following steps: Analyze the tonality of the audio scene described by the first set of spatial audio parameters and / or the second set of spatial audio parameters, or analyze the tonality of the transmitted downmixing belonging to the first set of spatial audio parameters and / or the second set of spatial audio parameters, to obtain a tonality value describing the tonality; and The random noise to be injected depends on the tone value.
10. The method (100) of claim 9, wherein the random noise is reduced proportionally by a factor that decreases along with the reciprocal of the tonality value, or the random noise is reduced proportionally if the tonality increases.
11. The method (100) of claim 1, wherein the method (100) includes the step of extrapolating the first arrival direction (azi1, ele1) information to obtain the replacement arrival direction information.
12. The method (100) of claim 11, wherein the extrapolation is based on one or more additional direction of arrival information belonging to one or more sets of spatial audio parameters.
13. The method (100) according to claim 11, wherein the extrapolation is performed if the first diffusion information and / or the second diffusion information (Ψ1, Ψ2) indicate low diffusion or if the first diffusion information and / or the second diffusion information (Ψ1, Ψ2) are below a predetermined threshold for diffusion information.
14. The method (100) according to claim 1, wherein the first set of spatial audio parameters belongs to a first time point and / or a first frame, and wherein the second set of spatial audio parameters belongs to a second time point and / or a second frame; or The first set of spatial audio parameters belongs to a first time point, and the second time point is after the first time point, or the second frame is after the first frame.
15. The method (100) according to claim 1, wherein the first set of spatial audio parameters includes a first subset of spatial audio parameters for a first frequency band and a second subset of spatial audio parameters for a second frequency band; and / or The second set of spatial audio parameters includes another first subset of spatial audio parameters for the first frequency band and another second subset of spatial audio parameters for the second frequency band.
16. A method (200) for decoding a DirAC-encoded audio scene, comprising the following steps: Decode the DirAC encoded audio scene, including downmixing, the first set of spatial audio parameters, and the second set of spatial audio parameters; Perform the method for loss concealment according to any one of claims 1 to 15.
17. A computer-readable digital storage medium having stored thereon a computer program having program code for performing the method (100, 200) according to any one of claims 1 to 15 or according to claim 16 when run on a computer.
18. A loss-concealing device (50) for concealing the loss of spatial audio parameters, the spatial audio parameters including at least direction of arrival information, the device comprising: The receiver (52) is configured to receive (100) a first set of spatial audio parameters including first direction of arrival (azi1, ele1) information, and to receive (120) a second set of spatial audio parameters including second direction of arrival (azi2, ele2) information. The processor (54) is configured to replace the second group of arrival direction (azi2, ele2) information with replacement arrival direction information derived from the first arrival direction (azi1, ele1) information in the event that at least the second arrival direction (azi2, ele2) information or a portion thereof is lost or damaged. The replacement includes causing the replacement to jitter the direction information; and / or The replacement includes injecting random noise into the first arrival direction (azi1, ele1) information to obtain the replacement arrival direction information.
19. A decoder (70) for DirAC encoded audio scenes, the decoder comprising the loss concealment device according to claim 18.
Citation Information
Patent Citations
Packet loss shielding device and method and audio processing system
CN104282309A