Improved Stability of Interchannel Time Difference (ITD) Estimators for Coincident Stereo Acquisition.
By identifying coincident microphone configurations and adapting ITD searches to favor time lags closer to zero, the method stabilizes ITD detection, addressing errors in spatial audio encoding and decoding, thus improving audio quality and stability.
Patent Information
- Application Number
- JP2023577407
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-15
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Existing spatial audio encoding and decoding technologies face challenges in accurately estimating inter-channel time differences (ITD) for coincident microphone configurations, leading to unstable audio impressions due to erroneous ITD detection, especially in noisy or reverberant environments.
A method is introduced to identify coincident microphone configurations and adapt the ITD search to favor time lags closer to zero, using generalized cross-correlation with phase transform (GCC-PHAT) analysis and low-pass filtering to stabilize ITD detection, ensuring accurate spatial audio rendering.
This approach stabilizes ITD detection, improving encoding quality and stability of reconstructed audio signals from coincident microphone configurations, enhancing the overall audio experience by reducing unstable energy fluctuations.
Smart Images

Figure 0007680574000061 
Figure 0007680574000062 
Figure 0007680574000063
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to communications, and more particularly to methods and related encoders and decoders supporting audio encoding and decoding. [Background technology]
[0002] Spatial audio or 3D audio is a general formulation to represent various kinds of multi-channel audio signals. Depending on the capture method and rendering method, the audio scene is represented by a spatial audio format. Typical spatial audio formats defined by the capture method (microphones) are represented as, for example, stereo, binaural, ambisonics, etc. A spatial audio rendering system (headphones or speakers) can render a spatial audio scene in stereo (left and right channels 2.0) or more advanced multi-channel audio signals (2.1, 5.1, 7.1, etc.).
[0003] Recent techniques for the transmission and manipulation of such audio signals allow end users to have an enhanced audio experience with higher spatial quality, often resulting in better intelligibility as well as augmented reality. Spatial audio coding techniques such as MPEG Surround or MPEG-H 3D Audio generate compact representations of spatial audio signals that are compatible with data-rate constrained applications, such as streaming over the Internet. However, the transmission of spatial audio signals is limited when data-rate constraints are strong, and therefore post-processing of decoded audio channels is also used to enhance spatial audio reproduction. A commonly used technique can, for example, blind upmix decoded mono or stereo signals to multi-channel audio (5.1 channels or more).
[0004] To efficiently render spatial audio scenes, spatial audio coding and processing techniques exploit the spatial properties of multi-channel audio signals. In particular, the time and level differences between channels of spatial audio capture are used to approximate the interaural cues that characterize the perception of directional sound in space. It is very important that the inter-channel time differences are relevant from a perceptual aspect, since they are only an approximation of what the auditory system can detect (i.e., interaural time differences and interaural level differences, ear entrances). Inter-channel time differences and inter-channel level differences (ICTD and ICLD) are commonly used to model the directional components of multi-channel audio signals, while inter-channel cross-correlation (ICC), which models the inter-aural cross-correlation (IACC), is used to characterize the width of the audio image. The stereo image can also be modeled with inter-channel phase differences (ICPD), especially for low frequencies.
[0005] It should be noted that the binaural cues relevant to spatial auditory perception are called interaural level difference (ILD), interaural time difference (ITD), and interaural coherence or interaural correlation (IC or IACC). Considering a typical multi-channel signal, the corresponding cues related to the channels are inter-channel level difference (ICLD), inter-channel time difference (ICTD), and inter-channel coherence or inter-channel correlation (ICC). Since spatial audio processing mostly operates on captured audio channels, the "C" is sometimes dropped, and the terms ITD, ILD and IC are also used when referring to the audio channels.
[0006] FIG. 1 shows a conventional setup using parametric spatial audio analysis. A stereo signal pair is input to a stereo encoder 110. A spatial analyzer 112 assists a downmixer 114, which generates a single channel representation of the two input channels. The downmix process aims to compensate for channel differences in time, correlation and phase, thereby maximizing the energy of the downmix signal. This achieves an efficient encoding of the stereo signal. The downmix signal is forwarded to a downmix encoder 116. Parameters from the spatial analysis are encoded by a parameter encoder 118 and transmitted to the decoder together with the encoded downmix. Typically, some of the stereo parameters are represented in spectral subbands on a perceptual frequency scale, such as the equivalent rectangular bandwidth (ERB) scale. The stereo decoder 120 performs stereo synthesis in a spatial synthesizer 126 based on the signal from the downmix decoder 124 and the parameters from the parameter decoder 122. The stereo synthesis operation aims to restore channel differences in time, level, correlation and phase, and to generate a stereo image similar to the input audio signal.
[0007] Since the encoded parameters are used to render spatial audio to the human auditory system, the inter-channel parameters may be extracted and encoded using perceptual considerations to maximize perceptual quality.
[0008] Stereo and multi-channel audio signals are complex signals that can be difficult to model, especially when the environment is noisy or reverberant, or when the various audio components of the mixture overlap in time and frequency, i.e., in cases of noisy speech, musical speech or simultaneous talkers.
[0009] When it comes to estimating ICTD, traditional parametric methods rely on the cross-correlation function (CCF), r, which is a measure of similarity between two waveforms x(n) and y(n). xy and is generally specified in the time domain as r xy (n, τ) = E[x(n)y(n+τ)] where τ is the time lag parameter and E[·] is the expectation operator. For a signal frame of length N, the cross-correlation is usually estimated as TIFF0007680574000001.tif13170
[0010] The ICC is conventionally obtained as the maximum of the CCF normalized by the signal energy according to: TIFF0007680574000002.tif12170
[0011] The time lag τ corresponding to the ICC is determined as the ICTD between channel x and channel y. The CCF can also be calculated using the discrete Fourier transform as follows: r xy (τ)=DFT -1 (X(k)Y * (k) where X[k] is the discrete Fourier transform (DFT) of the time-domain signal x[n], and Y * [k] is the complex conjugate of the discrete Fourier transform (DFT) of the time domain signal y[n], i.e. TIFF0007680574000003.tif17170, DFT -1 The DFT(·) or IDFT(·) is the Inverse Discrete Fourier Transform. Note, however, that the DFT replicates the analysis frame into a periodic signal, resulting in a circular convolution of x(n) and y(n). Based on this, the analysis frame is usually padded with zeros to match the true cross-correlation.
[0012] If y(n) is purely a delayed version of x(n), then the cross-correlation function is given by: TIFF0007680574000004.tif7170Here, * denotes convolution, and δ(τ-τ 0 ) is the Kronecker delta function, i.e., τ 0is equal to 1 in x(n) and equal to zero otherwise. This means that the cross-correlation function between x and y is the autocorrelation function for x(n), r xx This means that the cross-correlation is a delta function spread by convolution with (τ). For a signal frame with several delay components, e.g. several speakers, there will be a peak at each delay that exists between the signals, and the cross-correlation will be r xy (τ)=r xx (τ)*Σ i δ(τ-τ i )
[0013] The delta functions are then spread apart from one another, which can make it difficult to distinguish between some delays within a signal frame. However, there exists a generalized cross-correlation (GCC) function that does not have this spreading. GCC is generally defined as follows: TIFF0007680574000005.tif6170 where ψ[k] is the frequency weighting. In spatial audio, the phase transform (PHAT) has been utilized due to its robustness to reverberation in low noise environments. The phase transform is essentially the absolute value of each frequency coefficient, i.e. The file is TIFF0007680574000006.tif10170.
[0014] This weighting whitens the cross spectrum so that the power of each component is equal. With pure delay and uncorrelated noise in the signals x[n] and y[n], the phase-transformed GCC (GCC-PHAT) is simply the Kronecker delta function δ(τ-τ 0 ), i.e. The file is TIFF0007680574000007.tif15170.
[0015] FIG. 2 shows signal pairs with inter-channel time differences, their cross-correlations, and generalized cross-correlations by phase transform analysis for a pure delay situation.
[0016] In a real scenario of analyzing recorded stereo signals, the channels do not differ only by delay, but may for example have different noises, variations in the frequency response of the microphone and recording equipment, and have different reverberation patterns. In this case, the time lag τ is usually found by identifying the maximum of GCC-PHAT. In such situations, the analysis is even more likely to show frame-to-frame variations. This is a typical characteristic in short-term Fourier analysis, but also because the source signal may vary in level and spectral content, which is the case for example in voice recordings. For this reason, it is beneficial to apply a stabilization to the final analysis of the time lag. This can be done by slowing down or preventing the update of the time lag when the signal energy is low relative to the background noise.
[0017] In US 2020 / 0194013, ITD selection is stabilized by applying an adaptive low-pass filter in GCC-PHAT. Low-pass filtering is applied to the cross-correlation by adaptively filtering the cross-correlation of successive frames. A low-pass filter is also applied to the time-domain representation of the cross-correlation. For clean signals with a high estimated signal-to-noise ratio (SNR), more advanced low-pass filtering is used.
[0018] US Patent Application Publication No. 20200211575 describes a method for reusing previously stored ITD values depending on an SNR estimate, thereby achieving more stable ITD parameters over time.
[0019] The time lag between channels in stereo recordings is due to the physical distance between the microphones. As shown in Figure 3, the AB microphone configuration usually has a relatively large distance between the microphones, about 1 to 1.5 meters. Therefore, recordings using the AB configuration often have a time delay between the channels, depending on the location of the captured audio source. Some microphone configurations, such as XY and MS, try to place the microphone membranes as close to each other as possible, so-called coincident microphone configurations. These coincident microphone configurations usually have very small or zero time delay between the channels. The XY configuration captures the stereo image mainly through level differences. The MS setup, short for Mid-Side, has a front channel pointed forward and a microphone with a figure-of-eight pickup pattern to capture the surrounding environment in the side channels. The Mid-Side representation is converted to a Left-Right representation using the following relationship: TIFF0007680574000008.tif9170 The side channel S is added to the left and right channels with opposite sign. More generally, a stereo representation can be obtained by converting two or more mono signals into a stereo representation, where the time difference between the signals (related to the physical distance of capture) must be small. Another example of a suitable capture technique is the use of a tetrahedral microphone with four closely spaced cardioids, from which a stereo representation can be formed. Summary of the Invention
[0020] For an MS coincident microphone configuration (hereafter referred to as "coincident configuration" and abbreviated as "CC"), the time lag should ideally always be close to zero. However, due to reverberation and noise, an occasional time lag may be detected. When the time lag is encoded in the context of a stereo or multi-channel audio encoder, a sudden jump in the time lag caused by an incorrectly detected lag may give an unstable impression of the location of the audio source in the reconstructed audio signal. Furthermore, an inaccurate or unstable time lag may adversely affect the downmix signal, which may exhibit unstable energy as a result of these errors.
[0021] Even if low-pass filtering of GCC-PHAT is applied as proposed in US20200194013, detection of erroneous ITD in CC signals may occur. As outlined in US20200211575, the ability to reuse previously stored ITD values does not prevent erroneous ITD estimation in CC signals. In fact, added stabilization may cause erroneous decisions to persist even longer.
[0022] Certain aspects of the present disclosure and embodiments thereof may provide solutions to these and other problems. Various embodiments of the inventive concepts described herein detect coincident configurations, such as MS microphone configurations. If such a configuration (e.g., MS microphone configuration) is detected, the time lag detection may be adapted such that time lags closer to zero are preferred.
[0023] According to some embodiments of the inventive concepts, there is provided a method for identifying a coincident microphone configuration CC and adapting an inter-channel time difference ITD search in an encoder or decoder. The method includes generating a cross-correlation of a pair of channels of the multi-channel audio signal for each frame m of the multi-channel audio signal. The method includes determining a first ITD estimate based on the cross-correlation. The method includes determining whether the multi-channel audio signal is a CC signal. The method includes biasing an ITD search to favor ITDs closer to zero to obtain a final ITD in response to determining that the multi-channel audio signal is a CC signal.
[0024] Similar apparatus, computer programs, and computer program products are provided according to other embodiments of the inventive concept.
[0025] An advantage that can be achieved is that it allows for stabilization of the time lag or ITD detection, which improves the encoding quality and stability of the reconstructed audio of a stereo signal of a coincident configuration, e.g., from an MS configuration. Stabilizing the time lag or ITD detection improves the encoding quality and stability of the reconstructed audio of a stereo signal of a coincident configuration, e.g., from an MS configuration.
[0026] Configuration detection can be based on the GCC-PHAT spectrum, which is already computed to estimate the time lag, giving only a very small computational overhead compared to the baseline system.
[0027] The accompanying drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this specification, illustrate certain non-limiting embodiments of the inventive concepts. [Brief description of the drawings]
[0028] [Figure 1]FIG. 1 is a block diagram showing a stereo encoder and decoder system. [Diagram 2] 1 is a diagram of a signal pair with inter-channel time difference, their cross-correlation, and generalized cross-correlation by phase transform analysis. [Diagram 3] FIG. 1 is a diagram of microphone configurations and their capture patterns. [Figure 4] FIG. 1 is a diagram of possible antisymmetric configurations for CC signals. [Diagram 5] FIG. 13 is a diagram of an example mask for emphasizing ITDs near zero, in accordance with some embodiments of the inventive concepts. [Figure 6] 1 is a flowchart illustrating operations for identifying CC signals and adapting ITD searches in accordance with some embodiments of the inventive concepts. [Figure 7] 4 is a block diagram illustrating the operation of an encoder / decoder apparatus for identifying a CC signal and adapting an ITD search in accordance with some embodiments of the inventive concepts. [Figure 8] 1 is a flowchart illustrating operations for identifying MS constituent signals and adapting an ITD search in accordance with some embodiments of the inventive concepts. [Figure 9] 1 is a block diagram illustrating the operation of an encoder / decoder apparatus for identifying MS constituent signals and adapting an ITD search in accordance with some embodiments of the inventive concepts. [Figure 10] 1 is a block diagram illustrating an exemplary environment in which an encoder and / or decoder may operate, according to some embodiments of the inventive concept. [Figure 11] FIG. 1 is a block diagram of a virtualization environment in accordance with some embodiments. [Figure 12] 1 is a block diagram illustrating an encoder in accordance with some embodiments of the inventive concept. [Figure 13] FIG. 2 is a block diagram illustrating a decoder, in accordance with some embodiments of the inventive concepts. [Figure 14] 4 is a flowchart illustrating the operation of an encoder or decoder according to some embodiments of the inventive concept. [Figure 15] 4 is a flowchart illustrating the operation of an encoder or decoder according to some embodiments of the inventive concept. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0029] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. The embodiments are provided as examples to convey the scope of the subject matter to those skilled in the art, and examples of embodiments of the inventive concept are shown. However, the inventive concept may be embodied in many different forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided so that this disclosure will be comprehensive and complete, and will fully convey the scope of the inventive concept to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Elements from one embodiment may be implicitly assumed to be present / used in another embodiment.
[0030] Before describing the embodiments in further detail, Figure 10 illustrates an example of an operating environment for an encoder 110 that may be used to encode a bitstream as described herein. The encoder 110 receives audio from a network 1002 and / or a storage device 1004, encodes the audio into a bitstream as described below, and transmits the encoded audio to a decoder 120 over a network 1008. The storage device 1004 may be part of a storage depository of multi-channel audio signals, such as a storage repository of a store or streaming audio service, a separate storage component, a component of a mobile device, etc. The decoder 120 may be part of a device 1010 having a media player 1012. The device 1010 may be a mobile device, a set-top device, a desktop computer, etc.
[0031] FIG. 11 is a block diagram illustrating a virtualization environment 1100 in which functions implemented by some embodiments may be virtualized. In this context, virtualizing means creating a virtual version of an apparatus or device, which may include virtualizing a hardware platform, storage devices, and networking resources. As used herein, virtualization may apply to any device or component thereof described herein and relates to implementations in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functionality described herein may be implemented as virtual components, executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1100 hosted by one or more of the hardware nodes, such as a network node, a UE, a core network node, or a hardware computing device operating as a host. Furthermore, in embodiments in which the virtual node does not require wireless connectivity (e.g., a core network node or a host), the node may be fully virtualized.
[0032] An application 1102 (which may alternatively be referred to as a software instance, a virtual appliance, a network function, a virtual node, a virtual network function, etc.) is run in the virtualized environment 1100 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.
[0033] The hardware 1104 includes processing circuitry, memory that stores software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices described herein, such as network interfaces, input / output interfaces, etc. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 1106 (also referred to as a hypervisor or virtual machine monitor (VMM)), provide VMs 1108A and 1108B (one or more of which may be generally referred to as VMs 1108), and / or perform any of the functions, features and / or benefits described in connection with some embodiments described herein. The virtualization layer 1106 may present a virtual operating platform that appears to be networking hardware to the VMs 1108.
[0034] The VMs 1108 may comprise virtual processing, virtual memory, virtual networking or interfaces, and virtual storage, and may be run by a corresponding virtualization layer 1106. Different embodiments of instances of virtual appliances 1102 may be implemented in one or more of the VMs 1108, and the implementation may be done in different ways. Hardware virtualization is referred to in some contexts as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry-standard high-volume server hardware, physical switches, and physical storage that may be located in data centers and customer premises equipment.
[0035] In the context of NFV, the VMs 1108 may be software implementations of physical machines that run programs as if they were running on a physical non-virtual machine. Each of the VMs 1108, and the portion of the hardware 1104 on which each VM runs, forms a separate virtual network element, even if the hardware is dedicated to each VM and / or shared by each VM with other VMs. Furthermore, in the context of NFV, a virtual network function is responsible for handling a particular network function running within one or more VMs 1108 on the hardware 1104 and corresponds to the application 1102.
[0036] The hardware 1104 may be implemented in a standalone network node having general or specific components. The hardware 1104 may implement some functions through virtualization. Alternatively, the hardware 1104 may be part of a larger cluster of hardware (e.g., in a data center or CPE) where many hardware nodes work together and are managed via a management and orchestration 1110 that oversees, among other things, the lifecycle management of the application 1102. In some embodiments, the hardware 1104 may be coupled to one or more radio units, each including one or more transmitters and one or more receivers that may be coupled to one or more antennas. The radio units may communicate directly with the hardware nodes via one or more appropriate network interfaces or may be used in combination with virtual components to provide a virtual node with wireless capabilities, such as a radio access node or base station. In some embodiments, some signaling may be provided by using a control system 1112 that may alternatively be used for communication between the hardware nodes and the radio units.
[0037] 12 is a block diagram illustrating elements of an encoder 1000 configured to encode audio frames according to some embodiments of the inventive concepts. As shown, the encoder 1000 may include a network interface circuit 1205 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The encoder 1000 may also include a processor circuit 1201 (also referred to as a processor) coupled to the network interface circuit 1205, and a memory circuit 1203 (also referred to as a memory) coupled to the processor circuit. The memory circuit 1203 may include computer readable program code that, when executed by the processor circuit 1201, causes the processor circuit to perform operations according to embodiments disclosed herein.
[0038] According to other embodiments, the processor circuitry 1201 may be defined to include memory such that a separate memory circuitry is not required. As discussed herein, the operations of the encoder 1000 may be performed by the processor 1201 and / or the network interface 1205. For example, the processor 1201 may control the network interface 1205 to send communications to the decoder 1006 and / or receive communications from one or more other network nodes / entities / servers, such as other encoder nodes, depository servers, etc., via the network interface 1205. Furthermore, modules may be stored in the memory 1203, and these modules may provide instructions such that, when the instructions of the modules are executed by the processor 1201, the processor 1201 performs the respective operations.
[0039] 13 is a block diagram illustrating elements of a decoder 1006 configured to decode audio frames in accordance with some embodiments of the inventive concepts. As shown, the decoder 1006 may include a network interface circuit 1305 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The decoder 1006 may also include a processor circuit 1301 (also referred to as a processor) coupled to the network interface circuit 1305, and a memory circuit 1303 (also referred to as a memory) coupled to the processor circuit. The memory circuit 1303 may include computer readable program code that, when executed by the processor circuit 1301, causes the processing circuit to perform operations according to embodiments disclosed herein.
[0040] According to other embodiments, the processor circuitry 1301 may be defined to include memory such that a separate memory circuitry is not required. As discussed herein, the operations of the decoder 1006 may be performed by the processor 1301 and / or the network interface 1305. For example, the processor circuitry 1301 may control the network interface circuitry 1305 to receive communications from the encoder 1000. Additionally, modules may be stored in the memory 1303, and these modules may provide instructions such that, when the instructions of the modules are executed by the processor circuitry 1301, the processor circuitry 1301 performs respective operations.
[0041] Consider a system designated to obtain spatial representation parameters of an audio input consisting of two or more audio channels. The system may be part of a stereo encoding and decoding system or encoder / decoder as outlined in FIG. 1. The audio input is segmented into time frames m. In the case of multi-channel approaches, spatial parameters are usually obtained for a channel pair, in the case of a stereo setup, this pair is simply the left and right channels L and R. In an encoder, the method may be part of a spatial analysis to assist the downmix procedure and encode the spatial parameters to represent a spatial image. In a decoder, the method may complement the downmix procedure when the number of channels received is larger than can be handled by the decoder unit, for example in the case of a stereo decoder with mono audio playback capabilities. Hereafter, we focus on the inter-channel time difference (ITD) parameter as part of the set of spatial parameters derived by the spatial analyzer 112 for a single channel pair l(n,m) and r(n,m), where n represents the sample number and m represents the frame number. Hereafter, the index m is used to indicate the value calculated for frame m.
[0042] Referring to Figure 6, the system has a specified method that is activated for stereo signals coming from a coincident configuration. The spatial representation parameters include ITD parameters, which in some embodiments may be derived using a Generalized Cross-Correlation with Phase Transform (GCC-PHAT) analysis of the input channels in block 610. The analysis may include smoothing of cross-correlation between time frames, as proposed in US Patent Application Publication No. 20200194013. The ITD for frame m in these embodiments 0 The first estimate of the (m) parameter is the absolute maximum of GCC-PHAT in block 620. The first estimate may be determined according to: TIFF0007680574000009.tif8170, where ITD 0(m) is the first estimate of the ITD, τ is the time lag parameter, TIFF0007680574000010.tif7170 is GCC-PHAT.
[0043] It has been observed that the GCC-PHAT of an MS signal (i.e., a certain type of CC) may exhibit an anti-symmetric pattern, as shown in Figure 4. This structure comes from the time difference due to the small distance between the microphones in an MS setup, and the fact that the S signal is added to the left and right channels with opposite signs. This pattern may be exploited when forming the coincident configuration detection variable D(m) for frame m in computing the CC detection variable in block 630. TIFF0007680574000011.tif7170
[0044] Alternative detection variables that have been found to give a positive indication of coincident organization of some stereo representations are: TIFF0007680574000012.tif75170, where R is the search range and W is the symmetry-ITD 0 (m) defines a region around the first estimate of ITD that coincides at the time lag of ITD 0 ’ (m) is an ITD candidate limited to the search range [-R, R], and is determined, for example, as follows: For coincident configurations such as MS signals, where symmetry appears close to τ=0, a suitable search range may be R=10 or in the range R∈[5,20]. A suitable value defining the matching region is W=1 or in the range [0,5]. The embodiments described herein assume 32 kHz sampling of the audio signal, and suitable ranges for the parameters may depend on the sampling frequency.
[0045] To stabilize the detector, the decision variables, D LP (m) = αD(m) + (1-α)DLP (m-1) It may be desirable to low-pass filter where α is a low-pass filter coefficient. Suitable values for α may be α=0.1 or in the range of α∈(0,0.2). If the formation of D(m) does not include absolute values, the low-pass filter may include absolute values. D LP (m) = α | D(m) | + (1-α) D LP (m-1) Since the detector variables only give valid values when the source is active, it is beneficial to restrict the decision variable updates to this situation. The low-pass filtered decision variable equations become: where A(m) is TRUE if frame m is active, i.e., classified as containing an active source signal such as speech, and FALSE otherwise. A(m) can be, for example, the output of a voice activity detector (VAD), or the absolute maximum of GCC-PHAT compared to a threshold, TIFF0007680574000015.tif7170 indicates that the source is active. Here, C thr The appropriate value is C thr =0.5 or C thr ∈[0.3,0.9]. Another way to achieve this behavior is to use the activity measure A(m) to adapt the low-pass filter coefficient α, D LP (m) = α(m)D(m) + (1-α(m))D LP (m-1) TIFF0007680574000016.tif12170Here, suitable values for the filter coefficients are α high = 0.1 or α∈[α low ,0.5], and α low = 0.01 or α low ∈[0,α highIf the activity indicator is false, A(m)=FALSE, then the detector variable may be unreliable and it may be desirable to decay the detector variable towards a predetermined value; TIFF0007680574000017.tif12170, where D 0 D 0 =0 or D 0 =D THR and D THR is the decision threshold described below.
[0046] To determine whether the signal is a CC signal, the detector variable may be compared to a threshold in block 640 . TIFF0007680574000018.tif10170 The absolute value is D(m), and the result is D LP If not included in forming (m), the comparison to the threshold may include an absolute value. TIFF0007680574000019.tif11170
[0047] Note that indicating that a signal is a CC signal means that the signal comes from a coincident microphone configuration. If a CC signal is detected, the ITD search may be influenced to favor ITDs closer to zero. For example, as described in U.S. Patent Application Publication No. 20200194013, ITD stabilization may be applied and the stabilized ITD, ITD, may be selected in block 650. stab (m) is obtained. If a CC signal is detected, in some embodiments of the inventive concept, in block 660, the ITD with the smallest absolute value is selected. TIFF0007680574000020.tif18170, where ITD 1 (m) is the final ITD, ITD 0 (m) is the first ITD estimate, ITD stab(m) is the stabilized ITD. The stabilization procedure may result in a stabilized ITD that is the same as the first ITD estimate, which means that the ITD is the same even if the CC signal is not detected, i.e., CC detection=FALSE. 1 (m) is ITD 0 (m). In another embodiment, the switch to a smaller absolute value means that the absolute value goes from zero to [-R 1 ,R 1 ] is only performed if the TIFF0007680574000021.tif2417032kHz sampling frequency, R 1 The appropriate value of R 1 =10 or R 1 ∈[5,20].
[0048] Further stabilization can be applied, for example, taking into account the previous ITD value as described in US Patent Application Publication No. 20200211575. Again, if a CC signal is detected, in block 660, the stabilization result is accepted if the absolute value is close to zero. Again, the decision to retain the previously obtained ITD instead of the stabilized ITD is also determined by the fact that the previously obtained ITD has changed from zero to, for example, [-R 1 ,R 1 ] may depend on whether the
[0049] Another method of prioritizing ITDs closer to zero is the GCC-PHAT method, which complements the stabilization660 by giving more weight to values closer to zero. The weighting w(τ) is given by w(τ) = max(0,1-|τ(1+C) / ITD MAX |) can be obtained by
[0050] On the other hand, if no CC signal is detected, the weighting is omitted, which is equivalent to setting the weighting to one. TIFF0007680574000023.tif12170
[0051] This weighting function can be a suitable value for those constants for a sampling frequency of 32 kHz, C=5 and ITD MAX This effectively masks out the wedge of correlation values around zero, as shown in Figure 5 for = 200. In this case, the ITD estimate is the absolute maximum of the weighted GCC-PHAT. TIFF0007680574000024.tif8170
[0052] If CC detection = FALSE, the ITD already obtained 0 (m) may be used.
[0053] With reference to FIG. 7, the above-described embodiment may be implemented by a cross-correlation analyzer 710 that can generate a GCC-PHAT analysis of input signals L and R. A first ITD estimate is generated by an ITD analyzer 720. A CC detector 730 uses at least the output of the cross-correlation analyzer, and optionally the first ITD estimate, to detect low ITD signals, such as CC signals. The CC detector forms a CC detector variable that is compared to a threshold to determine whether a CC signal is present. If a CC signal is detected, it instructs an ITD stabilizer 740 to favor ITD values closer to zero.
[0054] 8 shows an embodiment where CC detection is based on analysis of the previous frame. During system startup, the MS detector variable memory and the MS detector flags are initialized in block 810. For each frame m, blocks 820 through 850 are performed.
[0055] At block 820, the cross-correlation TIFF0007680574000025.tif7170 is calculated. In block 830, the absolute maximum value of the weighted cross-correlation ITD 1 (m) Determined according to TIFF0007680574000026.tif8170.
[0056] The weighting may be the same as in block 640 above, but the decision is based on the CC detection from the previous frame. TIFF0007680574000027.tif12170
[0057] The identified maximum value may be further stabilized in optional block 840, similar to the stabilization performed above in block 660. A CC detection variable is derived in block 850, similar to the derivation described above in block 630. This value is then stored for use in the next frame. TIFF0007680574000028.tif11170 The absolute value is D(m), and the result is D LP If not included in forming (m), the comparison to the threshold may include an absolute value. TIFF0007680574000029.tif11170
[0058] In this case, the decision variables are the instantaneous estimates ITD, including the stabilization method that may be implemented in block 840. 0 (m) or the final ITD value ITD(m).
[0059] 9, the embodiment described in FIG. 8 may be implemented by a cross-correlation analyzer 910 capable of generating a GCC-PHAT analysis of the input signals L and R. A weighter and absolute maximum finder 920 weights the cross-correlation and determines the absolute maximum ITD of the weighted cross-correlation. An optional ITD stabilizer 930 stabilizes the final ITD 1 The identified maximum ITD is stabilized to obtain (m). The MS detector variable and CC detector flag updater 940 derives the CC detection variable and provides the CC detection variable to a CC detector variable and CC detector flag memory 950 for storing the CC detector variable for use in the next frame.
[0060] In the following description, the encoder may be either the stereo encoder 110, the encoder 1000, the virtualization hardware 1104 or the virtual machine 1108A, 1108B, but the encoder 1000 shall be used to describe the functionality of the operation of the encoder. Similarly, the decoder may be either the stereo decoder 120, the decoder 1006, the hardware 1104 or the virtual machine 1108A, 1108B, but the decoder 1006 shall be used to describe the functionality of the operation of the decoder. Next, the operation of the encoder 1000 (implemented using the structure of the block diagram of FIG. 12) or the decoder 1006 (implemented using the structure of the block diagram of FIG. 13) will be described with reference to the flowchart of FIG. 14 according to some embodiments of the inventive concept. For example, modules may be stored in the memory 1203 of FIG. 12 or the memory 1303 of FIG. 13, and these modules may provide instructions such that when the instructions of the modules are executed by the respective processing circuit 1201 / 1301, the processing circuit 1201 / 1301 performs the respective operations of the flowchart.
[0061] 14 shows a method for identifying coincident microphone configurations CC and adapting inter-channel time difference ITD search in an encoder or decoder. In the case of a decoder, this method is mainly used when the decoder receives a stereo signal but the audio device only has mono playback capability.
[0062] 14, the operations of blocks 1401 to 1409 are performed for each frame m of a multi-channel audio signal. In block 1401, the processing circuit 1201 / 1301 generates cross-correlations of pairs of channels of the multi-channel audio signal. The cross-correlations may be generated as described above in FIGs. 6 and 8. In some embodiments of the inventive concept, the cross-correlations are generalized cross-correlations with phase transforms (GCC-PHAT).
[0063] In block 1403, the processing circuit 1201 / 1301 determines a first ITD estimate based on the cross-correlation. The processing circuit 1201 / 1301 may determine the first ITD estimate by determining the first ITD estimate as an absolute maximum of the cross-correlation. In some embodiments, the processing circuit 1201 / 1301 determines the absolute maximum of the cross-correlation according to: TIFF0007680574000030.tif8170, where ITD 0 (m) is the first ITD estimate, TIFF0007680574000031.tif7170 is the cross-correlation and τ is the time lag parameter.
[0064] In block 1405, the processing circuit 1201 / 1301 determines whether the multi-channel audio signal is a CC signal.
[0065] In some embodiments of the inventive concept, the processing circuit 1201 / 1301 determines whether the multi-channel audio signal is a CC signal based on a CC detection variable. Figure 15 illustrates an embodiment of determining whether the multi-channel audio signal is a CC signal based on a CC detection variable. Referring to Figure 15, in block 1501, the processing circuit 1201 / 1301 calculates a CC detection variable. The calculation of the CC detection variable has been described above.
[0066] In block 1503, processing circuit 1201 / 1301 determines whether the CC detection variable is above a threshold. In some of these embodiments, processing circuit 1201 / 1301 determines whether the CC detection variable is above a threshold by determining whether the absolute value of the CC detection variable is above the threshold.
[0067] In block 1505, the processing circuit 1201 / 1301 determines that the multi-channel audio signal is a CC signal in response to determining that the CC detection variable is above the threshold. In block 1507, the processing circuit 1201 / 1301 determines that the multi-channel audio signal is not a CC signal in response to determining that the CC detection variable is not above the threshold.
[0068] In other embodiments, the processing circuit 1201 / 1301 determines whether the multi-channel audio signal is a CC signal by detecting one of an antisymmetric and symmetric pattern of cross-correlation in a channel pair of the multi-channel audio signal. In some embodiments, detecting the antisymmetric pattern in the components includes detecting the antisymmetric pattern according to: TIFF0007680574000032.tif7170 where D(m) is the CC detection variable, TIFF0007680574000033.tif7170 is GCC-PHAT and ITD 0 (m) is the first ITD estimate.
[0069] In another embodiment of the inventive concept, the processing circuit 1201 / 1301 detects one of the antisymmetric and symmetric patterns in the cross-correlation by detecting the antisymmetric pattern according to at least one of the following: TIFF0007680574000034.tif79170 where D(m) is the CC detection variable, TIFF0007680574000035.tif7170 is GCC-PHAT, R is the search range, W defines the area around the first estimate of the matching ITD, and ITD 0 ’ (m) is an ITD candidate limited to the search range [-R, R].
[0070] Returning to FIG. 14, in block 1407, in response to determining that the multi-channel audio signal is a CC signal, the processing circuit 1201 / 1301 biases the ITD search to favor ITDs closer to zero to obtain the final ITD.
[0071] In some embodiments, the processing circuit 1201 / 1301 biases the ITD search to favor ITDs closer to zero to obtain the final ITD by selecting the ITD with the smallest absolute value. In these embodiments, the processing circuit 1201 / 1301 selecting the ITD with the smallest absolute value includes selecting an ITD as the final ITD according to: TIFF0007680574000036.tif12170, where ITD 1 (m) is the final ITD, ITD 0 (m) is the first ITD estimate, ITD stab (m) is a stabilized ITD.
[0072] In another embodiment of the inventive concept, the processing circuitry 1201 / 1301 biases the ITD search in favor of ITDs closer to zero by selecting the final ITD from ITD candidates within a limited range around zero.
[0073] In a further embodiment of the inventive concept, the processing circuit 1201 / 1301 biases the ITD search in favor of ITDs closer to zero by applying cross-correlation weighting to assign greater weight to cross-correlation values closer to zero.
[0074] Returning to FIG. 14, in block 1409, in response to determining that the multi-channel audio signal is not a CC signal, the processing circuit 1201 / 1301 obtains a final ITD without prioritizing ITDs closer to zero.
[0075] In some other embodiments of the inventive concept, the processing circuit 1201 / 1301 applies stabilization to the selected ITD candidate to obtain a final ITD, the selected ITD candidate being selected from the at least one generated ITD candidate.
[0076] Various operations from the flowchart of Figure 14 may be optional with respect to some embodiments of the encoder / decoder and related methods. With respect to the method of example embodiment 1 (described below), for example, the operation of block 1409 of Figure 14 may be optional.
[0077] While the computing devices (e.g., UE, network node, host) described herein may include the depicted combination of hardware components, other embodiments may include computing devices having different combinations of components. It should be understood that these computing devices may include any suitable combination of hardware and / or software necessary to perform the tasks, features, functions, and methods disclosed herein. The determining, calculating, obtaining, or similar operations described herein may be performed by a processing circuit, which may process information, for example, by transforming the obtained information to other information, by comparing the obtained or transformed information to information stored in the network node, and / or by performing one or more operations based on the obtained or transformed information and as a result of said processing making a decision. Furthermore, while a component is shown as a single box located within a larger box or as a single box nested within multiple boxes, in reality the computing device may include multiple different physical components that make up a single illustrated component, and functionality may be divided among the separate components. For example, a communication interface may be configured to include any of the components described herein, and / or functionality of a component may be divided between a processing circuit and a communication interface. In another example, non-computationally intensive functions of any of such components may be implemented in software or firmware, and computationally intensive functions may be implemented in hardware.
[0078] In certain embodiments, some or all of the functionality described herein may be provided by a processing circuit executing instructions stored in a memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuit without executing instructions stored in a separate or distinct device-readable storage medium, such as in a hardwired manner. In any of these particular embodiments, the processing circuit may be configured to perform the above-mentioned functions, whether or not it executes instructions stored in a non-transitory computer-readable storage medium. Benefits provided by such functionality are not limited to the processing circuit alone or other components of the computing device, but are enjoyed by the computing device as a whole, and / or by end users and wireless networks in general.
[0079] Exemplary embodiments are described below. Embodiment 1. A method for identifying coincident microphone configurations CC and adapting inter-channel time difference ITD search in an encoder (110, 1000) or decoder (120, 1006), comprising: For each frame m of the multi-channel audio signal, Generating (1401) a cross-correlation of a channel pair of a multi-channel audio signal; determining (1403) a first ITD estimate based on the cross-correlation; Determining (1405) whether the multi-channel audio signal is a CC signal; In response to determining that the multi-channel audio signal is a CC signal, biasing (1407) the ITD search to favor ITDs closer to zero to obtain a final ITD; A method comprising: In response to determining that the multi-channel audio signal is not a CC signal, obtaining a final ITD without prioritizing ITDs closer to zero (1409). 2. The method of embodiment 1, further comprising: Embodiment 3. The method of embodiment 2, wherein obtaining a final ITD when the multi-channel audio signal is not a CC signal includes obtaining the final ITD by setting the final ITD to the first ITD estimate value. Embodiment 4. The method of embodiment 1 or 2, further comprising applying stabilization to the selected ITD candidates to obtain a final ITD. Embodiment 5. The method of embodiment 4, wherein applying stabilization further comprises generating at least one ITD candidate. Embodiment 6. A method as described in any one of embodiments 1 to 5, wherein biasing the ITD search to favor ITDs closer to zero to obtain a final ITD includes obtaining the final ITD by selecting an ITD having a smallest absolute value. Embodiment 7. Selecting the ITD with the smallest absolute value includes selecting an ITD as the final ITD according to: TIFF0007680574000037.tif12170, where ITD 1 (m) is the final ITD, ITD 0 (m) is the first ITD estimate, ITD stab (m) is a stabilized ITD; 7. The method of embodiment 6. Embodiment 8. The method of any one of embodiments 1 to 7, wherein biasing the ITD search to favor ITDs closer to zero includes selecting a final ITD from ITD candidates within a limited range around zero. Embodiment 9. A method as described in any one of embodiments 1 to 3, wherein biasing the ITD search to favor ITDs closer to zero to obtain a final ITD includes applying cross-correlation weighting to assign greater weight to cross-correlation values closer to zero. Embodiment 10. The method of any one of embodiments 1 to 9, wherein determining the first ITD estimate includes determining the first ITD estimate as an absolute maximum of the cross-correlation. Embodiment 11. Determining the first ITD estimate as an absolute maximum of the cross-correlation includes determining the absolute maximum according to: TIFF0007680574000038.tif8170, where ITD 0 (m) is the first ITD estimate, TIFF0007680574000039.tif7170 is the cross-correlation and τ is the time lag parameter, 11. The method of embodiment 10. Embodiment 12. The method of any one of embodiments 1 to 11, wherein the cross-correlation is generalized cross-correlation with phase transform (GCC-PHAT). Embodiment 13. Determining whether the multi-channel audio signal is a CC signal comprises: DETECTING ONE OF ANTISYMMETRIC AND SYMMETRIC PATTERNS OF CROSS-CORRELATION BETWEEN CHANNEL PAIRS OF A MULTI-CHANNEL AUDIO SIGNAL - Patent application The method according to any one of embodiments 1 to 12, comprising: Embodiment 14. Detecting an antisymmetric pattern in a component comprises detecting an antisymmetric pattern according to: TIFF0007680574000040.tif7170 where D(m) is the CC detection variable, TIFF0007680574000041.tif7170 is GCC-PHAT and ITD 0 (m) is the first ITD estimate; 14. The method of embodiment 13. Embodiment 15. Detecting one of the antisymmetric and symmetric patterns in the cross-correlation includes detecting the antisymmetric pattern according to at least one of the following: TIFF0007680574000042.tif79170 where D(m) is the CC detection variable, TIFF0007680574000043.tif6170 is GCC-PHAT, R is the search range, W defines the area around the first estimate of the matching ITD, and ITD 0 ’(m) is an ITD candidate limited to the search range [-R,R]. 14. The method of embodiment 13. Embodiment 16. Determining whether the multi-channel audio signal is a CC signal comprises: Calculating the CC detection variables (1501); determining 1503 whether a CC detection variable is above a threshold; determining (1505) that the multi-channel audio signal is a CC signal in response to determining that the CC detection variable is above the threshold; The method according to any one of embodiments 1 to 12, comprising: Embodiment 17. The method of embodiment 16, wherein determining whether the CC detection variable is above a threshold value comprises determining whether the absolute value of the CC detection variable is above a threshold value. Embodiment 18. The method of any one of embodiments 14 to 17, further comprising filtering the CC detection variables with low-pass filtering to stabilize the CC detection. Embodiment 19. The method of embodiment 18, wherein the low-pass filtering on the CC detection variables is adaptive depending on at least the output A(m) of the activity detector. Embodiment 20. Filtering the CC detection variables with low-pass filtering includes filtering with adaptive low-pass filtering according to: D LP (m) = α(m)D(m) + (1-α(m))D LP (m-1) TIFF0007680574000044.tif12170 where A(m) is the output of the activity detector and α high and α low are the filter coefficients, 20. The method of embodiment 19. Embodiment 21. An apparatus (110, 120, 1000, 1006), comprising: A processing circuit (1201, 1301); A memory (1205, 1305) coupled to the processing circuitry, which, when executed by the processing circuitry, causes the apparatus to: For each frame m of the multi-channel audio signal, Generating (1401) a cross-correlation of a channel pair of a multi-channel audio signal; determining (1403) a first ITD estimate based on the cross-correlation; determining (1405) whether the multi-channel audio signal is a CC signal; and In response to determining that the multi-channel audio signal is a CC signal, bias the ITD search to favor ITDs closer to zero to obtain a final ITD (1407). Memory and instructions An apparatus (110, 120, 1000, 1006) comprising:
[0046] Embodiment 22. In response to determining that the multi-channel audio signal is not a CC signal, obtaining a final ITD without prioritizing ITDs closer to zero (1409). 22. The apparatus (110, 120, 1000, 1006) of embodiment 21, further comprising: Embodiment 23. An apparatus (110, 120, 1000, 1006) as described in embodiment 22, wherein obtaining a final ITD when the multi-channel audio signal is not a CC signal includes obtaining a final ITD by setting the final ITD to the first ITD estimate value. Embodiment 24. An apparatus (110, 120, 1000, 1006) as described in embodiment 21 or 22, wherein the memory includes further instructions that, when executed by the processing circuit, cause the apparatus to apply stabilization to the selected ITD candidates to obtain a final ITD. Embodiment 25. An apparatus (110, 120, 1000, 1006) as described in embodiment 24, wherein applying stabilization further comprises generating at least one ITD candidate. Embodiment 26. An apparatus (110, 120, 1000, 1006) described in any one of embodiments 21 to 25, wherein biasing the ITD search to prioritize ITDs closer to zero to obtain a final ITD includes obtaining the final ITD by selecting an ITD having a smallest absolute value. Embodiment 27. Selecting the ITD with the smallest absolute value includes selecting an ITD as the final ITD according to: TIFF0007680574000045.tif12170, where ITD 1 (m) is the final ITD, ITD 0 (m) is the first ITD estimate, ITD stab (m) is a stabilized ITD; An apparatus (110, 120, 1000, 1006) as described in embodiment 26. Embodiment 28. An apparatus (110, 120, 1000, 1006) described in any one of embodiments 21 to 27, wherein biasing the ITD search to favor ITDs closer to zero includes selecting a final ITD from ITD candidates within a limited range around zero. Embodiment 29. An apparatus (110, 120, 1000, 1006) described in any one of embodiments 21 to 27, wherein biasing the ITD search to favor ITDs closer to zero to obtain a final ITD includes applying cross-correlation weighting to assign greater weight to cross-correlation values closer to zero. Embodiment 30. An apparatus (110, 120, 1000, 1006) described in any one of embodiments 21 to 29, wherein determining the first ITD estimate value includes determining the first ITD estimate value as an absolute maximum of the cross-correlation. Embodiment 31. Determining the first ITD estimate as an absolute maximum of the cross-correlation includes determining the absolute maximum according to: TIFF0007680574000046.tif8170, where ITD 0 (m) is the first ITD estimate, TIFF0007680574000047.tif7170 is the cross-correlation and τ is the time lag parameter, An apparatus (110, 120, 1000, 1006) as described in embodiment 30. Embodiment 32. An apparatus (110, 120, 1000, 1006) according to any one of embodiments 21 to 31, wherein the cross-correlation is generalized cross-correlation with phase transform (GCC-PHAT). Embodiment 33. Determining whether the multi-channel audio signal is a CC signal comprises: DETECTING ONE OF ANTISYMMETRIC AND SYMMETRIC PATTERNS OF CROSS-CORRELATION BETWEEN CHANNEL PAIRS OF A MULTI-CHANNEL AUDIO SIGNAL - Patent application The device (110, 120, 1000, 1006) according to any one of embodiments 21 to 31, comprising: Embodiment 34. Detecting an antisymmetric pattern in a component includes detecting an antisymmetric pattern according to: TIFF0007680574000048.tif7170 where D(m) is the CC detection variable, TIFF0007680574000049.tif7170 is GCC-PHAT and ITD 0 (m) is the first ITD estimate; The apparatus (110, 120, 1000, 1006) according to embodiment 33. Embodiment 35. Detecting one of the antisymmetric and symmetric patterns in the cross-correlation includes detecting the antisymmetric pattern according to at least one of the following: TIFF0007680574000050.tif79170 where D(m) is the CC detection variable, TIFF0007680574000051.tif6170 is GCC-PHAT, R is the search range, W defines the area around the first estimate of the matching ITD, and ITD 0 ’ (m) is an ITD candidate limited to the search range [-R,R]. An apparatus (110, 120, 1000, 1006) as described in embodiment 35. Embodiment 36. Determining whether the multi-channel audio signal is a CC signal comprises: Calculating the CC detection variables (1501); determining 1503 whether a CC detection variable is above a threshold; determining (1505) that the multi-channel audio signal is a CC signal in response to determining that the CC detection variable is above the threshold; The device (110, 120, 1000, 1006) according to any one of embodiments 21 to 32, comprising: Embodiment 37. An apparatus (110, 120, 1000, 1006) as described in embodiment 33, wherein determining whether the CC detection variable is above a threshold value includes determining whether an absolute value of the CC detection variable is above a threshold value. Embodiment 38. An apparatus (110, 120, 1000, 1006) described in any one of embodiments 34 to 37, wherein the memory includes further instructions that, when executed by the processing circuit, cause the apparatus to low-pass filter the CC detection variable to stabilize the CC detection. Embodiment 39. An apparatus (110, 120, 1000, 1006) as described in embodiment 38, wherein the low-pass filtering on the CC detection variable is adaptive depending on at least the output A(m) of the activity detector. Embodiment 40. Filtering the CC detection variables with low-pass filtering includes filtering with adaptive low-pass filtering according to: D LP (m) = α(m)D(m) + (1-α(m))D LP (m-1) TIFF0007680574000052.tif12170 where A(m) is the output of the activity detector and α high and α low are the filter coefficients, An apparatus (110, 120, 1000, 1006) as described in embodiment 39. For each frame m of a multi-channel audio signal, generating (1401) a cross-correlation of a channel pair of a multi-channel audio signal; determining (1403) a first ITD estimate based on the cross-correlation; determining (1405) whether the multi-channel audio signal is a CC signal; and In response to determining that the multi-channel audio signal is a CC signal, bias (1407) the ITD search to favor ITDs closer to zero to obtain a final ITD. The apparatus (110, 120, 1000, 1006) is adapted to Embodiment 42. The apparatus (110, 120, 1000, 1006) according to embodiment 41, adapted to perform according to embodiments 2 to 20. Embodiment 43. A computer program comprising a program code executed by a processing circuit (1201 / 1301) of an apparatus (110, 120, 1000, 1006), the execution of which causes the apparatus (110, 120, 1000, 1006) to: For each frame m of the multi-channel audio signal, Generating (1401) a cross-correlation of a channel pair of a multi-channel audio signal; determining (1403) a first ITD estimate based on the cross-correlation; determining (1405) whether the multi-channel audio signal is a CC signal; and In response to determining that the multi-channel audio signal is a CC signal, bias the ITD search to favor ITDs closer to zero to obtain a final ITD (1407). Computer program. Embodiment 44. A computer program as described in embodiment 43, wherein the program code includes further program code for causing an apparatus (110, 120, 1000, 1006) to perform according to any one of embodiments 2 to 20. Embodiment 45. A computer program product including a non-transitory storage medium including a program code executed by a processing circuit (1201 / 1301) of an apparatus (110, 120, 1000, 1006), the execution of which causes the apparatus (110, 120, 1000, 1006) to: For each frame m of the multi-channel audio signal, Generating (1401) a cross-correlation of a channel pair of a multi-channel audio signal; determining (1403) a first ITD estimate based on the cross-correlation; determining (1405) whether the multi-channel audio signal is a CC signal; and In response to determining that the multi-channel audio signal is a CC signal, bias the ITD search to favor ITDs closer to zero to obtain a final ITD (1407). Computer program products. Embodiment 46. A computer program as described in embodiment 45, wherein the non-transitory storage medium comprises further program code for causing an apparatus (110, 120, 1000, 1006) to perform according to any one of embodiments 2 to 20.
[0080] Explanations of various abbreviations / acronyms used in this disclosure are provided below. Abbreviation Explanation CC coincident microphone configuration ILD Interaural level difference or interchannel level difference ITD Interaural or Interchannel Time Difference IC or IACC Interaural coherence or correlation or interchannel coherence or correlation GCC general cross-correlation GCC-PHAT Generalized Cross-Correlation by Phase Transform
Claims
1. A method for identifying coincident microphone configurations (CC) and adapting an inter-channel time difference (ITD) search, the method being performed by a processor circuit included in an encoder (110, 1000) or a processor circuit included in a decoder (120, 1006), For each frame m of the multi-channel audio signal, generating cross-correlations of pairs of channels of the multi-channel audio signal (1401); determining 1403 a first ITD estimate based on the cross-correlation; determining (1405) whether the multi-channel audio signal is a CC signal; In response to determining that the multi-channel audio signal is a CC signal, biasing (1407) the ITD search to favor ITDs closer to zero to obtain a final ITD. A method comprising:
2. In response to determining that the multi-channel audio signal is not a CC signal, obtaining the final ITD without prioritizing ITDs closer to zero (1409). The method of claim 1 further comprising:
3. The method of claim 2 , wherein obtaining the final ITD when the multi-channel audio signal is not a CC signal comprises obtaining the final ITD by setting the final ITD to the first ITD estimate value.
4. The method of claim 1 or 2, further comprising applying a stabilization to the ITD to obtain the final ITD.
5. The method of claim 4 , wherein applying stabilization further comprises generating at least one ITD candidate.
6. 6. The method of claim 1, wherein biasing the ITD search to favor ITDs closer to zero to obtain the final ITD comprises obtaining the final ITD by selecting an ITD having a smallest absolute value.
7. Selecting the ITD having the smallest absolute value includes selecting the ITD as the final ITD according to: Here, I.T.D. 1 (m) is the final ITD, ITD 0 (m) is the first ITD estimate, ITD stab (m) is a stabilized ITD; The method according to claim 6.
8. 8. The method of claim 1, wherein biasing the ITD search to favor ITDs closer to zero comprises selecting the final ITD from ITD candidates within a limited range around zero.
9. 4. The method of claim 1, wherein biasing the ITD search to favor ITDs closer to zero to obtain the final ITD comprises applying cross-correlation weighting to assign greater weight to cross-correlation values closer to zero.
10. The method of any one of claims 1 to 9, wherein determining the first ITD estimate comprises determining the first ITD estimate as an absolute maximum of the cross-correlation.
11. Determining the first ITD estimate as the absolute maximum of the cross-correlation includes determining the absolute maximum according to: Here, I.T.D. 0 (m) is the first ITD estimate; is the cross-correlation and τ is a time lag parameter. The method of claim 10.
12. The method according to any one of claims 1 to 11, wherein the cross-correlation is a generalized cross-correlation with phase transform (GCC-PHAT).
13. determining whether the multi-channel audio signal is a CC signal, detecting one of an antisymmetric and a symmetric pattern of the cross-correlation in the channel pairs of the multi-channel audio signal; The method according to any one of claims 1 to 12, comprising:
14. Detecting the antisymmetric pattern in a component comprises detecting the antisymmetric pattern according to: where D(m) is the CC detection variable, is GCC-PHAT, and ITD 0 (m) is the first ITD estimate; The method of claim 13.
15. Detecting one of an antisymmetric pattern and a symmetric pattern in the cross-correlation includes detecting the antisymmetric pattern according to at least one of the following: where D(m) is the CC detection variable, is the GCC-PHAT, R is the search range, and W defines the area around the first ITD estimate that matches, and ITD 0 '(m) is an ITD candidate limited to the search range [-R, R]; The method of claim 13.
16. determining whether the multi-channel audio signal is a CC signal, Calculating CC detection variables (1501); determining 1503 whether the CC detection variable is above a threshold; determining (1505) that the multi-channel audio signal is a CC signal in response to determining that the CC detection variable is above the threshold; The method according to any one of claims 1 to 12, comprising:
17. The method of claim 16 , wherein determining whether the CC detection variable is above the threshold comprises determining whether an absolute value of the CC detection variable is above the threshold.
18. The method of any one of claims 14 to 17, further comprising filtering the CC detection variables with low-pass filtering to stabilize CC detection.
19. 20. The method of claim 18, wherein the low-pass filtering on the CC detection variable is adaptive depending on at least an output A(m) of an activity detector.
20. Filtering the CC detection variables with low pass filtering comprises adaptive low pass filtering according to: DLP(m)=α(m)D(m)+(1-α(m))DLP(m-1) where A(m) is the output of the activity detector, and α high and α low are the filter coefficients, 20. The method of claim 19.
21. For each frame m of the multi-channel audio signal, generating (1401) a cross-correlation of a pair of channels of the multi-channel audio signal; determining 1403 a first ITD estimate based on the cross-correlation; determining 1405 whether the multi-channel audio signal is a CC signal; and In response to determining that the multi-channel audio signal is a CC signal, bias (1407) an ITD search to favor ITDs closer to zero to obtain a final ITD. The apparatus (110, 120, 1000, 1006) is adapted to:
22. An apparatus (110, 120, 1000, 1006) according to claim 21, adapted to carry out the method according to any one of claims 2 to 20.
Citation Information
Patent Citations
Multichannel audio encoder and method for encoding multichannel audio signals
JP2015514234A
Apparatus and method for estimating inter-channel time difference
JP2019502966A
Method and apparatus for increasing stability of inter-channel time difference parameter
JP2020065283A
Determining the Inter-Channel Time Difference of a Multi-Channel Audio Signal
US20180301154A1
Method and appparatus for increasin stability of an inter-channel time difference parameter
US20200286495A1