Efficient Time Delay Synthesis
The method efficiently adjusts inter-channel time differences by sharing transition lengths between channels, addressing computational complexity and artifacts in spatial audio rendering, ensuring smooth transitions across various playback systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2026-03-12
AI Technical Summary
Existing spatial audio rendering methods face challenges in efficiently transitioning inter-channel time differences (ITDs) across zero boundaries, leading to computational complexity and artifacts during transitions, especially when adapting audio for different playback systems.
A method and apparatus that efficiently adjust inter-channel time differences (ITDs) by sharing the total transition length between output channels, using time shifting operations based on the current and previous frame ITDs, minimizing computational complexity and transition artifacts.
The method maintains consistent speed and low computational complexity while updating time delays, reducing transition artifacts and ensuring smooth audio transitions across different playback systems.
Smart Images

Figure 2026508733000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to communications, and more particularly to communication methods and associated devices and nodes that support audio encoding and decoding. [Background technology]
[0002] Spatial audio is a representation of a sound field that immerses the listener. There are several formats of spatial audio. The most common is the stereo format, in which the sound field is rendered through either two speakers or a set of headphones. In scenarios where playback is over a larger set of loudspeakers, such as 5.1, 7.1+4, or 22.2, spatial audio is often referred to as multichannel audio. There are also spatial audio formats that do not rely on the layout of the loudspeaker system but rather represent the sound field itself. Such representations include wave field synthesis (WFS), in which the sound field is captured by an array of microphones and reproduced symmetrically by an array of loudspeakers. Another popular format is ambisonics, which relies on spherical harmonics captured using a compact microphone array. Ambisonics have recently become more popular because they are well-suited for listener-centric rendering, such as virtual reality (VR) and augmented reality (AR) audio rendering, and because they are inherently suitable for rotation. Ambisonics can also be combined with 360 video capture for reconstruction of the experienced scene.
[0003] Multichannel audio formats can be played directly on the loudspeaker setup for which they were designed. However, without a loudspeaker setup, the audio cannot be played without adapting the audio. This adaptation is often referred to as rendering spatial audio for the playback system. If you have a 2.2 multichannel signal or an Ambisonics signal, it can be rendered for playback on, for example, a 5.1 system or a set of headphones. When rendering for headphones, the audio reaching the ears is generally modeled using head-related filters (HRFs) or head-related transfer functions (HRTFs). These filters model the direction of arrival (DoA) of the sound source, so the listener perceives the sound as coming from this direction. This is achieved by a time difference caused by spectral coloration, level differences between the two ears, and differences in path lengths to the left and right ears. This time difference is often called the interaural time difference or interchannel time difference (ITD). The time difference between channels can be created by filtering one or both channels with a Dirac pulse, as follows: h=δ(tt Δ ) However, transitions between different time shifts, eg for moving sources, need to be handled. Summary of the Invention
[0004] When modeling the HRF, spectral coloration can be performed using filters, and time differences can be generated by time shifting. The present disclosure applies time shifting in an efficient manner when crossing the zero boundary due to the shift.
[0005] When changing the sign of the time delay parameter, a shift operation needs to be performed on both output channels. To limit the complexity of the shift operation, the total transition length is shared between the channels in proportion to the size of the shift on each side of the zero point.
[0006] According to a first aspect, a method for adjusting timing of output audio signals to achieve a desired inter-channel time difference (ITD) between the output audio signals is presented. The method includes receiving a current ITD value and an audio frame, and determining transition times t1, t2 for implementing a time shift to apply to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame. The time shift within the determined transition times t1, t2 is applied in generating the first output signal and the second output signal. The method further includes storing at least a portion of the audio frame for use in synthesizing the ITD in a subsequent frame.
[0007] According to a second aspect, an apparatus for adjusting timing of output audio signals to achieve a desired inter-channel time difference (ITD) between the output audio signals is presented. The apparatus is adapted to receive a current ITD value and an audio frame, and determine transition times t1, t2 for implementing a time shift to apply to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame. The apparatus is adapted to apply the time shift within the determined transition times t1, t2 in generating the first output signal and the second output signal.
[0008] According to a third aspect, an apparatus is presented, comprising: a processing circuit; and a memory coupled to the processing circuit, the memory including instructions that, when executed by the processing circuit, cause the apparatus to perform operations, the operations including receiving a current inter-channel time difference (ITD) value and an audio frame; determining transition times t1, t2 for performing a time shift to apply to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame; and applying the time shift within the determined transition times t1, t2 in generating the first output signal and the second output signal.
[0009] According to a fourth aspect, there is presented a computer program comprising program code that is executed by processing circuitry of an apparatus, the execution of which causes the apparatus to perform the operations of the first aspect.
[0010] According to a fifth aspect, there is presented a computer program product comprising a non-transitory storage medium containing program code that is executed by processing circuitry of an apparatus, the execution of the program code causing the apparatus to perform the operations of the first aspect.
[0011] Some embodiments may provide one or more of the following technical advantages: The speed of adjustment is kept consistent for switching across the zero boundary, and computational complexity is kept low. The method aims to create two channels with a time delay, which can be updated for each frame. The updates to the time delay can be done with minimal transition artifacts.
[0012] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate several non-limiting embodiments of the inventive concepts. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a block diagram illustrating an environment in which various embodiments of the present disclosure may be implemented. [Figure 2] FIG. 2 is a block diagram of an audio object renderer according to some embodiments of the present disclosure. [Figure 3] FIG. 2 is a block diagram of a parametric stereo decoder according to some embodiments of the present disclosure. [Figure 4] 10 is a flowchart illustrating the operation of an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 5] FIG. 1 is a block diagram of an ITD combiner according to some embodiments of the present disclosure. [Figure 6] 10 is a flowchart illustrating the operation of an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 7] FIG. 10 is a diagram of a buffer operation performed by an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 8] 10 is a flowchart illustrating the operation of an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 9] FIG. 10 is a diagram of a buffer operation performed by an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 10] FIG. 10 is a diagram of a buffer operation performed by an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 11] FIG. 10 is a diagram of a buffer operation performed by an ITD combiner in accordance with some embodiments of the present disclosure. [Figure 12] FIG. 1 is a diagram of a sinc resampling function for handling compressing and expanding segments of a signal. [Figure 13] FIG. 2 is a block diagram of an audio object renderer according to some embodiments. [Figure 14] FIG. 1 is a block diagram of a host computer in communication with an encoder and / or decoder, according to some embodiments. [Figure 15] FIG. 1 is a block diagram of a virtualized environment, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0014] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Examples of embodiments of the inventive concepts are shown. The embodiments are provided as examples to convey the scope of the subject matter to those skilled in the art. However, the inventive concepts may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concepts to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. It may be implicitly assumed that a component from one embodiment is present / used in another embodiment.
[0015] FIG. 1 illustrates an example of an operating environment in which various embodiments of the present disclosure may be implemented. Referring to FIG. 1 , in the exemplary operating environment 100, an encoder 102 receives data, such as an audio file, to be encoded from an entity, such as a host 106, and / or from storage 108 over a network 104. In some embodiments, the host 106 may communicate directly with the encoder 102. The encoder 102 either encodes the audio file as described herein and stores the encoded audio file in storage 108, or transmits the encoded audio file over a network 110 to a decoder 112 having an audio object renderer 114. The audio object renderer 114 in the decoder 112 renders the decoded audio file and transmits the rendered decoded audio file to an audio player 116 for playback. For example, the audio player 116 may play the rendered decoded audio file for a spatial audio presentation, such as a virtual reality conference or a computer game. The audio player 116 may be or be included within a user device, terminal, mobile phone, or the like. In other embodiments, the host 106 may transmit the encoded audio file to the audio object renderer 114 over the network 110. In some embodiments, the audio object renderer 114 may be a standalone device between the decoder 112 and the audio player 116 in scenarios where the decoded audio does not directly "fit" the audio player 116, such as rendering the decoded audio file for headphones.
[0016] As shown previously, when changing the sign of the time delay parameter, a shift operation needs to be performed on both output channels. To limit the complexity of the shift operation, the total transition length is shared between the channels. The sharing is done in proportion to the size of the shift on each side of the zero point according to the following formula: TIFF2026508733000002.tif22170 where ITD(m) and ITD(m-1) are the inter-channel time differences, m is the subframe index, and t tot is the total transition length, and t1 and t2 are transition lengths for performing a time expansion or compression operation. A time expansion or compression operation is sometimes called a time shifting operation or a resampling operation. The transition length may be expressed in seconds or in number of samples for a discretely sampled audio signal. The transition length is sometimes called the transition time.
[0017] The above formula can be simplified (with integer rounding) to: TIFF2026508733000003.tif16170
[0018] In some embodiments of the present disclosure, the method operates in an ITD synthesizer implemented within the audio object renderer 114, as shown in Figure 2. In other embodiments, the ITD synthesizer may be implemented within a parametric stereo decoder. The method operates on segments of audio, called frames, where each frame or subframe m consists of N samples. x(m,n), n=0,1,2,...,N-1
[0019] A frame may represent a segment of audio, for example from a decoded audio object, a mono downmix channel in a parametric stereo decoder, or an input channel to an audio object renderer. Here, the audio object renderer 114 receives an audio object comprising an audio signal and position metadata representing the position of the audio object. The position can be absolute or relative to the listener position. The position metadata is input to the HR filter module 210, which provides ITD values and a set of HR filters for the left and right channels. The time delay parameter for frame m is in the range ITD(m) = [-ITD MAX,ITD MAX ] is an integer within
[0020] If the frame comes from a parametric stereo decoder, ITD(m) can be found by analyzing the input channels to the stereo encoder. Preferably, the input channels are aligned by compensating for ITD(m) before creating the downmix channels. The downmix channels are encoded with stereo parameters including ITD(m) and are decoded and reconstructed in the parametric stereo decoder. The parametric stereo decoder reconstructs the downmix signal, the stereo parameters including at least the reconstruction of ITD(m), and combines the two output channels with the corresponding ITD(m).
[0021] The audio and positional metadata may come from, for example, an audio object decoder or may be generated by a 3D audio engine for spatial audio presentation, such as for virtual reality conferencing or computer games. The HR filter module 210 may be, for example, a database of stored filters and ITD values, or it may be a model-based database that generates filters and ITD values for given positional data. The ITD values are input to the ITD synthesizer 220, which generates two output signals based on the input audio frame, where the output signals have the desired ITDs. The two channels are filtered through a left filter 230 and a right filter 240 to generate combined left and right channels. The audio object may be summed with one or more additional objects. The output left and right channels may be forwarded to an audio device for playback.
[0022] This is illustrated in the flowchart of Fig. 4 of the operation of a method 400 that the ITD synthesizer 220 implements in some embodiments. In block 401, the ITD synthesizer 220 receives a current ITD and an audio frame, where each frame m comprises N samples. Upon receipt of the audio frame, or at any time during the processing of the frame, the ITD synthesizer 220 may store at least a portion of the current input audio frame in block 403 to be used for processing the ITD in a subsequent audio frame.
[0023] In block 405, the ITD synthesizer 220 determines transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on the sign of the inter-channel time difference (ITD) of the current audio frame and the ITD of the previous audio frame. The ITD, in the context of FIG. 2, comes from the HRF filter module 210. In other embodiments in which the ITD synthesizer is implemented as part of a parametric stereo decoder, the ITD is reconstructed from the bitstream. Here, the time shift refers to an operation to smoothly transition to the target ITD.
[0024] In block 407, the ITD combiner 220 applies a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1.
[0025] Before describing the ITD synthesizer 220 in further detail, FIG. 3 illustrates an embodiment in which the ITD synthesizer may be implemented within a parametric stereo decoder 300. In the parametric stereo decoder 300, stereo parameters, including ITD parameters, are decoded by a parameter decoder 310, and optionally, a reconstructed residual signal is produced by a residual decoder 320. A downmix decoder 330 is configured to decode and reconstruct the encoded downmix signal to output a reconstructed downmix signal to which a time shift may be applied. The reconstructed downmix, reconstructed stereo parameters, and optionally, the reconstructed residual signal are fed to a stereo upmixer 340 to produce a reconstructed stereo signal. The ITD synthesizer 220 is part of the stereo upmixer 340.
[0026] The ITD synthesizer 220 performs the operations described in more detail in Figure 5 and illustrated in Figure 6. Referring to Figure 6, in step 601, a processing buffer 510 is populated using the current input audio frame x(m,n) and a signal memory 520. The length of the memory is N mem is at least the maximum time shift ITD MAX and the look-back / look-ahead memory rs required for the resampling function. LA It should be a harmony with. N mem =ITD MAX +rs LA
[0027] The processing buffer 510 is shown in FIG. 7, where the middle plot represents the processing buffer x buf (n) indicates x buf(0) corresponds to the first value of the current frame. The input frame is also provided to the signal memory 520 to be used in the next frame. In step 603, the ITD value ITD(m) of the current frame, along with ITD(m-1), is input from the ITD memory 540 to the transition length calculator 530. The transition length is calculated based on ITD(m) and ITD(m-1). First, the total transition length is calculated as the frame length N plus the look-ahead rs needed by the resamplers 570, 580. LA It is calculated based on ITD(m). <rs LA If , a small portion of the processing buffer must be kept to accommodate the look-ahead without introducing processing delays due to resampling. The transition can be divided into three parts, t1, t2, and t3. The transition length t3 can be seen as the buffer length to avoid reading from the middle of memory in the resampling operation. The resampler must keep at least rs at the end of the buffer to perform the resampling or filtering operation. LA samples must be left. The transition length t3 is calculated according to: t3=max(0,rs LA -|ITD(m)|)
[0028] In that case, the total transition time t tot is as follows: t tot =N max -t3 where N max is the maximum allowable transition length. N max is N max = N, which means the full frame time is allowed for the transition. However, if N is large, N can be set to 0 to achieve faster transitions and potentially lower complexity. max It may be desirable to limit the maximum allowable transition length using ≦N.
[0029] 8 shows the operations performed by the ITD combiner 220 in determining the total transition length. In block 801, the ITD combiner 220 determines the frame length N of frame m and the look-ahead memory rs required for resampling. LA and calculate the total transition length based on the total transition length, and the total transition is divided into two parts comprising t1 and t2.
[0030] In block 803, the ITD combiner 220 stores the look-ahead memory rs LA In block 805, the ITD combiner 220 determines the transition time t3 based on the maximum allowable transition length N max and transition time t3 to determine the total transition length.
[0031] The subsequent time shift operations can be divided into two groups. 1. ITD(m) and ITD(m-1) have the same sign, or one of them is 0. 2. The sign of ITD is non-zero and changing, i.e., ITD(m)·ITD(m-1)<0.
[0032] Case 1 - The signs of the ITDs are the same or one of them is 0
[0033] If ITD(m) and ITD(m-1) have the same sign or one of them is 0, the shift can be handled by processing only one of the channels, which means processing step 605, where resampler 570 adjusts processing buffer 510 to populate output buffer A 550. This can be achieved by allocating the entire transition length to t1 and setting t2 to 0, i.e. TIFF2026508733000004.tif34170 where n1, n2, and n3 indicate the start index of each time shift segment, assuming that the current input subframe starts at n=0, and L in,1 , L in,2is the length of resampling segments 1 and 2. In this case, the input signal portion of the processing buffer is simply copied to the output buffer B560.
[0034] When the time delay of the current frame is the same as the previous frame, i.e., ITD(m) = ITD(m-1), an output time delay synthesis is created by pointing to the corresponding starting point in the processing buffer. The sign of ITD(m) determines which of the two channels the delay should be applied to. For example, a positive ITD(m) may indicate that the left channel of a stereo pair is ahead of the right channel; in that case, the right channel should be delayed and the left channel should be output without delay. This situation is shown in FIG. 7. In this case, the resampling operation on output buffer A550 has the same input and output lengths and is equivalent to a copy operation.
[0035] When the absolute value of the time delay of the current frame is greater than that of the previous frame, i.e., |ITD(m)| > |ITD(m-1)|, a transition is generated to allow a smooth transition between delay values. This situation is illustrated in Figure 9. The transition is generated by dividing the length of the frame by x of length t1 + |ITD(m-1)| - |ITD(m)|. buf (n), n = ITD(m-1),...,N-1-ITD(m), to an output frame of length t1. If the transition time t3 is greater than 0 (i.e., t3 > 0), the last t3 samples of the output channel are simply copied from the processing buffer and arrive with a delay of |ITD(m)|, x buf (n), n=N-1-t3-|ITD(m)|,...N-1-|ITD(m)|. Here, it may also be noted that if |ITD(m-1)|=|ITD(m)|, the input length is the same as the output length, and resampling will be equivalent to a copy operation.
[0036] When the absolute value of the time delay is reduced, i.e., |ITD(m)|<|ITD(m-1)|, the formula for the input frame length remains the same. However, the length t1 + |ITD(m-1)| - |ITD(m)| will now be greater than the resulting length t1, and resampling corresponds to shortening the length of the frame. This is shown in Figure 10. In this example, ITD(m) = 0, which means that t3 = rs LA means that the last t3 samples of output buffer A are copied from processing buffer 510.
[0037] Case 2 - ITDs have different signs and are non-zero If the sign of ITD is changing, i.e., ITD(m) ITD(m-1)<0, then shift operations must be performed on both channels. In this case, the total transition length t tot is split into two parts according to the following: TIFF2026508733000005.tif40170 where [·] indicates rounding to the nearest integer. Next, output buffers A and B are assembled using time-shifting the signal using transition times t1 and t2, as shown in Figure 11. First, a signal of length L in,1 The resampling starts from n1 in the output buffer A up to the first t1 samples. sf −t1 samples are populated by copying the remainder of the Processing Buffer to Output Buffer A. Then the first sample of the Processing Buffer, starting from index 0, is copied to the first t1 samples of Output Buffer B. The next t2 samples of Output Buffer B are copied to a buffer of length L in,2 Finally, the last t3 samples of the processing buffer are copied to the output buffer B. The last L ITDmemsamples are stored in memory for processing the next subframe. Output buffers A and B are allocated to output channels 0 and 1 for left and right HRIR filtering, respectively. The allocation of output channels is based on the signs of ITD(m) and ITD(m-1) as follows: TIFF2026508733000006.tif11170 where ∧ denotes logical AND and ∨ denotes inclusive OR. The resampling operation is implemented using a polyphase filter with a sinc function from a lookup table. The benefit of splitting the transition length between the two channels is that the computational complexity of the resampler is proportional to the transition length, and the total transition length is reduced to t tot By constraining , the total complexity is kept under a certain bound. Furthermore, the transition rates on the left and right channels are kept generally the same (since integer rounding is generally performed). This keeps transition artifacts to a minimum, since artifacts from transitions are lower for lower transition rates.
[0038] The resampling and copying operations can also be described with reference to the buffer index as follows: When the ITD codes are different, the first transition should move from ITD(m-1) to 0 on output buffer A 550, and then shift from 0 to ITD(m) on buffer B 560. An example of this process is shown in Figure 11. In step 605, resampler 570 fills the first t1 samples of output buffer A 550. This is done by filling samples x of length t1 + |ITD(m-1)|. buf (n), n=-|ITD(m-1)|,...,t1-1 to fit samples n=0,...,t1-1 in output buffer A550. In the same step, samples x buf (n), n=0,...,t1-1 are copied to the corresponding index n=0,...,t1-1 in output buffer B 560. In step 607, resampler 580 copies samples x buf(n), n=t1,...,t1+t2-1-|ITD(m)|, is adapted to fit the length t2 samples n=t1,...,t1+t2-1 in output buffer B 560. If t3>0, the last t3 samples, n=N-1-t3-|ITD(m)|,...,N-1-|ITD(m)|, are copied from processing buffer 510 to output buffer B 560 in step 609. Note that output buffer 550 and output buffer 560 may have different alignments such that the indices are shifted depending on the processing buffers. In that case, the indices above relating to output buffers 550 and 560 would be shifted by that amount, but segments would still be added in the same manner as described herein. For example, output buffer A 550 in FIGS. 7 and 9 could be offset by -|ITD(m)| to be aligned with the processing buffer index. Furthermore, resampler 570 and resampler 580 may be implemented using the same resampling function operating on different inputs.
[0039] In step 611, which is common to both Case 1 and Case 2 above, output buffer A 550 and output buffer B 560 are allocated to output 0 and output 1. For intermediate buffers A and B, A always corresponds to the channel that currently has a non-0 ITD and is delayed, while buffer B corresponds to the channel with a 0 ITD. The intermediate buffers use these assumptions to simplify processing, and output allocation is a simple step that can be performed last to allocate processed buffers to the correct output channels. The allocation of output buffers depends on the signs of ITD(m-1) and ITD(m), according to the following pseudocode: ● IF ITD(m-1)=0 ○ IF ITD(m)>0 Output buffer A550->Output 1, Output buffer B560->Output 0 ○ ELSE Output buffer A550->Output 0, Output buffer B560->Output 1 ● ELSE ○ IF ITD(m-1)>0 Output buffer A550->Output 1, Output buffer B560->Output 0 ○ ELSE Output buffer 550 -> output 0, output buffer 560 -> output 1 The allocation of the output buffer can also be simplified to: TIFF2026508733000007.tif11170 where ∧ denotes logical AND, ∨ denotes inclusive OR, (A,B)->(1,0) means allocating output buffer A 550 to output 1 and output buffer B 560 to output 0, and (A,B)->(1,0) means allocating output buffer A 550 to output 0 and output buffer B 560 to output 1. Output buffers 0 and 1 may correspond to binaural channels left and right, respectively. The numbering may also be done differently, for example, output buffers 1 and 2.
[0040] In other words, if ITD(m-1) is 0, the current ITD, ITD(m), is used to determine which buffer to delay. If ITD(m) is positive, output 0 precedes output 1, and output 1 should be delayed. If ITD(m) is negative, output 1 precedes output 0, and output 0 should be delayed. If ITD(m-1) is not 0, the previous ITD, ITD(m-1), determines which buffer to shift first. If ITD(m-1) is positive, output 1 is shifted first in buffer A 550, followed by shifting output 0 in buffer B 560. If ITD(m-1) is negative, output 0 is shifted first in buffer A 550, followed by shifting output 1 in buffer B 560. Note that the sign convention of ITD(m) can be reversed, in which case output 0 and output 1 would exchange places as described above.
[0041] Note that step 611 may occur before step 605 by having already allocated Output Buffer A 550 and Output Buffer B 560 to designated Outputs 1 and 2 so that the outputs are populated during steps 605-611. In one embodiment, Output 0 and Output 1 may correspond to the left and right channels, respectively.
[0042] Resampling using the sinc function
[0043] The described method relies on a resampling function to handle the compression and expansion of segments of the signal. This can be achieved using a sinc resampling function, as shown in Figure 12. in of input signal y(n) and output length L out Assuming that, the input signal follows these steps: TIFF2026508733000008.tif8170 can be resampled. For each n, Calculate TIFF2026508733000009.tif7170, Here, [·] represents a round-down operation.
[0044] Because sinc functions are computationally complex to calculate, it is sometimes desirable to store the sinc function in a table with a predefined resolution. For example, R = 1, which means that there are 64 samples between zero crossings of the sinc function. sinc A resolution of ∑ = 64 may be suitable (see Figure 12). The corresponding index in the sinc table is then found as follows: TIFF2026508733000010.tif5170
[0045] The output value z(k) is the sum It can be found by TIFF2026508733000011.tif11170, where: TIFF2026508733000012.tif28170
[0046] While the above embodiments have been described using an audio object renderer (e.g., a decoder), it should be noted that the various embodiments described above may also be performed in an encoder, with the shifted outputs (i.e., output 0 and output 1) being shifted in the encoder rather than in the audio object renderer.
[0047] 13 illustrates an audio object renderer 114 (e.g., a decoder) according to some embodiments, where the audio object renderer 114 is implemented as a standalone device. As used herein, an audio object renderer refers to a device capable of, configured to, and / or operable to decode encoded objects and to communicate with a network node, encoder, and / or decoder. Examples of audio object renderers include, but are not limited to, smartphones, mobile phones, cell phones, voice-over-IP (VoIP) phones, wireless local loop phones, desktop computers, personal digital assistants (PDAs), wireless cameras, gaming consoles or devices, storage devices, playback appliances, wearable terminal devices, wireless endpoints, mobile stations, tablets, laptop computers, laptop embedded appliances (LEEs), laptop mounted appliances (LMEs), smart devices, wireless customer premises equipment (CPEs), vehicle-mounted or vehicle-embedded / integrated wireless devices, etc.
[0048] The audio object renderer may support device-to-device (D2D) communications, for example, by implementing 3GPP standards for sidelink communications, dedicated short-range communications (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-everything (V2X). In other examples, a decoder may not necessarily have a user in the sense of a human user who owns and / or operates the associated device.
[0049] The audio object renderer 114 includes a processing circuit 1302 operably coupled via a bus 1304 to an input / output interface 1306, a power supply 1308, a memory 1310, a communication interface 1312, and / or any other components, or any combination thereof. Some decoders may utilize all or a subset of the components shown in FIG. 13. The level of integration between components may vary from decoder to decoder. Additionally, some decoders may include multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0050] The processing circuit 1302 is configured to process instructions and data and may be configured to implement any sequential state machine operable to execute instructions stored in the memory 1310 as a machine-readable computer program. The processing circuit 1302 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), programmable logic together with appropriate firmware, one or more stored computer programs such as a microprocessor or digital signal processor (DSP) together with appropriate software, a general-purpose processor, or any combination of the above. For example, the processing circuit 1302 may include multiple central processing units (CPUs).
[0051] In this example, the input / output interface 1306 may be configured to provide one or more interfaces to an input device, an output device, or one or more input and / or output devices. Examples of output devices include a speaker, a sound card, a video card, a display, a monitor, an actuator, an emitter, a smart card, another output device, or any combination thereof. An input device may allow a user to capture information to the audio object renderer 114. Examples of input devices include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a webcam, etc.), a microphone, a sensor, a directional pad, a trackpad, a scroll wheel, a smart card, etc. A presence-sensitive display may include a capacitive or resistive touch sensor for detecting input from a user. The sensor may be, for example, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, a light sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as the input device. For example, a universal serial bus (USB) port may be used to accommodate input and output devices.
[0052] In some embodiments, the power source 1308 is structured as a battery or battery pack. Other types of power sources may be used, such as an external power source (e.g., an electrical outlet), a photovoltaic device, or a battery. The power source 1308 may further include power circuitry for delivering power to various portions of the audio object renderer 114 from the power source 1308 itself and / or from an external power source via an interface such as an input circuit or a power cable. Delivering power may be for charging the power source 1308, for example. The power circuitry may perform any formatting, conversion, or other modification on the power from the power source 1308 to make it suitable for the respective components of the audio object renderer 114 being powered.
[0053] The memory 1310 may be or be configured to include memory, such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disk, optical disk, hard disk, removable cartridge, flash drive, etc. In one example, the memory 1310 includes one or more application programs 1314, such as an operating system, a web browser application, a widget, a gadget engine, or other applications, and corresponding data 1316. The memory 1310 may store any of a variety of different operating systems or combinations of operating systems for use by the audio object renderer 114.
[0054] The memory 1310 may be configured to include several physical drive units, such as a redundant array of independent disks (RAID), flash memory, a USB flash drive, an external hard disk drive, a thumb drive, a pen drive, a key drive, a high-density digital versatile disc (HD-DVD) optical disc drive, an internal hard disk drive, a Blu-ray optical disc drive, a holographic digital data storage (HDDS) optical disc drive, an external mini dual in-line memory module (DIMM), a synchronous dynamic random access memory (SDRAM), an external micro-DIMM SDRAM, a smart card memory such as a tamper-resistant module in the form of a universal integrated circuit card (UICC) containing one or more subscriber identity modules (SIMs), such as a USIM and / or ISIM, other memory, or any combination thereof. The UICC may be, for example, an embedded UICC (eUICC), an integrated UICC (iUICC), or a removable UICC, commonly known as a "SIM card." The memory 1310 may enable the audio object renderer 114 to access, offload, or upload data, instructions, application programs, and the like stored on a temporary or non-transitory memory medium. An article of manufacture, such as an article of manufacture utilizing a communication system, may be tangibly embodied as or in the memory 1310, which may be or comprise a device-readable storage medium.
[0055] The processing circuit 1302 may be configured to communicate with an access network or other networks using a communication interface 1312. The communication interface 1312 may comprise one or more communication subsystems and may include or be communicatively coupled to an antenna 1322. The communication interface 1312 may include one or more transceivers used to communicate, such as by communicating with one or more remote transceivers of another device capable of wireless communication (e.g., another UE or network node in the access network). Each transceiver may include a transmitter 1318 and / or a receiver 1320 suitable for providing network communication (e.g., optical, electrical, frequency allocation, etc.). Moreover, the transmitter 1318 and receiver 1320 may be coupled to one or more antennas (e.g., antenna 1322) and may share circuit components, software, or firmware, or alternatively, be implemented separately.
[0056] In the illustrated embodiment, the communication capabilities of communication interface 1312 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communication such as Bluetooth, near-field communication, location-based communication such as using a Global Positioning System (GPS) to determine location, another similar communication capability, or any combination thereof. Communication may be implemented in accordance with one or more communication protocols and / or standards such as IEEE 802.11, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Synchronous Optical Networking (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), etc.
[0057] Regardless of the type of sensor, the audio object renderer may provide the output of the decoded data to a network node via a wireless connection through the audio object renderer's communication interface 1312.
[0058] The audio object renderer, when in the form of an Internet of Things (IoT) device, can be a device for use in one or more application domains, including, but not limited to, urban wearable technology, augmented industrial applications, and healthcare. Non-limiting examples of such IoT devices are or are incorporated into a connected refrigerator or freezer, a TV, a connected lighting device, an energy meter, a robotic vacuum cleaner, a voice-controlled smart speaker, a home security camera, a thermostat, an electric door lock, a connected doorbell, an autonomous vehicle, a surveillance system, a weather monitoring device, a vehicle parking monitoring device, an electric vehicle charging station, a smartwatch, a fitness tracker, a head-mounted display for augmented reality (AR) or virtual reality (VR), or a wearable for haptic augmentation or sensory augmentation. A decoder in the form of an IoT device comprises circuitry and / or software depending on the intended application of the IoT device, in addition to the other components described with respect to the audio object renderer 114 shown in FIG. 13 .
[0059] 14 is a block diagram of a host 1400 in accordance with various aspects described herein. As used herein, a host 1400 may be or comprise various combinations of hardware and / or software, including a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, a container, or processing resources in a server farm. The host 1400 may provide one or more services to one or more UEs.
[0060] Host 1400 includes a processing circuit 1402 operably coupled to an input / output interface 1406, a network interface 1408, a power supply 1410, and a memory 1412 via a bus 1404. In other embodiments, other components may be included. Features of these components may be substantially similar to those described with respect to the devices of previous figures, such as FIG. 13, and therefore those descriptions are generally applicable to the corresponding components of host 1400.
[0061] Memory 1412 may include one or more computer programs, including one or more host application programs 1414 and data 1416, which may include user data, e.g., data generated by a UE for host 1400 or data generated by host 1400 for the UE. Embodiments of host 1400 may utilize only a subset or all of the shown components. Host application program 1414 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG, VP9) and audio codecs (e.g., EVS, IVAS, FLAC, Advanced Audio Coding (AAC), MPEG, G.711), including transcoding for multiple different classes, types, or implementations of UE (e.g., handsets, desktop computers, wearable display systems, heads-up display systems). The host application program 1414 may also provide user authentication and license checks, and may periodically report health, route, and content availability to a central node, such as a device in the core network or a device on the edge of the core network. Thus, the host 1400 may select and / or indicate a different host for over-the-top services for the UE. The host application program 1414 may support various protocols, such as HTTP Live Streaming (HLS) protocol, Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Dynamic Adaptive Streaming over HTTP (MPEG-DASH), etc.
[0062] FIG. 15 is a schematic block diagram illustrating a virtualization environment 1500 in which functionality implemented by some embodiments of the audio object renderer 114 or components of the audio object renderer 114 may be virtualized. In this context, virtualizing means creating a virtual version of an apparatus or device, which may include virtualizing a hardware platform, storage devices, and networking resources. Virtualization, as used herein, may apply to any device described herein, or components thereof, and relates to implementations in which at least a portion of functionality is implemented as one or more virtual components. Some or all of the functionality described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1500 hosted by one or more hardware nodes, such as a decoder, encoder, network node, UE, core network node, or hardware computing device acting as a host. Furthermore, in embodiments in which the virtual node does not require wireless connectivity (e.g., to a core network node or host), the node may be fully virtualized.
[0063] An application 1502 (which may alternatively be referred to as a software instance, a virtual appliance, a network function, a virtual node, a virtual network function, etc.) is run in the virtualized environment 1500 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.
[0064] Hardware 1504 includes processing circuitry, memory that stores software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices described herein, such as network interfaces, input / output interfaces, etc. Software is executed by the processing circuitry to instantiate one or more virtualization layers 1506 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 1508A and 1508B (one or more of which may be referred to generically as VMs 1508), and / or implement any of the functions, features, and / or benefits described with respect to some embodiments described herein. Virtualization layer 1506 may present to VMs 1508 a virtual operating platform that appears to be networking hardware.
[0065] The VMs 1508 may comprise virtual processing, virtual memory, virtual networking or interfaces, and virtual storage, and may be run by a corresponding virtualization layer 1506. Different embodiments of the virtual appliance 1502 instance may be implemented on one or more of the VMs 1508, and the implementation may be done in different ways. Hardware virtualization is referred to in some contexts as network functions virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry-standard high-volume server hardware, physical switches, and physical storage, which may be located in data centers and customer premises equipment.
[0066] In the context of NFV, a VM 1508 may be a software implementation of a physical machine that runs programs as if the programs were running on a physical, non-virtualized machine. Each VM 1508 and the portion of the hardware 1504 on which it runs, whether hardware dedicated to that VM and / or hardware shared by that VM with other VMs, form a separate virtual network element. Further, in the context of NFV, a virtual network function is responsible for handling a particular network function running in one or more VMs 1508 on the hardware 1504 and corresponds to the application 1502.
[0067] The hardware 1504 may be implemented in a standalone network node with general or specific components. The hardware 1504 may implement some functions via virtualization. Alternatively, the hardware 1504 may be part of a larger cluster of hardware (e.g., as in a data center or CPE) where many hardware nodes cooperate and are managed via a management and orchestration 1510 that, among other things, oversees the lifecycle management of the application 1502. In some embodiments, the hardware 1504 is coupled to one or more radio units, each including one or more transmitters and one or more receivers, which may be coupled to one or more antennas. The radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with virtual components to provide a virtual node with wireless capabilities, such as a wireless access node or base station. In some embodiments, some signaling may be provided using a control system 1512, which may alternatively be used for communication between the hardware nodes and the radio units.
[0068] While the computing devices described herein (e.g., decoders, audio object renderers, encoders, hosts) may include the depicted combinations of hardware components, other embodiments may comprise computing devices with different combinations of components. It should be understood that these computing devices may comprise any suitable combination of hardware and / or software required to perform the tasks, features, functions, and methods disclosed herein. The determining, calculating, obtaining, or similar operations described herein may be performed by processing circuitry, which may process information by, for example, transforming the obtained information to other information, comparing the obtained or transformed information to information stored in a network node, and / or performing one or more operations based on the obtained or transformed information and as a result of the processing making a decision. Moreover, while a component is depicted as a single box located within a larger box or nested within multiple boxes, in reality the computing device may comprise multiple different physical components that make up the single depicted component, and functionality may be partitioned among the separate components. For example, a communications interface may be configured to include any of the components described herein, and / or the functionality of those components may be partitioned between the processing circuitry and the communications interface. In another example, non-computationally intensive functionality of any of such components may be implemented in software or firmware, and computationally intensive functionality may be implemented in hardware.
[0069] In some embodiments, some or all of the functionality described herein may be provided by a processing circuit executing instructions stored in a memory, which in some embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuit without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hardwired manner. In any of these particular embodiments, the processing circuit may be configured to perform the described functionality, regardless of whether or not it executes instructions stored on a non-transitory computer-readable storage medium. Benefits provided by such functionality are not limited to the processing circuit alone or to other components of the computing device, but are enjoyed by the computing device as a whole and / or by end users and wireless networks generally.
[0070] Illustrative Embodiments 1. A method in an inter-channel time difference (ITD) combiner (220, 340, 1502), comprising: Receiving (401) a current ITD and an audio frame, where each frame m comprises N samples; storing (403) in a signal memory at least a portion of the currently input audio frame; determining (405) transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on an inter-channel time difference (ITD) of a current input audio frame and an ITD of a previous input audio frame; applying (407) a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1; A method comprising: 2. The method of embodiment 1, wherein the audio frame is a portion of an object audio signal having position metadata describing a position relative to the listener, and the method further comprises obtaining the time shift from the position metadata. 3. The frame length of frame m and the look-ahead memory rs required for resampling LA Calculating a total transition length based on (801), wherein the total transition is divided into two parts comprising t1 and t2; Look-ahead memory rs LA determining (803) a buffer length t3 based on determining (805) a total transition length based on the maximum allowable transition length and the buffer length t3; 3. The method of embodiment 1 or 2, further comprising: 4. Determining the buffer length t3 is performed by: t3=max(0,rs LA -|ITD(m-1)|) and The total transition length is t tot =N max -t3, N max ≦N and determining according to 5. Determining the transition times t1 and t2 is In response to the current ITD and the previous ITD having the same sign, allocating the total transition length to one of the transition times t1, t2 and setting the other to 0. 5. The method of any one of embodiments 1 to 4, comprising: 6. Determining the transition times t1 and t2 is applying a shift operation to both output signal 0 and output signal 1 by splitting the total transition length into two parts to determine transition times t1 and t2 in response to the code of the current ITD being different from the code of the previous ITD; 5. The method of any one of embodiments 1 to 4, comprising: 7. Splitting the total transition length into two parts to determine the transition times t1 and t2 can be done by dividing the total transition length by Including splitting according to TIFF2026508733000013.tif14170, where [·] represents the round-to-nearest integer operation, 7. The method of embodiment 6. 8. Populating a processing buffer (510) using the current input audio frame of the object audio and the signal memory. further comprising Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 is In response to the sign of the current ITD and the sign of the previous ITD being the same, or in response to one of the current ITD and the previous ITD being 0, adjusting the processing buffer (510) to populate the first output buffer (550) by allocating the total transition length to t1 and setting t2 to 0; Copying the input signal portion of the processing buffer (510) to a second output buffer (560); Including, 8. The method of any one of embodiments 3 to 7. 9. Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 and in response to ITD(m)=ITD(m−1) and the sign of one of the current ITD and the previous ITD being negative, thereby indicating that one of output signal 0 and output signal 1 is ahead of the other of output signal 0 and output signal 1, either the first output buffer (550) or the second output buffer (560) delaying the output buffer associated with the other of output signal 0 and output signal 1 by a total transition length. 9. The method of embodiment 8, further comprising: 10. Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 In response to |ITD(m)|>|ITD(m-1)| The length of the frame in the processing buffer (510) is expressed as x of length t1 + |ITD(m-1)| - |ITD(m)| buf (n), n=|ITD(m-1)|,...,N-1-|ITD(m)| to an output frame of length t1; In response to the buffer length t3 being greater than 0, x buf (n), n=N-1-t3-|ITD(m)|,...N-1-|ITD(m)| by adding the last t3 samples of the output channel. Generating transitions by 10. The method of embodiment 8 or 9, further comprising: 11. Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 adding the last t_3 samples of the first output buffer (550) by copying them from the processing buffer (510) in response to |ITD(m)|<|ITD(m-1)|; 11. The method of any one of embodiments 8 to 10, further comprising: 12. Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 is In response to ITD(m)·ITD(m-1)<0, further comprising splitting the total transition length according to TIFF2026508733000014.tif16170; where [·] represents the rounding to nearest integer operation, and splitting the total transition length is Sample x of length t1 + |ITD(m-1)| buf (n), n=-|ITD(m-1)|,...,t1-1 to match the samples n=0,...,t1-1 of the first output buffer (550); Sample x buf(n), n=0,...,t1-1 to the corresponding index in the second output buffer (560); Sample x of length t2-|ITD(m)| buf (n), n=t1,...,t1+t2-1-|ITD(m)| to fit the samples n=t1,...,t1+t2-1 of length t2 in the second output buffer (560); 12. The method of any one of embodiments 8 to 11, comprising: 13. Applying the determined transition times t1, t2 in generating the output signal 0 and the output signal 1 allocating a first output buffer (550) to output signal 1 and a second output buffer (560) to output signal 0 in response to ITD(m-1)=0 and ITD(m)>0; allocating a first output buffer (550) to output signal 0 and a second output buffer (560) to output signal 1 in response to ITD(m-1)=0 and ITD(m)≦0; In response to ITD(m-1)>0, allocating a first output buffer (550) to output signal 1 and a second output buffer (560) to output signal 0; In response to ITD(m-1)<0, allocating the first output buffer (550) to output signal 0 and the second output buffer (560) to output signal 1; 13. The method of any one of embodiments 8 to 12, further comprising: 14. An apparatus (114, 1502) having an ITD combiner, the ITD combiner comprising: Receiving (401) a current ITD and an audio frame, where each frame m comprises N samples; storing (403) in a signal memory at least a portion of the currently input audio frame; determining (405) transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on an inter-channel time difference (ITD) of a current input audio frame and an ITD of a previous input audio frame; applying (407) a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1; An apparatus (114, 1502) adapted to perform the steps of: 15. The apparatus (114, 300, 1502) of embodiment 14, wherein the ITD combiner (220, 340, 1502) is further adapted to perform according to any one of embodiments 2 to 12. 16. An apparatus (114, 300, 1502) having an inter-channel time difference (ITD) combiner (220, 340, 1502), wherein the ITD combiner (220, 340, 1502) comprises: A processing circuit (1202); a memory (1210) coupled to the processing circuit; wherein the memory includes instructions that, when executed by the processing circuit, cause the ITD synthesizer (220, 340, 1502) to perform operations, the operations including: Receiving (401) a current ITD and an audio frame, where each frame m comprises N samples; storing (403) in a signal memory at least a portion of the currently input audio frame; determining (405) transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on an inter-channel time difference (ITD) of a current input audio frame and an ITD of a previous input audio frame; applying (407) a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1; An apparatus (114, 300, 1502) including: 17. The apparatus (114, 300, 1502) of embodiment 16, wherein the memory includes further instructions that, when executed by the processing circuit, cause the ITD synthesizer (220, 340, 1502) to perform according to any one of embodiments 2 to 13. 18. A computer program comprising program code executed by a processing circuit (1202) of an apparatus (112, 300, 1502) having an inter-channel time difference (ITD) combiner (220, 340, 1502), wherein execution of the program code causes the ITD combiner (220, 340, 1502) to perform an operation, the operation being: Receiving (401) a current ITD and an audio frame, where each frame m comprises N samples; storing (403) in a signal memory at least a portion of the currently input audio frame; determining (405) transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on an inter-channel time difference (ITD) of a current input audio frame and an ITD of a previous input audio frame; applying (407) a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1; a computer program comprising: 19. The computer program of embodiment 18, further comprising program code, the execution of which causes the ITD synthesizer (220, 340, 1502) to operate according to any one of embodiments 2 to 13. 20. A computer program product comprising a non-transitory storage medium containing program code executed by a processing circuit (1202) of an apparatus (114, 300, 1502) having an inter-channel time difference (ITD) combiner (220, 340, 1502), the execution of the program code causing the ITD combiner (220, 340, 1502) to perform operations, the operations including: Receiving (401) a current ITD and an audio frame, where each frame m comprises N samples; storing (403) in a signal memory at least a portion of the currently input audio frame; determining (405) transition times t1, t2 for implementing a time shift to apply to at least one of output signal 0 and output signal 1 based on an inter-channel time difference (ITD) of a current input audio frame and an ITD of a previous input audio frame; applying (407) a time shift within the determined transition times t1, t2 in generating output signal 0 and output signal 1; a computer program product, 21. The computer program of embodiment 19, wherein the non-transitory storage medium includes further program code, the execution of which causes the ITD synthesizer (220, 340, 1502) to operate according to any one of embodiments 2 to 13.
Claims
1. 1. A method for adjusting timing of output audio signals to achieve a desired inter-channel time difference (ITD) between the output audio signals, the method comprising: receiving (401) a current ITD value and an audio frame; a transition time t for implementing a time shift to apply to at least one of the first output signal and the second output signal based on the ITD of the current frame and the ITD of the previous frame; 1 , t 2 determining (405) In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 applying (407) said time shift in A method comprising:
2. 2. The method of claim 1, wherein at least a portion of the audio frame is stored in a memory for use in synthesizing an ITD in a subsequent frame (403).
3. 3. The method of claim 1, wherein the audio frame is part of an audio object comprising an audio signal and position metadata describing an object position, the method further comprising obtaining the time shift from the position metadata.
4. The frame length of the current frame m and the look-ahead memory rs required for resampling LA and calculating (801) a total transition length based on t 1 and 2 Calculating (801) a total transition length, which is divided into two parts comprising: The look-ahead memory rs LA Based on the transition length t 3 determining (803) The maximum allowable transition length and the transition length t 3 determining (805) the total transition length based on The method of any one of claims 1 to 3, further comprising:
5. The transition length t 3 Determining the transition length t 3 of, t 3 =max(0,rs LA -|ITD(m-1)|) where ITD(m-1) is the ITD of the previous audio frame comprising N samples. 3 and determining the total transition length as t tot = N max -t 3 , N max determining according to N≦N, N max determining said total transition length, where is the maximum allowable transition length; The method of claim 4, comprising:
6. The transition time t 1 , t 2 To determine In response to the current ITD and the previous ITD having the same sign, a total transition length is calculated based on the transition time t 1 , t 2 and set the other to 0.
6. The method of claim 1, comprising:
7. The transition time t 1 , t 2 To determine In response to the sign of the current ITD being different from the sign of the previous ITD, the transition time t 1 , t 2 applying a shift operation to both the first output signal and the second output signal by splitting a total transition length into two parts to determine 6. The method of claim 1, comprising:
8. The transition time t 1 , t 2 Splitting the total transition length into two parts to determine splitting according to where ITD(m) is the current ITD and [·] represents a round-to-nearest integer operation. The method of claim 7.
9. populating a processing buffer (510) with said audio frames; further comprising The transition time t determined in generating the first output signal and the second output signal 1 , t 2 Applying in response to the current ITD and the previous ITD having the same sign, or in response to one of the current ITD and the previous ITD being 0, The total transition length is t 1 Assign to t 2 adjusting said processing buffer (510) to populate a first output buffer (550) by setting copying the input signal portion of said processing buffer (510) to a second output buffer (560); Including, 9. The method according to any one of claims 4 to 8.
10. In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 Applying and in response to ITD(m)=ITD(m−1) and the sign of one of the current ITD and the previous ITD being negative, thereby indicating that one of the first output signal and the second output signal is ahead of the other of the first output signal and the second output signal, either the first output buffer (550) or the second output buffer (560) delaying an output buffer associated with the other of the first output signal and the second output signal by the total transition length.
10. The method of claim 9, further comprising:
11. In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 Applying In response to |ITD(m)|>|ITD(m-1)| The length of the frame in the processing buffer (510) is defined as length t 1 + |ITD(m-1)|-|ITD(m)|'s x buf (n), n=|ITD(m-1)|, . . . , N-1-|ITD(m)|, length t 1 and extending it to the output frame of The transition length t 3 is greater than 0, and buf (n), n=N-1-t3-|ITD(m)|,... N-1-|ITD(m)| is copied to the last t of the output channel. 3 Adding samples and Generating transitions by 11. The method of claim 9 or 10, further comprising:
12. In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 Applying In response to |ITD(m)|<|ITD(m-1)|, the last t 3 Adding samples 12. The method of any one of claims 9 to 11, further comprising:
13. In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 Applying In response to ITD(m)·ITD(m−1)<0, and further comprising splitting the total transition length according to where [·] represents a rounding operation to the nearest integer, and splitting the total transition length is Length t 1 + |ITD(m-1)| sample x buf (n), n=-|ITD(m-1)|,. .. .. ,t 1 -1 to the samples n=0, . . . , t of the first output buffer (550). 1 Resampling to fit −1; Sample x buf (n), n=0, . .. .. ,t 1 copying −1 to the corresponding index in said second output buffer (560); Length t 2 - Sample x of |ITD(m)| buf (n), n = t 1 ,... ,t 1 +t 2 -1-|ITD(m)| to the length t 2 Sample n=t 1 ,... ,t 1 +t 2 Resampling to fit -1 13. The method of any one of claims 9 to 12, comprising:
14. generating the second output signal; The transition length t 3 is greater than 0, and buf (n), n=N-1-t3-|ITD(m)|,...N-1-|ITD(m)| 3 Adding samples 14. The method of any one of claims 9 to 13, further comprising:
15. In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 Applying allocating the first output buffer (550) to the second output signal and the second output buffer (560) to the first output signal in response to ITD(m-1)=0 and ITD(m)>0; allocating the first output buffer (550) to the first output signal and the second output buffer (560) to the second output signal in response to ITD(m-1)=0 and ITD(m)≦0; allocating the first output buffer (550) to the second output signal and the second output buffer (560) to the first output signal in response to ITD(m-1)>0; allocating the first output buffer (550) to the first output signal and the second output buffer (560) to the second output signal in response to ITD(m-1)<0; 15. The method of any one of claims 9 to 14, further comprising:
16. 1. An apparatus (112, 300, 1502) for adjusting timing of output audio signals to achieve a desired inter-channel time difference (ITD) between the output audio signals, said apparatus comprising: receiving a current ITD value and an audio frame; a transition time t for implementing a time shift to apply to at least one of the first output signal and the second output signal based on the ITD of the current frame and the ITD of the previous frame; 1 , t 2 and determining In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 applying the time shift within An apparatus (112, 300, 1502) adapted to perform the steps of:
17. 17. The apparatus (112, 300, 1502) of claim 16, wherein the apparatus is further adapted to perform a method according to any one of claims 2 to 15.
18. An apparatus (112, 300, 1502) comprising: A processing circuit (1202); a memory (1210) coupled to said processing circuit; wherein the memory contains instructions that, when executed by the processing circuitry, cause the device to perform operations, the operations including: receiving a current inter-channel time difference (ITD) value and an audio frame; a transition time t for implementing a time shift to apply to at least one of the first output signal and the second output signal based on the ITD of the current frame and the ITD of the previous frame; 1 , t 2 and determining In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 applying (407) said time shift in An apparatus (112, 300, 1502) including:
19. 19. The apparatus (112, 300, 1502) of claim 18, wherein the memory includes further instructions that, when executed by the processing circuitry, cause the apparatus to perform the operations of any one of claims 2 to 15.
20. A computer program comprising program code that is executed by a processing circuit (1202) of an apparatus (112, 300, 1502), the execution of the program code causing the apparatus to perform operations, the operations comprising: receiving a current ITD value and an audio frame; a transition time t for implementing a time shift to apply to at least one of the first output signal and the second output signal based on the ITD of the current frame and the ITD of the previous frame; 1 , t 2 and determining In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 applying the time shift within a computer program comprising:
21. 21. The computer program of claim 20, comprising further program code, execution of which causes the device (112, 300, 1502) to operate in accordance with any one of claims 2 to 15.
22. 1. A computer program product comprising a non-transitory storage medium containing program code that is executed by a processing circuit (1202) of a device (112, 300, 1502), the execution of the program code causing the device (112, 300, 1502) to perform operations, the operations comprising: receiving a current ITD value and an audio frame; a transition time t for implementing a time shift to apply to at least one of the first output signal and the second output signal based on the ITD of the current frame and the ITD of the previous frame; 1 , t 2 and determining In generating the first output signal and the second output signal, the determined transition time t 1 , t 2 applying the time shift within a computer program product,
23. 23. The computer program product of claim 22, wherein the non-transitory storage medium comprises further program code, execution of which causes the device (112, 300, 1502) to operate in accordance with any one of claims 2 to 15.
Citation Information
Patent Citations
Method and apparatus for signal reconstruction during stereo signal encoding
JP2020531912A
Stereo Signal Processing Method and Apparatus
US20200082834A1