Efficient time delay synthesis
By receiving the current inter-channel time difference value and audio frame, the time shift conversion time suitable for the output signal is determined, which solves the problem that multi-channel audio signals cannot be effectively played back without speaker configuration, and realizes efficient inter-channel time difference adjustment and effective processing of mobile sound sources.
Patent Information
- Application Number
- CN202380077164.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-09
- Filing Date
- 2023-11-08
- Publication Date
- 2025-06-24
AI Technical Summary
Without a speaker configuration, multi-channel audio signals cannot be played back effectively, especially when processing mobile sound sources need to be processed, and prior art is difficult to efficiently adjust the time difference between channels.
By receiving the current inter-channel time difference (ITD) value and audio frame, the time shift conversion time applicable to the output signal is determined, and the time shift is applied when generating the output signal to achieve the desired inter-channel time difference.
It is realized that the timing of the output audio signal is efficiently adjusted when crossing the zero boundary, ensuring that the time difference between the output audio signals meets the expected value, reducing the computational complexity and improving the processing efficiency of the mobile sound source.
Smart Images

Figure CN120202682A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to communications, and more particularly to communication methods, related devices, and nodes that support audio encoding and decoding. Background Art
[0002] Spatial audio is a description of the sound field for an immersed listener. There are several formats of spatial audio. The most common format is the stereo format, where the sound field is rendered through two speakers or a set of headphones. In scenarios where the playback is on a larger set of speakers such as 5.1, 7.1+4, or 22.2, spatial audio is typically referred to as multichannel audio. There are also spatial audio formats that describe the sound field itself regardless of the layout of the speaker system. Such descriptions include wave field synthesis (WFS), where the sound field is captured by a microphone array and symmetrically reproduced by a speaker array. Another popular format is ambisonics, which depends on spherical harmonics captured using a compact microphone array. Currently, ambisonics has recently become increasingly popular as they are suitable for listener-centric rendering, such as virtual reality (VR) and augmented reality (AR) audio rendering, and they are inherently suitable for rotation. They can also be coupled with 360 video capture for reconstructing an experienced scene.
[0003] Multichannel audio formats can be directly played back on the speaker settings they are designed for. However, if there is no speaker configuration, the audio cannot be played back in the case of misaligned audio. This adaptation is typically referred to as rendering the spatial audio of the playback system. If there is a 22.2 multichannel signal or an ambisonic signal, it can be rendered, for example, to be played back on a 5.1 system or a set of headphones. When rendering for headphones, the audio reaching the ears is typically modeled using head-related filters (HRFs) or head-related transfer functions (HRTFs). The filters model the direction of arrival (DoA) of the sound source so that the listener perceives the sound coming from that direction. This is achieved through spectral coloring, the horizontal difference between the ears, and the time difference caused by the difference in the lengths of the paths to the left and right ears. This time difference is typically referred to as the interaural time difference or inter-channel time difference (ITD). The time difference between channels can be created by filtering one or both channels with a Dirac impulse: h = δ(t - t Δ ). However, it is necessary to handle the conversion between different time shifts, for example, for a moving source. Summary of the Invention
[0004] When modeling the HRF, filters can be used to complete spectral coloring, and the time difference can be generated by time shifting. The present disclosure applies time shifting in an efficient manner when crossing the zero boundary for shifting.
[0005] When changing the sign of the time delay parameter, a shift operation needs to be performed on the two output channels. To limit the complexity of the shift operation, the total conversion length is shared between the channels. The sharing is done proportionally to the size of the shift on each side of the zero point.
[0006] According to a first aspect, a method is provided for adjusting the timing of an output audio signal to achieve a desired inter-channel time difference (ITD) between output audio signals. The method includes: receiving a current ITD value and an audio frame, and determining conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame. When generating the first output signal and the second output signal, the time shift is applied within the determined conversion times t1, t2.
[0007] According to a second aspect, an apparatus is provided for adjusting the timing of an output audio signal to achieve a desired inter-channel time difference ITD between output audio signals. The apparatus is adapted to: receive a current ITD value and an audio frame, and determine conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame. The apparatus is adapted to: apply the time shift within the determined conversion times t1, t2 when generating the first output signal and the second output signal.
[0008] According to a third aspect, an apparatus is provided that includes: a processing circuit and a memory coupled to the processing circuit, where the memory includes instructions that, when executed by the processing circuit, cause the apparatus to perform operations including: receiving a current inter-channel time difference ITD value and an audio frame; determining conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal based on the ITD of the current frame and the ITD of a previous frame; and applying the time shift within the determined conversion times t1, t2 when generating the first output signal and the second output signal.
[0009] According to a fourth aspect, a computer program is provided that includes program code to be executed by a processing circuit of an apparatus, whereby execution of the program code causes the apparatus to perform the operations of the first aspect.
[0010] According to a fifth aspect, a computer program product is provided that includes a non-transitory storage medium that includes program code to be executed by a processing circuit of an apparatus, whereby execution of the program code causes the apparatus to perform the operations of the first aspect
[0011] Certain embodiments may provide one or more of the following technical advantages. For switching across the zero boundary, the speed of adjustment remains consistent and the computational complexity remains low. The method is designed to utilize time delay to generate two channels, where the time delay can be updated per frame. The update of the time delay can be accomplished with minimal conversion artifacts. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this application, illustrate certain non-limiting embodiments of the inventive concept. In the drawings:
[0013] Figure 1 is a block diagram showing an environment in which various embodiments of the disclosure may be implemented;
[0014] Figure 2 is a block diagram of an audio object renderer according to some embodiments of the disclosure;
[0015] Figure 3 is a block diagram of a parametric stereo decoder according to some embodiments of the disclosure;
[0016] Figure 4 is a flowchart showing the operation of an ITD synthesizer according to some embodiments of the disclosure;
[0017] Figure 5 is a block diagram of an ITD synthesizer according to some embodiments of the disclosure;
[0018] Figure 6 is a flowchart showing the operation of an ITD synthesizer according to some embodiments of the disclosure;
[0019] Figure 7 is an illustration of buffer operations performed by an ITD synthesizer according to some embodiments of the disclosure;
[0020] Figure 8 is a flowchart showing the operation of an ITD synthesizer according to some embodiments of the disclosure;
[0021] Figures 9 to 11 is an illustration of buffer operations performed by an ITD synthesizer according to some embodiments of the disclosure;
[0022] Figure 12 is an illustration of a sinc resampling function for processing the compression and expansion segments of a signal;
[0023] Figure 13 is a block diagram of an audio object renderer according to some embodiments;
[0024] Figure 14Block diagram of a host computer communicating with an encoder and / or decoder according to some embodiments; and
[0025] Figure 15 Block diagram of a virtualized environment according to some embodiments. DETAILED DESCRIPTION
[0026] Some embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. The embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art, where examples of embodiments of the inventive concept are shown. However, the inventive concept may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concept to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment may be suitably assumed to be present / used in another embodiment.
[0027] Figure 1 An example of an operating environment in which various embodiments of the present disclosure may be implemented is shown. Turning to Figure 1 , in an example operating environment 100, an encoder 102 receives data to be encoded (such as an audio file) from an entity and / or from a storage device 108 via a network 104 (such as a host 106). In some embodiments, the host 106 may communicate directly with the encoder 102. The encoder 102 encodes the audio file as described herein and stores the encoded audio file in the storage device 108 or transmits the encoded audio file via a network 110 to a decoder 112 having an audio object renderer 114. The audio object renderer 114 within the decoder 112 renders the decoded audio file and transmits the rendered decoded audio file to an audio player 116 for playback. For example, the audio player 116 may play the rendered decoded audio file for a spatial audio representation such as a virtual reality conference or a computer game. The audio player 116 may be a user device, a terminal, a mobile phone, etc. or may be included in a user device, a terminal, a mobile phone, etc. In other embodiments, the host 106 may transmit the encoded audio file to the audio object renderer 114 via the network 110. In some embodiments, where the decoded audio does not directly "fit" the audio player 116, the audio object renderer 114 may be a separate device between the decoder 112 and the audio player 116, e.g., to render the decoded audio file for headphones.
[0028] As previously mentioned, when changing the sign of the time delay parameter, a shift operation needs to be performed on two output channels. To limit the complexity of the shift operation, the total conversion length is shared between the channels. The sharing is done proportionally to the size of the shift on each side of the zero according to the following formula: where ITD(m) and ITD(m - 1) are the inter-channel time differences, m is the sub-frame index, t tot is the total conversion length, and t1 and t2 are the conversion lengths for performing time stretching or compression operations. The time stretching or compression operation can also be referred to as a time shift operation or a resampling operation. The conversion length can be expressed in seconds or the number of samples of a discrete sampled audio signal. The conversion length can also be referred to as the conversion time.
[0029] The above formula can be simplified (rounding integers) to:
[0030] In some embodiments of the present disclosure, as Figure 2 shown, the method operates in the ITD synthesizer implemented within the audio object renderer 114. In other embodiments, the ITD synthesizer can be implemented within the parametric stereo decoder. The method operates on audio segments called frames, where each frame or sub-frame m consists of N samples. x(m,n), n = 0, 1, 2, …, N - 1
[0031] The frames can, for example, constitute audio segments from decoded audio objects, the mono downmix channel in the parametric stereo decoder, or the input channels to the audio object renderer. Here, the audio object renderer 114 receives an audio object, which includes an audio signal and position metadata describing the position of the audio object. The position can be absolute or relative to the listener's position. The position metadata is input to the HR filter module 210, which provides ITD values and a set of HR filters for the left and right channels. The time delay parameter for frame m is an integer within the range ITD(m) = [-ITD MAX , ITD MAX .
[0032] In the case where the frame comes from the parametric stereo decoder, ITD(m) can be found by analyzing the input channels to the stereo encoder. Preferably, the input channels are aligned by compensating for ITD(M) before generating the downmix channel. The downmix channel will be encoded together with the stereo parameters, including ITD(m) reconstructed and decoded in the parametric stereo decoder. The parametric stereo decoder will reconstruct the downmix signal, and the stereo parameters will at least include the reconstruction of ITD(m) and synthesize two output channels with the corresponding ITD(m).
[0033] The audio and position metadata may, for example, come from an audio object decoder or be generated by a 3D audio engine for spatial audio representation, such as for virtual reality conferences or computer games. The HR filter module 210 may, for example, be a database of stored filters and ITD values, or it may be a model-based database of filters and ITD values that produce given position data. The ITD values are input to an ITD synthesizer 220 that produces two output signals based on an input audio frame, wherein the output signals have the desired ITD. The two channels are filtered by a left filter 230 and a right filter 240 to produce synthesized left and right channels. The audio object may be added with one or more additional objects. The output left and right channels may be forwarded to an audio device for playback.
[0034] This is the case in the operation of method 400. Figure 4 As shown in the flowchart of , the ITD synthesizer 220 performs in some embodiments. In block 401, the ITD synthesizer 220 receives a current ITD and an audio frame, wherein each frame m includes N samples. Upon receiving the audio frame or at any time during the processing of the frame, the ITD synthesizer 220 may store at least a portion of the current input audio frame to be used for processing the ITD in a subsequent audio frame in block 403.
[0035] In block 405, the ITD synthesizer 220 determines a transition time t1, t2 for performing a time shift applied to at least one of the output signal 0 and the output signal 1 based on the sign of the inter-channel time difference ITD of the current audio frame and the ITD of the previous audio frame. Figure 2 The HRF filter module 210 in the context of . In other embodiments where the ITD synthesizer is implemented as part of a parametric stereo decoder, the ITD is reconstructed from the bitstream. Here, time shifting refers to the operation of smoothly switching to a target ITD.
[0036] In block 407 , the ITD synthesizer 220 applies a time shift within the determined transition times t1 , t2 when generating OUTPUT SIGNAL 0 and OUTPUT SIGNAL 1 .
[0037] Before describing further details of the ITD synthesizer 220, Figure 3An embodiment is shown in which the ITD synthesizer can be implemented within the parametric stereo decoder 300. In the parametric stereo decoder 300, stereo parameters including the ITD parameter are decoded by the parameter decoder 310, and optionally, a reconstructed residual signal is generated by the residual decoder 320. The downmix decoder 330 is configured to decode and reconstruct the encoded downmix signal to output a reconstructed downmix signal, where time shifting can be applied. The reconstructed downmix, the reconstructed stereo parameters, and optionally the reconstructed residual signal are fed to the stereo upmixer 340 to generate a reconstructed stereo signal. The ITD synthesizer 220 is part of the stereo upmixer 340.
[0038] The ITD synthesizer 220 is described in further detail in Figure 5 and also performs the operations shown in Figure 6 . Turning to Figure 6 , in step 601, the processing buffer 510 is filled using the current input audio frame x(m,n) and the signal memory 520. The length of the memory N men should be at least the sum of the maximum time shift ITD MAX and the lookback / lookahead memory rs LA required for the resampling function. N mem = ITD MAX + rs LA
[0039] Figure 7 The processing buffer 510 is shown in buf , where the middle curve shows the processing buffer x buf (n), where x LA The total conversion length is calculated based on the frame length N and the lookahead rs LA required by the resamplers 570, 580. If ITD(m) < rs LA , then a small portion of the processing buffer must be reserved to accommodate the lookahead without introducing processing delay due to resampling. The conversion can be divided into three parts, t1, t2, and t3. The conversion length t3 can be considered the buffer length to avoid reading the memory during the resampling operation. The resampler must leave at least rs t3 = max(0, rs LA - |ITD(m)|)
[0040] Then the total conversion time t tot is t tot = N max - t3 where N max is the maximum allowable conversion length. It can be set to N max = N, meaning that the full frame time is allowed for performing the conversion. However, if N is large, then it may be necessary to use N max ≤ N to limit the maximum allowable conversion length to achieve a faster conversion and possibly lower complexity.
[0041] Figure 8 FIG. shows the operations performed by the ITD synthesizer 220 in determining the total conversion length. In block 801, the ITD synthesizer 220 calculates the total conversion length based on the frame length N of frame m and the look-ahead memory rs required for resampling LA and the total conversion is divided into two parts including t1 and t2.
[0042] In block 803, the ITD synthesizer 220 determines the conversion time t3 based on the look-ahead memory rs LA In block 805, the ITD synthesizer 220 determines the total conversion length based on the maximum allowable conversion length N max and the conversion time t3.
[0043] The following time shift operations can be divided into two groups: 1. The signs of ITD(m) and ITD(m - 1) are the same or one of them is zero. 2. The sign of ITD is non-zero and changes, i.e., ITD(m)·ITD(m - 1) < 0.
[0044] Case 1 - The signs of ITD are the same or one of them is zero
[0045] If the signs of ITD(m) and ITD(m - 1) are the same or one of them is zero, then the shift can be processed by only processing one of the channels, which means processing step 605, where the resampler 570 adjusts the processing buffer 510 to fill the output buffer A 550. This can be achieved by allocating the entire conversion length to t1 and setting t2 to zero, i.e., where n1, n2, n3 represent the starting indices of each time shift segment assuming the current input sub-frame starts at n = 0, L in,1 , L i,2 is the length of resampling segments 1 and 2. In this case, the input signal portion of the processing buffer is simply copied to the output buffer B 560.
[0046] When the time delay of the current frame is the same as that of the previous frame, i.e., ITD(m) = ITD(m - 1), output time delay synthesis is generated by pointing to the corresponding starting point in the processing buffer. The sign of ITD(m) determines in which of the two channels the delay is applied. For example, a positive ITD(m) may indicate that the left channel of a stereo pair is in front of the right channel, in which case the right channel should be delayed and the left channel should be output without delay. Figure 7 This situation is shown in. In this case, the resampling operation on the output buffer A550 has the same input and output lengths and is equivalent to a copy operation.
[0047] When the absolute value of the time delay of the current frame is greater than the absolute value of the previous frame, |ITD(m)| > |ITD(m - 1)|, a conversion is generated to allow for a smooth transition between the delay values. Figure 9 This situation is shown in. The conversion is completed by extending the length of the frame from x of length t1 + |ITD(m - 1)| - |ITD(m)|, n = ITD(m - 1), …, N - 1 - ITD(m) to the output frame of length t1. If the conversion time t3 is greater than zero (i.e., t3 > 0), then simply copy the last t3 samples of the output channel from the processing buffer to achieve a delay of |ITD(m)|, x(n), n = N - 1 - t3 - |ITD(m)|, … N - 1 - |ITD(m)|. Here, it can also be noted that if |ITD(m - 1)| = |ITD(m)|, then the input length is the same as the output length and the resampling will be equivalent to a copy operation. buf (n), n = ITD(m - 1), …, N - 1 - ITD(m) to the output frame of length t1. If the conversion time t3 is greater than zero (i.e., t3 > 0), then simply copy the last t3 samples of the output channel from the processing buffer to achieve a delay of |ITD(m)|, x(n), n = N - 1 - t3 - |ITD(m)|, … N - 1 - |ITD(m)|. Here, it can also be noted that if |ITD(m - 1)| = |ITD(m)|, then the input length is the same as the output length and the resampling will be equivalent to a copy operation. buf (n), n = N - 1 - t3 - |ITD(m)|, … N - 1 - |ITD(m)|. Here, it can also be noted that if |ITD(m - 1)| = |ITD(m)|, then the input length is the same as the output length and the resampling will be equivalent to a copy operation.
[0048] When the absolute value of the time delay decreases, i.e., |ITD(m)| < |ITD(m - 1)|, the expression for the input frame length remains the same. However, the length t1 + |ITD(m - 1)| - |ITD(m)| will now be greater than the resulting length t1, and the resampling corresponds to shortening the length of the frame. This is shown in Figure 10 In this example, ITD(m) = 0, which means that the last t3 samples of the output buffer A are copied from the processing buffer 510.
[0049] Case 2 - The signs of ITD are different and non - zero
[0050] If the sign of ITD changes, i.e., if ITD(m)·ITD(m - 1) < 0, then a shift operation must be performed on both channels. In this case, the total conversion length t tot is divided into two parts according to the following formula: where [·] represents rounding to the nearest integer. Next, as Figure 11 shown, the signals are time-shifted using the conversion times t1 and t2 to assemble output buffer A and output buffer B. First, the buffer is resampled from n1 of length L in ,1 to the first t1 samples of output buffer A. The last L sf -t1 samples are filled by copying the remaining part of the processing buffer to output buffer A. Then, the first sample of the processing buffer starting from index 0 is copied to the first t1 samples of output buffer B. The next t2 samples of output buffer B are created by resampling the samples starting from n2 in the processing buffer of length L in,2 . Finally, the last t3 samples of the processing buffer are copied to output buffer B. The last L ITDmem samples of the input frame are stored in the memory for processing the next sub-frame. Output buffer A and output buffer B are assigned to output channel 0 and output channel 1 respectively for left HRIR filtering and right HRIR filtering. The assignment of the output channels is done based on the signs of ITD(m) and ITD(m - 1) as follows:
[0051] The resampling and copying operations can also be described with reference to the indices of the buffers as follows. When the ITD signs are different, the first transition is a 550 shift from ITD(m - 1) to 0 turning on output buffer A followed by a 560 shift from 0 to ITD(m) turning on output buffer B. An example of this process is shown in Figure 11 . In step 605, resampler 570 fills the first t1 samples of output buffer A 550. This is done by resampling samples of length t1 + |ITD(m - 1)| of x buf (n), n = -|ITD(m - 1)|, …, t1 - 1 to fit the samples of output buffer A 550 of x buf (n), n = 0, …, t1 - 1. In the same step, the samples x buf (n), n = 0, …, t1 - 1 are copied to the corresponding indices n = 0, …, t1 - 1 in output buffer B560. In step 607, resampler 580 adapts samples of length t2 - |ITD(m)| of x buf(n), where n = t1, …, t1 + t2 - 1 - |ITD(m)|, to match the length t2 of the samples in output buffer B 560, i.e., n = t1, …, t1 + t2 - 1. In the case where t3 > 0, in step 609, the last t3 samples n = N - 1 - t3 - |ITD(m)|, …, N - 1 - |ITD(m)| are copied from processing buffer 510 to output buffer B 560. It should be noted that output buffer 550 and output buffer 560 may have different alignments, such that the index is shifted according to the processing buffer. In this case, the above-mentioned indices associated with output buffer 550 and output buffer 560 will be shifted by that amount, but the segment is still appended in the same manner as described here. For example, Figure 7 and Figure 9 output buffer A 550 in may be offset by -|ITD(m)| to align with the processing buffer index. Additionally, resampler 570 and resampler 580 can be implemented using the same resampling function for different input operations.
[0052] In step 611, similar to the above cases 1 and 2, output buffer A 550 and output buffer B 560 are assigned to output 0 and output 1. In intermediate buffer A and intermediate buffer B, A always corresponds to the currently non-zero ITD and delayed channel, while buffer B corresponds to the channel with zero ITD. The intermediate buffer simplifies the processing using these assumptions, and the output assignment is a simple step that can be done at the end to assign the processed buffer to the correct output channel. The assignment of the output buffer depends on the signs of ITD(m - 1) and ITD(m) following this pseudocode: · If ITD(m - 1) = 0 ○ If ITD(m) > 0 ■ Output buffer A 550 outputs 1, and output buffer B 560 outputs 0 ○ Otherwise ■ Output buffer A 550 outputs 0, and output buffer B 560 outputs 1 · Otherwise ○ If ITD(m - 1) > 0 ■ Output buffer A 550 outputs 1, and output buffer B 560 outputs 0 ○ Otherwise ■ Output buffer 550 outputs 0, and output buffer 560 outputs 1 can also be simplified to: Where ∧ represents logical AND and ∨ represents inclusive OR, (A,B)→(1,0) means that output buffer A 550 is assigned to output 1 and output buffer B 560 is assigned to output 0, and (A,B)→(0,1) means that output buffer A 550 is assigned to output 0 and output buffer B 560 is assigned to output 1. Output buffer 0 and output buffer 1 can correspond to the left and right earphone channels respectively. They can also be numbered differently, such as output buffer 1 and output buffer 2.
[0053] In other words, if ITD(m - 1) is zero, then the current ITD ITD(m) is used to determine which buffer to delay. If ITD(m) is positive, output 0 is before output 1, and output 1 should be delayed. If ITD(m) is negative, output 1 is before output 0, and output 0 should be delayed. If ITD(m - 1) is not zero, then the previous ITD ITD(m - 1) decides which buffer to shift first. If ITD(m - 1) is positive, output 1 is shifted first in buffer A 550, and then output 0 is shifted in buffer B 560. If ITD(m - 1) is negative, output 0 is shifted first in buffer A 550, and then output 1 is shifted in buffer B 560. It should be noted that the definition of the sign of ITD(m) can be reversed, in which case output 0 and output 1 will switch the above positions.
[0054] It should be noted that step 611 can occur before step 605 by already assigning output buffer A 550 and output buffer B 560 to the designated outputs 1 and 2, such that the outputs are filled during steps 605 to 611. In an embodiment, output 0 and output 1 can correspond to the left and right channels respectively.
[0055] Resampling with a sinc function
[0056] The method described depends on a resampling function to process the compressed and expanded segments of the signal. This can be implemented using a sinc resampling function, as Figure 12 shown. Given an input signal y(n) of length L in and an output length L out the input signal can be resampled at fractional indices after these steps For each n, calculate where, represents the floor operation.
[0057] Since the sinc function is computationally complex, it may be desirable to store it in a table with a predefined resolution. For example, a resolution of R sinc = 64, meaning that there are 64 samples between the zero crossings of the sinc function may be appropriate (see Figure 12 ). The corresponding index in the sinc table can then be found at the following.
[0058] The output value z(k) can be found by the following sum where
[0059] Note that while the above embodiments have been described using an audio object renderer (e.g., a decoder), the various embodiments above can also be done at the encoder, where the shifted outputs (i.e., output 0 and output 1) are shifted at the encoder rather than at the audio object renderer.
[0060] Figure 13 An audio object renderer 114 (e.g., a decoder) according to some embodiments is shown, where the audio object renderer 114 is implemented as a stand-alone device. As used herein, an audio object renderer refers to a device that is capable of, configured to, arranged to, and / or operable to decode an encoded object and communicate with a network node, an encoder, and / or a decoder. Examples of audio object renderers include, but are not limited to, smart phones, mobile phones, cellular phones, Internet Protocol voice (VoIP) phones, wireless local loop phones, desktop computers, personal digital assistants (PDAs), wireless cameras, game consoles or devices, storage devices, playback appliances, wearable terminal devices, wireless endpoints, mobile stations, tablets, laptop computers, laptop embedded devices (LEE), laptop mobile devices (LME), smart devices, wireless customer premise equipment (CPE), in-vehicle or vehicle embedded / integrated wireless devices, etc.
[0061] The audio object renderer may support device-to-device (D2D) communication, such as by implementing 3GPP standards for sidelink communication, dedicated short range communication (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-everything (V2X). In other examples, the decoder may not necessarily have a user in the sense of a human user who owns and / or operates the relevant device.
[0062] The audio object renderer 114 includes processing circuitry 1302 that is operably coupled via a bus 1304 to an input / output interface 1306, a power supply 1308, a memory 1310, a communication interface 1312, and / or any other components or any combination thereof. Some decoders may utilize Figure 13 all or a subset of the components shown therein. The level of integration between components may vary from one decoder to another. Additionally, some decoders may include multiple instances of components, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0063] The processing circuitry 1302 is configured to process instructions and data and may be configured to implement any sequential state machine operable to execute instructions stored as a machine-readable computer program in the memory 1310. The processing circuitry 1302 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.); programmable logic along with appropriate firmware; one or more stored computer programs, a general-purpose processor (such as a microprocessor or a digital signal processor (DSP)) along with appropriate software; or any combination of the above. For example, the processing circuitry 1302 may include multiple central processing units (CPUs).
[0064] In this example, the input / output interface 1306 may be configured to provide one or more interfaces to an input device, an output device, or one or more input and / or output devices. Examples of output devices include speakers, sound cards, video cards, displays, monitors, actuators, transmitters, smart cards, another output device, or any combination thereof. Input devices may allow a user to collect information into the audio object renderer 114. Examples of input devices include touch-sensitive or presence-sensitive displays, cameras (e.g., digital cameras, digital video cameras, webcams, etc.), microphones, sensors, direction pads, touchpads, rollers, smart cards, etc. A presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. The sensors may be, for example, accelerometers, gyroscopes, tilt sensors, force sensors, magnetometers, optical sensors, proximity sensors, biometric sensors, etc., or any combination thereof. Output devices may use the same type of interface port as input devices. For example, a universal serial bus (USB) port may be used to provide both input devices and output devices.
[0065] In some embodiments, power supply 1308 is configured as a battery or battery pack. Other types of power supplies may be used, such as an external power supply (e.g., a power outlet), a photovoltaic device, or a battery. Power supply 1308 may also include a power circuit for delivering power from power supply 1308 itself and / or an external power supply to various parts of the audio object renderer 114 via an input circuit or an interface such as a power cable. The delivered power may be, for example, for charging power supply 1308. The power circuit may perform any formatting, transformation, or other modification of the power from power supply 1308 to make the power suitable for the corresponding components of the audio object renderer 114 to which the power is supplied.
[0066] Memory 1310 may be or be configured to include a memory such as random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, etc. In one example, memory 1310 includes one or more applications 1314 (such as an operating system, a web browser application, a widget, a widget engine, or other applications) and corresponding data 1316. Memory 1310 may store any one of a variety of operating systems or combinations of operating systems for use by the audio object renderer 114.
[0067] Memory 1310 may be configured to include multiple physical drive units such as redundant array of independent disks (RAID), flash memory, USB flash drives, external hard disk drives, thumb drives, pen drives, key drives, high density digital versatile disc (HD-DVD) optical disc drives, internal hard disk drives, Blu-ray disc drives, holographic digital data storage (HDDS) optical disc drives, external micro dual in-line memory modules (DIMMs), synchronous dynamic random access memory (SDRAM), external micro DIMM SDRAM, smart card memory (such as a tamper-resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs) such as USIM and / or ISIM), other memories, or any combination thereof. The UICC may be, for example, an embedded UICC (eUICC), an integrated UICC (iUICC), or a removable UICC commonly referred to as a "SIM card". Memory 1310 may allow the audio object renderer 114 to access instructions, applications, etc. stored on a transient or non-transient memory medium to offload data or upload data. An article of manufacture such as a communication system may be tangibly embodied as memory 1310 or in memory 1310, and memory 1310 may be or include a device-readable storage medium.
[0068] The processing circuit 1302 may be configured to communicate with an access network or other network using the communication interface 1312. The communication interface 1312 may include one or more communication subsystems and may include or be communicatively coupled to the antenna 1322. The communication interface 1312 may include one or more transmitters for communicating, such as by communicating with one or more remote transmitters of another device capable of wireless communication (e.g., another UE or a network node in an access network). Each transceiver may include a transmitter 1318 and / or a receiver 1320 adapted to provide network communication (e.g., optical, electrical, frequency allocation, etc.). Additionally, the transmitter 1318 and the receiver 1320 may be coupled to one or more antennas (e.g., the antenna 1322) and may share circuit components, software, or firmware, or alternatively be implemented separately.
[0069] In the illustrated embodiment, the communication functions of the communication interface 1312 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communication such as Bluetooth, near-field communication, location-based communication such as using the Global Positioning System (GPS) to determine location, another similar communication function, or any combination thereof. The communication may be implemented according to one or more communication protocols and / or standards, such as IEEE 802.11, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Synchronous Optical Network (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), etc.
[0070] Regardless of the type of sensor, the audio object renderer may provide an output of the decoded data via a wireless connection to a network node through its communication interface 1312.
[0071] When in the form of an Internet of Things (IoT) device, the audio object renderer may be a device for use in one or more application domains, including but not limited to urban wearable technology, extended industrial applications, and healthcare. Non-limiting examples of such IoT devices are devices embedded in: connected refrigerators or freezers, TVs, connected lighting devices, electricity meters, robotic vacuum cleaners, voice-controlled smart speakers, home security cameras, thermostats, electric door locks, connected doorbells, autonomous vehicles, surveillance systems, weather monitoring devices, vehicle parking monitoring devices, electric vehicle charging stations, smart watches, fitness trackers, head-mounted displays for augmented reality (AR) or virtual reality (VR), wearable devices for tactile or sensory augmentation. The decoder in the form of an IoT device, in addition to regarding Figure 13In addition to the other components described by the audio object renderer 114 shown, it also includes circuits and / or software depending on the intended application of the IoT device.
[0072] Figure 14 is a block diagram of a host 1400 according to various aspects described herein. As used herein, the host 1400 can be or include various combinations of hardware and / or software, including standalone servers, blade servers, cloud-implemented servers, distributed servers, virtual machines, containers, or processing resources in a server farm. The host 1400 can provide one or more services to one or more UEs.
[0073] The host 1400 includes processing circuitry 1402 that is operably coupled via a bus 1404 to an input / output interface 1406; a network interface 1408; a power supply 1410, and a memory 1412. Other components may be included in other embodiments. The characteristics of these components may be substantially similar to the characteristics described for the devices in the previous figures (such as Figure 13 ) such that their description generally applies to the corresponding components of the host 1400.
[0074] The memory 1412 may include one or more computer programs that include one or more host applications 1414 and data 1416, which may include user data, e.g., data generated by the UE for the host 1400 or data generated by the host 1400 for the UE. Embodiments of the host 1400 may utilize only a subset or all of the shown components. The host applications 1414 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG, VP9) and audio codecs (e.g., Enhanced Voice Services (EVS), Immersive Voice Audio Service (IVAS), Free Lossless Audio Codec (FLAC), Advanced Audio Coding (AAC), MPEG, G.711), including transcoding for multiple different categories, types, or implementations of UEs (e.g., mobile phones, desktop computers, wearable display systems, head-up display systems). The host applications 1414 may also provide user authentication and license checking and may periodically report health, routing, and content availability to a central node (such as a device in or on the edge of the core network). Thus, the host 1400 can select and / or indicate different hosts for the over-the-top services for the UE. The main application 1414 may support various protocols, such as the HTTP Live Streaming (HLS) protocol, Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Adaptive Streaming over HTTP (MPEG-DASH), etc.
[0075] Figure 15FIG. is a block diagram showing a virtualized environment 1500 in which functions implemented by some embodiments of an audio object renderer 114 or components of an audio object renderer 114 may be virtualized. In this context, virtualization means creating a virtual version of a device or apparatus, which may include a virtualized hardware platform, storage devices, and network resources. As used herein, virtualization may be applied to any device or its components described herein and involves an implementation in which at least a portion of the functions are implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) that are implemented in one or more virtualized environments 1500 hosted by one or more hardware nodes, such as a hardware computing device operating as a decoder, encoder, network node, UE, core network node, or host. Additionally, in embodiments where the virtual node does not require a radio connection (e.g., a core network node or host), the node may be fully virtualized.
[0076] An application 1502 (which may alternatively be referred to as a software instance, virtual appliance, network function, virtual node, virtual network function, etc.) runs in the virtualized environment 1500 to implement some of the features, functions, and / or benefits of some embodiments disclosed herein.
[0077] The hardware 1504 includes a processing circuit, a memory storing software and / or instructions executable by the hardware processing circuit, and / or other hardware devices as described herein, such as a network interface, an input / output interface, etc. The software may be executed by the processing circuit to instantiate one or more virtualization layers 1506 (also referred to as a hypervisor or virtual machine monitor (VMM)), provide VMs 1508A and 1508B (one or more of which may generally be referred to as VM 1508), and / or perform any of the functions, features, and / or benefits described with respect to some embodiments described herein. The virtualization layer 1506 may present a virtual operating platform that appears as networked hardware to the VMs 1508.
[0078] The VMs 1508 include virtual processing, virtual memory, virtual networking or interfaces, and virtual storage, and may be run by the corresponding virtualization layer 1506. Different embodiments of instances of the virtual appliance 1502 may be implemented on one or more of the VMs 1508 and may be implemented in different ways. Virtualization of hardware is referred to as network function virtualization (NFV) in some contexts. NFV may be used to consolidate many network device types onto industry-standard high-volume server hardware, physical switches, and physical storage that may be located in data centers and customer premise equipment.
[0079] In the context of NFV, VM 1508 can be a software implementation of a physical machine running programs as if they were executed on a physical non-virtualized machine. Each of the VMs 1508 and that part of the hardware 1504 that executes the VM, which is the hardware dedicated to the VM and / or shared by the VM with other VMs in the VM, form separate virtual network elements. Still in the context of NFV, the virtual network function is responsible for handling specific network functions that run in one or more VMs 1508 above the hardware 1504 and correspond to the application 1502.
[0080] The hardware 1504 can be implemented in an independent network node with general or specific components. The hardware 1504 can implement some functions via virtualization. Alternatively, the hardware 1504 can be part of a larger hardware cluster (e.g., such as in a data center or CPE), where many hardware nodes work together and are managed via the management and orchestration 1510, which in particular supervises the lifecycle management of the application 1502. In some embodiments, the hardware 1504 is coupled to one or more radio units, each radio unit including one or more transmitters and one or more receivers that can be coupled to one or more antennas. The radio units can communicate directly with other hardware nodes via one or more suitable network interfaces and can be used in combination with virtual components to provide virtual nodes with radio capabilities, such as radio access nodes or base stations. In some embodiments, a control system 1512 can be used to provide some signaling, which can alternatively be used for communication between the hardware node and the radio unit.
[0081] Although the computing devices (e.g., decoder, audio object renderer, encoder, host) described herein may include the shown combinations of hardware components, other embodiments may include computing devices having different combinations of components. It should be understood that these computing devices may include any suitable combination of hardware and / or software required to perform the tasks, features, functions, and methods disclosed herein. The determination, calculation, obtaining, or similar operations described herein may be performed by a processing circuit that may process information by, for example, converting the obtained information into other information, comparing the obtained information or the converted information with information stored in a network node, and / or performing one or more operations based on the obtained information or the converted information and making a determination as a result of such processing. Additionally, although a component is depicted as a single box located within a larger box or nested within multiple boxes, in practice, a computing device may include multiple different physical components that make up a single shown component, and the functionality may be divided among separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of a component may be divided between a processing circuit and a communication interface. In another example, the non-computationally intensive functionality of any such component may be implemented in software or firmware, and the computationally intensive functionality may be implemented in hardware.
[0082] In some embodiments, some or all of the functionality described herein may be provided by a processing circuit that executes instructions stored in a memory, which in some embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by a processing circuit without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hard-wired manner. In any of those particular embodiments, whether or not instructions stored on a non-transitory computer-readable storage medium are executed, the processing circuit may be configured to perform the described functionality. The benefits provided by such functionality are not limited to the separate processing circuit or other components of a computing device, but are generally enjoyed by the computing device and / or typically by an end user and a wireless network.
[0083] Example embodiments 1. A method in an inter-channel time difference (ITD) synthesizer (220, 340, 1502), the method comprising: Receiving (401) a current ITD and an audio frame, where each frame m includes N samples; Storing (403) at least a portion of the current input audio frame in a signal memory; Determining (405) transition times t1, t2 for performing a time shift to be applied to at least one of output signal 0 and output signal 1 based on the inter-channel time difference (ITD) of the current input audio frame and the ITD of a previous input audio frame; and When generating output signal 0 and output signal 1, apply (407) time shift within the determined conversion times t1, t2. 2. The method according to embodiment 1, wherein the audio frame is a part of an object audio signal having position metadata describing the position relative to the listener, and the method further comprises obtaining a time shift from the position metadata. 3. The method according to any one of embodiments 1 to 2, further comprising: Based on the frame length of frame m and the look-ahead memory rs required for resampling LA calculate (801) the total conversion length, which is divided into two parts including t1 and t2; Based on the look-ahead memory rs LA determine (803) the buffer length t3; and Based on the maximum allowable conversion length and the buffer length t3, determine (805) the total conversion length. 4. The method according to embodiment 3, wherein determining the buffer length t3 includes determining the buffer length t3 according to the following formula: t3 = max(0, rs LA - |ITD(m - 1)|), and determine the total conversion length according to the following formula: t tot = N max - t3 where N max ≤ N. 5. The method according to any one of embodiments 1 to 4, wherein determining the conversion times t1, t2 includes: In response to the signs of the current ITD and the previous ITD being the same, assign the total conversion length to one of the conversion times t1, t2 and set the other conversion time to zero. 6. The method according to any one of embodiments 1 to 4, wherein determining the conversion times t1, t2 includes: In response to the signs of the current ITD and the previous ITD being different, apply a shift operation to both output signal 0 and output signal 1 by splitting the total conversion length into two parts to determine the conversion times t1, t2. 7. The method according to embodiment 6, wherein splitting the total conversion length into two parts to determine the conversion times t1, t2 includes splitting the total conversion length according to the following formula: where [·] represents the rounding operation to the nearest integer. 8. The method according to any one of embodiments 3 to 7, further comprising: Fill a processing buffer (510) using a current input audio frame of the object audio and a signal memory; Wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 includes: In response to the sign of the current ITD being the same as the sign of the previous ITD or in response to one of the current ITD and the previous ITD being zero: Adjust the processing buffer (510) to fill a first output buffer (550) by allocating the total transition length to t1 and setting t2 to zero; and Copy the input signal portion of the processing buffer (510) to a second output buffer (560). 9. The method according to embodiment 8, wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 further includes: In response to ITD(m)=ITD(m - 1) and the sign of one of the current ITD and the previous ITD being negative, thereby indicating that one of output signal 0 and output signal 1 is before the other of output signal 0 and output signal 1, delay an output buffer associated with either one of output signal 0 and output signal 1 in either the first output buffer (550) or the second output buffer (560) by the total transition length. 10. The method according to any one of embodiments 8 to 9, wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 further includes: In response to |ITD(m)|>|ITD(m - 1)|, generate a transition by: Extend the length of the frames in the processing buffer (510) from length t1 + |ITD(m - 1)| - |ITD(m)| of x buf (n), n = |ITD(m - 1)|, …, N - 1 - |ITD(m)| to output frames of length t1; and In response to the buffer length t3 being greater than zero, add the last t3 samples of the output channel by copying from the processing buffer (510), x buf (n), n = N - 1 - t3 - |ITD(m)|, … N - 1 - |ITD(m)|. 11. The method according to any one of embodiments 8 to 10, wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 further includes: In response to |ITD(m)|<|ITD(m - 1)|, add the last t3 samples of the first output buffer (550) by copying from the processing buffer (510). 12. According to the method according to any one of embodiments 8 to 11, wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 further comprises: in response to ITD(m)·ITD(m - 1)<0, splitting the total transition length according to the following formula: where [·] represents rounding to the nearest integer, and splitting the total transition length includes: resampling the samples x of length t1 + |ITD(m - 1)| buf (n), n = -|ITD(m - 1)|, …, t1 - 1 to fit the samples n = 0, …, t1 - 1 of the first output buffer (550); copying the samples x buf (n), n = 0, …, t1 - 1 to the corresponding indices in the second output buffer (560); resampling the samples x of length t2 - |ITD(m)| buf (n), n = t1, …, t1 + t2 - 1 - |ITD(m)| to fit the samples of length t2 in the second output buffer (560) n = t1, …, t1 + t2 - 1. 13. According to the method according to any one of embodiments 8 to 12, wherein applying the determined transition times t1, t2 when generating output signal 0 and output signal 1 further comprises: in response to ITD(m - 1) = 0 and ITD(m)>0, assigning the first output buffer (550) to output signal 1 and assigning the second output buffer (560) to output signal 2; in response to ITD(m - 1) = 0 and ITD(m)≤0, assigning the first output buffer (550) to output signal 0 and assigning the second output buffer (560) to output signal 1; in response to ITD(m - 1)>0, assigning the first output buffer (550) to output signal 1 and assigning the second output buffer (560) to output signal 0; and in response to ITD(m - 1)<0, assigning the first output buffer (550) to output signal 0 and assigning the second output buffer (560) to output signal 1. 14. An apparatus (114, 1502) having an ITD synthesizer adapted to: receive (401) a current ITD and an audio frame, where each frame m includes N samples; store (403) at least a portion of the current input audio frame in a signal memory; determining (405) a conversion time t1, t2 for performing a time shift applied to at least one of output signal 0 and output signal 1 based on the inter-channel time difference ITD of the current input audio frame and the ITD of the previous input audio frame; and When generating output signal 0 and output signal 1, a time shift is applied (407) within the determined transition times t1, t2. 15. The apparatus (114, 300, 1502) of embodiment 14, wherein the ITD synthesizer (220, 340, 1502) is further adapted to perform according to any one of embodiments 2 to 12. 16. An apparatus (114, 300, 1502) having an inter-channel time difference (ITD) synthesizer (220, 340, 1502), comprising: processing circuit (1202); and A memory (1210) coupled to the processing circuit, wherein the memory includes instructions that, when executed by the processing circuit, cause the ITD synthesizer (220, 340, 1502) to perform operations including: Receiving (401) a current ITD and an audio frame, wherein each frame m includes N samples; storing (403) at least a portion of a current input audio frame in a signal memory; Based on the inter-channel time difference ITD of the current input audio frame and the ITD of the previous input audio frame, a method for performing at least one of the output signal 0 and the output signal 1 is determined (405). the applied time-shifted transition times t1, t2; and When generating output signal 0 and output signal 1, a time shift is applied (407) within the determined transition times t1, t2. 17. The apparatus (114, 300, 1502) of embodiment 16, wherein the memory includes further instructions that, when executed by the processing circuit, cause the ITD synthesizer (220, 340, 1502) to perform operations according to any one of embodiments 2 to 13. 18. A computer program comprising program code to be executed by a processing circuit (1202) of an apparatus (112, 300, 1502) having an inter-channel time difference (ITD) synthesizer (220, 340, 1502), whereby execution of the program code causes the ITD synthesizer (220, 340, 1502) to perform operations comprising: Receiving (401) a current ITD and an audio frame, wherein each frame m includes N samples; storing (403) at least a portion of a current input audio frame in a signal memory; Determine (405) conversion times t1, t2 for a time shift to be applied to at least one of output signal 0 and output signal 1 based on the inter-channel time difference ITD of the current input audio frame and the ITD of a previous input audio frame; and Apply (407) the time shift within the determined conversion times t1, t2 when generating output signal 0 and output signal 1. 19. The computer program according to claim 18, comprising additional program code, whereby execution of the program code causes the ITD synthesizer (220, 340, 1502) to perform according to any one of embodiments 2 to 13. 20. A computer program product comprising a non-transitory storage medium including program code to be executed by a processing circuit (1202) of a device (114, 300, 1502), the device (112, 300, 1502) having an inter-channel time difference ITD synthesizer (220, 340, 1502), whereby execution of the program code causes the ITD synthesizer (220, 340, 1502) to perform operations including: Receiving (401) a current ITD and an audio frame, where each frame m includes N samples; Storing (403) at least a portion of the current input audio frame in a signal memory; Determine (405) conversion times t1, t2 for a time shift to be applied to at least one of output signal 0 and output signal 1 based on the inter-channel time difference ITD of the current input audio frame and the ITD of a previous input audio frame; and Apply (407) the time shift within the determined conversion times t1, t2 when generating output signal 0 and output signal 1. 19. The computer program product according to claim 19, wherein the non-transitory storage medium includes additional program code, whereby execution of the program code causes the ITD synthesizer (220, 340, 1502) to perform according to any one of embodiments 2 to 13.
Claims
1. A method for adjusting the timing of an output audio signal to achieve a desired inter-channel time difference ITD between output audio signals, the method comprising: Receiving (401) a current ITD value and an audio frame; Determining (405) conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal, based on the ITD of the current frame and the ITD of a previous frame; And Applying (407) the time shift within the determined conversion times t1, t2 when generating the first output signal and the second output signal.
2. The method according to claim 1, wherein at least a portion of the audio frame is stored (403) in a memory for synthesizing the ITD in subsequent frames.
3. The method according to any one of claims 1 to 2, wherein the audio frame is part of an audio object, the audio object comprising an audio signal and position metadata describing the position of the object, the method further comprising obtaining the time shift from the position metadata.
4. The method according to any one of claims 1 to 3, further comprising: Based on the frame length of the current frame m and the look-ahead memory rs required for resampling LA to calculate (801) the total conversion length, which is divided into two parts including t1 and t2; Based on the aforementioned advanced memory rs LA to determine (803) the conversion length t3; and Determining (805) the total conversion length based on a maximum allowable conversion length and the conversion length t3.
5. The method according to claim 4, wherein determining the conversion length t3 comprises determining the conversion length t3 according to the following formula: t3 = max(0, rs LA - |ITD(m - 1)|), where ITD(m - 1) is the ITD of the previous audio frame including N samples, and the total conversion length is determined according to the following formula: t tot = N max - t3, where N max ≤ N, where N max is the maximum allowable conversion length.
6. The method according to any one of claims 1 to 5, wherein determining the conversion times t1, t2 comprises: In response to the signs of the current ITD and the previous ITD being the same, allocating the total conversion length to one of the conversion times t1, t2 and setting the other conversion time to zero.
7. The method according to any one of claims 1 to 5, wherein determining the conversion times t1, t2 comprises: In response to the signs of the current ITD and the previous ITD being different, applying a shift operation to both the first output signal and the second output signal by splitting the total conversion length into two parts to determine the conversion times t1, t2.
8. The method according to claim 7, wherein splitting the total conversion length into two parts to determine the conversion times t1, t2 comprises splitting the total conversion length according to the following formula: where ITD(m) is the current ITD and [·] represents rounding to the nearest integer.
9. The method according to any one of claims 4 to 8, further comprising: Using the audio frame to fill a processing buffer (510); wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal comprises: In response to the signs of the current ITD and the previous ITD being the same or in response to one of the current ITD and the previous ITD being zero: Adjusting the processing buffer (510) to fill a first output buffer (550) by allocating the total conversion length to t1 and setting t2 to zero; and Copying an input signal portion of the processing buffer (510) to a second output Buffer (560).
10. The method according to claim 9, wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal further comprises: In response to ITD(m) = ITD(m - 1) and the sign of one of the current ITD and the previous ITD being negative, thereby indicating that one of the first output signal and the second output signal is before the other of the first output signal and the second output signal, delaying the output buffer associated with either one of the first output buffer (550) or the second output buffer (560) that is associated with the other of the first output signal and the second output signal by the total conversion length.
11. The method according to any one of claims 9 to 10, wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal further comprises: In response to |ITD(m)| > |ITD(m - 1)|, generating a conversion by: Extending the length of the frame in the processing buffer (510) from length t1 + x of |ITD(m - 1)| - |ITD(m)| buf (n), n = |ITD(m - 1)|, …, N - 1 - |ITD(m)| to an output frame of length t1; and In response to the conversion length t3 being greater than zero, the last t3 samples of the output channel are added by copying from the processing buffer (510), x buf (n), where n = N - 1 - t3 - |ITD(m)|, … N - 1 - |ITD(m)|.
12. The method according to any one of claims 9 to 11, wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal further comprises: In response to |ITD(m)| < |ITD(m - 1)|, adding the last t3 samples of the first output buffer (550) by copying from the processing buffer (510).
13. The method according to any one of claims 9 to 12, wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal further comprises: In response to ITD(m)·ITD(m - 1) < 0, splitting the total conversion length according to the following formula: where [·] represents rounding to the nearest integer, and splitting the total conversion length includes: Samples x of length t1 + |ITD(m - 1)| buf (n), where n = -|ITD(m - 1)|, …, t1 - 1 are resampled to fit samples n = 0, …, t1 - 1 of the first output buffer (550); Copy the sample x buf (n), where n = 0, …, t1 - 1 to the corresponding index in the second output buffer (560); Samples x of length t2 - |ITD(m)| buf (n), where n = t1, …, t1 + t2 - 1 - |ITD(m)| are resampled to fit samples of length t2 in the second output buffer (560) for n = t1, …, t1 + t2 - 1.
14. The method according to any one of claims 9 to 13, wherein generating the second output signal further comprises: In response to the conversion length t3 being greater than zero, the last y3 samples of the output channel are added by copying from the processing buffer (510), x buf (n), where n = N - 1 - y3 - |ITD(m)|, … N - 1 - |ITD(m)|.
15. The method according to any one of claims 9 to 14, wherein applying the determined conversion times t1, t2 when generating the first output signal and the second output signal further comprises: In response to ITD(m - 1) = 0 and ITD(m) > 0, assigning the first output buffer (550) to the second output signal and assigning the second output buffer (560) to the first output signal; In response to ITD(m - 1) = 0 and ITD(m) ≤ 0, assigning the first output buffer (550) to the first output signal and assigning the second output buffer (560) to the second output signal; In response to ITD(m - 1) > 0, assigning the first output buffer (550) to the second output signal and assigning the second output buffer (560) to the first output signal; And In response to ITD(m - 1) < 0, assign the first output buffer (550) to the first output signal and assign the second output buffer (560) to the second output signal.
16. An apparatus (112, 300, 1502) for adjusting the timing of output audio signals to achieve a desired inter-channel time difference ITD between the output audio signals, the apparatus being adapted to: Receive a current ITD value and an audio frame; Based on the ITD of the current frame and the ITD of the previous frame, determine conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal; And When generating the first output signal and the second output signal, apply the time shift within the determined conversion times t1, t2.
17. The apparatus (112, 300, 1502) according to claim 16, wherein the apparatus is further adapted to perform the method according to any one of claims 2 to 15.
18. An apparatus (112, 300, 1502) comprising: A processing circuit (1202); And A memory (1210) coupled to the processing circuit, wherein the memory includes instructions that, when executed by the processing circuit, cause the apparatus to perform operations, the operations including: Receive a current inter-channel time difference ITD value and an audio frame; Based on the ITD of the current frame and the ITD of the previous frame, determine conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal; and When generating the first output signal and the second output signal, apply (407) the time shift within the determined conversion times t1, t2.
19. The apparatus (112, 300, 1502) according to claim 18, wherein the memory includes additional instructions that, when executed by the processing circuit, cause the apparatus to perform the operations according to any one of claims 2 to 15.
20. A computer program comprising program code to be executed by a processing circuit (1202) of an apparatus (112, 300, 1502), whereby execution of the program code causes the apparatus to perform operations, the operations including: Receive a current ITD value and an audio frame; Based on the ITD of the current frame and the ITD of the previous frame, determine conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal; And When generating the first output signal and the second output signal, apply the time shift within the determined conversion times t1, t2.
21. The computer program according to claim 20, including additional program code, whereby execution of the program code causes the apparatus (112, 300, 1502) to perform according to any one of claims 2 to 15.
22. A computer program product, comprising a non-transitory storage medium including program code to be executed by a processing circuit (1202) of a device (112, 300, 1502), whereby execution of the program code causes the device (112, 300, 1502) to perform operations, the operations including: Receiving a current ITD value and an audio frame; Determining conversion times t1, t2 for performing a time shift to be applied to at least one of a first output signal and a second output signal, based on the ITD of the current frame and the ITD of a previous frame; And Applying the time shift within the determined conversion times t1, t2 when generating the first output signal and the second output signal.
23. The computer program product according to claim 22, wherein the non-transitory storage medium includes additional program code, whereby execution of the program code causes the device (112, 300, 1502) to perform according to any one of claims 2 to 15.