Post-processing of binaural signals
By transforming binaural audio into the frequency domain and applying HRTF-based separation, the method addresses the limitations of linear gain in existing systems, enabling precise manipulation and improved spatial processing of binaural audio.
Patent Information
- Application Number
- JP2023536843
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-02
- Filing Date
- 2021-12-16
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-12-16
AI Technical Summary
Existing audio post-processing systems for binaural audio struggle with frequency-dependent level and time differences, failing to effectively separate and process binaural audio signals due to their reliance on linear gain, which is inadequate for binaural audio's frequency-dependent characteristics.
A method for binaural audio processing that involves transforming binaural signals into a frequency domain, estimating head-related transfer functions (HRTFs), and separating sources based on these attributes to apply frequency-dependent levels and time delays to principal and residual components separately.
This approach enhances the listener's experience by allowing precise manipulation of binaural audio objects and improved spatial processing, maintaining audio quality and spatial accuracy during head movements.
Smart Images

Figure 0007778789000034 
Figure 0007778789000035 
Figure 0007778789000036
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 155,471, filed March 2, 2021, and Spanish Patent Application No. P202031265, filed December 17, 2020, which are incorporated herein by reference in their entireties.
[0002] [Technical field] The present disclosure relates to audio processing, and in particular to post-processing of binaural audio signals. [Background technology]
[0003] Unless otherwise noted, the approaches described in this section are not prior art to the claims herein and are not admitted to be prior art by inclusion in this section.
[0004] Audio source separation generally refers to extracting specific components from an audio mix in order to isolate or manipulate the level, location, or other attributes of an object present in a mixture of other sounds. Source separation methods can be based on algebraic derivation, use machine learning, etc. After extraction, some manipulation can be applied to mix the separated components with background audio. Also, in stereo or multi-channel audio, many models exist on how to isolate or manipulate objects present in the mix from specific spatial locations. These models are based on linear real-valued mixing models, e.g., the object to be extracted or manipulated is assumed to be present in the mix signal with a linear, frequency-independent gain. In other words, an object signal x with object index i is extracted from the object signal x with a specific spatial location. i , and the mixed signal s j , the assumed model is based on the unknown linear gain g according to equation (1). ij Use:
number
[0005] Binaural audio content is becoming widely available, including stereo signals intended for playback over headphones. Sources of binaural audio include rendered and captured binaural audio.
[0006] Rendered binaural audio generally represents computationally generated audio. For example, object-based audio, such as Dolby Atmos™ audio, can be rendered for headphones using head-related transfer functions (HRTFs) that introduce inter-aural time difference (ITD) and inter-aural level difference (ILD) as well as reflections occurring at the human ear. If done correctly, the perceived location of objects can be manipulated anywhere around the listener. Additionally, room reflections and delayed reverberation can be added to create a sense of perceived distance. One product with a binaural renderer that positions source objects around the listener is the Dolby Atmos Production Suite™ (DAPS) system.
[0007] Captured binaural audio generally refers to audio produced by capturing microphone signals at the ears. One way to capture binaural audio is by placing microphones at the ears of a dummy head. Another method is made possible by the strong growth of the wireless earphone market. Earphones may also contain microphones, for example for making phone calls, making binaural audio capture more accessible to consumers.
[0008] Both rendered and captured binaural audio typically requires some form of post-processing. Examples of such post-processing include reorienting or rotating the scene to compensate for head movement, readjusting the levels of specific objects relative to the background (e.g., boosting speech or dialogue levels or attenuating background sounds or room reverberation), equalization or dynamic range processing of specific objects in the mix or only from specific directions such as in front of the listener, etc. Summary of the Invention
[0009] Existing audio post-processing systems have many problems. One problem is that many existing signal decomposition and upmixing processes use linear gain. While linear gain works well for channel-based signals such as stereo audio, it does not work well for binaural audio, which has frequency-dependent level and time differences. There is a need for an improved upmix process that works well for binaural audio.
[0010] While methods exist for reorienting and rotating binaural signals, these methods generally operate on the full mix or only on coherent elements to achieve relative rotational changes. It is necessary to separate binaurally rendered objects from the mix and perform different processing based on different objects.
[0011] Embodiments relate to a method for extracting and processing one or more objects from a binaural rendition or capture, the method focusing on (1) estimating attributes of HRTFs used during the rendering or present in the capture, (2) separating sources based on the estimated HRTF attributes, and (3) processing the separated source or sources.
[0012] According to an embodiment, a computer-implemented method for audio processing includes performing a signal transformation on a binaural signal, including transforming the binaural signal from a first signal domain to a second signal domain and generating a transformed binaural signal, where the first signal domain is the time domain and the second signal domain is the frequency domain. The method further includes performing spatial analysis on the transformed binaural signal, where performing the spatial analysis includes generating estimated rendering parameters, where the estimated rendering parameters include a level difference and a phase difference. The method further includes extracting estimated objects from the transformed binaural signal using at least a first subset of the estimated rendering parameters, where extracting the estimated objects includes generating a left principal component signal, a right principal component signal, a left residual component signal, and a right residual component signal. The method further includes performing object processing on the estimated objects using at least a second subset of the estimated rendering parameters, where performing object processing includes generating processed signals based on the left principal component signal, the right principal component signal, the left residual component signal, and the right residual component signal.
[0013] As a result, the listener's experience is improved as the system can apply different frequency-dependent levels and time delays to the binaural signals.
[0014] Generating the processed signals includes generating a left principal processed signal and a right principal processed signal from the left principal component signal and the right principal component signal using a first set of object processing parameters, and generating a left residual processed signal and a right residual processed signal from the left residual component signal and the right residual component signal using a second set of object processing parameters, the second set of object processing parameters being different from the first set of object processing parameters. In this manner, the principal components and the residual components can be processed separately.
[0015] According to another embodiment, an apparatus includes a processor configured to control the apparatus to implement one or more of the methods described herein. The apparatus may further include similar details to one or more of the methods described herein.
[0016] According to another embodiment, a non-transitory computer readable medium stores a computer program that, when executed by a processor, controls an apparatus to perform processes including the methods described herein.
[0017] The following detailed description and accompanying drawings provide a further understanding of the features and advantages of various implementations. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a block diagram of an audio processing system 100.
[0019] [Figure 2] FIG. 2 is a block diagram of an object processing system 208.
[0020] [Figure 3A] 1 illustrates an embodiment of the object processing system 108 (see FIG. 1) for re-rendering. [Figure 3B] 1 illustrates an embodiment of the object processing system 108 (see FIG. 1) for re-rendering.
[0021] [Figure 4] FIG. 4 is a block diagram of an object processing system 408.
[0022] [Figure 5] FIG. 5 is a block diagram of an object processing system 508.
[0023] [Figure 6] 6 illustrates a device architecture 600 for implementing the features and processes described herein, according to an embodiment.
[0024] [Figure 7] 7 is a flowchart of a method 700 of audio processing. DETAILED DESCRIPTION OF THE INVENTION
[0025] This specification describes techniques related to audio processing. Throughout the following detailed description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present disclosure, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may also include modifications and equivalents of the features and concepts that may be described herein.
[0026] In the following description, various methods, processes, and procedures are set forth. While certain steps may be described in a particular order, such order is primarily for convenience and clarity. Certain steps may be repeated more than once, may occur before or after other steps, and may occur in parallel with other steps, even if those steps are described in a different order. A second step need only follow a first step if the first step must be completed before the second step can begin. Such situations will be specifically pointed out when not clear from the context.
[0027] The terms "and," "or," and "and / or" are used herein. Such terms should be construed as having an inclusive meaning. For example, "A and B" may mean at least the following: "both A and B," "at least both A and B." For example, "A or B" may mean at least the following: "at least A," "at least B," "both A and B," "at least both A and B." For example, "A and / or B" may mean at least the following: "A and B," "A or B." When an exclusive or is intended, such should be particularly noted. For example, "either A or B," "at most one of A and B," etc.
[0028] This specification describes various processing functions that can be associated with structures such as blocks, elements, components, circuits, etc. Generally, these structures may be implemented by a processor controlled by one or more computer programs.
[0029] 1. Binaural post-processing system
[0030] As described in more detail below, embodiments describe methods for extracting one or more components from a binaural mixture and estimating their position or rendering parameters, including (1) frequency dependence and (2) relative time differences, which enables one or more of the following: precise manipulation of the position of one or more objects in a binaural performance or capture, processing of one or more objects in a binaural performance or capture where the processing depends on the estimated position of each object, and source separation, including estimation of the position of each source from the binaural performance or capture.
[0031] FIG. 1 is a block diagram of an audio processing system 100. The audio processing system 100 can be implemented by one or more computer programs executed by one or more processors. The processor may be a component of a device that implements the functionality of the audio processing system 100, such as a headset, headphones, a mobile phone, or a laptop computer. The audio processing system 100 includes a signal transformation system 102, a spatial analysis system 104, an object extraction system 106, and an object processing system 108. The audio processing system 100 may include other components and functions not discussed in detail (for brevity). Generally, in the audio processing system 100, a binaural signal is first processed by the signal transformation system 102 using a time-frequency transform. The spatial analysis system 104 then estimates rendering parameters, such as binaural rendering parameters, including level and time differences applied to one or more objects. These one or more objects are then extracted by the object extraction system 106 and / or processed by the object processing system 108. The following paragraphs provide a detailed description of each component.
[0032] The signal transformation system 102 receives the binaural signal 120 and performs a signal transformation on the binaural signal 120 to generate a transformed binaural signal 122. The signal transformation includes transforming the binaural signal 120 from a first signal domain to a second signal domain. The first signal domain may be the time domain, and the second signal domain may be the frequency domain. The signal transformation may be one of many time-to-frequency transformations, including a Fourier transform, such as a fast Fourier transform (FFT) or a discrete Fourier transform (DFT), a quadrature mirror filter (QMF) transform, a complex QMF (CQMF) transform, a hybrid CQMF (HCQMF) transform, etc. The signal transformation may result in a complex-valued signal.
[0033] Generally, the signal conversion system 102 provides some time / frequency separation to the binaural signal 120, resulting in the converted binaural signal 122. For example, the signal conversion system 102 may convert blocks or frames of the binaural signal 120, e.g., 10-100 ms blocks, such as 20 ms blocks. The converted binaural signal 122 then corresponds to a set of time-frequency tiles for each converted block of the binaural signal 120. The number of tiles depends on the number of frequency bands implemented by the signal conversion system 102. For example, the signal conversion system 102 may be implemented by a filter bank with between 10 and 100 bands, such as 20 bands, in which case the converted binaural signal 122 will have the same number of time-frequency tiles.
[0034] The spatial analysis system 104 receives the transformed binaural signal 122 and performs spatial analysis on the transformed binaural signal 122 to generate a number of estimated rendering parameters 124. Generally, the estimated rendering parameters 124 correspond to parameters such as head-related transfer functions (HRTFs), head-related impulse responses (HRIRs), and binaural room impulse responses (BRIRs). The estimated rendering parameters 124 include a number of level differences (parameter h), as described in more detail below, and a number of phase differences (parameter φ), as described in more detail below.
[0035] The object extraction system 106 receives the transformed binaural signal 122 and the estimated rendering parameters 124 and performs object extraction on the transformed binaural signal 122 using the estimated rendering parameters 124 to generate a number of estimated objects 126. Generally, the object extraction system 106 generates one object for each time-frequency tile of the transformed binaural signal 122. For example, for 100 tiles, the number of estimated objects will be 100.
[0036] Each estimated object can be represented as a principal component signal, hereafter denoted as x, and a residual component signal, hereafter denoted as d. The principal component signal is the left principal component signal x l and the right principal component signal x r The residual component signal can include a left residual component signal d l and the right residual component signal d r The estimated object 126 then includes four component signals for each time-frequency tile.
[0037] The object processing system 108 receives the estimated objects 126 and the estimated rendering parameters 124 and performs object processing on the estimated objects 126 using the estimated rendering parameters 124 to generate a processed signal 128. The object processing system 108 may use a different subset of the estimated rendering parameters 124 than that used by the object extraction system 106. The object processing system 108 may implement many different object processing processes, as described in more detail below.
[0038] 2. Spatial analysis and object extraction
[0039] The audio processing system 100, as implemented by the spatial analysis system 104 and the object extraction system 106, may perform many calculations as part of performing spatial analysis and object extraction. These calculations may include one or more of HRTF estimation, phase unwrapping, object estimation, object separation, and phase alignment.
[0040] 2.1.HRTF Estimation
[0041] In the following, we assume that the signals exist in subbands and time frames using a time-frequency transform that provides complex-valued signals (e.g., DFT, CQMF, HCQMF, etc.). We assume that within each time / frequency tile, we can model a complex-valued binaural signal pair (l[n], r[n]) with n frequency or time indices according to equations (2a)-(2b).
number
[0042] Complex phase angle φ l and φ r represents the phase shift introduced by the HRTF within a narrow subband. l and h rrepresents the magnitude of the HRTF applied to the principal component signal x, and d r are the two unknown residual signals. In most cases, the HRTF φ l and φ r We are not interested in the absolute phase of . Instead, we can use the inter-aural phase difference (IPD) φ. By pushing the IPD φ to the right channel signal, our signal model can be expressed as equations (3a)-(3b):
number
[0043] Similarly, we may be primarily interested in estimating the head shadow effect (e.g., inter-aural level difference (ILD)). Therefore, we can write a model using a real-valued head shadow attenuation h, as in equations (4a)-(4b).
number
[0044] Assume that the expected value of the dot product of the residual signal is 0, as in equation (5):
number
[0045] Furthermore, we assume that the expectation of the dot product of signal x and any residual signal is also zero, as in equation (6):
number
[0046] Finally, we also require that the two residual signals have equal energy, as in equation (7):
number
[0047] Next, the relative IPD phase angle φ is calculated directly as in equation (8):
number
[0048] That is, the phase difference for each tile is calculated as the phase angle of the dot product of the left component l of the transformed binaural signal (eg, 122 in FIG. 1) and the right component r* of the transformed binaural signal.
[0049] Next, create a modified right channel signal r' by applying the relative phase angle, as in equation (9):
number
[0050] The principal component x^' is estimated from l[n] and r'[n] according to a weighted combination as shown in equation (10):
number
[0051] In equation (10), the caret or hat symbol ^ represents the estimate, and the weight w'r can be calculated according to equation (11):
number
[0052] The cost function Ex can be formulated as in equation (12):
number
[0053] Setting the partial derivatives of:
number
number
[0054] Then, equations (14a) to (14c) can be written as follows:
number
[0055] Substitution leads to equations (15a)-(15i):
number
[0056] Like equation (16), equations (15a)-(15i) provide a solution for the level difference h present in the HRTFs:
number
[0057] That is, the level difference for each tile is calculated according to a quadratic equation based on the left component of the transformed binaural signal, the right component of the transformed binaural signal, and the phase difference. An example of the left component of the transformed binaural signal is the left component of 122 in FIG. 1, represented by the variables l and l* in Equations A, B, and C. An example of the right component of the transformed binaural signal is the right component of 122 in FIG. 1, represented by the variables r' and r'* in Equations A, B, and C. An example of the phase difference is the phase difference information of the estimated rendering parameter 124, represented by the IPD phase angle φ in Equation (8), which is used to calculate r' according to Equation (9).
[0058] As a specific example, the spatial analysis system 104 (see FIG. 1 ) can estimate the HRTFs by manipulating the transformed binaural signal 122 using equations (1)-(16), in particular equation (8) to generate the IPD phase angle φ and equation (16) to generate the level difference h as part of generating the estimated rendering parameters 124.
[0059] 2.2. Phase Unwrapping
[0060] In the previous section, the estimated IPDφ is always wrapped to a 2π interval according to equation (8). To accurately determine the location of a given object, the phase needs to be unwrapped. In general, unwrapping refers to using neighboring bands to determine the most likely location given multiple possible locations indicated by the wrapped IPD. Different strategies can be employed to unwrap the phase: evidence-based unwrapping and model-based unwrapping.
[0061] 2.2.1. Evidence-Based Unwrapping
[0062] Evidence-based phase unwrapping can use information from neighboring bands to derive an optimal estimate of the unwrapped IPD. Suppose we have three IPD estimates for neighboring subbands b-1, b, and b+1, and φ b-1 , φ b , φ b+1 The unwrapped phase candidate φ^ for band b is b is given by the following equation (17):
number
[0063] Each candidate φ^ b,Nb is expressed as ITDτ^ in the following equation (18). b,N Having:
number
[0064] In equation (18), f b represents the center frequency of band b. 2 b There is also an estimate of the total energy of the principal components of , given by equation (19):
number
[0065] Therefore, the principal component x of band b b The cross-correlation function Rb(τ) of band b as a function of the ITD τ can be modeled as in equation (20):
number
[0066] Now, for each unwrapped IPD candidate, we can accumulate the energy across neighboring bands v and take the maximum as the estimate that occupies most of the energy in a single ITD across bands, as in equation (21):
number
[0067] That is, the system estimates the total energy of the left and right principal component signals in each band, calculates the cross-correlation based on each band, and selects an appropriate phase difference for each band according to the energy between neighboring bands based on the cross-correlation.
[0068] 2.2.2. Model-Based Unwrapping
[0069] In model-based unwrapping, given estimates of the head shadow parameters, e.g., as in Eq. (16), a simple HRTF model (e.g., a spherical head model) is used to calculate N^ given the value of h in band b. b The optimal value of h can be found, i.e., the optimal unwrapped phase that matches the magnitude of a given head shadow magnitude. This unwrapping can be performed computationally given the model and the values of h for the various bands. That is, the system selects the appropriate phase difference for a given band from many candidate phase differences depending on the level difference for that band applied to the head-related transfer function.
[0070] As a specific example, for both types of unwrapping, the spatial analysis system 104 (see FIG. 1) can perform phase unwrapping as part of generating the estimated rendering parameters 124 .
[0071] 2.3. Main Object Estimation
[0072] <xx*> ,<dd*> , and the weights w according to the estimates of h (according to equations (15a), (15b), and (16)). l , w' r We can calculate the equations (10) to (11) as well. We can then repeat equations (13a) to (13b) from above as equations (22a) to (22b):
number
[0073] Next, the weight w l , w' r can be calculated as:
number
[0074] As a specific example, the spatial analysis system 104 (see FIG. 1) (see FIG. 1) may perform the estimation of the dominant object by generating weights as part of generating the estimated rendering parameters 124.
[0075] 2.4. Separation of the main object and residual
[0076] The system can estimate two binaural signal pairs: one for the rendered principal components and one for the residuals. The rendered principal component pair can be expressed as Equations (24a)-(24b):
number
[0077] In equations (24a)-(24b), the signal l x [n] corresponds to the left principal component signal (e.g., 220 in Figure 2), and the signal r x [n] corresponds to the right principal component signal (e.g., 222 in Figure 2). Equations (24a)-(24b) can be expressed in terms of the upmix matrix M as in Equation (25):
number
[0078] residual signal l d [n] and r d [n] can be estimated as in equation (26):
number
[0079] In equation (26), the signal l d Signal [n] corresponds to the left residual component signal (eg, 224 in FIG. 2), and signal [n] corresponds to the right residual component signal (eg, 226 in FIG. 2).
[0080] The perfect reconstruction requirement gives the expression for D according to equation (27):
number
[0081] In equation (27), I corresponds to the identity matrix.
[0082] As a specific example, object extraction system 106 (see FIG. 1) can perform main object estimation as part of generating estimated object 126. Estimated object 126 can then be provided to an object processing system (e.g., 108 in FIG. 1, 208 in FIG. 2, etc.), e.g., as component signals 220, 222, 224, and 226 (see FIG. 2).
[0083] 2.5. Global phase matching
[0084] Up to now, all phase matching has been applied to the right channel and the right channel prediction coefficients. See, for example, equation (9). To obtain a more balanced distribution, one strategy is to align the phase of the extracted principal components and residuals to the downmix m as in equation m = l + r. The phase shift θ applied to the two prediction coefficients is given by equation (28):
number
[0085] Next, the weighting equations in (10) and (23a)-(23b) are modified using the phase shift θ, and our signal x̂ is θ gives the final predicted coefficients of:
number
[0086] This modifies equation (25) to equation (30):
number
[0087] Therefore, the submix extraction matrix M does not change as a result of θ, but as in Eq. (31), x̂ θ The prediction coefficients for calculating depend on θ:
number
[0088] Finally, x^ θ The re-rendering of is given by equation (32):
number
[0089] As a specific example, the spatial analysis system 104 (see FIG. 1 ) may perform part of the global phase matching as part of generating the weights as part of generating the estimated rendering parameters 124, and the object extraction system 106 may perform part of the global phase matching as part of generating the estimated objects 126.
[0090] 3. Object Processing
[0091] As previously mentioned, the object processing system 108 can implement a number of different object processing processes. These object processes include one or more of rearrangement, level adjustment, equalization, dynamic range adjustment, dessing, multiband compression, immersion enhancement, enveloping, upmixing, conversion, channel remapping, storage, and archiving. Rearrangement generally refers to moving one or more identified objects within a perceived audio scene, such as by adjusting HRTF parameters of the left and right component signals of a processed binaural signal. Level adjustment generally refers to adjusting the level of one or more identified objects within a perceived audio scene. Equalization generally refers to adjusting the timbre of one or more identified objects by applying a frequency-dependent gain. Dynamic range adjustment generally refers to adjusting the loudness of one or more identified objects to fit within a defined loudness range. For example, adjusting the audio so that a nearby speaker is not perceived as too loud and a distant speaker is not perceived as too quiet. De-essing generally refers to sibilance reduction, such as reducing the listener's perception of harsh consonants like "s," "sh," "x," "ch," "t," and "th." Multiband compression generally refers to applying different loudness adjustments to different frequency bands of one or more identified objects. For example, reducing the loudness and loudness range of the noise band and increasing the loudness of the speech band. Immersion enhancement generally refers to adjusting the parameters of one or more identified objects to match other sensory information, such as a video signal. For example, matching a moving sound to a moving, three-dimensional collection of video pixels or adjusting the wet / dry balance so that echoes correspond to the perceived visual size of the room. Enveloping generally refers to adjusting the position of one or more identified objects to enhance the perception that sounds are emanating from all around the listener.Upmixing, conversion, and channel remapping generally refer to changing one type of channel arrangement to another. Upmixing generally refers to increasing the number of channels in an audio signal, for example, upmixing a two-channel signal such as binaural audio to a 12-channel signal such as 7.1.4-channel surround sound. Conversion generally refers to reducing the number of channels in an audio signal, for example, converting a six-channel signal such as 5.1-channel surround sound to a two-channel signal such as stereo audio. Channel remapping generally refers to an operation that includes both upmixing and conversion. Storage and archiving generally refers to storing a binaural signal as one or more extracted objects with associated metadata and a binaural residual signal.
[0092] Various audio processing systems and tools may be used to perform the object processing process, including, for example, the Dolby Atmos Production Suite™ (DAPS) system, the Dolby Volume™ system, the Dolby Media Enhance™ system, and the Dolby™ Mobile Capture Audio Processing System.
[0093] The following figures provide details of object processing in various embodiments of the audio processing system 100.
[0094] 2 is a block diagram of an object processing system 208. The object processing system 208 can be used as the object processing system 108 (see FIG. 1).
[0095] The object processing system 208 receives a left principal component signal 220, a right principal component signal 222, a left residual component signal 224, a right residual component signal 226, a first set of object processing parameters 230, a second set of object processing parameters 232, and estimated rendering parameters 124 (see FIG. 1). Component signals 220, 222, 224, and 226 are component signals corresponding to an estimated object 126 (see FIG. 1). The estimated rendering parameters 124 include level and phase differences calculated by the spatial analysis system 104 (see FIG. 1).
[0096] The object processing system 208 uses object processing parameters 230 to generate a left main processed signal 240 and a right main processed signal 242 from the left main component signal 220 and the right main component signal 222. The object processing system 208 uses object processing parameters 232 to generate a left residual processed signal 244 and a right residual processed signal 246 from the left residual component signal 224 and the right residual component signal 226. The processed signals 240, 242, 244, and 246 correspond to the processed signal 128 (see FIG. 1). The object processing system 208 can perform direct-feed processing, such as generating a left (or right) main (or residual) processed signal from only the left (or right) main (or residual) component signal. The object processing system 208 can perform cross-feed processing, such as generating a left (or right) main (or residual) processed signal from both the left and right main (or residual) component signals.
[0097] Depending on the specific type of processing performed, the object processing system 208 can use one or more level differences and one or more phase differences of the estimated rendering parameters 124 when generating one of the processed signals 240, 242, 244, and 246. As an example, the rearrangement uses at least some, e.g., all, of the level differences and at least some, e.g., all, of the phase differences. As another example, the level adjustment uses at least some, e.g., all, of the level differences and fewer, e.g., none, of the phase differences. As another example, the rearrangement uses fewer, e.g., none, of the level differences and at least some, e.g., low frequencies below 1.5 kHz. Using only low frequencies is acceptable because inter-channel phase differences above these frequencies do not contribute significantly to the perceived location of the source, but changing the phase can introduce audible artifacts. Therefore, adjusting only low-frequency phase differences and leaving high-frequency phase differences intact may result in a better tradeoff between audio quality and perceived location.
[0098] The object processing parameters 230 and 232 enable the object processing system 208 to use one set of parameters to process the principal component signals 220 and 222 and another set of parameters to process the residual component signals 224 and 226. This allows for differential processing of the principal and residual components when performing the different object processing processes described above. For example, in a rearrangement, the principal components may be rearranged as determined by the object processing parameters 230, but the object processing parameters 232 are such that the residual components are unchanged. As another example, in a multi-band compression, a band of principal components may be compressed using the object processing parameters 230, and a band of residual components may be compressed using different object processing parameters 232.
[0099] The object processing system 208 may include additional components to perform additional processing steps. One of the additional components is an inverse transform system. The inverse transform system performs an inverse transform on the processed signals 240, 242, 244, and 246 to generate time-domain processed signals. The inverse transform is the inverse of the transform performed by the signal transformation system 102 (see FIG. 1).
[0100] Another additional component is a time-domain processing system. Several audio processing techniques work well in the time domain, such as delay effects, echo effects, reverberation effects, pitch shifting, timbre modification, etc. By implementing a time-domain processing system after the inverse transform system, the object processing system 208 can perform time-domain processing on the processed signal to generate a modified time-domain signal.
[0101] The details of object processing system 208 may otherwise be similar to the details of object processing system 108 .
[0102] 3A-3B illustrate an embodiment of object processing system 108 (see FIG. 1) for re-rendering. FIG. 3A is a block diagram of object processing system 308, which can be used as object processing system 108. Object processing system 308 receives left principal component signal 320, right principal component signal 322, left residual component signal 324, right residual component signal 326, and sensor data 330. Component signals 320, 322, 324, and 326 are component signals corresponding to estimated object 126 (see FIG. 1). Sensor data 330 corresponds to data generated by sensors such as gyroscopes or other head-tracking sensors located on a device such as a headset, headphones, earphones, or microphone.
[0103] The object processing system 308 uses the sensor data 330 to generate a left main processed signal 340 and a right main processed signal 342 based on the left main component signal 320 and the right main component signal 322. The object processing system 308 generates a left residual processed signal 344 and a right residual processed signal 346 from the sensor data 330 without modification. The object processing system 308 can use direct-feed processing or cross-feed processing in a manner similar to the object processing system 208 (see FIG. 2). The object processing system 308 can use binaural panning to generate the main processed signals 340 and 342. That is, the main component signals 320 and 322 are treated as objects to which binaural panning is applied, and the diffuse sound of the residual component signals 324 and 326 is not modified.
[0104] Alternatively, the object processing system 308 may generate a mono object from the left principal component signal 320 and the right principal component signal 322 and perform binaural panning on the mono object using the sensor data 330. The object processing system 308 may also generate the mono object using a phase-aligned downmix.
[0105] Furthermore, as head-tracking systems become a common feature in high-end earphone and headphone products, it becomes possible to know the listener's orientation in real time and rotate the scene accordingly, for example in virtual reality, augmented reality, or other immersive media applications. However, unless an object-based presentation is available, the effectiveness and quality of rotation methods in rendered binaural presentations are limited. To address this issue, the object extraction system 106 (see FIG. 1) isolates and estimates the principal components, and the object processing system 308 treats the principal components as objects and applies binaural panning while leaving the remaining diffuse sound untouched. This enables applications such as:
[0106] One application example is an object processing system 308 that rotates the audio scene according to the listener's viewpoint while preserving the precise position conveyed by the objects, without compromising the spatiality of the audio scene conveyed by the ambience in the afterimage.
[0107] Another application is an object processing system 308 that compensates for unwanted head rotations that occur during recording with binaural earphones or microphones. Head rotations can be inferred from the positions of the principal components. For example, assuming the principal components are stationary, any detected changes in position can be corrected. Head rotations can also be inferred by acquiring head tracking data synchronously with the audio recording.
[0108] 3B is a block diagram of an object processing system 358, which can be used as the object processing system 108 (see FIG. 1). The object processing system 358 receives a left principal component signal 370, a right principal component signal 372, a left residual component signal 374, a right residual component signal 376, and configuration information 380. The component signals 370, 372, 374, and 376 are component signals corresponding to the estimated object 126 (see FIG. 1). The configuration information 380 corresponds to a channel layout for upmixing, conversion, or channel remapping.
[0109] The object processing system 358 uses the configuration information 380 to generate a multi-channel output signal 390. The multi-channel output signal 390 then corresponds to the particular channel layout specified in the configuration information 380. For example, if the configuration information 380 specifies upmixing to 5.1 channel surround sound, the object processing system performs upmixing from the component signals 370, 372, 374, and 376 to generate six channels of a 5.1 channel surround sound channel signal.
[0110] More specifically, reproducing binaural recordings with a loudspeaker layout poses several challenges if one wishes to preserve the recording's spatial characteristics. Typical solutions involve crosstalk cancellation, which tends to be effective only in very small listening areas in front of the loudspeakers. By using the separation of the principal and residual components and estimating the location of the principal components, the object processing system 358 can treat the principal components as dynamic objects with relative positions over time, which can be accurately rendered across various loudspeaker layouts. The object processing system 358 can process the diffuse components using a 2-to-N-channel upmixer to form an immersive channel-based bed. Together, the dynamic objects resulting from the principal components and the channel-based bed resulting from the residual components result in an immersive presentation of the original binaural recording with any set of loudspeakers. An example of a system for generating an upmix of diffuse content is one in which the diffuse content is decorrelated and dispersed according to an orthogonal matrix, as described in Mark Vinton, David McGrath, Charles Robinson and Phillip Brown, “Next Generation Surround Decoding and Upmixing for Consumer and Professional Applications”, in 57th International Conference: The Future of Audio Entertainment Technology - Cinema, Television and the Internet (March 2015).
[0111] The advantage of this time-frequency decomposition over many existing systems is that re-panning can be different for each object, rather than rotating the entire sound field to accommodate head movement. Additionally, many existing systems add excessive interaural time delay (ITD) to the signal, potentially resulting in a delay greater than natural. The object processing system 358 helps overcome these issues compared to these existing systems.
[0112] 4 is a block diagram of an object processing system 408, which can be used as the object processing system 108 (see FIG. 1). The object processing system 408 receives a left principal component signal 420, a right principal component signal 422, a left residual component signal 424, a right residual component signal 426, and configuration information 430. The component signals 420, 422, 424, and 426 are component signals corresponding to the estimated object 126 (see FIG. 1). The configuration information 430 corresponds to configuration settings for the audio improvement process.
[0113] The object processing system 408 uses the configuration information 430 to generate a left main processed signal 440 and a right main processed signal 442 based on the left main component signal 420 and the right main component signal 422. The object processing system 408 generates a left residual processed signal 444 and a right residual processed signal 446 without modification from the configuration information 430. The object processing system 408 can use direct-feed processing or cross-feed processing in a manner similar to the object processing system 208 (see FIG. 2). The object processing system 408 can use manual audio improvement processing parameters provided by the configuration information 430, or the configuration information 430 can correspond to settings for automatic processing by an audio improvement processing system such as that described in International Publication WO 2020 / 014517. That is, the main component signals 420 and 422 are treated as objects to which audio improvement processing is applied, and the diffuse sound of the residual component signals 424 and 426 is not altered.
[0114] Specifically, binaural recordings of audio content, such as podcasts or vlogs, often contain contextual environmental sounds, such as crowd noise, nature sounds, and urban noise, alongside the speech. It is often desirable to improve speech quality, such as level, tonality, and dynamic range, without affecting the background sounds. Separation into dominant and residual components allows the object processing system 408 to perform independent processing. Level, equalization, sibilance reduction, and dynamic range adjustments can be applied to the dominant components based on configuration information 430. After processing, the object processing system 408 recombines the signals into processed signals 440, 442, 444, and 446 to form an enhanced binaural presentation.
[0115] 5 is a block diagram of an object processing system 508, which can be used as the object processing system 108 (see FIG. 1). The object processing system 508 receives a left principal component signal 520, a right principal component signal 522, a left residual component signal 524, a right residual component signal 526, and configuration information 530. The component signals 520, 522, 524, and 526 are component signals corresponding to the estimated object 126 (see FIG. 1). The configuration information 530 corresponds to the configuration settings of the level adjustment process.
[0116] The object processing system 508 generates a left main processed signal 540 and a right main processed signal 542 based on the left main component signal 520 and the right main component signal 522 using a first set of level adjustment values in the configuration information 530. The object processing system 508 generates a left residual processed signal 540 and a right residual processed signal 542 based on the left residual component signal 520 and the right residual component signal 522 using a second set of level adjustment values in the configuration information 530. The object processing system 508 can use direct-feed processing or cross-feed processing in a manner similar to the object processing system 208 (see FIG. 2).
[0117] More specifically, recordings made in reverberant environments, such as large indoor spaces or rooms with reflective surfaces, may contain a significant amount of reverberation, especially if the sound source of interest is not close to the microphone. Excessive reverberation can reduce the intelligibility of the sound source. In binaural recordings, reverberant sounds and ambient sounds, such as non-localized noise from nature or machinery, tend to be uncorrelated in the left and right channels and therefore remain primarily in the residual signal after decomposition is applied. This characteristic allows the object processing system 508 to control the amount of ambient sound in the recording, e.g., the perceived amount of reverberation, by controlling the relative levels of the dominant and residual components and adding them to the modified binaural signal. The modified binaural signal can then be, for example, reduced in residual content for increased intelligibility or reduced in dominant content for increased perceived immersion.
[0118] The desired balance between the dominant and residual components set in the configuration information 530 can be defined manually, such as by manipulating a fader or a "balance" knob, or it can be determined automatically based on an analysis of their relative levels and a definition of the desired balance between them. In one embodiment, such analysis is a comparison of the root-mean-square (RMS) levels of the dominant and residual components over the entire recording. In another embodiment, the analysis is performed adaptively over time, adjusting the relative levels of the dominant and residual signals accordingly in a time-varying manner. For speech content, this process can be preceded by content analysis such as voice activity detection, which can modify the relative balance between the dominant and residual components during speech or non-speech portions in different ways.
[0119] 4. Hardware and Software Details
[0120] The following paragraphs describe various hardware and software details related to the aforementioned binaural post-processing.
[0121] 6 illustrates a device architecture 600 for implementing the features and processes described herein, according to an embodiment. Architecture 600 can be implemented in any electronic device, including, but not limited to, a desktop computer, consumer audio / visual (AV) equipment, wireless broadcast equipment, mobile devices such as smartphones, tablet computers, laptop computers, wearable devices, etc. In the exemplary embodiment shown, architecture 600 is for a laptop computer and includes a processor 601, a peripherals interface 602, an audio subsystem 603, a speaker 604, a microphone 605, sensors 606 such as an accelerometer, gyro, barometer, magnetometer, camera, etc., a position processor 607, such as a GNSS receiver, a wireless communication subsystem 608 such as Wi-Fi, Bluetooth, cellular, etc., and an I / O subsystem 609 including a touch controller 610 and other input controllers 611, a touch surface 612, and other input / control devices 613. Other architectures having more or fewer components can also be used to implement the disclosed embodiments.
[0122] Memory interface 414 couples to processor 601, peripherals interface 602, and memory 615, e.g., Flash, RAM, ROM, etc. Memory 615 stores computer program instructions and data, including, but not limited to, operating system instructions 616, communications instructions 617, GUI instructions 618, sensor processing instructions 619, telephony instructions 620, electronic messaging instructions 621, web browsing instructions 622, audio processing instructions 623, GNSS / navigation instructions 624, and applications / data 625. Audio processing instructions 623 include instructions for performing the audio processing described herein.
[0123] According to an embodiment, architecture 600 may correspond to a computer system, such as a laptop computer, implementing audio processing system 100 (see FIG. 1), one or more object processing systems described herein (e.g., 208 of FIG. 2, 308 of FIG. 3A, 358 of FIG. 3B, 408 of FIG. 4, 508 of FIG. 5, etc.), etc.
[0124] According to an embodiment, architecture 600 can correspond to multiple devices. The multiple devices can communicate via wired or wireless connections, such as IEEE 802.15.1 standard connections. For example, architecture 600 can correspond to a computer system or mobile phone implementing processor 601, a headset implementing audio subsystem 603 such as a speaker, one or more sensors 606 such as a gyroscope or other head tracking sensor, etc. For example, architecture 600 can correspond to earphones implementing computer system or mobile phone implementing processor 601, an audio subsystem 603 such as a microphone and speaker, etc.
[0125] Figure 7 is a flowchart of an audio processing method 700. Method 700 can be performed by a device, such as a laptop computer, a mobile phone, or the like, having components of architecture 600 of Figure 6 to implement the functionality of, for example, audio processing system 100 (see Figure 1), one or more object processing systems described herein (e.g., 208 of Figure 2, 308 of Figure 3A, 358 of Figure 3B, 408 of Figure 4, 508 of Figure 5, etc.), etc., by executing one or more computer programs.
[0126] At 702, a signal transformation is performed on the binaural signal. Performing the signal transformation includes transforming the binaural signal from a first signal domain to a second signal domain and generating a transformed binaural signal. The first signal domain may be the time domain and the second signal domain may be the frequency domain. For example, the signal transformation system 102 (see FIG. 1) may transform the binaural signal 120 to generate the transformed binaural signal 122.
[0127] At 704, spatial analysis is performed on the transformed binaural signal. Performing the spatial analysis includes generating estimated rendering parameters, where the estimated rendering parameters include level differences and phase differences. For example, the signal transformation system 104 (see FIG. 1) can perform spatial analysis on the transformed binaural signal 122 to generate estimated rendering parameters 124.
[0128] At 706, an estimated object is extracted from the transformed binaural signal using at least a first subset of the estimated rendering parameters. Extracting the estimated object includes generating a left principal component signal, a right principal component signal, a left residual component signal, and a right residual component signal. For example, the object extraction system 106 (see FIG. 1) may perform object extraction on the transformed binaural signal 122 using one or more of the estimated rendering parameters 124 to generate an estimated object 126. The estimated object 126 may correspond to component signals such as the left principal component signal 220, the right principal component signal 222, the left residual component signal 224, the right residual component signal 226 (see FIG. 2), or component signals 320, 322, 324, and 326 of FIG. 3.
[0129] At 708, object processing is performed on the estimated object using at least a second subset of the plurality of estimated rendering parameters. Performing the object processing includes generating a processed signal based on the left principal component signal, the right principal component signal, the left residual component signal, and the right residual component signal. For example, object processing system 108 (see FIG. 1) can perform object processing on estimated object 126 using one or more of estimated rendering parameters 124 to generate processed signal 128. As another example, processing system 208 (see FIG. 2) can perform object processing on component signals 220, 222, 224, and 226 using one or more of estimated rendering parameters 124 and object processing parameters 230 and 232.
[0130] Method 700 may include additional steps corresponding to other functions of audio processing system 100, one or more of object processing systems 108, 208, 308, etc., as described herein. For example, method 700 may include receiving sensor data, head tracking data, etc., and performing processing based on the sensor data or head tracking data. As another example, object processing (see 708) may include processing principal components using one set of processing parameters and processing residual components using another set of processing parameters. As another example, method 700 may include performing an inverse transform, performing time-domain processing on the inverse transformed signal, etc.
[0131] Implementation details
[0132] The embodiments may be implemented in hardware, executable modules stored on a computer-readable medium, or a combination of both, e.g., a programmable logic array, etc. Unless otherwise specified, steps performed by the embodiments may relate to any particular computer or other apparatus, although they may be inherent in a particular embodiment. In particular, various general-purpose mechanisms may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus, e.g., integrated circuits, etc., to perform the required method steps. Thus, the embodiments may be implemented in one or more computer programs executing on one or more programmable computer systems, each including at least one processor, at least one data storage system including volatile and non-volatile memory and / or storage elements, at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.
[0133] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device, such as a solid-state memory or medium, or a magnetic or optical medium, so as to configure and operate the computer to perform the procedures described herein when the storage medium or device is read by a computer system. The system of the present invention is also contemplated to be implemented as a computer-readable storage medium and configured with a computer program, where the storage medium is configured to operate a computer system to perform the functions described herein in a specific and predetermined manner. Software itself, and intangible or transitory signals, are excluded insofar as they are non-patentable subject matter.
[0134] Any of the systems described herein may be implemented in any suitable computer-based audio processing network environment that processes digital or digitized audio files. Portions of the adaptive audio system may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0135] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that various functions disclosed herein may be described in terms of their operation as hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media, using any number of combinations of register transfers, logical components, and / or other characteristics. The computer-readable media on which such formatted data and / or instructions are embodied include various forms of physical non-transitory, non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.
[0136] The foregoing description has described various embodiments of the present disclosure, along with examples of how aspects of the disclosure may be implemented. The above examples and embodiments should not be considered the only embodiments, but have been presented to illustrate the flexibility and advantages of the present disclosure as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be utilized without departing from the spirit and scope of the present disclosure as defined by the claims.
Claims
1. 1. A computer-implemented method of audio processing, comprising: performing a signal transformation on a binaural signal, the binaural signal being a binaural performance or a binaural capture, the performing a signal transformation comprising: converting the binaural signal from a first signal domain to a second signal domain; generating a transformed binaural signal, wherein the first signal domain is the time domain and the second signal domain is the frequency domain, the signal transformation is a time-to-frequency transformation, and the transformed binaural signal comprises a plurality of time-frequency tiles transformed over a given time period; and performing a spatial analysis on each of the plurality of time-frequency tiles of the transformed binaural signal, the performing of the spatial analysis comprising generating a plurality of estimated rendering parameters, a given time-frequency tile of the plurality of time-frequency tiles being associated with a given subset of the plurality of estimated rendering parameters, the plurality of estimated rendering parameters comprising a plurality of level differences and a plurality of phase differences, the plurality of estimated rendering parameters corresponding to at least one of head-related transfer functions, head-related impulse responses, and binaural room impulse responses used during binaural performance or present in binaural capture; generating a plurality of objects from the transformed binaural signal using at least a first subset of the estimated rendering parameters, the objects being represented by a left principal component signal, a right principal component signal, a left residual component signal, and a right residual component signal for each time-frequency tile of the transformed binaural signal; performing object processing on the plurality of objects using at least a second subset of the plurality of estimated rendering parameters, wherein the performing object processing comprises generating processed signals based on the left principal component signal, the right principal component signal, the left residual component signal, and the right residual component signal; Including, The method, wherein the object processing includes at least one of rearrangement, level adjustment, equalization, dynamic range adjustment, dessing, multi-band compression, immersion enhancement, enveloping, upmixing, conversion, channel remapping, storage, and archiving.
2. The step of generating a processed signal comprises: generating a left principal processed signal and a right principal processed signal from the left principal component signal and the right principal component signal using a first set of object processing parameters; generating a left residual processed signal and a right residual processed signal from the left residual component signal and the right residual component signal using a second set of object processing parameters, the second set of object processing parameters being different from the first set of object processing parameters; Including, The object processing is performed by the left main processed signal, the right main processed signal, The method of claim 1 , comprising using the left residual processed signal and the right residual processed signal.
3. receiving sensor data from a sensor, the sensor being a component of at least one of a headset, headphones, earphones, and a microphone; The method of claim 1 , wherein performing the object processing comprises generating the processed signal based on the sensor data.
4. The step of executing the object processing includes: applying binaural panning to the left and right principal component signals based on sensor data, wherein applying binaural panning comprises generating a left principal processed signal and a right principal processed signal; generating a left residual processed signal and a right residual processed signal from the left residual component signal and the right residual component signal without applying the binaural panning; The method of claim 1 , comprising:
5. The step of executing the object processing includes: generating a mono object signal from the left principal component signal and the right principal component signal; applying binaural panning to the mono object based on sensor data; generating a left residual processed signal and a right residual processed signal from the left residual component signal and the right residual component signal without applying the binaural panning; The method of claim 1 , comprising:
6. The step of executing the object processing includes: generating a multi-channel output signal from the left principal component signal, the right principal component signal, the left residual component signal, and the right residual component signal; 2. The method of claim 1, wherein the multi-channel output signal includes at least one left channel and at least one right channel, the at least one left channel including at least one of a front left channel, a side left channel, a rear left channel, and a left elevation channel, and the at least one right channel including at least one of a front right channel, a side right channel, a rear right channel, and a right elevation channel.
7. The step of executing the object processing includes: applying audio enhancement processing to the left principal component signal and the right principal component signal, wherein applying the audio enhancement processing comprises generating a left principal processed signal and a right principal processed signal; generating a left residual processed signal from the left residual component signal and a right residual processed signal from the right residual component signal without applying the audio enhancement processing; The method of claim 1 , comprising:
8. The step of generating a processed signal comprises: applying a level adjustment to the left principal component signal and the right principal component signal using a first level adjustment value, wherein applying the level adjustment includes generating a left principal processed signal and a right principal processed signal; applying a level adjustment to the left residual component signal and the right residual component signal using a second level adjustment value, wherein applying the level adjustment includes generating a left residual processed signal and a right residual processed signal, and the second level adjustment value is different from the first level adjustment value; Including, The method of claim 1 , wherein the object processing includes using the left main processed signal, the right main processed signal, the left residual processed signal, and the right residual processed signal.
9. 9. The method of claim 1, wherein the plurality of phase differences are a plurality of unwrapped phase differences, and the plurality of unwrapped phase differences are unwrapped by performing at least one of evidence-based unwrapping and model-based unwrapping.
10. The step of performing evidence-based unwrapping comprises: estimating the total energy of the left and right principal component signals in each band; calculating a cross-correlation based on each band; selecting the plurality of unwrapped phase differences from a plurality of candidate phase differences according to energy across neighboring bands based on the cross-correlation; 10. The method of claim 9, comprising:
11. The step of performing model-based unwrapping comprises:
10. The method of claim 9, comprising selecting the plurality of unwrapped phase differences from a plurality of candidate phase differences according to a given level difference applied to a head-related transfer function of a given band.
12. 12. The method according to claim 1, wherein a given phase difference of the plurality of phase differences is calculated as the phase angle of a dot product of a left component of the transformed binaural signal and a right component of the transformed binaural signal for a given index in the second signal domain.
13. 13. The method according to claim 1, wherein a given level difference among the plurality of level differences is calculated according to a quadratic equation based on a left component of the transformed binaural signal, a right component of the transformed binaural signal, and a given phase difference among the plurality of phase differences.
14. performing an inverse signal transform on the left main processed signal, the right main processed signal, the left residual processed signal, and the right residual processed signal to generate processed signals, the processed signals being in the first signal domain; 9. The method of any one of claims 2, 4, 7, and 8, further comprising:
15. performing time domain processing on the processed signal, wherein performing time domain processing comprises generating a modified time domain signal; 15. The method of any one of claims 1 to 14, further comprising:
16. A non-transitory computer readable medium storing a computer program which, when executed by a processor, controls an apparatus to perform processes including the method of any one of claims 1 to 15.
17. 1. An apparatus for audio processing, said apparatus comprising: a processor and an optional sensor, the processor being configured to control the device to perform a process comprising the method of any one of claims 1 to 15; Equipment including.
Citation Information
Patent Citations
Audio signal filter generation method and parameterization device therefor
JP2017505039A
Binaural integration cross-correlation autocorrelation mechanism
JP2017530579A
Binaural rendering for headphones using metadata processing
US20160266865A1
Device and method for processing audio signal
US20180324542A1
Audio Encoding and Decoding Using Presentation Transform Parameters
US20200227052A1