Audio filterbank with decorrelating components
The audio filter bank with decorrelation components addresses high latency in conventional systems by integrating decorrelation processing into a single mixer, enhancing efficiency in audio signal transformation.
Patent Information
- Application Number
- JP2025077740
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-03
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2040-09-02
AI Technical Summary
Conventional audio signal processing systems require multiple linear mixers for decorrelation processing, leading to high latency.
An audio filter bank with decorrelation components that integrates decorrelation processing into a single linear mixer, using a composite frequency-domain gain vector formed by augmenting component vectors with modified frequency responses to achieve decorrelation.
Reduces latency by enabling decorrelation processing with a single linear mixer, improving efficiency in audio signal transformation.
Smart Images

Figure 2025134682000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 895,096, filed September 3, 2019, which is incorporated herein by reference. [Technical field]
[0002] The present disclosure relates generally to audio signal processing, and more particularly to processing a set of one or more frequency-domain input audio signals to generate a new set of one or more frequency-domain output audio signals. [Background technology]
[0003] Audio signal processing generally involves transforming a set of input audio signals into a new set of audio output signals, the number of which may be equal to or greater than the number of input audio signals. For example, a surround sound system may use linear matrix operations to convert two input audio signals (eg, a stereo audio signal) into five output audio signals. A linear matrix operation applies a matrix containing coefficients that can vary as a function of time or frequency to the input audio signal. The linear matrix operation can also determine the covariance of the output audio signal when the input audio signal has undergone decorrelation processing. Summary of the Invention
[0004] The multi-input, multi-output audio process is implemented as a linear system for use in an audio filter bank to transform a set of frequency-domain input audio signals into a set of frequency-domain output audio signals. The transfer function from one input to one output is defined as a frequency-dependent gain function. In some implementations, the transfer function includes a direct component substantially defined as a frequency-dependent gain and one or more decorrelated components having a frequency-varying group phase response. The transfer function is formed from a set of subband functions, each subband function formed from a corresponding set of component transfer functions including the direct component and one or more decorrelated components.
[0005] In some implementations, a method for converting a set of frequency-domain input audio signals into a set of frequency-domain output audio signals includes calculating, using one or more processors, each frequency-domain output audio signal as a sum of filtered frequency-domain input audio signals, wherein each filter used to filter the frequency-domain input audio signal is characterized by a complex gain function over a respective sub-band frequency range of the frequency-domain input audio signal, and wherein the contribution of the frequency-domain input audio signals to the frequency-domain output audio signal is determined by a composite frequency-domain gain vector, the composite frequency-domain gain vector being obtained by: calculating, using the one or more processors, a set of component frequency-domain gain vectors, at least one of the component frequency-domain gain vectors being a decorrelated component frequency-domain gain vector formed by augmenting the component frequency-domain gain vector with an additional component frequency-domain gain vector having a modified frequency response to produce a decorrelation effect; and adding, using the one or more processors, the component frequency-domain gain vectors to form the composite frequency-domain gain vector.
[0006] In some implementations, the decorrelated component frequency domain gain vector is formed by scaling at least one of the component frequency domain vectors by a component gain value.
[0007] In some embodiments, one or more of the component frequency-domain gain vectors include a phase response that varies across the subband frequency range, thereby providing a group delay that is substantially constant across the subband frequencies, the group delay being substantially constant when the variation in the group delay is sufficiently small so as to be perceptually inconsequential to a listener.
[0008] In some embodiments, one or more of the component frequency domain gain vectors include a phase response that varies across the subband frequency range, thereby providing a group delay that varies across the subband frequency range to provide the decorrelation effect.
[0009] In some implementations, the decorrelated component frequency domain gain vectors are formed by multiplying the component frequency domain gain vectors by a decorrelation function.
[0010] In some implementations, the audio filter bank with decorrelation components comprises a converter configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals, and a linear mixer configured to convert the set of frequency-domain input audio signals into a set of frequency-domain output audio signals, each frequency-domain output audio signal being a sum of filtered frequency-domain input audio signals, each filter used to filter the frequency-domain input audio signals being characterized by a complex gain function over a respective sub-band frequency range of the frequency-domain input audio signals, and the contribution of the frequency-domain input audio signals to the frequency-domain output audio signals being determined by a composite frequency-domain gain vector.
[0011] In some embodiments, the composite frequency domain gain vector is obtained by calculating a set of component frequency domain gain vectors, at least one of which is a decorrelated component frequency domain gain vector formed by augmenting the component frequency domain gain vector with an additional component frequency domain gain vector having a modified frequency response to produce a decorrelation effect, and summing the component frequency domain gain vectors to form the composite frequency domain gain vector.
[0012] In some implementations, the decorrelated component frequency domain gain vector is formed by scaling at least one of the component frequency domain vectors by a component gain value.
[0013] In some embodiments, one or more of the component frequency-domain gain vectors include a phase response that varies across the subband frequency range, thereby providing a group delay that is approximately constant across the subband frequencies, the group delay being approximately constant when the group delay variations are small enough to be perceptually inconsequential to a listener.
[0014] In some embodiments, one or more of the component frequency domain gain vectors include a phase response that varies across the subband frequency range, thereby providing a group delay that varies across the subband frequency range to provide the decorrelation effect.
[0015] In some implementations, the decorrelated component frequency domain gain vectors are formed by multiplying the component frequency domain gain vectors by a decorrelation function.
[0016] In some embodiments, a filter bank-based audio system comprises: a converter configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals; and a linear mixer configured to convert the set of frequency-domain input signals into a set of frequency-domain output signals, the linear mixer including weighting coefficients to provide a frequency-dependent gain function comprising a direct component defined as a frequency-dependent gain and one or more decorrelated components having a frequency-varying group phase response, the frequency-dependent gains being formed from a set of subband functions, each subband function being formed from a corresponding set of component transfer functions comprising the direct component and one or more decorrelated components.
[0017] Other embodiments disclosed herein are directed to systems, devices, and computer-readable media. Details of the disclosed implementations are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0018] Particular embodiments disclosed herein provide one or more of the following advantages: The disclosed implementations integrate decorrelation processing into the audio filterbank, thereby enabling an input audio signal to be mapped to an output audio signal using a single linear mixer, resulting in lower latency than conventional audio filterbanks that use multiple linear mixers to perform decorrelation processing. [Brief explanation of the drawings]
[0019] The figures show a particular arrangement and order of schematic elements, such as those representing devices, units, instruction blocks, and data elements, for ease of explanation. However, it should be understood that the particular order and arrangement of the schematic elements in the figures does not imply a particular order or sequence of operations, or separation of operations, is required. Furthermore, the inclusion of a schematic element in a figure does not imply that such element is required in all embodiments, or that features represented by such element cannot be included in or combined with other elements in some embodiments.
[0020] Furthermore, when a connecting element, such as a solid or dashed line or arrow, is used in the drawings to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of such a connecting element does not mean that the connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the disclosure. Furthermore, for ease of explanation, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, it is understood that such element represents one or more signal paths necessary to affect the communication.
[0021] [Figure 1] 1 illustrates a set of input audio signals being filtered using an array of filters to produce a set of audio output signals, according to one or more embodiments.
[0022] [Figure 2] 1 illustrates a desired frequency response curve according to one or more embodiments.
[0023] [Figure 3] 1 illustrates a set of filter bank frequency responses according to one or more embodiments.
[0024] [Figure 4]1 illustrates a bandpass response of an exemplary component frequency domain gain vector in accordance with one or more embodiments.
[0025] [Figure 5] 1 illustrates the frequency response of a subband filter having a group delay that varies significantly across frequency in accordance with one or more embodiments.
[0026] [Figure 6] 1 illustrates a known method for mixing input signals using a direct mixing matrix and one or more decorrelation mixing matrices to generate an output signal, according to one or more embodiments.
[0027] [Figure 7] FIG. 2 is a flow diagram of an exemplary process for converting a set of frequency-domain input audio signals into a set of frequency-domain output audio signals, according to one or more embodiments.
[0028] [Figure 8] FIG. 8 illustrates a block diagram of a system suitable for implementing the features and processes described with reference to FIGS. 1-7, according to one or more embodiments.
[0029] The use of the same reference symbols in the various drawings indicates like elements. DETAILED DESCRIPTION OF THE INVENTION
[0030] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of various described embodiments. It will be apparent to those skilled in the art that various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Although several features are described below, each can be used independently of each other or in any combination with the other features. term
[0031] As used herein, the term "comprises" and variations thereof should be construed as open-ended, meaning "including, but not limited to." The term "or" should be construed as "and / or" unless the context clearly indicates otherwise. The term "based on" should be interpreted as "based at least in part on." "One implementation" and "one example" should be interpreted as "at least one implementation." "Another example" should be interpreted as "at least one other example." "Determined," "determine," or "determining" can be interpreted as obtaining, receiving, calculating, estimating, predicting, or deriving. Moreover, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. System Overview
[0032] FIG. 1 illustrates a linear mixing system 100 in which a set of input audio signals are filtered to produce a set of audio output signals, according to one or more embodiments. System 100 may be implemented, for example, within an audio filter bank. An audio filter bank includes an array of bandpass filters that separates the input audio signals into multiple frequency subbands of the input audio signal. In the illustrated example, linear mixing system 100 includes a bank of filters 101 and a summer 102. N input signals (X1...X) are filtered to produce a set of audio output signals. N ) are processed by a bank of filters 101 and added together by summers 102 to produce M output signals (Y1...Y M ) The linear mixing system 100 may be defined in terms of the frequency-domain input signal and the frequency-domain output signal as follows:
number
number
number
[0033] According to equation [3], the frequency domain output audio signal [Outside 1] TIFF2025134682000005.tif8170 is formed as a sum of filtered frequency-domain input audio signals Xn(f), where the contribution of the frequency-domain input audio signals Xn(f), n∈[1…N], to Ym(f) is determined by a composite frequency-domain vector Gm,n(f)
number
[0034] For purposes of the following discussion, G(f) will be referred to as an exemplary composite frequency-domain gain vector, and this term should be understood to refer to any one of the composite frequency-domain gain vectors G(f) used in equations [3] and [4].
[0035] 2 illustrates a desired frequency response curve for a filter, according to one or more embodiments. The desired frequency response of an exemplary composite frequency-domain gain vector can be generated by a process that generates a smoothed function, as shown in FIG. 2, where the filter gain 20 is a function of frequency, such as the control frequency f c1 , f c2 ... and a corresponding set of predefined component gain values w1, w2, .... For example, at a frequency f c2 The gain 21 of the filter at is set by the component gain value w2, as shown in Figure 2. The frequency response shown in Figure 2 is achieved by the weighted addition of a number of predefined component frequency domain gain vectors.
[0036] Figure 3 shows the reference frequency band 2, H 0,23 shows a set of filter bank frequency responses according to one or more embodiments, with response 300 in (f). The frequency responses of these predefined component frequency domain gain vectors are hereafter referred to as component frequency domain gain vectors H 0,b (f) for b∈[1...B], where B is the number of bands (e.g., B=5 in the example of Figure 3), each of the component frequency-domain gain vectors has an alternative representation h0,b(n) in the form of a time-domain impulse response.
[0037] In one embodiment, the desired filter response (see FIG. 2) may be formed from a weighted sum of predetermined filter bank responses, which can be expressed as a time-domain or frequency-domain sum:
number
[0038] In some implementations, the set of component frequency-domain gain vectors may be further modified by additional component frequency-domain gain vectors H that modify their frequency responses to produce a decorrelation effect. 0,b The augmented set of component frequency-domain gain vectors is hereafter referred to as the decorrelated component frequency-domain gain vectors and is denoted by the following terms:
number
[0039] The augmented set of component frequency domain gain vectors is given by Equation [7]
number
[0049] , it can be used in a filterbank-based audio processing system to generate a composite frequency-domain gain vector.
[0040] 4 illustrates a bandpass response of an exemplary component frequency domain gain vector in accordance with one or more embodiments. In the illustrated example, the component frequency domain gain vector H 0,b (f) has an amplitude response 401 that is generally dominant over a particular subband range of the total frequency range, and a group delay 402 that is substantially constant over the subband range. A group delay is considered to be substantially constant if the variation in group delay is small enough that it is not perceptually significant to a listener when the filter is used to process an audio signal.
[0041] FIG. 5 illustrates the frequency response of a subband filter having a group delay that varies significantly with frequency, in accordance with one or more embodiments. 1,b The frequency response of the decorrelated component frequency domain gain vector, such as (f) (l≠0), exhibits a group delay 502 that varies across the subband frequency range, and the variation in group delay is reflected by the decorrelated component frequency domain gain vector H 1,b (f) The input audio signal filtered by (l≠0) is expressed as the component frequency-domain gain vector H 0,b (f) is perceived as being decorrelated with the input audio signal filtered by (f).
[0042] As is known, methods are known in the art for generating frequency responses with group delay variations that vary over a wide frequency range for the purpose of creating a perceived decorrelation effect. In one embodiment, a known decorrelation frequency response may be adapted by applying magnitude responses 501 to form a decorrelated component frequency domain gain vector. In one embodiment, a known decorrelation function D l(f) Compute a set of B decorrelated component frequency domain gain vectors using (l∋[1...L]):
number
[0043] 6 illustrates a system 600 for mixing input signals to generate an output signal using a direct mixing matrix and one or more decorrelation mixing matrices, according to one or more embodiments. l Given (f)(l∋[l...L]), an N-channel input signal (X) is processed by system 600 to produce an M-channel output signal (Y). In this example, processing of one subband (e.g., band b) is shown, where N-channel input 601 (X) is applied directly to linear mixing matrix 602 (C) (e.g., an MxN matrix) to produce M-channel direct signal 603. N-channel input 601 also applies to linear mixer 610 (Ql) (e.g., K L ×N matrix) and is processed by a KL decorrelation filter 612 (D l ) through a bank of K L 611 channels, each of which has a frequency response D L (f) to generate the KL channel signal 613, which is then fed to a linear mixer 614 (P l ) (e.g., M × K L matrix) to generate an M-channel decorrelated component signal 615. The M-channel direct signal 603 is then added to the M-channel decorrelated component signal (e.g., decorrelated component signal (615)) to generate an M-channel output 602(Y).
[0044] In this embodiment, instead of the process shown in FIG. 6, the linear mixing matrices C,Q1...Q L and P1...P L function as a set of weighting coefficients [Outside 2] TIFF2025134682000011.tif9170. According to one embodiment, referring back to equation [4], the output channel Y m (f) is
number
[0045] Equation [9] can be implemented in a filter bank-based audio processing system, where the number of filters is (L+1)×B instead of the B filters known in the art. This set of extended filters can further be thought of as the previously known B filters with the addition of L×B filters corresponding to L different decorrelation functions.
[0046] In some implementations, equation [9] is implemented as an audio filter bank including converters (e.g., Fast Fourier Transforms) configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals Xn(f), and a linear mixer (performing a matrix multiplication operation) [Outside 3] The present invention is configured to implement TIFF2025134682000013.tif15170 to transform a set of frequency-domain input audio signals Xn(f) into a set of frequency-domain output audio signals Ym(f). Each frequency-domain output audio signal is a sum of filtered frequency-domain input audio signals, and each filter used to filter the frequency-domain input audio signals is characterized by a complex gain function over a respective subband frequency range of the frequency-domain input audio signals. The contributions of the frequency-domain input audio signals to the frequency-domain output audio signals are determined by a composite frequency-domain gain vector.
[0047] In some embodiments, equation [9] is implemented as an audio filter bank system including a transformer (e.g., a Fast Fourier Transform) configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals Xn(f), and a linear mixer (software or hardware implementing a sum of products operation) [Outside 4] The system is configured to implement TIFF2025134682000014.tif16170 and convert a set of frequency-domain input audio signals Xn(f) into a set of frequency-domain output audio signals Ym(f). The linear mixer includes weighting coefficients (elements of Gm,n(f)) that provide a frequency-dependent gain function that includes a direct component defined as a frequency-dependent gain and one or more decorrelated components having a frequency-varying group phase response. The frequency-dependent gain is formed from a set of subband functions, each formed from a corresponding set of component transfer functions that include the direct component and one or more decorrelated components. Process Example
[0048] 7 is a flow diagram of an exemplary process 700 for converting a set of frequency-domain input audio signals into a set of frequency-domain output audio signals, according to one or more embodiments. Process 700 can be implemented, for example, by system 800 described with reference to FIG.
[0049] The process 700 then calculates each frequency-domain output audio signal as a sum of filtered frequency-domain input audio signals that define a complex gain function over the respective subband frequency range, with the contribution of the frequency-domain input audio signals to the frequency-domain output audio signal being determined by a composite frequency-domain gain vector (701).
[0050] Process 700 continues by obtaining a composite frequency-domain gain vector by calculating 702 a set of component frequency-domain gain vectors, at least one of which is a decorrelated component frequency-domain gain vector that is augmented with additional component frequency-domain gain vectors having modified frequency responses to form a component frequency-domain gain vector, producing a decorrelation effect.
[0051] The process 700 continues by adding the component frequency-domain gain vectors to form a composite frequency-domain gain vector (703). System Architecture Example
[0052] 8 shows a block diagram of an exemplary system 800 suitable for implementing exemplary embodiments of the present disclosure. System 800 includes one or more server computers or any client devices, including, but not limited to, call servers, user devices, conference room systems, home theater systems, virtual reality (VR) gear, and immersive content ingestion devices. System 800 also includes any consumer device, including, but not limited to, smartphones, tablet computers, wearable computers, vehicle computers, gaming consoles, surround systems, kiosks, etc.
[0053] As shown, the system 800 includes a central processing unit (CCU) 801 that can execute various processes according to programs stored in, for example, a read-only memory (ROM) 802 or programs loaded from, for example, a storage unit 808 into a random access memory (RAM) 803. The RAM 803 also stores data required by the CPU 801 when executing the various processes, as needed. The CPU 801, the ROM 802, and the RAM 803 are connected to one another via a bus 804. An input / output interface 805 is also connected to the bus 804.
[0054] The following components are connected to the I / O interface 805: an input unit 806 including a keyboard, mouse, etc.; an output unit 807 including a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 808 including a hard disk or other suitable storage device; and a communication unit 809 including a network interface card such as a network card.
[0055] In some implementations, input unit 806 includes one or more microphones at different locations (depending on the host device) to enable capturing audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0056] In some implementations, output unit 807 includes a system with a variable number of speakers. The output unit 807 can render audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats) (depending on the capabilities of the host device).
[0057] The communication unit 809 is configured to communicate with other devices (eg, via a network). Optionally, a drive 810 is also connected to the I / O interface 805. A removable medium 811, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is mounted on the drive 810, and, as necessary, a computer program read therefrom is installed in the storage unit 808. It will be appreciated by those skilled in the art that, although the system 800 is described as including the above-mentioned components, in actual applications, some of these components may be added, removed, and / or substituted, and all such modifications or variations are within the scope of the present disclosure.
[0058] According to exemplary embodiments of the present disclosure, the above-described processes may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded from a network via a communications unit 809, mounted, and / or installed from a removable medium 811, as shown in FIG. 8.
[0059] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or special-purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the above-described units may be executed by a control circuit (e.g., a CPU in combination with other components of FIG. 8 ), which may then perform the operations described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). While various aspects of the exemplary embodiments of the present disclosure have been illustrated and described using block diagrams, flowcharts, or some other pictorial representations, it should be understood that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, special-purpose circuitry or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof, as non-limiting examples.
[0060] Furthermore, the various blocks illustrated in the flowcharts may be viewed as method steps and / or as operations resulting from the operations of computer program code and / or as a plurality of coupled logic circuit elements configured to perform the associated functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the above-described methods.
[0061] In the context of this disclosure, a machine / computer-readable medium may be any tangible medium that includes an instruction execution system, apparatus, or device, or that can store a program for use in connection with an instruction execution system, apparatus, or device. The machine / computer-readable medium may be a machine / computer-readable signal medium or a machine / computer-readable storage medium. The machine / computer-readable medium may be non-transitory and include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine / computer-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or Flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0062] Computer program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. The computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that the program code, when executed by the processor of the computer or other programmable data processing apparatus, causes the computer to perform the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on the computer, partially on the computer as a stand-alone software package, partially on the computer and partially on a remote computer, entirely on a remote computer or server, or entirely on a computer distributed across one or more remote computers and / or servers.
[0063] While this document contains many specific implementation details, these should not be construed as limiting the scope of the claims, but rather as descriptions of features specific to specific embodiments. Specific features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a specific combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be carved out of that combination, and the claimed combination may be directed to subcombinations or variations of subcombinations. The logic flows depicted in the figures do not require the specific order shown to achieve desired results. Additionally, other steps may be provided or steps may be removed from the described flows, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the claims.
[0064] [Appendix 1] 1. A method for converting a set of frequency domain input audio signals into a set of frequency domain output audio signals, comprising: calculating, using one or more processors, each frequency-domain output audio signal as a sum of the filtered frequency-domain input audio signals; Each filter used to filter the frequency-domain input audio signal is characterized by a complex gain function over a respective subband frequency range of the frequency-domain input audio signal, and the contribution of the frequency-domain input audio signal to the frequency-domain output audio signal is determined by a composite frequency-domain gain vector, which is calculating, using one or more processors, a set of component frequency domain gain vectors, at least one of the component frequency domain gain vectors being a decorrelated component frequency domain gain vector formed by augmenting a component frequency domain gain vector with an additional component frequency domain gain vector having a modified frequency response to produce a decorrelation effect; adding, using one or more processors, the component frequency domain gain vectors to form a composite frequency domain gain vector; The method is obtained by [Appendix 2] 2. The method of claim 1, wherein the decorrelated component frequency domain gain vector is formed by scaling at least one of the component frequency domain vectors by the component gain value. [Appendix 3] 3. The method of any one of claims 1 or 2, wherein one or more of the component frequency-domain gain vectors include a phase response that varies across the subband frequency range to provide a group delay that is substantially constant across the subband frequencies, and wherein the group delay is substantially constant when variations in group delay are sufficiently small that they are not perceptually significant to a listener. [Appendix 4] 4. The method of any one of claims 1 to 3, wherein one or more of the component frequency-domain gain vectors include a phase response that varies across the subband frequency range and provides a group delay that varies across the subband frequency range to provide a decorrelation effect. [Appendix 5] 5. The method of any one of claims 1 to 4, wherein the decorrelated component frequency domain gain vectors are formed by multiplying the component frequency domain gain vectors by a decorrelation function. [Appendix 6] one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method of any one of claims 1 to 5; A system having: [Appendix 7] A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method of any one of claims 1 to 5. [Appendix 8] 1. An audio filterbank with a decorrelation component, comprising: a converter configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals; 1. A linear mixer configured to convert a set of frequency-domain input audio signals into a set of frequency-domain output audio signals, comprising: a linear mixer, wherein each frequency-domain output audio signal is a sum of filtered frequency-domain input audio signals, each filter used to filter the frequency-domain input audio signals is characterized by a complex gain function over a respective sub-band frequency range of the frequency-domain input audio signal, and the contribution of the frequency-domain input audio signals to the frequency-domain output audio signal is determined by a composite frequency-domain gain vector; Audio filter bank. [Appendix 9] The composite frequency domain gain vector is calculating a set of component frequency-domain gain vectors, at least one of the component frequency-domain gain vectors being a decorrelated component frequency-domain gain vector formed by augmenting the component frequency-domain gain vector with an additional component frequency-domain gain vector having a modified frequency response to produce a decorrelation effect; adding the component frequency domain gain vectors to form a composite frequency domain gain vector; The audio filter bank according to claim 8, obtained by: [Appendix 10] 10. The audio filterbank of claim 8 or 9, wherein the decorrelated component frequency domain gain vector is formed by scaling at least one of the component frequency domain vectors by the component gain value. [Appendix 11] 11. The audio filterbank of any one of clauses 8 to 10, wherein one or more of the component frequency-domain gain vectors include a phase response that varies across the subband frequency range, provide a group delay that is substantially constant across the subband frequencies, and the group delay is substantially constant when variations in group delay are sufficiently small so as to be perceptually insignificant to a listener. [Appendix 12] 12. The audio filterbank of any one of claims 8 to 11, wherein one or more of the component frequency domain gain vectors include a phase response that varies across the subband frequency range and provide a group delay that varies across the subband frequency range to provide a decorrelation effect. [Appendix 13] 13. The audio filterbank of any one of claims 8 to 12, wherein the decorrelated component frequency domain gain vectors are formed by multiplying the component frequency domain gain vectors by a decorrelation function. [Appendix 14] 1. A filter bank based audio system, comprising: a converter configured to convert a set of time-domain input audio signals into a set of frequency-domain input audio signals; 1. A linear mixer configured to convert a set of frequency domain input signals into a set of frequency domain output signals, comprising: a linear mixer including weighting coefficients that provide a frequency-dependent gain function including a direct component defined as a frequency-dependent gain and one or more decorrelated components having a frequency-varying group phase response, the frequency-dependent gain being formed from a set of subband functions, each subband function being formed from a corresponding set of component transfer functions including the direct component and one or more decorrelated components; have Audio system.
Claims
1. 1. A method for converting a set of frequency domain input audio signals into a set of frequency domain output audio signals, comprising: calculating, using one or more processors, each frequency-domain output audio signal as a sum of the filtered frequency-domain input audio signals; Each filter used to filter the frequency-domain input audio signal is characterized by a complex gain function over a respective subband frequency range of the frequency-domain input audio signal, and the contribution of the frequency-domain input audio signal to the frequency-domain output audio signal is determined by a composite frequency-domain gain vector, which is using the one or more processors to define a filter gain value for the filter according to a set of component gain values; calculating, using the one or more processors, a set of component frequency domain gain vectors, at least one of the component frequency domain gain vectors being a decorrelated component frequency domain gain vector formed by at least one of the component frequency domain gain vectors based on the determined filter gain values to produce a decorrelation effect; using the one or more processors to form a composite frequency domain gain vector based on the component frequency domain gain vectors; The method is obtained by
2. 10. The method of claim 1, wherein one or more of the component frequency-domain gain vectors include a phase response that varies across a subband frequency range to provide a group delay that is substantially constant across the subband frequencies, and wherein the group delay is substantially constant if variations in the group delay are sufficiently small to be perceptually inconsequential to a listener.
3. 3. The method of claim 1, wherein one or more of the component frequency domain gain vectors includes a phase response that varies across the subband frequency range and provides a group delay that varies across the subband frequency range to provide a decorrelation effect.
4. 4. The method of claim 1, wherein the decorrelated component frequency domain gain vectors are formed by multiplying the component frequency domain gain vectors by a decorrelation function.
Citation Information
Patent Citations
Acoustic signal processing apparatus, acoustic signal processing method, and program
JP2013102389A
Improvement of audio signals in FM stereo wireless receivers using parametric stereo
JP2013504908A
Improved multichannel upmixing using multichannel decorrelation
JP2013517687A
Adaptive Diffusivity Signal Generation in an Upmixer
JP2016537855A
Decorrelator structure for parametric reconstruction of audio signals
JP2016539358A