Method, system, and non-transitory computer readable medium for encoding spatial cues along a sagittal plane into a monaural signal to generate a plurality of resulting channels
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2026-08-01
AI Technical Summary
Existing audio encoding technologies struggle to effectively incorporate spatial cues into mono audio signals, limiting the ability to create immersive and spatially aware audio experiences, particularly in conferencing, video playback, and entertainment applications.
The use of all-pass filter networks, specifically single-input multiple-output (SIMO) all-pass filters, to encode spatial cues into mono audio signals, transforming them into multiple channels with targeted amplitude and phase shifts, enhancing the perceived spatial quality and immersion.
This approach allows for the creation of immersive audio experiences by effectively encoding spatial cues, improving speech intelligibility and overall spatial awareness in audio content, reducing the need for multiple amplifiers and speakers, and enhancing the sense of space in sound fields.
Smart Images

Figure TWG2TB001903315_001 
Figure TWG2TB001903315_002 
Figure TWG2TB001903315_003
Abstract
Description
[Technical Field]
[0001] This invention generally relates to audio processing, and more specifically, to encoding spatial cues into audio content. [Previous Technology]
[0002] Audio content can be encoded to include the spatial properties of a sound field to allow a user to perceive a sense of space within the sound field. For example, the audio from a specific sound source (e.g., a voice or musical instrument) can be mixed into audio content in a way that produces a sense of space associated with the audio, such as the perception that the audio arrives at the user from a specific direction or is located in a specific type of location (e.g., a small room, a large auditorium, etc.). [Summary of the Invention]
[0003] Some embodiments include a method for encoding a spatial cue along a sagittal plane into a mono signal to generate a plurality of resulting channels. The method includes: determining, by a processing circuitry system, a target amplitude response of one of the middle or side components of the plurality of resulting channels based on a spatial cue associated with a frequency-dependent phase shift; converting the target amplitude response of the middle or side component into a transfer function of a single-input multiple-output (SMILE) all-pass filter; and processing the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
[0004] Some embodiments include a system for generating a plurality of channels from a single mono channel, wherein the plurality of channels are encoded using one or more spatial cues. The system includes one or more computing units configured to determine a target amplitude response of one of the middle or side components of the plurality of channels based on a spatial cue associated with a frequency-dependent phase shift. The one or more computers are further configured to convert the target amplitude response of the middle or side component into a transfer function of a single-input multiple-output (SIMO) all-pass filter and to process the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
[0005] Some embodiments include a non-transitory computer-readable medium comprising stored instructions for generating a plurality of channels from a mono channel, wherein the plurality of channels are encoded with one or more spatial cues, the instructions configuring the at least one processor, when executed by at least one processor, to: determine a target amplitude response of one of the middle or side components of the plurality of resulting channels based on a spatial cue associated with a frequency-dependent phase shift; convert the target amplitude response of the middle or side component into a transfer function of a single-input multiple-output (SMILE) all-pass filter; and process the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
[0006] Some embodiments relate to using a series of Hilbert transforms to spatially shift a portion of audio content (e.g., speech). Some embodiments include one or more processors and a non-transitory computer-readable medium. The computer-readable medium includes stored program code that, when executed by the one or more processors, configures the one or more processors to: separate an audio channel into a low-frequency component and a high-frequency component; apply a first Hilbert transform to the high-frequency component to generate a first left branch component and a first right branch component, the first left branch component being 90 degrees out of phase with the first right branch component; apply a second Hilbert transform to the first right branch component to generate a second left branch component and a second right branch component, the second left branch component being 90 degrees out of phase with the second right branch component; combine the first left branch component with the low-frequency component to generate a left channel; and combine the second right branch component with the low-frequency component to generate a right channel.
[0007] Some embodiments include a non-transitory computer-readable medium containing stored program code. When executed by one or more processors, the program code configures the one or more processors to: separate an audio channel into a low-frequency component and a high-frequency component; apply a first Hilbert transform to the high-frequency component to generate a first left branch component and a first right branch component, the first left branch component being 90 degrees out of phase with the first right branch component; apply a second Hilbert transform to the first right branch component to generate a second left branch component and a second right branch component, the second left branch component being 90 degrees out of phase with the second right branch component; combine the first left branch component with the low-frequency component to generate a left channel; and combine the second right branch component with the low-frequency component to generate a right channel.
[0008] Some embodiments include a method executed by one or more processors. The method includes: separating an audio channel into a low-frequency component and a high-frequency component; applying a first Hilbert transform to the high-frequency component to generate a first left branch component and a first right branch component, the first left branch component being 90 degrees out of phase with the first right branch component; applying a second Hilbert transform to the first right branch component to generate a second left branch component and a second right branch component, the second left branch component being 90 degrees out of phase with the second right branch component; combining the first left branch component with the low-frequency component to generate a left channel; and combining the second right branch component with the low-frequency component to generate a right channel.
Implementation Method
[0037] The figures and the following description are for illustrative purposes only and are preferred embodiments. It should be noted that, from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily regarded as feasible alternatives that can be adopted without departing from the claimed principles.
[0038] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that similar or identical element symbols may be used in the drawings and may indicate similar or identical functions, wherever feasible. The embodiments of the disclosed systems (or methods) depicted in the drawings are for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.
[0039] In various applications involving multiple simultaneous streams of audible content, it is desirable to encode spatial perceptual cues into a single mono audio source. Examples of such applications include: • Conferencing use cases, where adding spatial perceptual cues applied to one or more telephony devices can help improve overall speech intelligibility and enhance the overall immersion of the listener. • Video and music playback / streaming use cases, where one or more audio channels or the signal components of one or more audio channels can be enhanced by adding spatial perceptual cues to improve the intelligibility or spatiality of other elements of the speech or mix. • Co-watching entertainment use cases, where the stream consists of individual content channels, such as one or more telephony devices and entertainment material, which must be mixed together to form an immersive experience, and applying spatial perceptual cues to one or more elements can increase the perceptual difference between the mix elements to broaden the listener's perceptual bandwidth.
[0040] The embodiments relate to an audio system that modifies the perceived spatial quality (e.g., sound field and overall position of a target listener's head) of one or more audio channels. In some embodiments, modifying the perceived spatial quality of an audio channel can be used to separate the coloration of a particular source from its perceived spatial position and / or reduce the number of amplifiers and speakers required to encode this effect.
[0041] Audio signal processing performed by an audio system is referred to as Perceptual Sound Field Modification (PSM) processing. The perceptual result of PSM processing is referred to herein as a spatial displacement. Psychoacoustic effects are generally perceived by the user as a general displacement of a sound source above, around, or near the head to perceptually distinguish the sound source from other parts of the audio content. This psychoacoustic effect arises from the phase and time relationship between the left and right channels, enhanced by a network of all-pass filters and delays. In some embodiments, this filter and delay network may be implemented as one or more second-order all-pass sections (such as a series of Hilbert transforms) or using a first-order non-orthogonal rotational decorrelation (FNORD) filter network, each of which will be described in more detail below. The perceptual result of PSM processing can vary depending on the listening configuration (e.g., headphones or speakers, etc.). For some content and algorithm configurations, the result may also produce the impression that the perceptual signal propagates (e.g., diffuses) around the listener's head. For mono input signals (such as non-spatial audio signals), the diffusion effect of PSM processing can be used for mono to stereo upmixing.
[0042] In some embodiments, the audio system may isolate a target portion of an audio signal from a remaining portion of the audio signal, apply various configurations of PSM processing to perceptually shift the target portion, and remix the processed result with the remaining portion (e.g., which may be unprocessed or differently processed). This system can be perceived to clarify, enhance, or otherwise distinguish the target portion within the entire audio mix. In some embodiments, PSM processing is used to perceptually shift a portion of an audio signal containing singing or speaking sounds. Conventionally, speech in television, film, or music audio streams is typically located at the center of the sound field and thus at a portion of the middle component (also referred to as a non-spatial or non-correlated component) of a stereo or multi-channel audio signal. Therefore, PSM processing can be applied to a middle component of an audio signal or a super-middle component containing the spectral energy of a side component (also referred to as a spatial or non-correlated component) with spectral energy removed from the middle component.
[0043] PSM processing can be combined with other types of processing. For example, an audio system may apply processing to a shifted portion of an audio signal to perceptually transform and distinguish the shifted portion from other components in the mix. These additional types of processing may include one or more of the following: single or multi-band equalization, single or multi-band dynamic processing (e.g., limiting, compression, expansion, etc.), single or multi-band gain or delay, crosstalk processing (e.g., crosstalk cancellation and / or crosstalk simulation processing), or crosstalk compensation. In some embodiments, PSM processing may be performed in conjunction with mid / side processing, such as sub-band spatial processing, wherein sub-bands of the mid and side components of an audio signal generated by PSM processing are gain-adjusted to enhance the spatial feel of the sound field.
[0044] Isolation of an audio channel used for PSM processing can be achieved in various ways. In some embodiments, PSM processing can be performed on spectral quadrature sound components (such as the super-mid component of an audio signal). In other embodiments, PSM processing is performed on an audio channel associated with a sound source (e.g., a speech), and the processed channel is then mixed with other audio content (e.g., background music).
[0045] Although the following discussion focuses primarily on upmixing a mono signal to stereo (i.e., two output channels), given the large percentage of audio presentation devices are stereo, it should be understood that the techniques discussed can be easily extended to include more channels. Stereo embodiments can be discussed from a center / side processing perspective, where the phase difference between the left and right channels becomes a complementary region of amplification and attenuation in the center / side space. Example Audio Processing System
[0046] Figure 1 is a block diagram of an audio processing system 100 according to one or more embodiments. System 100 uses PSM processing to spatially shift an audio signal and applies other types of spatial (e.g., center / side) processing. Some embodiments of system 100 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0047] System 100 includes a PSM module 102, an L / R to M / S converter module 104, a component processor module 106, an M / S to L / R converter module 108, and a crosstalk processor module 108. The PSM module 102 receives input audio 120 and generates spatially shifted left and right channels 122 and 124. The operation of the PSM 102 according to various embodiments will be described in more detail below with reference to Figures 6 to 11.
[0048] The L / R to M / S converter module 104 receives the left channel 122 and the right channel 124 and generates a middle component 126 (e.g., a non-spatial component) and a side component 128 (e.g., a spatial component) from channels 122 and 124. In some embodiments, the middle component 126 is generated based on the sum of the left channel 122 and the right channel 122, and the side component 128 is generated based on the difference between the left channel 122 and the right channel 124. In some embodiments, the transformation of a point in the L / R space into a point in the M / S space can be expressed according to equation (1) as follows: (1)
[0049] The inverse transformation can be expressed by equation (2) as follows: (2)
[0050] It should be understood that in other embodiments, other L / R to M / S type transformations can be used to generate the middle component 126 and the side component 128. In some embodiments, the transformations shown in equations (1) and (2) can be used instead of the true intersection form, where, due to reduced computational complexity, both the forward and inverse transformations are converted by √2. For ease of discussion, regardless of the particular transformation used, the convention of transforming the coordinates of a column vector by right multiplication and the notation of the transformed coordinates carrying its base as one of its labels will be used, as shown in the following equation (3): (3)
[0051] The component processor module 106 processes the mid component 126 to generate a processed mid component 130 and processes the side component 128 to generate a processed side component 314. The processing of each of the components 126 and 128 may include various types of filtering, such as spatial cue processing (e.g., amplitude- or delay-based translation, stereo processing, etc.), single or multi-band equalization, single or multi-band dynamic processing (e.g., compression, expansion, limiting, etc.), single or multi-band gain or delay stages, adding audio effects, or other types of processing. In some embodiments, the component processor module 106 uses the mid component 126 and the side component 128 to perform sub-band spatial processing and / or crosstalk compensation processing. Sub-band spatial processing is the processing of spatially enhanced audio signals performed on the frequency sub-bands of the mid and side components. Crosstalk compensation processing is the processing of adjusting spectral artifacts caused by crosstalk processing, such as crosstalk compensation for speakers or crosstalk simulation for headphones. The various components that may be included in the component processor module 106 will be further described with respect to Figures 12A to 13.
[0052] The M / S to L / R converter module 108 receives the processed middle component 130 and the processed side component 132 and generates a processed left component 134 and a processed right component 136. In some embodiments, the M / S to L / R converter module 108 transforms the processed middle and side components 130 and 132 based on an inverse transformation performed by the L / R to M / S converter module 104. For example, the processed left component 134 is generated based on the sum of the processed middle component 130 and the processed side component 132, and the processed right component 136 is generated based on the difference between the processed middle component 130 and the processed side component 132. Other M / S to L / R type transformations can be used to generate the processed left component 134 and the processed right component 136.
[0053] Crosstalk processor module 110 receives and performs crosstalk processing on processed left component 134 and processed right component 136. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is the processing performed on an audio signal (e.g., via headphone output) to simulate the effect of a speaker. Crosstalk cancellation is the processing performed on an audio signal (e.g., via speaker output) to reduce crosstalk caused by the speaker. Crosstalk processor module 110 outputs a left channel 138 and a right output channel 140. In some embodiments, crosstalk processing (e.g., simulation or cancellation) may be performed before component processing, such as before the left channel 122 and right channel 124 are converted into center and side components. Various components that may be included in crosstalk processor module 110 will be further described with respect to Figures 15 and 16.
[0054] In some embodiments, the PSM module 100 is incorporated into the component processor module 106. The L / R to M / S converter module 104 receives a left channel and a right channel, which may represent (e.g., stereo) inputs of the audio processing system 100. The L / R to M / S converter module 104 uses the left and right input channels to generate a center component and a side component. The PSM module 100 of the component processor module 106 processes the center component and / or side component as inputs (as discussed herein with respect to input audio 102) to generate a left and right channel. The component processor module 106 may also perform other types of processing on the center and side components, and the M / S to L / R converter module 108 generates left and right channels from the processed center and side components. The left channel generated by the PSM module 100 and the left channel generated by the M / S to L / R converter module 108 are combined to generate a processed left component. The right channel generated by PSM module 100 is combined with the right channel generated by M / S to L / R converter module 108 to generate the processed right component.
[0055] System 100 provides a left channel 138 to a left speaker 112 and a right channel 140 to a right speaker 114. Speakers 112 and 114 may be components of a smartphone, tablet, smart speaker, laptop, desktop computer, fitness equipment, etc. Speakers 112 and 114 may be part of a device including system 100 or may be detached from system 100, such as being connected to system 100 via a network. The network may include wired and / or wireless connections. The network may include a local area network, a wide area network (e.g., including the Internet), or a combination thereof.
[0056] Figure 2 is a block diagram of a computing system environment 200 according to one embodiment. The computing system 200 may include an audio system 202 connected to user devices 210a and 210b via a network 208, which may include one or more computing devices (e.g., servers). The audio system 202 provides audio content to user devices 210a and 210b (also individually referred to as user device 210) via the network 208. The network 208 facilitates communication between the system 202 and the user device 210. The network 106 may include various types of networks, including the Internet.
[0057] The audio system 202 includes one or more processors 204 and computer-readable media 206. One or more processors 204 execute program modules that cause the one or more processors 204 to perform functions (such as generating multiple output channels from a single mono channel). Several processors 204 may include a central processing unit (CPU), a graphics processing unit (GPU), a controller, a state machine, other types of processing circuitry systems, or one or more combinations thereof. A processor 204 may further include local memory storing program modules, operating system data, etc.
[0058] Computer-readable media 206 is a non-transitory storage medium storing program code for PSM module 102, component processor module 106, crosstalk processor module 110, L / R and M / S conversion modules 104 and 108, and a channel summing module 212. PSM module 102 generates multiple output channels from a mono channel, which can be further processed using component processor module 106, crosstalk processor module 110, and / or L / R and M / S conversion modules 104 and 108. System 202 provides multiple output channels to user device 210a, which includes multiple speakers 214 to display each of the output channels.
[0059] The channel summing module 212 generates a mono output channel by summing together multiple output channels generated by the PSM module 102 and / or other modules. The system 202 provides the mono output channel to the user device 210b, which includes a single speaker 216 to present the mono output channel. In some embodiments, the channel summing module 212 is located at the user device 210b. The audio system 202 provides multiple output channels to the user device 210b, which converts the multiple channels into a mono output channel for the speaker 216. A user device 210 presents audio content to a user. The user device 210 may be a computing device of a user, such as a music player, smart speaker, smartphone, wearable device, tablet computer, laptop, desktop computer, or the like. (Center / Side Space Shading)
[0060] In some embodiments, spatial cues are encoded into an audio signal by generating a coloring effect in the center / side space while avoiding it in the left / right space. In some embodiments, this is achieved by applying an all-pass filter in the left / right space, the all-pass filter having the property of being explicitly selected to cause a target coloring in the center / side. For example, in a two-channel system, the relationship between the left / right phase angle and the center / side gain can be expressed by the following equation (4): (4)
[0061] wherein are two-dimensional column vectors composed of the target gain factors (in decibels) at a specific frequency ω for the center and side, respectively, and are one of the target functions of the phase relationship between the left and right channels. According to the following equations (5) and (6), the desired frequency-dependent phase difference for application in the left / right space is obtained by solving equation (4): (5) (6)
[0062] It should be noted that if a colorless constraint is imposed on the system in the left / right space, only the transfer function of the central or side component can be specified. Therefore, the system of equations (5) and (6) is overdetermined, where only one of the above equations can be solved without violating the desired symmetry. In some embodiments, a specific equation is selected to generate control over the central or side component. An additional degree of freedom can be achieved if the colorless constraint on the system in the left / right space is not considered. In systems with more than two channels, different techniques such as pairwise or hierarchical sum and difference transformations can be used instead of central and side components. An example all-pass filter implementation for encoding elevation cues.
[0063] In some embodiments, spatial perception cues can be encoded into an audio signal by embedding frequency-dependent amplitude cues (i.e., coloring) into the central / lateral space, while constraining the left / right signals to be colorless. For example, elevation cues (e.g., spatial perception cues located along a sagittal plane) can be encoded using this framework because the left / right cues for elevation are theoretically color-symmetrical.
[0064] In some embodiments, a significant feature of the elevation cue based on the Head Related Transfer Function (HRTF) is a notch that starts at about 8 kHz and monotonically increases to about 16 kHz according to elevation, which can be used to derive an appropriate coloring for encoding one of the channels in the elevation. Using this encoding cue, a corresponding frequency-dependent phase shift can be derived, which can be further used to derive a function implemented via a filter network (e.g., PSM module 100), as described below. In some embodiments, the HRTF-based elevation cue can be characterized as a notch that starts at about 8 kHz and monotonically increases to about 12 kHz according to elevation.
[0065] For ease of discussion, the following exemplary filter framework is discussed in relation to encoding the same perceptual cue, based on some embodiments, wherein the target elevation angle is 60 degrees (e.g., spatially shifting the audio content to 60 degrees above the horizontal in the sagittal plane). However, it should be understood that similar techniques can be used to encode perceptual cues with different elevation angles in other embodiments. Figure 3 illustrates a graph of a sampled HRTF measured at a 60-degree elevation angle according to some embodiments. Figure 4 illustrates a graph of an example of a perceptual cue characterized by a target amplitude function corresponding to an infinitely decaying region at approximately 11 kHz, according to some embodiments. Such cues can be used to generate elevation perception in most individuals across various presentation scenarios. Although the graph in Figure 4 illustrates a simplified sampled HRTF, it should be understood that more complex cues can also be derived based on the framework described herein. Design using second-order all-pass sections.
[0066] In some embodiments, the PSM module 100 is implemented using two independently cascaded second-order all-pass filters plus a delay element to achieve the desired phase shift in the left / right space to encode perceptual cues, such as the perceptual cues described above with respect to Figure 4. In some embodiments, the second-order segment is implemented as a dual second-order segment, wherein coefficients are applied to the feedback and feedforward taps for up to two delayed samples. As discussed herein, the convention of naming the feedback coefficients for one and two samples A1 and A2, respectively, and the convention of naming the feedforward coefficients for zero, one, and two samples B0, B1, and B2, respectively, are used.
[0067] In some embodiments, the PSM module 100 is implemented using a second-order all-pass filter configured to perform pole and zero elimination to allow the magnitude components of the transfer function to remain flat as the phase response changes. By using all-pass filter segments on both channels in the left / right space, a specific phase shift in the spectrum can be guaranteed. This has the additional benefit of allowing a given phase shift between the left and right, which will result in an increased sense of spatial extension and desired zeros in the middle / side space.
[0068] Table 1 below illustrates a set of exemplary bisecond-order coefficients that can be used in a second-order all-pass filter framework with an additional 2-sample delay on the right channel, according to some embodiments. The bisecond-order coefficients shown in Table 1 can be designed for a 44.1 kHz sampling rate, but can also be used in systems with other sampling rates (e.g., 48 kHz). Table 1
[0069] A filter network having the coefficients shown in Table 1 can produce an appropriate phase response in the left / right space resulting in a significant zero / amplification in the middle / side space at 11 kHz. Figure 5 illustrates a frequency response generated according to some embodiments by driving a second-order all-pass filter segment having the coefficients shown in Table 1 with white noise, showing the output frequency response of one of the multiple channels (middle) and 502 and one of the multiple channels (side) and 504.
[0070] In some embodiments, the PSM module 100 implemented using a second-order all-pass filter segment can be further amplified by using a cross-network to exclude processing of frequency regions where it is not needed. The use of a cross-network can increase the flexibility of the embodiment by allowing further processing of perceptually important cues (excluding unnecessary auditory data).
[0071] In some embodiments, the PSM module 100 implemented using a second-order all-pass filter segment can be implemented using a serially linked Hilbert transform network, as will be described in more detail below. Example Hilbert Transform Perceptual Sound Field Modification (HPSM) Module
[0072] Figure 6 is a block diagram of a PSM module implemented using Hilbert transform according to one or more embodiments. PSM module 600 (also referred to as a Hilbert Transform Perceptual Sound Field Modification (HPSM) module) applies a serially connected Hilbert transform network to an input audio 602 (which may correspond to the input audio 120 shown in Figure 1) to perceptually shift the input audio 602.
[0073] Module 600 includes a cross-connect module 604, a gain unit 610, a gain unit 612, a Hilbert transform module 614, a Hilbert transform module 620, a delay unit 626, a gain unit 628, a delay unit 630, a gain unit 632, an adder unit 634, and an adder unit 636. Some embodiments of module 600 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0074] The cross-connect module 604 receives input audio 602 and generates a low-frequency component 606 and a high-frequency component 608. The low-frequency component includes a sub-band of the input audio 602 having a frequency lower than that of the sub-band of the high-frequency component 608. In some embodiments, the low-frequency component 606 includes a first portion of the input audio containing low frequencies, and the high-frequency component 608 includes the remaining portion of the input audio containing high frequencies.
[0075] As will be discussed in more detail below, the high-frequency component 608 is processed using a series of Hilbert transforms, while the low-frequency component 606 bypasses the Hilbert transform series, and is then recombined with the processed high-frequency component 608. The crossover frequency between the frequency component 606 and the high-frequency component 608 is adjustable. For example, more frequencies can be included in the high-frequency component 608 to increase the perceived intensity of spatial displacement of the HPSM module 600, while more frequencies can be included in the low-frequency component 606 to decrease the perceived intensity of displacement. In another example, the crossover frequency is set such that the frequency of a sound source of interest (e.g., a speech) is included in the high-frequency component 608.
[0076] Input audio 602 may comprise a mono channel or may be a mixture of a stereo signal or other multi-channel signals (e.g., surround sound, high-fidelity stereo, etc.). In some embodiments, input audio 602 is audio content associated with a sound source to be incorporated into an audio mix. For example, input audio 602 may be speech processed by module 600 and the processing result combined with other audio content (e.g., background music) to generate an audio mix.
[0077] Gain unit 610 applies a gain to low-frequency component 606 and gain unit 612 applies a gain to high-frequency component 608. Gain units 610 and 612 can be used to adjust the overall level of low-frequency component 606 and high-frequency component 608 relative to each other. In some embodiments, gain unit 610 or gain unit 612 may be omitted from module 600.
[0078] Hilbert transform modules 614 and 620 apply a series of Hilbert transforms to the high-frequency component 608. Hilbert transform module 614 applies a Hilbert transform to the high-frequency component 608 to generate a left branch component 616 and a right branch component 618. The left branch component 616 and the right branch component 618 are audio components that are 90 degrees out of phase with each other. In some embodiments, the left branch component 616 and the right branch component 618 are out of phase with each other by an angle other than 90 degrees, such as between 20 degrees and 160 degrees.
[0079] The Hilbert transform module 620 applies a Hilbert transform to the right branch component 618 generated by the Hilbert transform module 614 to generate a left branch component 122 and a right branch component 624. The left branch component 622 and the right branch component 624 are audio components that are 90 degrees out of phase with each other. In some embodiments, the Hilbert transform module 620 generates the right branch component 624 but not the left branch component 122. In some embodiments, the left branch component 622 and the right branch component 624 are out of phase with each other by an angle other than 90 degrees, such as between 20 degrees and 160 degrees.
[0080] In some embodiments, each of the Hilbert transform modules 614 and 620 is implemented in the time domain and includes a cascaded all-pass filter and a delay, as discussed in more detail below with reference to FIG7. In other embodiments, the Hilbert transform modules 614 and 620 are implemented in the frequency domain.
[0081] Delay unit 626, gain unit 628, delay unit 630, and gain unit 632 provide tuning control to manipulate the perceptual outcome of the program in module 600. Delay unit 626 applies a time delay to the left branch component 616 generated by Hilbert transform module 614. Gain unit 628 applies a gain to the left branch component 616. In some embodiments, delay unit 626 or gain unit 628 may be omitted from module 600.
[0082] Delay unit 630 applies a time delay to the right branch component 624 generated by Hilbert transform module 620. Gain unit 632 applies a gain to right branch component 624. In some embodiments, delay unit 630 or gain unit 632 may be omitted from module 600.
[0083] Adder 634 combines low-frequency component 606 with left-branch component 616 to generate left channel 642. Left-branch component 616 is derived from the output of one of the first Hilbert transform modules 614 in the series. Left-branch component 616 may include a delay applied by delay unit 626 and a gain applied by gain unit 628.
[0084] Adder 636 combines low-frequency component 606 with right-branch component 624 to generate right channel 644. Right-branch component 624 is derived from the output of one of the second Hilbert transform modules 620 in the series. Right-branch component 624 may include a delay applied by delay unit 626 and a gain applied by gain unit 628.
[0085] Figure 7 is a block diagram of a Hilbert transform module 700 according to one or more embodiments. The Hilbert transform module 700 is an example of either a Hilbert transform module 614 or a Hilbert transform module 620. The Hilbert transform module 700 receives an input component 702 and uses the input component 702 to generate a left branch component 712 and a right branch component 724. Some embodiments of the Hilbert transform module 700 have components different from those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.
[0086] The Hilbert transform module 700 includes an all-pass filter cascade module 740 for generating the left branch component 712 and a delay unit 714 and an all-pass filter cascade module 742 for generating the right branch component 724. The all-pass filter cascade module 714 includes a series of all-pass filters 704, 706, 708 and 710. The delay unit 714 applies a time delay to the input component 702. The all-pass filter cascade module 742 includes a series of all-pass filters 716, 718, 720 and 722. Each of the all-pass filters 704 to 710 and 716 to 722 transmits frequencies with equal gain while changing the phase relationship between different frequencies. In some embodiments, each of the all-pass filters 704 to 710 and 716 to 722 is a double second-order filter defined by equation (7): (7)
[0087] where z is a complex variable, and the coefficients a0, a1, a2, b0, b1, and b2 are filter coefficients. Different bi-second-order filters may contain different coefficients applying different phase transitions.
[0088] All-pass filter cascade modules 740 and 742 may each contain a different number of all-pass filters. The Hilbert transform module 700 has one of eight all-pass filters, an 8th-order filter, with four left-branch components 712 and four right-branch components 724. In other embodiments, the Hilbert transform module 700 may be an 8th-order filter (e.g., four all-pass filters in all-pass filter cascade modules 740 and 742) or a 6th-order filter (e.g., three all-pass filters in all-pass filter cascade modules 740 and 742).
[0089] As discussed above in conjunction with Figure 6, module 600 includes a series of Hilbert transform modules 614 and 620. Using Hilbert transform module 700 for each of Hilbert transform modules 614 and 620, the left branch component 616 is generated by a cascaded all-pass filter module 740 applied to the high-frequency component 608. The right branch component 624 is generated by two cascaded modules 742 passing through Hilbert transform module 700, two delay units 714, and two all-pass filter modules. In some embodiments, Hilbert transform modules 614 and 620 may be different. For example, Hilbert transform modules 614 and 620 may include filters of different orders, such as an 8th-order filter for one of the Hilbert transform modules and a 6th-order filter for the other.
[0090] When Hilbert transform module 700 is used in Hilbert transform modules 614 and 620, the right branch component 624 includes a phase and delay relationship with the right branch component 618 generated by the all-pass filter and delay of Hilbert transform module 620. The right branch component 624 also includes a phase and delay relationship with the high-frequency component 608 generated by the all-pass filter and delay of Hilbert transform modules 614 and 620. In some embodiments, Hilbert transform module 620 uses left branch component 616 instead of right branch component 618 to generate left branch component 622 and right branch component 624. This results in the right branch component 624 having a phase and delay relationship with the high-frequency component 608 generated by the all-pass filter (e.g., without delay) of Hilbert transform 614 and the delay and all-pass filter of Hilbert transform module 620.
[0091] Figure 8 illustrates a frequency response generated by a white noise-driven HPSM module (as described in Figure 6) according to some embodiments, showing the output frequency response of one of the multiple channels (center) and 802 and one of the multiple channels (side) and 804.
[0092] As shown in Figure 8, while this filter does generate the desired perceptual cue in the region of approximately 11 kHz, it also imparts additional coloration to the lower frequencies in the middle and sides. In some embodiments, this can be corrected by applying a cross-network (such as the cross-network module 604 illustrated in Figure 6) to the input audio so that the HPSM module processes only the audio data within the desired frequency range (e.g., a high-frequency component) or by directly removing pole / zero pairs corresponding to that spectral transform region. A first-order nonorthogonal rotational decorrelation (FNORD) design is used.
[0093] In some embodiments, a similar perceptual effect can be achieved using a first-order nonorthogonal rotational decorrelation (FNORD) filter network. Figure 9 is a block diagram of a PSM module 900 implemented using an FNORD filter network according to some embodiments. The PSM module 900 (which may correspond to the PSM module 102 illustrated in Figure 1) provides decorrelation of a mono channel into multiple channels and includes an amplitude response module 902, an all-pass filter configuration module 904, and an all-pass filter module 906. The PSM module 900 processes a mono input channel x(t) 912 to generate multiple output channels, such as a channel ya(t) provided to a speaker 910a and a channel yb(t) provided to a speaker 910b (which may correspond to the left speaker 112 and the right speaker 114 illustrated in Figure 1). Although two output channels are shown, the PSM module 900 can generate any number of output channels (each referring to a channel y(t)). The PSM module 900 may be implemented as part of a computing device, such as a music player, speaker, smart speaker, smartphone, wearable device, tablet computer, laptop computer, desktop computer, or the like. Although Figure 9 illustrates the PSM module 900 as including an amplitude response module 902, a filter configuration module 904, and an all-pass filter module 906, in some embodiments, the PSM module 900 may include the all-pass filter module 906, wherein the amplitude response module 902 and / or the filter configuration module 904 are implemented separately from the PSM module 900.
[0094] The amplitude response module 902 determines and defines a target amplitude response for one or more spatial cues encoded into the output channel y(t) (e.g., encoded into the middle and side components of the output channel y(t)). The target amplitude response is defined by the relationship between the amplitude value and frequency value of the channel (e.g., the middle and side components of the channel), such as an amplitude that varies with frequency. In some embodiments, the target amplitude response defines one or more spatial cues on the channel, which may include a target broadband attenuation, a target subband attenuation, a critical point, a filter characteristic, or a sound field position. The amplitude response module 902 may receive data 914 and a mono channel x(t) 912 and use these inputs to determine the target amplitude response. The data 914 may include information such as the characteristics of the spatial cue to be encoded, the characteristics of a presentation device (e.g., one or more speakers), the expected content of the audio data, or the listener's perceptual ability in the background. In some embodiments, the mono channel x(t) 912 may correspond to the audio input 120 illustrated in FIG1 or a portion of the audio input (e.g., a high-frequency component of the input audio, such as the high-frequency component 608 of the input audio 602 illustrated in FIG6). In embodiments where the mono channel x(t) 912 corresponds to a portion of the audio input, the output channel y(t) may be combined with a channel corresponding to the remaining portion of the audio input (e.g., a low-frequency component illustrated in FIG6) to generate a combined output channel.
[0095] Target broadband attenuation is a specification of attenuation across all frequencies. Target subband attenuation is a specification of amplitude within a frequency range defined by a subband. Target amplitude response may include one or more target subband attenuation values for a different subband.
[0096] A critical point is a specification of the curvature of the target amplitude response of a filter, described as a frequency value at which the gain of one of the output channels (e.g., a side component of the output channel) is at a predetermined value (such as -3 dB or -∞ dB). The placement of this point can have a global effect on the curvature of the target amplitude response. One instance of a critical point corresponds to a frequency at which the target amplitude response is -∞ dB. Because the behavior of the target amplitude response is to invalidate the signal at frequencies close to this point, this critical point is a zero point. Another instance of a critical point corresponds to a frequency at which the target amplitude response is -3 dB. Because the behavior of the target amplitude response of the sum and difference channels (e.g., the side components of the channel) intersects at this point, this critical point is an intersection point.
[0097] Filter characteristics are parameters that specify how the in-channel and side components are filtered. Examples of filter characteristics include a high-pass filter characteristic, a low-pass characteristic, a band-pass characteristic, or a band-reject characteristic. The shape of the filter characteristic description is as if it were the result of an equalization filter. Equalization filtering can be described in terms of which frequencies can be passed through the filter or which frequencies are rejected. Thus, a low-pass characteristic allows frequencies below a certain inflection point to pass through and attenuates frequencies above the inflection point. A high-pass characteristic, conversely, allows frequencies above a certain inflection point to pass through and attenuates frequencies below the inflection point. A band-pass characteristic allows frequencies within a frequency band around a certain inflection point to pass through while attenuating other frequencies. A band-reject characteristic rejects frequencies within a frequency band around a certain inflection point while allowing other frequencies to pass through.
[0098] The target amplitude response can define more than one spatial thread encoded into the output channel y(t). For example, the target amplitude response can define a spatial thread specified by the critical point and one of the filter characteristics of one or more components of an all-pass filter. In another instance, the target amplitude response can define a spatial thread specified by the target broadband attenuation, the critical point, and the filter characteristics. Although discussed as independent specifications, the specifications of most regions of the parameter space can be interdependent. This result can be caused by the nonlinearity of the system in terms of phase. To address this, additional higher-order descriptors for the target amplitude response can be designed, which are nonlinear functions of the target amplitude response parameters.
[0099] The filter configuration module 904 determines the properties of a single-input multiple-output (SMO) all-pass filter based on the target amplitude response received by the self-amplitude response module 902. Specifically, the filter configuration module determines a transfer function of the all-pass filter based on the target amplitude response and determines the coefficients of the all-pass filter based on the transfer function. The all-pass filter is a decorrelation filter whose encoding is based on a spatial cue described by a target amplitude response and applied to a mono input channel x(t) to generate output channels ya(t) and yb(t).
[0100] An all-pass filter may include different configurations and parameters based on spatial cues and / or constraints defined by the target amplitude response. One of the filters with a target amplitude response having encoded spatial cues may be colorless, for example, preserving the spectral content (e.g., all) of individual output channels (e.g., left / right output channels). Thus, the filter can be used to encode elevation cues by embedding coloring in the center / side space in the form of frequency-dependent amplitude cues, while preserving the spectral content of the left and right signals. Because the filter is colorless, the mono content can be placed at a specific location in the sound field (e.g., specified by the target elevation angle), where the spatial location of the audio signal is decoupled from its overall coloring.
[0101] Figures 10A and 10B are block diagrams of an example PSM module based on a first-order non-orthogonal rotating decorrelation (FNORD) technique according to some embodiments. Figure 10A shows a detailed view of a PSM module 900 according to some embodiments, while Figure 10B provides a more detailed view of a broadband phase rotator 1004 within an all-pass filter module 906 of the PSM module 900 according to some embodiments.
[0102] As shown in Figure 10A, the all-pass filter module 906 receives information in the form of a mono input audio signal x(t) 912, a rotation control parameter θbf 1048, and a first-order coefficient βbf 1050. The input audio signal x(t) 912 and the rotation control parameter θbf 1048 are utilized by a wideband phase rotator 1004, which processes the input audio signal 912 using the rotation control parameter θbf 1048 to generate a left wideband rotation component 1020 and a right wideband rotation component 1022. According to some embodiments, the left wideband rotation component 1020 is then provided to a narrowband phase rotator 1024 for further processing, while the right wideband rotation component 1022 is output as the output channel yb(t) of the PSM module 900 (e.g., as the right output channel). Narrow-band phase rotator 1024 receives a left wideband rotation component 1020 from wideband phase rotator 1004 and a first-order coefficient βbf 1050 from filter configuration module 904 to generate a narrow-band rotation component 1028, which is then provided as the output channel ya(t) of PSM module 900 (e.g., as the left output channel).
[0103] According to some embodiments, the control data 914 for configuring the amplitude response module 902 may include a critical point fc 1038, a filter characteristic θbf 1036, and a sound field position Γ 1040. This data is provided to the PSM module 900 via the amplitude response module 902, which determines it to be an intermediate representation of data in the form of a critical point (in radians) ωc 1044, a filter characteristic θbf 1042, and a quadratic term φ 1046. In some embodiments, the amplitude response module 902 modifies one or more of the parameters of the control data 914 (e.g., the critical point fc 1038, the filter characteristic θbf 1036, and / or the sound field position Γ 1040) based on one or more parameters of the input audio signal x(t) 912. In some embodiments (such as the one shown in FIG. 10A), the filter characteristic θbf 1042 is equivalent to the filter characteristic θbf 1036. These intermediate representations 1042, 1044, and 1046 are provided to a filter configuration module 904, which generates filter configuration data that may include at least a first-order coefficient βbf 1050 and a rotation control parameter 1048. The first-order coefficient βbf 1050 is provided to an all-pass filter module 906 via a first-order all-pass filter 1026. In some embodiments, the rotation control parameter θbf 1048 may be equivalent to filter characteristics θbf 1036 and 1042, while in other embodiments, this parameter may be scaled for convenience. For example, in some embodiments, the filter characteristic is associated with a parameter range (e.g., 0 to 0.5) having a meaningful center point, and the rotation control parameter is scaled relative to the filter characteristic to change the parameter range, for example, from 0 to 1. In some embodiments, the filter characteristics are linearly scaled (e.g., preserving the increased resolution at the poles relative to the center point), while in other embodiments, a non-linear mapping may be used (e.g., for the increased numerical resolution around the center point). In the following equations, the rotation control parameter θbf is considered unscaled, but it should be understood that the same principle applies when scaling the rotation control parameter. The rotation control parameter θbf 1048 is provided to the all-pass filter module 906 via a wideband phase rotator 1004.
[0104] FIG10B describes in detail one exemplary embodiment of a broadband phase rotator 1004 according to some embodiments. The broadband phase rotator 1004 receives information in the form of a mono input audio signal x(t) 912 and rotation control parameters θbf 1048. The input audio signal x(t) 912 is first processed by a Hilbert transform module 1006 to generate a left branch component 1008 and a right branch component 1010. According to some embodiments, the Hilbert transform module 1006 can be implemented using the configuration shown in FIG7, but it should be understood that other embodiments of the Hilbert transform module 1006 can be used in other embodiments. The left branch component 1008 and the right branch component 1010 are provided to a 2D orthogonal rotation module 1012. The left branch component 1008 is also provided to the output of the broadband phase rotator 1004 as a right broadband rotation component 1022. When the broadband phase rotator 1004 is configured to rotate the left and right branch signals relative to each other, in some embodiments this is achieved by keeping the left branch component 1008 constant as the right broadband rotation component 1022 and rotating the left and right branch components to form the left broadband rotation component 1020.
[0105] According to some embodiments, in addition to the left branch component 1008 and the right branch component 1010, the 2D orthogonal rotation module 1012 may also receive a rotation control parameter θbf 1048 from the filter configuration module 904, as shown in FIG10A. The 2D orthogonal rotation module 1012 uses this data to generate the left rotation component 1014 and the right rotation component 1016. The projection module 1018 then receives the left rotation component 1014 and the right rotation component 1016, which are combined (e.g., added) to form the left broadband rotation component 1020. As shown in Figure 10A, the wideband phase rotator 1004 outputs the left wideband rotation component 1020 to the narrowband phase rotator 1024 to generate the narrowband rotation component 1028 as the left output channel ya(t) of the PSM module and to generate the right wideband rotation component 1022 as the right output channel yb(t) of the PSM module (which bypasses or passes through the narrowband phase rotator 1024 indefinitely). In other embodiments, the narrowband rotation component 1028 and the left branch component 1008 (which acts as the right wideband rotation component 1022 in the embodiments shown in Figures 10A and 10B) are instead mapped to the right and left output channels yb(t) and ya(t), respectively.
[0106] In some embodiments, the PSM module 900 may be formally described by the following equation (8): (8)
[0107] In some embodiments, this single-input multiple-output all-pass filter consists of several parts, each of which will be interpreted sequentially. According to some embodiments, these components may include Af, Ab, and H2.
[0108] According to some embodiments, Af may correspond to the narrow-band phase rotator 1024 in FIG10A. Af is a first-order all-pass filter having the output of one channel in the form of equation (9): (9) where βf is a coefficient of the filter in the range of -1 to +1. The second output of the filter may pass through the input unchanged. Therefore, according to some embodiments, the filter Af implementation may be defined by equation (10): (10)
[0109] The transfer function of Af is expressed as a differential phase shift from one output to another. This differential phase shift is a function of the angular frequency ω defined by equation (11): (11) where the target amplitude response can be derived by replacing θ in equation (5) or (6), depending on whether the response is placed in the middle (equation (5)) or the side (equation (6)).
[0110] The summation gain αf = 3 dB can be used as the frequency fc of the tuning critical point, which is defined as follows: (12) and (13)
[0111] By normalizing the target amplitude response to 0 dB, this critical point corresponds to the parameter fc, which can be -3 dB. In equation (8), the output of Af is indicated by a subscript: according to some embodiments, only the output of the first channel is used.
[0112] In equation (8), Ab is a single-input multiple-output full-pass filter, which corresponds to the broadband phase rotator 1004 in Figure 10A. Ab can be formally defined as in equation (14): (14) where H2(x(t)) is a discrete form of the filter, implemented using a pair of orthogonal full-pass filters, and defined using a continuous-time prototype according to equation (15): (15)
[0113] In some embodiments, the all-pass filter provides constraints on the 90-degree phase relationship between the two output signals and the unit amplitude relationship between the input and the two output signals, but does not necessarily guarantee a specific phase relationship between the input (mono) signal and either of the two (stereo) output signals. The discrete form of
[0114] is represented by the symbol and defined by its effect on the mono signal x(t). The result is a two-dimensional vector defined by equation (16): (16)
[0115] According to some embodiments, the discrete single-input multiple-output all-pass filter may correspond to the Hilbert transform module 1006 in FIG10B and also to the Hilbert transform module 700 in FIG7. According to some embodiments, in equation (14), θ determines the rotation angle of the first output of Ab relative to the second output.
[0116] Finally, according to some embodiments, the parameters supplied to the complete system Abf in equation (8) can be determined as follows. These parameters may include βbf and θbf, which correspond to the rotation control parameter θbf1048 and the first-order coefficient βbf1050 in Figure 10A. In some embodiments, βbf can be determined from a center angular frequency ωc as follows: (17) where ωc can be calculated from a desired center frequency fc using equation (12). In Figure 10A, ωc corresponds to the critical point ωc1044, fc corresponds to the critical point fc1038, and the action of equation (17) is partially executed within the filter configuration module 904, resulting in the first-order coefficient βbf. In some embodiments, the quadratic term φ1046 can be derived from θbf and a Boolean sound field position parameter Γ via equation (18): (18)
[0117] This quadratic term 1046 is provided to the filter configuration module 904 by the amplitude response module 902 in Figure 10A.
[0118] In some embodiments, higher-order parameters fc, θbf, and Γ are sufficient for intuitive and convenient tuning of the system. According to these embodiments, the center frequency fc determines an inflection point (in Hz) where the target amplitude response asymptotically approaches -∞ dB. The parameter θbf allows control of the filter characteristics around the inflection point fc. For 0 < θbf < 1 / 4, the characteristics are low-pass and have a zero at fc, and one spectral slope in the target amplitude function smoothly interpolates from a tendency towards low frequencies to a flat surface as θbf increases. For 1 / 4 < θbf < 1 / 2, the characteristics smoothly interpolate from a flat surface with a zero at fc to a high-pass surface as θbf increases. At the point θbf = 1 / 4, the target amplitude function is pure band-resistant and has a zero at fc. The parameter Γ is a Boolean value that places the target amplitude function determined by fc and θbf into the middle channel (i.e., L+R) or the side channel (i.e., LR). Due to the full-pass constraint on the two outputs of the filter network, the effect of Γ is to switch between complementary target amplitude responses.
[0119] In some embodiments, to achieve an amplitude response of one vertical criterion at 60 degrees, the FNORD filter network described above can be configured using parameters fc = 11 kHz, θbf = 0.13, and Γ = 1. Figure 11 illustrates a frequency response diagram showing the output frequency response of an FNORD filter network configured to achieve an amplitude response of one vertical criterion at 60 degrees according to some embodiments. Figure 11 illustrates the output frequency response in the mid-component 1110 and the lateral component 1120, wherein the FNORD filter network is driven by white noise. In some embodiments, the filter parameters fc, θbf, and / or Γ are selected based on an analysis of the HRTF-based elevation criterion at the desired angle.
[0120] In some embodiments, the PSM module 900 uses a frequency domain specification for the all-pass filter. For example, in some cases, a more complex spatial cue is required, such as a spatial cue sampled from a human factors engineering dataset. Within certain limitations, the above techniques can be used to embed an arbitrary cue into the phase difference of an audio stream based on the frequency domain representation of the cue's magnitude. For example, the filter configuration module 904 can use an equation of the form of equation (5) or (6) to vectorize one of the K narrow-band attenuation coefficients in the middle or side of the target amplitude response and to determine one of the K phase angles as a vectorized transfer function:
[0121] The phase angle vector θ is generated by equation (19), which defines a finite impulse response filter: (19) where DFT-1 represents the inverse discrete Fourier transform (idft) and the vector of 2(K-1) FIR filter coefficients Bn(θ) can then be applied to x(t), as defined by equation (20): (20) where convolution operation is represented.
[0122] To reproduce the effect from the aforementioned example and achieve a target amplitude response corresponding to a height cue of 60 degrees, an observed HRIR can be sampled and applied to a DFT of length 2 (K-1) such that it can be used to determine a target amplitude response vector using the following operation: (21) where and are operations that return the real and imaginary parts of a complex number, respectively, and all operations are applied to the vector component by component. This target amplitude response inserted into the middle or side can now be applied to one of equations (5) or (6) to determine a vector of K phase angles from which an FIR filter B can be derived. This filter is then inserted into equation (19) to derive a single-input multiple-output all-pass filter.
[0123] Although equations (19) and (20) provide an effective way to constrain the target amplitude response, their implementation will generally rely on a relatively high-order FIR filter generated by an inverse DFT operation. This may not be suitable for resource-constrained systems. In such cases, a low-order infinite impulse response (IIR) implementation can be used, such as that discussed in conjunction with equation (8).
[0124] The all-pass filter module 906 applies the all-pass filter configured by the filter configuration module 904 to the mono channel x(t) to generate output channels ya(t) and yb(t). The application of the all-pass filter to channel x(t) can be performed as defined by equations (8), (20) or as depicted in FIG9 or FIG10A. The all-pass filter module 906 provides each output channel to a respective loudspeaker, such as providing channel ya(t) to loudspeaker 910a and channel yb(t) to loudspeaker 910b. Although not shown in FIG9, it should be understood that output channels ya(t) and yb(t) can be provided to loudspeakers 910a and 910b via one or more intermediate components (e.g., component processor module 106, crosstalk processor module 110 and / or L / R to M / S converter module and M / S to L / R converter modules 104 and 108, as shown in FIG1).
[0125] In some embodiments, PSM processing may be performed on a target portion of a received audio signal, such as a mid-component or a super-mid-component of the audio signal. FIG12 is a block diagram of an audio processing system 1200 according to one or more embodiments. System 1200 generates a super-mid-component to isolate a target portion (e.g., speech) of an audio signal, and performs PSM processing on the super-mid-component to spatially shift the target portion. Some embodiments of system 1200 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0126] System 1200 includes an L / R to M / S converter module 1206, an orthogonal component generator module 1212, an orthogonal component processor module 1214 including a PSM module 102, and a serial audio processor module 1224.
[0127] The L / R to M / S converter module 1206 receives a left channel 1202 and a right channel 1204 and generates a middle component 1208 and a side component 1210 from channels 1202 and 1204. The discussion concerning the L / R to M / S converter module 104 can be applied to the L / R to M / S converter module 1206.
[0128] The quadrature component generator module 1212 processes the center component 1208 and the side component 1210 to generate at least one of the following: a super-center component M1, a super-side component S1, a residual center component M2, and a residual side component S2. The super-center component M1 is the spectral energy of the center component 1208 with the spectral energy of the side component 1210 removed. The super-side component S1 is the spectral energy of the side component 1210 with the spectral energy of the center component 1208 removed. The residual center component M2 is the spectral energy of the center component 1208 with the spectral energy of the super-center component M1 removed. The residual side component S2 is the spectral energy of the side component 1210 with the spectral energy of the super-side component S1 removed. The system 1200 generates the left channel 1242 and the right output channel 1244 by processing at least one of the super-center component M1, the super-side component S1, the residual center component M2, and the residual side component S2. The orthogonal component generator module 1212 is further described relative to Figures 13A, 13B and 13C.
[0129] The quadrature component processor module 1214 processes one or more of the super-center component M1, super-side component S1, residual center component M2, and / or residual side component S2, and converts the processed components into a processed left component 1220 and a processed right component 1222. The discussion concerning component processor module 106 can be applied to quadrature component processor module 1214, except that processing is performed on the super-center component M1, super-side component S1, residual center component M2, and / or residual side component S2, but not on the center and side components. For example, the processing of components M1, M2, S1, and S2 can include various types of processing such as spatial cue processing (e.g., amplitude or delay-based translation, stereo processing, etc.), single or multi-band equalization, single or multi-band dynamic processing (e.g., compression, expansion, limiting, etc.), single or multi-band gain or delay stages, adding audio effects, or other types of processing. In some embodiments, the quadrature component processor module 1214 uses the super-center component M1, the super-side component S1, the residual center component M2, and / or the residual side component S2 to perform sub-band spatial processing and / or crosstalk compensation processing. The quadrature component processor module 1214 may further include an L / R to M / S converter to convert components M1, S2, S1, and S2 into a processed left component 1220 and a processed right component 1222.
[0130] The quadrature component processor module 1214 further includes a PSM module 102, which can operate on one or more of the super-center component M1, the super-side component S1, the residual center component M2, and / or the residual side component S2. For example, the PSM module 102 can receive the super-center component M1 as input and generate spatially shifted left and right channels. For example, the super-center component M1 may include an isolated portion of the audio signal representing speech, and can therefore be selected for HPSM processing. The left channel generated by the PSM module 102 is used to generate the processed left component 1020, and the right channel generated by the PSM module 102 is used to generate the processed right component 1222. The quadrature component processor module 1214 is further described with respect to FIG12.
[0131] Crosstalk processor module 1224 receives processed left component 1220 and processed right component 1222 and performs crosstalk processing on processed left component 1220 and processed right component 1222. Crosstalk processor module 1224 outputs left channel 1242 and right channel 1244. Discussions relating to crosstalk processor module 1224 are applicable to crosstalk processor module 1224. In some embodiments, crosstalk processing (e.g., analog or cancellation) may be performed prior to quadrature component processing, such as before converting left channel 1202 and right channel 1204 into center and side components. Left channel 1242 is available to left speaker 112 and right channel 1244 is available to right speaker 114. Indicative quadrature component generator
[0132] Figures 13A to 13C are block diagrams of orthogonal component generator modules 1313, 1323, and 1343 according to one or more embodiments. Orthogonal component generator modules 1313, 1323, and 1343 are examples of orthogonal component generator module 1212. Some embodiments of modules 1313, 1323, and 1343 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0133] Referring to Figure 13A, the orthogonal component generator module 1313 includes a subtraction unit 1305, a subtraction unit 1309, a subtraction unit 1315, and a subtraction unit 1319. As described above, the orthogonal component generator module 1313 receives the center component 1208 and the side component 1210, and outputs one or more of the super-center component M1, the super-side component S1, the remaining center component M2, and the remaining side component S2.
[0134] Subtraction unit 1305 removes the spectral energy of side component 1210 from the spectral energy of mid component 1208 to generate super-mid component M1. For example, subtraction unit 1305 subtracts a value of side component 1210 in the frequency domain from a value of mid component 1208 in the frequency domain while retaining only the phase to generate super-mid component M1. Frequency domain subtraction can be performed on the time domain signal using a Fourier transform to generate a signal in the frequency domain and then the signals in the frequency domain are subtracted. In other instances, frequency domain subtraction can be performed in other ways, such as using a wavelet transform instead of a Fourier transform. Subtraction unit 1309 generates a residual mid component M2 by removing the spectral energy of super-mid component M1 from the spectral energy of mid component 1208. For example, subtraction unit 1309 subtracts a value of super-mid component M1 in the frequency domain from a value of mid component 1208 in the frequency domain while retaining only the phase to generate residual mid component M2. However, in the time domain, subtracting the side component from the center results in the original right channel of the signal. In the frequency domain, the above operation isolates and distinguishes a portion of the spectral energy of the middle component that is different from the side component (referred to as M1 or super-middle) from a portion of the spectral energy of the middle component that is the same as the side component (referred to as M2 or residual middle).
[0135] In some embodiments, additional processing may be used when the spectral energy of the middle component 1006 minus the spectral energy of the side component 1210 results in a negative value for the super-middle component M1 (e.g., for one or more frequency grids in the frequency domain). In some embodiments, the super-middle component M1 is clamped to a value of 0 when the spectral energy of the middle component 1208 minus the spectral energy of the side component 1210 results in a negative value. In some embodiments, the super-middle component M1 is wrapped around by taking the absolute value of the negative value as the value of the super-middle component M1. Other types of processing may be used when the spectral energy of the middle component 1208 minus the spectral energy of the side component 1210 results in a negative value for M1. Similar additional processing, such as clamping to 0, wrapping, or other processing, may be used when the subtraction of the super-side component S1, the remaining side component S2, or the remaining middle component M2 results in a negative value. Clamping the super-middle component M1 to 0 when the subtraction results in a negative value will provide spectral orthogonality between M1 and the two side components. Similarly, when subtraction results in a negative value, clamping the super-side component S1 to 0 provides spectral orthogonality between S1 and the two mid-components. By generating orthogonality between the super-mid and side components and their appropriate mid / side counterparts (i.e., the side component of the super-mid, the mid component of the super-side), the derived residual mid-component M2 and residual side component S2 contain spectral energies that are not orthogonal (i.e., common) to their appropriate mid / side counterparts. That is, when clamping to 0 is applied to the super-mid and the residual mid-component is derived using this M1 component, a super-mid component without spectral energy common to the side components and a residual mid-component with spectral energy completely common to the side components are generated. The same relationship applies to the super-side and residual sides when the super-side is clamped to 0. When frequency domain processing is applied, there is usually a resolution trade-off between frequency and time information. As frequency resolution increases (i.e., as the FFT window size and the number of frequency grids increase), time resolution decreases, and vice versa. The aforementioned spectral subtraction occurs per frequency grid, and therefore in some cases (such as when removing acoustic energy from the super-mid component), a large FFT window size (e.g., 8192 samples, resulting in 4096 frequency grids for a real-valued input signal) may be preferable. Other cases may require higher time resolution and therefore lower total delay and lower frequency resolution (e.g., a 512-sample FFT window size, resulting in 256 frequency grids for a real-valued input signal). In the latter case, the low-frequency resolution when subtracting the mid and side components from each other to derive the super-mid component M1 and the super-side component S1 can produce audible spectral artifacts because the spectral energy of each frequency grid is an average representation of energy over an excessively wide frequency range. In this case, obtaining the absolute value of the difference between the mid and side components when deriving the super-mid component M1 or the super-side component S1 can help mitigate perceptual artifacts by allowing the true orthogonality of the components to diverge per frequency grid.Instead of surrounding 0, we can apply a coefficient to the subtracted value to scale the value between 0 and 1 and thus provide a method of interpolation between the full orthogonality of the super and remaining middle / side components at one extreme (i.e., with a value of 1) and the super-middle M1 and super-side S1, which are equal to their corresponding original middle and side components, at the other extreme (i.e., with a value of 0).
[0136] Subtraction unit 1315 removes the spectral energy of the mid-component 1208 in the frequency domain from the spectral energy of the side component 1210 in the frequency domain, while retaining only the phase, to generate a super-side component S1. For example, subtraction unit 1315 subtracts a value of the mid-component 1208 in the frequency domain from a value of the side component 1210 in the frequency domain, while retaining only the phase, to generate a super-side component S1. Subtraction unit 1319 removes the spectral energy of the super-side component S1 from the spectral energy of the side component 1210 to generate a residual side component S2. For example, subtraction unit 1319 subtracts a value of the super-side component S1 in the frequency domain from a value of the side component 1210 in the frequency domain, while retaining only the phase, to generate a residual side component S2.
[0137] In Figure 5B, the quadrature component generator module 1323 is similar to the quadrature component generator module 1313 because it receives the center component 1006 and the side component 1210 and generates the super-center component M1, the remaining center component M2, the super-side component S1, and the remaining side component S2. The difference between the quadrature component generator module 1323 and the quadrature generator module 1313 is that the super-center component M1 and the super-side component S1 are generated in the frequency domain and then these components are converted back to the time domain to generate the remaining center component M2 and the remaining side component S2. The orthogonal component generator module 1323 includes a forward FFT unit 1320, a bandpass unit 1322, a subtraction unit 1324, a superprocessor 1325, an inverse FFT unit 1326, a delay unit 1328, a subtraction unit 1330, a forward FFT unit 1332, a bandpass unit 1334, a subtraction unit 1336, a superprocessor 1337, an inverse FFT unit 1340, a delay unit 1342, and a subtraction unit 1344.
[0138] The forward Fast Fourier Transform (FFT) unit 1320 applies a forward FFT to the mid-component 1208 to transform the mid-component 1208 to a frequency domain. The transformed mid-component 1208 in the frequency domain includes a magnitude and a phase. The bandpass unit 1322 applies a bandpass filter to the frequency-domain mid-component 1208, wherein the bandpass filter specifies a frequency in the super-mid-component M1. For example, to isolate a typical vocal range, the bandpass filter may specify frequencies between 300 Hz and 8000 Hz. In another instance, to remove audio content associated with a typical vocal range, the bandpass filter may maintain lower frequencies (e.g., generated by a bass guitar or drum) and higher frequencies (e.g., generated by cymbals) in the super-mid-component M1. In other embodiments, in addition to and / or instead of the bandpass filter applied by the bandpass unit 1322, the quadrature component generator module 1323 applies various other filters to the frequency-domain mid-component 1208. In some embodiments, the quadrature component generator module 1323 does not include a bandpass unit 1322 and does not apply any filters to the frequency-domain component 1208. In the frequency domain, the subtraction unit 1324 subtracts the side component 1210 from the filtered mid-component to generate the super-intermediate component M1. In other embodiments, instead of the subsequent processing applied to the super-intermediate component M1 performed by a quadrature component processor module (e.g., the quadrature component processor module of FIG. 12), the quadrature component generator module 1323 applies various audio enhancements to the frequency-domain super-intermediate component M1. The super-intermediate processor 1325 performs processing on the super-intermediate component M1 in the frequency domain before converting it to the time domain. The processing may include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of the processing that may be performed by the quadrature component processor module 1214, the super-intermediate processor 1325 performs processing on the super-intermediate component M1. Inverse FFT unit 1326 applies an inverse FFT to the super-intermediate component M1 to transform it back to the time domain. The super-intermediate component M1 in the frequency domain contains a magnitude of M1 and the phase of the intermediate component 1208. Inverse FFT unit 1326 transforms the intermediate component 1208 to the time domain. Delay unit 1328 applies a delay to the intermediate component 1208, causing both the intermediate component 1208 and the super-intermediate component M1 to arrive at subtraction unit 1330 simultaneously. Subtraction unit 1330 subtracts the super-intermediate component M1 from the delayed intermediate component 1208 in the time domain to generate the remaining intermediate component M2. In this example, the spectral energy of the super-intermediate component M1 is removed from the spectral energy of the intermediate component 1208 using time-domain processing.
[0139] Forward FFT unit 1332 applies a forward FFT to side component 1210 to transform side component 1210 to the frequency domain. The transformed side component 1210 in the frequency domain includes a magnitude and a phase. Bandpass unit 1334 applies a bandpass filter to the frequency domain side component 1210. The bandpass filter specifies the frequency in the super-side component S1. In other embodiments, instead of a bandpass filter, quadrature component generator module 1323 applies various other filters to the frequency domain side component 1210. In the frequency domain, subtraction unit 1336 subtracts the mid-component 1208 from the filtered side component 1210 to generate the super-side component S1. In other embodiments, instead of post-processing applied to the super-side component S1 performed by a quadrature component processor (e.g., quadrature component processor module 1214), quadrature component generator module 1323 applies various audio enhancements to the frequency domain super-side component S1. The super-side processor 1337 performs processing on the super-side component S1 in the frequency domain before converting it to the time domain. This processing may include sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or with additional processing that may be performed by the quadrature component processor module 1214, the super-side processor 1337 processes the super-side component S1. The inverse FFT unit 1340 applies an inverse FFT to the super-side component S1 in the frequency domain to generate the super-side component S1 in the time domain. The super-side component S1 in the frequency domain includes a magnitude of S1 and the phase of the side component 1210. The inverse FFT unit 1326 converts the side component 1210 to the time domain. The delay unit 1342 delays the side component 1210 so that the side component 1210 and the super-side component S1 arrive at the subtraction unit 1344 simultaneously. Subtraction unit 1344 then subtracts the time-domain super-side component S1 from the time-domain time-delay side component 1210 to generate the remaining side component S2. In this example, the spectral energy of the super-side component S1 is removed using the spectral energy of the time-domain processing side component 1210.
[0140] In some embodiments, if the execution of the super-center processor 1325 and the super-side processor 1337 is performed by the orthogonal component processor module 1214, these components may be omitted.
[0141] In Figure 13C, the quadrature component generator module 1343 is similar to the quadrature component generator module 1323 because it receives the center component 1208 and the side component 1210 and generates the super-center component M1, the remaining center component M2, the super-side component S1 and the remaining side component S2. The only difference is that the quadrature component generator module 1343 generates each of the components M1, M2, S1 and S2 in the frequency domain and then converts these components to the time domain. The orthogonal component generator module 1343 includes a forward FFT unit 1347, a bandpass unit 1349, a subtraction unit 1351, a superprocessor 1352, a subtraction unit 1353, a residual processor 1354, an inverse FFT unit 1355, an inverse FFT unit 1357, a forward FFT unit 1361, a bandpass unit 1363, a subtraction unit 1365, a superprocessor 1366, a subtraction unit 1367, a residual processor 1368, an inverse FFT unit 1369, and an inverse FFT unit 1371.
[0142] Forward FFT unit 1347 applies a forward FFT to the mid-component 1208 to convert the mid-component 1208 to the frequency domain. The converted mid-component 1208 in the frequency domain includes a magnitude and a phase. Forward FFT unit 1361 applies a forward FFT to the side-component 1210 to convert the side-component 1210 to the frequency domain. The converted side-component 1210 in the frequency domain includes a magnitude and a phase. Bandpass unit 1349 applies a bandpass filter to the frequency-domain mid-component 1208, the bandpass filter specifying the frequency of the super-mid-component M1. In some embodiments, in addition to and / or instead of a bandpass filter, quadrature component generator module 1343 applies various other filters to the frequency-domain mid-component 1208. Subtraction unit 1351 subtracts the frequency-domain side-component 1210 from the frequency-domain mid-component 1208 to generate the super-mid-component M1 in the frequency domain. The super-mid component processor 1352 performs processing on the super-mid component M1 in the frequency domain before converting it to the time domain. In some embodiments, the super-mid component processor 1352 performs sub-band spatial processing and / or crosstalk compensation processing. In some embodiments, instead of and / or with additional processing that can be performed by the quadrature component processor module 1214, the super-mid component processor 1352 processes the super-mid component M1. The inverse FFT unit 1357 applies an inverse FFT to the super-mid component M1 to convert it back to the time domain. The super-mid component M1 in the frequency domain includes a magnitude of M1 and the phase of the mid component 1208. The inverse FFT unit 1357 converts the mid component 1208 to the time domain. The subtraction unit 1353 subtracts the super-mid component M1 from the mid component 1208 in the frequency domain to generate the remaining mid component M2. The remaining mid component processor 1354 performs processing on the remaining mid component M2 in the frequency domain before converting it to the time domain. In some embodiments, the residual processor 1354 performs sub-band spatial processing and / or crosstalk compensation processing on the residual intermediate component M2. In some embodiments, instead of and / or with additional processing that can be performed by the quadrature component processor module 1214, the residual processor 1354 processes the residual intermediate component M2. The inverse FFT unit 1355 applies an inverse FFT to transform the residual intermediate component M2 to the time domain. The residual intermediate component M2 in the frequency domain includes a magnitude of M2 and the phase of the intermediate component 1208, and the inverse FFT unit 1355 transforms the intermediate component 1208 to the time domain.
[0143] Bandpass unit 1363 applies a bandpass filter to the frequency-domain side component 1210. The bandpass filter specifies the frequency in the super-side component S1. In other embodiments, instead of a bandpass filter, quadrature component generator module 1343 applies various other filters to the frequency-domain side component 1210. In the frequency domain, subtraction unit 1365 subtracts the median component 1208 from the filtered side component 1210 to generate the super-side component S1. Super-side processor 1366 performs processing on the super-side component S1 in the frequency domain before converting it to the time domain. In some embodiments, super-side processor 1366 performs sub-band spatial processing and / or crosstalk compensation processing on the super-side component S1. In some embodiments, instead of processing that can be performed by quadrature component processor module 1214, super-side processor 1366 performs processing on the super-side component S1. Inverse FFT unit 1371 applies an inverse FFT to convert the super-side component S1 back to the time domain. The super-side component S1 in the frequency domain includes a magnitude of S1 and the phase of the side component 1210. The inverse FFT unit 1371 transforms the side component 1210 to the time domain. The subtraction unit 1367 subtracts the super-side component S1 from the side component 1210 in the frequency domain to generate the remaining side component S2. The remaining side processor 1368 performs processing on the remaining side component S2 in the frequency domain before transforming it to the time domain. In some embodiments, the remaining side processor 1368 performs sub-band spatial processing and / or crosstalk compensation processing on the remaining side component S2. In some embodiments, instead of and / or with additional processing that can be performed by the quadrature component processor module 1214, the remaining side processor 1368 performs processing on the remaining side component S2. The inverse FFT unit 1369 applies an inverse FFT to the remaining side component S2 to transform it to the time domain. The remaining side component S2 in the frequency domain includes a value of S2 and the phase of side component 1210. The inverse FFT unit 1369 converts side component 1210 to the time domain.
[0144] In some embodiments, if the execution of the processing by the super-mid processor 1352, the super-side processor 1366, the residual mid processor 1354, or the residual side processor 1368 is performed by the quadrature component processor module 1214, then these components may be omitted. Example Quadrature Component Processor
[0145] Figure 14A is a block diagram of an orthogonal component processor module 1417 according to one or more embodiments. The orthogonal component processor module 1417 is an example of an orthogonal component processor module 1412. Some embodiments of module 1417 have components different from those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.
[0146] The quadrature component processor module 1417 includes a component processor module 1420, a PSM module 102, an adder unit 1422, an M / S to L / R converter module 1424, an adder unit 1426 and an adder 1428.
[0147] Except for using the super-center component M1, super-side component S1, residual center component M2, and / or residual side component S2 instead of the center and side components, the component processor module 1420 performs the same processing as the component processor module 106. For example, the component processor module 1420 performs sub-band spatial processing and / or crosstalk compensation processing on at least one of the super-center component M1, residual center component M2, super-side component S1, and residual side component S2. Due to the sub-band spatial processing and / or crosstalk compensation of the component processor module 1420, the quadrature component processor module 1417 outputs at least one of the processed M1, processed M2, processed S1, and processed S2. In some embodiments, one or more of the components M1, M2, S1, or S2 may bypass the component processor module 1420.
[0148] In some embodiments, the quadrature component processor module 1417 performs subband spatial processing and / or crosstalk compensation processing on at least one of the super-center component M1, the remaining center component M2, the super-side component S1, and the remaining side component S2 in the frequency domain. The quadrature component generator module 410 may provide the frequency domain components M1, M2, S1, or S2 to the quadrature component processor module 1417 without performing an inverse FFT. After generating the processed M1, processed M2, and processed side component 1442, the quadrature component processor module 1417 may perform an inverse FFT to convert these components back to the time domain. In some embodiments, the quadrature component processor module 1417 performs an inverse FFT on the processed M1, processed M2, processed S1, and processed S1, and generates the processed side component 1446 in the time domain.
[0149] Indicative components of the quadrature component processor module 1417 are shown in Figures 15 and 16. In some embodiments, the quadrature component processor module 1417 performs both sub-band spatial processing and crosstalk compensation processing. The processing performed by the quadrature component processor module 1417 is not limited to sub-band spatial processing or crosstalk compensation processing. Any type of spatial processing using center / side space can be performed by the quadrature component processor module 1417, such as replacing the center component with a super-center component or replacing the side component with a super-side component. Some other types of processing may include gain application, amplitude- or delay-based translation, stereo processing, reverberation, dynamic range processing (such as compression and limiting), and other linear or nonlinear audio processing techniques and effects ranging from choral or flange effects to machine learning-based methods to vocal or instrumental style shifting, conversion, or resynthesis.
[0150] PSM module 102 receives processed M1 and applies PSM processing to spatially shift processed M1 to generate a left channel 1432 and a right channel 1434. Although PSM module 102 is shown as being applied to super-intermediate component M1, the PSM module may be applied to one or more of components M1, M2, S1, or S2. In some embodiments, a component processed by PSM module 102 bypasses the processing of component processor module 1420. For example, PSM module 102 may process super-intermediate component M1 instead of processed M1.
[0151] Adding unit 1422 adds processed S1 and processed S2 to generate a processed side component 1442. M / S to L / R converter module 1424 uses processed M2 and processed side component 1442 to generate a processed left component 1444 and a processed right component 1446. In some embodiments, processed left component 1444 is generated based on the sum of processed M2 and processed side component 1442, and processed right component 1446 is generated based on the difference between processed M2 and processed side component 1442. Other M / S to L / R type transformations can be used to generate processed left component 1444 and processed right component 1446.
[0152] Adding unit 1426 adds the left channel 1432 from PSM module 102 to the processed left component 1444 to generate left channel 1452. Adding unit 1428 adds the right channel 1434 from PSM module 102 to the processed right component 1446 to generate right channel 1454. More generally, one or more left channels from PSM module 102 can be added to a left component from M / S to L / R converter module 1424 (e.g., generated using over / residual components not processed by PSM module 102) to generate left channel 1452, and one or more right channels from PSM module 102 can be added to a right component from M / S to L / R converter module 1424 (e.g., generated using over / residual components not processed by PSM module 102) to generate right channel 1454.
[0153] Therefore, the quadrature component processor module 1417 applies PSM processing to one of the super-center components M1 of an audio signal, as isolated by the L / R to M / S converter module 1206 and the quadrature component generator module 1212. The PSM-enhanced stereo signal (containing the left channel 1432 and the right channel 1434) can then be added to the remaining left / right signals (e.g., the processed left component 1444 and the processed right component 1446, without super-center component generation). Alternatively or in place of this example, other methods of isolating the components of the input signal used for PSM processing can be used, including audio source separation based on machine learning.
[0154] In some embodiments, the quadrature component processor module 1417 applies PSM processing to a mid-component M of an audio signal instead of the super-mid-component M1. FIG14B illustrates a block diagram of a quadrature component processor module 1419 according to one or more embodiments. In some embodiments, the quadrature component processor module 1419 of FIG14B may be implemented as a part of an audio processing system similar to system 1200 illustrated in FIG12 but without the quadrature component generator module 1212, such that the quadrature component processor module receives mid-component and side-component signals (e.g., mid-component 1208 and side-component 1210) instead of super-mid-component, super-side-component, residual mid-component, and residual side-component. In some embodiments, the quadrature component processor module 1410 includes a component processor module similar to component processor module 106 to generate processed mid-component and processed side-component from the received mid-component and side-component (not shown in the figure). PSM module 102 receives the intermediate component M (or processed intermediate), and applies PSM processing to spatially shift the received intermediate signal to generate a PSM-processed left channel 1432 and a PSM-processed right channel 1434. These are combined with the side component S (or processed side) by M / S to L / R converter module 1424 to generate left channel 1452 and right channel 1454. For example, as shown in FIG14B, M / S to L / R converter module 1424 uses an adder unit 1460 to generate left channel 1452 as a sum of PSM-processed left channel 1432 and side component S, and uses a subtractor unit 1462 to generate right channel 1454 as a difference between PSM-processed right channel 1434 and side component S. In other words, the M / S to L / R converter module 1424 is used to mix a signal from one side (based on the left and right sides being located in a subspace defined by the left component, which is the inverse of the right component) into a PSM-processed stereo signal in the left and right spaces by combining the combined signal with the left channel and the inverse of the combined signal with the right channel. Example of a sub-band spatial processor.
[0155] Figure 15 is a block diagram of a sub-band spatial processor module 1510 according to one or more embodiments. The sub-band spatial processor module 1510 is an example of a component of a component processor module 106 or 1520. The sub-band spatial processor module 1510 includes a middle EQ filter 1504(1), a middle EQ filter 1504(2), a middle EQ filter 1504(3), a middle EQ filter 1504(4), a side EQ filter 1506(1), a side EQ filter 1506(2), a side EQ filter 1506(3), and a side EQ filter 1506(4). Some embodiments of the sub-band spatial processor module 1510 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0156] Subband spatial processor module 1510 receives a non-spatial component Ym and a spatial component Ys, and adjusts the gain of one or more of these components in the subband to provide spatial enhancement. When subband spatial processor module 1510 is part of component processor module 1420, the non-spatial component Ym may be an ultra-intermediate component M1 or a residual intermediate component M2. The spatial component Ys may be an ultra-side component S1 or a residual side component S2. When subband spatial processor module 1510 is part of component processor module 106, the non-spatial component Ym may be an intermediate component 126 and the spatial component Ys may be a side component 128.
[0157] The sub-band spatial processor module 1510 receives the non-spatial component Ym and applies intermediate EQ filters 1504(1) to 1504(4) to different sub-bands of Ym to generate an enhanced non-spatial component Em. The sub-band spatial processor module 1510 also receives the spatial component Ys and applies intermediate EQ filters 1506(1) to 1506(4) to different sub-bands of Ys to generate an enhanced spatial component Es. The sub-band filters may include various combinations of peak filters, notch filters, low-pass filters, high-pass filters, low-profile filters, high-profile filters, band-pass filters, band-stop filters and / or all-pass filters. The sub-band filters may also apply gain to their respective sub-bands. More specifically, the sub-band spatial processor module 1510 includes one sub-band filter for each of the n frequency sub-bands of the non-spatial component Ym and one sub-band filter for each of the n sub-bands of the spatial component Ys. For example, for n=4 sub-bands, the sub-band spatial processor module 1510 includes a series of sub-band filters for the non-spatial component Ym, including an equalization (EQ) filter 1504(1) for one sub-band (1), an EQ filter 1504(2) for one sub-band (2), an EQ filter 1504(3) for one sub-band (3), and an EQ filter 1504(4) for one sub-band (4). Each EQ filter 1504 applies a filter to a frequency sub-band portion of the non-spatial component Ym to generate an enhanced non-spatial component Em.
[0158] The sub-band spatial processor module 1510 further includes a series of sub-band filters for the frequency sub-band of the spatial component Ys, including a side equalization (EQ) filter 1506(1) for sub-band (1), a side EQ filter 1506(2) for sub-band (2), a side EQ filter 1506(3) for sub-band (3), and a side EQ filter 1506(4) for sub-band (4). Each side EQ filter 1506 applies a filter to a frequency sub-band portion of the spatial component Ys to generate an enhanced spatial component Es.
[0159] Each of the n frequency sub-bands of the non-spatial component Ym and the spatial component Ys can correspond to a frequency range. For example, frequency sub-band (1) can correspond to 0 Hz to 300 Hz, frequency sub-band (2) can correspond to 300 Hz to 510 Hz, frequency sub-band (3) can correspond to 510 Hz to 2700 Hz, and frequency sub-band (4) can correspond to 2700 Hz to the Nyquist frequency. In some embodiments, each of the n frequency sub-bands is a group of combined critical bands. The critical bands can be determined using a corpus of audio samples from various music genres. The long-term average energy ratio of the middle and side components on the critical bands is determined from the samples using 24 Bark measurements. Adjacent bands with similar long-term average ratios are then grouped together to form a critical band group. The range and number of frequency sub-bands can be adjusted.
[0160] In some embodiments, the sub-band spatial processor module 1510 processes the remaining middle component M2 into a non-spatial component Ym and uses one of the side component, the super-side component S1 or the remaining side component S2 as the spatial component Ys.
[0161] In some embodiments, the sub-band spatial processor module 1510 processes one or more of the super-intermediate component M1, the super-side component S1, the residual intermediate component M2, and the residual side component S2. The filters applied to the sub-bands of these components may be different. The super-intermediate component M1 and the residual intermediate component M2 may each be processed as discussed for the non-spatial component Ym. The super-side component S1 and the residual side component S2 may each be processed as discussed for the spatial component Ys. Example crosstalk compensation processor
[0162] FIG16 is a block diagram of a crosstalk compensation processor module 1610 according to one or more embodiments. The crosstalk compensation processor module 1610 is an example of a component of a component processor module 106 or 1420. Some embodiments of the crosstalk compensation processor module 1610 have components different from those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.
[0163] The crosstalk compensation processor module 1610 includes a mid-component processor 1620 and a side-component processor 1630. The crosstalk compensation processor module 1610 receives a non-spatial component Ym and a spatial component Ys and applies a filter to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or previous) crosstalk processing. When the crosstalk compensation processor module 1610 is part of the component processor module 1420, the non-spatial component Ym may be a super-mid-component M1 or a residual mid-component M2. The spatial component Ys may be a super-side-component S1 or a residual-side-component S2. When the crosstalk compensation processor module 1610 is part of the component processor module 106, the non-spatial component Ym may be a mid-component 126 and the spatial component Ys may be a side-component 128.
[0164] Crosstalk compensation processor module 1610 receives the non-spatial component Ym, and intermediate component processor 1620 applies a set of filters to generate an enhanced non-spatial crosstalk compensation component Zm. Crosstalk compensation processor module 1610 also receives the spatial sub-band component Ys and applies one of the sets of filters from one of the intermediate component processors 1630 to generate an enhanced spatial sub-band component Es. Intermediate component processor 1620 includes a plurality of filters 1640, such as m intermediate filters 1640(a), 1640(b) to 1640(m). Here, each of the m intermediate filters 1640 processes one of the m frequency bands of the non-spatial component Xm. Intermediate component processor 1620 thus generates an intermediate crosstalk compensation channel Zm by processing the non-spatial component Xm. In some embodiments, the intermediate filter 1640 is configured using a frequency response diagram of the non-spatial Xm and crosstalk processing through analog. Furthermore, by analyzing the frequency response plot, any spectral defects, such as peaks or troughs in the frequency response plot exceeding a predetermined threshold (e.g., 10 dB), can be estimated as artifacts in crosstalk processing. These artifacts primarily originate from the sum of delayed and potentially out-of-phase signals on the opposite side and their corresponding signals on the same side during crosstalk processing, thereby effectively incorporating a comb-filter-like frequency response into the final presentation. The intermediate crosstalk compensation channel Zm can be generated by the intermediate component processor 1620 to compensate for estimated peaks or troughs, where each of the m frequency bands corresponds to a peak or trough. Specifically, based on the specific delay, filter frequency, and gain applied in crosstalk processing, peaks and troughs are shifted upwards and downwards in the frequency response to cause variable amplification and / or attenuation of energy in specific regions of the spectrum. The intermediate filters 1640 can be configured to adjust one or more of the peaks and troughs.
[0165] The side component processor 1630 includes a plurality of filters 1650, such as m side filters 1650(a), 1650(b) to 1650(m). The side component processor 1630 generates a side crosstalk compensation channel Zs by processing the spatial component Xs. In some embodiments, a frequency response diagram of the spatial Xs with crosstalk processing can be obtained by simulation. By analyzing the frequency response diagram, any spectral defects, such as peaks or troughs in the frequency response diagram exceeding a predetermined threshold (e.g., 10 dB), can be estimated as peaks or troughs where crosstalk artifacts occur. The side crosstalk compensation channel Zs can be generated by the side component processor 1630 to compensate for the estimated peaks or troughs. Specifically, based on specific delays, filtering frequencies, and gains applied in the crosstalk processing, peaks and troughs are shifted up and down in the frequency response to cause variable amplification and / or attenuation of energy in specific regions of the spectrum. Each of the side filters 1650 can be configured to adjust one or more of the peaks and troughs. In some embodiments, the mid-component processor 1620 and the side-component processor 1630 may include different numbers of filters.
[0166] In some embodiments, the middle filter 1640 and the side filter 1650 may comprise a double second-order filter having a transfer function defined by equation (7). One way to implement such a filter is in the direct form I topology defined by equation (22): (22)
[0167] where X is the input vector and Y is the output. Other topologies can be used, depending on their maximum word length and saturation behavior. Next, a bi-second order can be used to implement a second-order filter with real-valued inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and then transformed into discrete time via a bilinear transform. Furthermore, frequency warping can be used to compensate for the resulting shift in center frequency and bandwidth.
[0168] For example, a peak filter may have an S-surface transfer function defined by equation (23): (23)
[0169] where s is a complex variable, A is the amplitude of the peak value, and Q is the filter "quality". The digital filter coefficients are defined by the following equation (24): (24) where is the center frequency of the filter (in radians) and. In addition, the filter quality Q can be defined by equation (25): (25)
[0170] where Δf is the bandwidth and fc is the center frequency. The intermediate filter 1640 is shown in series, and the side filter 1650 is shown in series. In some embodiments, the intermediate filter 1640 is applied in parallel to the intermediate component Xm, and the side filters are applied in parallel to the side component Xs.
[0171] In some embodiments, the crosstalk compensation processor module 1610 processes each of the super-center component M1, the super-side component S1, the remaining center component M2, and the remaining side component S2. The filters applied to each of these components may be different. Example Crosstalk Processor
[0172] Figure 17 is a block diagram of a crosstalk analog processor module 1700 according to one or more embodiments. The crosstalk analog processor module 1700 is an example of a crosstalk processor module 110 or a crosstalk processor module 1224. Some embodiments of the crosstalk analog processor module 1700 have components different from those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.
[0173] The crosstalk analog processor module 1700 generates the opposite sound component for output to stereo headphones, thereby providing a speaker-like listening experience on the headphones. The left input channel XL can be a processed left component 134 / 1220 and the right input channel XR can be a processed right component 136 / 1222.
[0174] The crosstalk analog processor module 1700 includes a left headshadow low-pass filter 1702, a left headshadow high-pass filter 1724, a left crosstalk delay 1704, and a left headshadow gain 1710 to process the left input channel XL. The crosstalk analog processor module 1700 further includes a right headshadow low-pass filter 1706, a right headshadow high-pass filter 1726, a right crosstalk delay 1708, and a right headshadow gain 1712 to process the right input channel XR. The crosstalk analog processor module 1500 further includes an adder unit 1714 and an adder unit 1716.
[0175] The left headshadow low-pass filter 1702 and the left headshadow high-pass filter 1724 apply modulation to the left input channel XL, modeling the frequency response of the signal after passing through the listener's head. The output of the left headshadow high-pass filter 1724 is provided to the left crosstalk delay 1704, which applies a time delay. The time delay represents the transaural distance by which a pair of side sound components are lateralized relative to a common side sound component. The left headshadow gain 1710 applies a gain to the output of the left crosstalk delay 1704 to generate the right-left analog channel WL.
[0176] Similarly, for the right input channel XR, the right headshot low-pass filter 1706 and the right headshot high-pass filter 1726 apply modulation to the right input channel XR, which models the frequency response of the listener's head. The output of the right headshot high-pass filter 1726 is provided to the right crosstalk delay 1708, which applies a delay. The right headshot gain 1712 applies a gain to the output of the right crosstalk delay 1708 to generate the right crosstalk analog channel WR.
[0177] The head shadow low-pass filter, head shadow high-pass filter, crosstalk delay and head shadow gain can be applied to the left and right channels in different orders.
[0178] Adder unit 1714 adds the right audio analog channel WR to the left input channel XL to generate a left output channel OL. Adder unit 1716 adds the left audio analog channel WL to the right input channel XR to generate a left output channel OR.
[0179] Figure 18 is a block diagram of a crosstalk cancellation processor module 1800 according to one or more embodiments. The crosstalk cancellation processor module 1800 is an example of a crosstalk processor module 110 or a crosstalk processor module 1224. Some embodiments of the cancellation processor module 1800 have components different from those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.
[0180] The crosstalk cancellation processor module 1800 receives a left input channel XL and a right input channel XR, and performs crosstalk cancellation on channels XL and XR to generate a left output channel OL and a right output channel OR. The left input channel XL can be a processed left component 134 / 1220 and the right input channel XR can be a processed right component 136 / 1222.
[0181] The crosstalk cancellation processor module 1800 includes an in-band to out-of-band crossover 1810, inverters 1820 and 1822, opposite-side estimators 1830 and 1840, combiners 1850 and 1852, and an in-band to out-of-band combiner 1860. These components operate together to split the input channels TL and TR into in-band and out-of-band components, and perform crosstalk cancellation on the in-band components to generate output channels OL and OR.
[0182] By dividing the input audio signal T into different frequency bands and performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed on a specific frequency band while avoiding degradation in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal T into different frequency bands, the audio signal after crosstalk cancellation may exhibit significant attenuation or amplification of non-spatial and spatial components in low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12000 Hz), or both. By selectively performing crosstalk cancellation on in-band components (e.g., between 250 Hz and 14000 Hz) where the majority of influential spatial cues reside, a balanced total energy can be preserved across the spectrum of the mixture, especially in the non-spatial components.
[0183] The in-band-out-of-band crossover 1810 separates the input channels TL and TR into in-band channels TL,In and TR,In, and out-of-band channels TL,Out and TR,Out, respectively. Specifically, the in-band-out-of-band crossover 1810 splits the left enhancement compensation channel TL into a left in-band channel TL,In and a left out-of-band channel TL,Out. Similarly, the in-band-out-of-band crossover 1810 splits the right enhancement compensation channel TR into a right in-band channel TR,In and a right out-of-band channel TR,Out. Each in-band channel may cover a portion of its respective input channel corresponding to a frequency range including, for example, 250 Hz to 14 kHz. For example, the frequency range can be adjusted according to speaker parameters.
[0184] Inverter 1820 and opposite-side estimator 1830 operate together to generate a left opposite-side cancellation component SL to compensate for an opposite-side sound component attributable to the left in-band channel TL,In. Similarly, inverter 1822 and opposite-side estimator 1840 operate together to generate a right opposite-side cancellation component SR to compensate for an opposite-side sound component attributable to the right in-band channel TR,In.
[0185] In one method, inverter 1820 receives an in-band channel TL,In and inverts the polarity of the received in-band channel TL,In to generate an inverted in-band channel TL,In'. Counter-side estimator 1830 receives the inverted in-band channel TL,In' and extracts a portion of the inverted in-band channel TL,In' corresponding to a pair of counter-side sound components through filtering. Because filtering is performed on the inverted in-band channel TL,In', the portion extracted by counter-side estimator 1830 becomes an inverted portion of the in-band channel TL,In attributable to the counter-side sound components. Therefore, the portion extracted by counter-side estimator 1830 becomes a left counter-side cancellation component SL, which can be added to a corresponding in-band channel TR,In to reduce the counter-side sound components attributable to the in-band channel TL,In. In some embodiments, inverter 1820 and counter-side estimator 1830 are implemented in a different sequence.
[0186] The inverter 1822 and the opposite-side estimator 1840 perform similar operations relative to the in-band channel TR,In to generate the right opposite-side canceled component SR. Therefore, for the sake of brevity, their detailed description is omitted here.
[0187] In one exemplary embodiment, the contralateral estimator 1830 includes a filter 1832, an amplifier 1834, and a delay unit 1836. The filter 1832 receives an inverted input channel TL,In' and extracts a portion of the inverted in-band channel TL,In' corresponding to a pair of contralateral sound components through a filtering function. An exemplary filter embodiment has a notch or overhead filter with a center frequency selected between 5000 Hz and 10000 Hz and a Q selected between 0.5 and 1.0. The gain (GdB) in decibels can be derived from equation (26): GdB = -3.0 - log1.333(D) (26) where D is a delay amount of the delay unit 1836 of the sample, for example, at a sampling rate of 48 kHz. An alternative embodiment has a low-pass filter with a corner frequency selected between 5000 Hz and 10000 Hz and a Q selected between 0.5 and 1.0. Furthermore, amplifier 1834 amplifies the extracted portion by a corresponding gain coefficient GL,In, and delay unit 1836 delays the amplified output from amplifier 1834 according to a delay function D to generate the left-side cancellation component SL. The opposite-side estimator 1840 includes a filter 1842, an amplifier 1844, and a delay unit 1846, which perform similar operations on the inverted in-band channel TR,In' to generate the right-side cancellation component SR. In one example, opposite-side estimators 1830 and 1840 generate the left and right opposite-side cancellation components SL and SR according to the following equations: SL=D[GL,In*F[TL,In']] (27) SR=D[GR,In*F[TR,In']] (28) where F[] is a filter function and D[] is a delay function.
[0188] The crosstalk cancellation configuration can be determined by speaker parameters. In one example, the filter center frequency, delay, amplifier gain, and filter gain can be determined based on an angle formed between two speakers relative to a listener. In some embodiments, values between speaker angles are used to interpolate other values.
[0189] Combiner 1850 combines the right-side cancelling component SR to the left in-band channel TL,In to generate a left in-band crosstalk channel UL, and combineer 1852 combines the left-side cancelling component SL to the right in-band channel TR,In to generate a right in-band crosstalk channel UR. In-band-out-band combiner 1860 combines the left in-band crosstalk channel UL with the out-of-band channel TL,Out to generate a left output channel OL, and combines the right in-band crosstalk channel UR with the out-of-band channel TR,Out to generate a right output channel OR.
[0190] Therefore, the left output channel OL contains an inverted right opposite-side cancellation component SR corresponding to a portion of the in-band channel TR,In that corresponds to the opposite-side sound, and the right output channel OR contains an inverted left opposite-side cancellation component SL corresponding to a portion of the in-band channel TL,In that corresponds to the opposite-side sound. In this configuration, a wavefront of an opposite-side sound component output by a left speaker via the left output channel OL can be canceled by a right speaker based on a wavefront of a same-side sound component output by the right output channel OR that reaches the right ear. Similarly, a wavefront of an opposite-side sound component output by a right speaker via the right output channel OR can be canceled by a left speaker based on a wavefront of a same-side sound component output by the left output channel OL that reaches the left ear. Therefore, opposite-side sound components can be reduced to enhance spatial detectability. Example PSM procedure flow
[0191] Figure 19 is a flowchart of a procedure 1900 for PSM processing according to one or more embodiments. Procedure 1900 may include fewer or more steps, and the steps may be performed in different orders. In some embodiments, PSM processing may be performed using a Hilbert transform perceptual sound field modification (HPSM) module.
[0192] An audio processing system (e.g., PSM module 102 of audio processing system 100 or 1200) splits an input channel 1905 into a low-frequency component and a high-frequency component. The crossover frequency defining the boundary between the low-frequency component and the high-frequency component can be adjusted to, for example, ensure that the frequencies of interest to the PSM processing are included in the high-frequency component. In some embodiments, the audio processing system applies a gain to the low-frequency component and / or the high-frequency component.
[0193] An input channel may be a specific portion of an audio signal extracted for PSM processing. In some embodiments, an input channel is a mid-component or a side component of an audio signal (e.g., stereo or multi-channel). In some embodiments, an input channel is a super-mid-component, a super-side component, a residual mid-component, or a residual side component of an audio signal. In some embodiments, an input channel is associated with a sound source such as a speech or musical instrument, which will be combined with other sounds to form an audio mix.
[0194] The audio processing system applies a first Hilbert transform to the high-frequency components to generate a first left branch component and a first right branch component, wherein the first left branch component is out of phase with the first right branch component by 90 degrees.
[0195] The audio processing system applies a second Hilbert transform to the first right branch component to generate a second left branch component and a second right branch component, wherein the first left branch component is out of phase with the first right branch component by 90 degrees.
[0196] In some embodiments, the audio processing system applies a delay and / or gain to a first left-branch component. The audio processing system may apply a delay and / or gain to a second right-branch component. These gains and delays can be used to manipulate the perceptual results of PSM processing.
[0197] The audio processing system combines the first left branch component and the low-frequency component 1920 to generate a left channel. The audio processing system combines the second right branch component and the low-frequency component 1925 to generate a right channel. The left channel can be provided to a left speaker and the right channel can be provided to a right speaker.
[0198] Figure 20 is a flowchart of another procedure 2000 for PSM processing using a first-order non-orthogonal rotatable decorrelation (FNORD) filter network, according to some embodiments. The procedure shown in Figure 20 can be executed by a component of an audio system (e.g., system 100, 202, or 1200). In other embodiments, other entities can perform some or all of the steps in Figure 20. Embodiments may include different and / or additional steps or perform steps in a different order.
[0199] The audio system determination 2005 defines a target amplitude response encoded into a mono audio signal to generate one or more spatial cues in a plurality of resulting channels, wherein one or more spatial cues are associated with one or more frequency-dependent amplitude cues encoded into the resulting channels in the middle / side space without altering the overall coloring of the resulting channels. One or more spatial cues may include at least one elevation cue associated with a target elevation angle. Each elevation cue may correspond to one or more frequency-dependent amplitude cues encoded into the middle / side space of the audio signal, such as a target magnitude function corresponding to a narrow region of infinite attenuation at one or more specific frequencies. On the other hand, since the left / right cues used for elevation are typically color-symmetrical, the left / right signals may be constrained to be colorless. In some embodiments, a spatial cue may be based on a sampled HRTF.
[0200] In some embodiments, the target amplitude response may further define one or more parametric spatial cues, which may include a target broadband attenuation, a target subband attenuation, a critical point, a filter characteristic, and / or a sound field location in which the cues are embedded. The critical point may be an inflection point at 3 dB. The filter characteristic may include one of a high-pass filter characteristic, a low-pass characteristic, a band-pass characteristic, or a band-block characteristic. The sound field location may include a center or side channel, or, in cases where the number of output channels is greater than two, other subspaces within the output space, such as subspaces determined by pairwise and / or hierarchical sums and / or differences. One or more spatial cues may be determined based on the characteristics of the presentation device (e.g., the frequency response of the speaker, the speaker location), the intended content of the audio data, the listener's perceptual ability in the background, or the minimum quality expected by the audio presentation system involved. For example, if the speaker cannot adequately reproduce frequencies below 200 Hz, spatial cues in this range should be avoided. Similarly, if the expected audio content is speech, the audio system can select a target amplitude response that only affects the frequencies most sensitive to the ear (starting within the expected bandwidth of speech). If the listener will receive audible cues from other sources in the background (such as another speaker array in the location), the audio system can determine a target amplitude response that complements these simultaneous cues.
[0201] The audio system uses a target amplitude response determination 2010 for a transfer function of a single-input multiple-output all-pass filter. The transfer function defines the relative rotation of the phase angle of the output channel. The transfer function describes the effect of a filter network on its inputs, with respect to each output, based on a phase angle rotation that varies according to frequency.
[0202] The audio system determines the coefficients of the 2015 all-pass filter based on the transfer function. These coefficients are selected and applied to the incoming audio stream in a manner best suited to the type of cue and / or constraint. Some instances of the coefficient set are defined in equations (12), (13), (17), and (19). In some embodiments, determining the coefficients of the all-pass filter based on the transfer function involves using an inverse discrete Fourier transform (idFT). In this case, the coefficient set can be determined as defined by equation (19). In some embodiments, determining the coefficients of the all-pass filter based on the transfer function involves using a phase vocoder. In this case, the coefficient set can be determined as defined by equation (19), except that these are applied in the frequency domain before resynthesizing the time-domain data. In some embodiments, the coefficients include at least a rotation control parameter and first-order coefficients, which are determined based on the received critical point parameter, filter characteristic parameters, and sound field position parameters.
[0203] The audio system processes a 2020 mono channel using the coefficients of an all-pass filter to generate a plurality of channels. For example, in some embodiments, the all-pass filter module receives a mono audio channel and performs a wideband phase rotation on the mono audio channel based on a rotation control parameter to generate a plurality of wideband rotated component channels (e.g., left and right wideband rotated components), and performs a narrowband phase rotation on at least one of the plurality of wideband rotated component channels based on first-order coefficients to determine a narrowband rotated component channel, which, together with one or more of the remaining wideband rotated component channels, forms a plurality of channels output by the audio system.
[0204] In some embodiments, if the system operates in the time domain, an IIR implementation such as Equation (8) is used, with coefficients scaled to accommodate appropriate feedback and feedforward delays. If an FIR implementation such as Equation (19) is used, only the feedforward delay can be used. If the coefficients are determined and applied in the frequency domain, they can be applied as a complex multiplication to the spectral data before resynthesis. The audio system can provide multiple output channels to a presentation device, such as a user device connected to the audio system via a network.
[0205] The above-described exemplary PSM processing flow each utilizes an all-pass filter network to encode spatial cues by perceptually placing the mono content into a specific location in the sound field (e.g., a location associated with a target elevation angle). Because the all-pass filter networks described herein are colorless, these filters allow the user to decouple the spatial location of the audio signal from its overall coloring. Orthogonal component spatial processing
[0206] Figure 21 is a flowchart of a procedure 2100 for spatial processing using at least one of a super-center, residual center, super-side, or residual side component, according to one or more embodiments. Spatial processing may include gain application, amplitude- or delay-based translation, stereo processing, reverberation, dynamic range processing (such as compression and limiting), linear or nonlinear audio processing techniques and effects, chorus effects, flange effects, machine learning-based methods to vocal or instrumental style shifting, conversion or resynthesis, and other techniques. The executable procedure provides spatially enhanced audio to a user's device. The procedure may contain fewer or more steps and may be performed in different orders.
[0207] An audio processing system (e.g., audio processing system 1200) receives 2110 an input audio signal (e.g., left channel 1202 and right input channel 1204). In some embodiments, the input audio signal may be a multichannel audio signal comprising multiple left and right channel pairs. Each left and right channel pair may be processed as discussed herein with respect to the left and right input channels.
[0208] The audio processing system generates a non-spatial mid-component (e.g., mid-component 1208) and a spatial side-component (e.g., side-component 1210) from the input audio signal. In some embodiments, an L / R to M / S converter (e.g., L / R to M / S converter module 1206) performs the conversion of the input audio signal to the mid- and side-components.
[0209] The audio processing system generates at least one of a super-mid component (e.g., super-mid component M1), a super-side component (e.g., super-side component S1), a residual mid component (e.g., residual mid component M2), and a residual side component (e.g., residual side component S2). The audio processing system can generate at least one and / or all of the above components. The super-mid component contains the spectral energy of the side component removed from the spectral energy of the mid component. The residual mid component contains the spectral energy of the super-mid component removed from the spectral energy of the mid component. The super-side component contains the spectral energy of the mid component removed from the spectral energy of the side component. The residual side component contains the spectral energy of the super-side component removed from the spectral energy of the side component. The processing for generating M1, M2, S1, or S2 can be performed in the frequency domain or the time domain.
[0210] The audio processing system filters at least one of the super-intermediate component, residual intermediate component, super-side component, and residual side component 2140 to enhance the audio signal. The filtering may include HPSM processing, wherein a series of Hilbert transforms are applied to one of the high-frequency components of the super-intermediate component, residual intermediate component, super-side component, or residual side component. In one example, the super-intermediate component receives HPSM processing, while one or more of the residual intermediate component, super-side component, or residual side component receives other types of filtering.
[0211] Filtering may include PSM processing, in which spatial cues are color-coded by means of the parametric specifications of the spatial cues (discussed in more detail above in conjunction with Figures 10A and 10B) or by human factors sampling of HRTF data (discussed above in conjunction with Equation 20). In one instance, the super-mid component receives PSM processing, while one or more of the remaining mid component, super-side component, or remaining side component do not receive filtering or other types of filtering.
[0212] Filtering may include other types of filtering, such as spatial cue processing. Spatial cue processing may include adjusting the frequency-dependent amplitude or frequency-dependent delay of one of the super-mid component, residual mid component, super-side component, or residual side component. Some examples of spatial cue processing include translation or stereo processing based on amplitude or delay.
[0213] Filtering may include dynamic range processing, such as compression or limiting. For example, when a critical level for compression is exceeded, the super-mid component, residual mid component, super-side component, or residual side component may be compressed according to a compression ratio. In another instance, when a critical level for limiting is exceeded, the super-mid component, residual mid component, super-side component, or residual side component may be limited to a maximum level.
[0214] Filtering may include machine learning-based modifications to the super-mid component, residual mid component, super-side component, or residual side component. Some instances include machine learning-based vocal or instrumental style shifts, transformations, or resynthesis.
[0215] Filtering of the super-mid component, residual mid component, super-side component, or residual side component may include gain application, reverberation, and other linear or nonlinear audio processing techniques and effects within the processing range, such as self-chorus and / or flange effects or other types of processing. In some embodiments, filtering may include filtering for sub-band spatial processing and crosstalk compensation, as discussed in more detail below with reference to Figure 22.
[0216] Filtering can be performed in the frequency domain or the time domain. In some embodiments, the intermediate and side components are transformed from the time domain to the frequency domain, super- and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are transformed to the time domain. In other embodiments, the super- and / or residual components are transformed to the time domain, and filtering is performed on these components in the time domain.
[0217] The audio processing system uses one or more of the filtered super-center / residual components to generate a left output channel (e.g., left output channel 1242) and a right output channel (e.g., right output channel 1244). For example, the M / S to L / R conversion can be performed using a center component or a side component generated from at least one of the filtered super-center component, the filtered residual center component, the filtered super-side component, or the filtered residual side component. In another example, the filtered super-center component or the filtered residual center component can be used as the center component for M / S to L / R conversion, or the filtered super-side component or the residual side component can be used as the side component for M / S to L / R conversion. Quadrature component sub-band spatial and crosstalk processing
[0218] Figure 22 is a flowchart of a procedure 2200 for subband spatial processing and crosstalk compensation processing using at least one of super-center, residual center, super-side, or residual side components, according to one or more embodiments. Crosstalk processing may include crosstalk cancellation or crosstalk simulation. Subband spatial processing can be performed to provide audio content with enhanced spatial detectability, for example by directing sound from a large area rather than a specific point in space corresponding to a speaker location to the listener's perception (e.g., sound field enhancement) and thereby creating a more immersive listening experience for the listener. Crosstalk simulation can be used for audio output to headphones to simulate a speaker experience with crosstalk. Crosstalk cancellation can be used for audio output to speakers to eliminate the effects of crosstalk interference. Crosstalk compensation compensates for spectral defects caused by crosstalk cancellation or crosstalk simulation. The procedure may include fewer or more steps and may be performed in different orders. Super and residual center / side components may be manipulated in different ways for different purposes. For example, in the case of crosstalk compensation, the target subband filter can be applied only to the super-mid component M1 (where most of the audio dialogue energy in a movie occurs) to attempt to remove spectral artifacts caused by crosstalk processing only in this component. In the case of sound field enhancement with or without crosstalk processing, the target subband gain can be applied to the remaining mid component M2 and the remaining side component S2. For example, the remaining mid component M2 can be attenuated and the remaining side component S2 can be amplified in reverse to increase the distance between these components from a gain angle (which, if done well, can increase spatial detectability) without producing a drastic overall change in the perceived loudness in the final L / R signal, while also avoiding attenuation of the super-mid component M1 (e.g., the portion of the signal that typically contains most of the sound energy).
[0219] The audio processing system receives an input audio signal 2210, which includes left and right channels. In some embodiments, the input audio signal may be a multi-channel audio signal comprising one of multiple left and right channel pairs. Each left and right channel pair may be processed as discussed herein with respect to the left and right input channels.
[0220] The audio processing system applies crosstalk processing 2220 to the received input audio signal. Crosstalk processing includes at least one of crosstalk simulation and crosstalk cancellation.
[0221] In steps 2230 to 2260, the audio processing system uses one or more of the super-center, super-side, residual center, or residual side components to perform sub-band spatial processing and crosstalk compensation for crosstalk processing. In some embodiments, crosstalk processing may be performed after the processing in steps 2230 to 2260.
[0222] The audio processing system generates a middle component and a side component from (for example, after crosstalk processing) the audio signal.
[0223] The audio processing system generates at least one of a super-intermediate component, a residual intermediate component, a super-side component, and a residual side component. The audio processing system can generate at least one and / or all of the above components.
[0224] An audio processing system applies sub-band filtering 2250 to at least one of the super-mid component, residual mid component, overside component, and residual side component to an audio signal, thereby applying sub-band spatial processing. Each sub-band may comprise a frequency range, such as that defined by a critical band set. In some embodiments, the sub-band spatial processing further includes delaying the sub-band of at least one of the super-mid component, residual mid component, overside component, and residual side component. In some embodiments, the filtering includes applying HPSM processing.
[0225] The audio processing system filters at least one of the super-mid component, residual mid component, overside component, and residual side component 2260 to compensate for spectral defects in crosstalk processing from the input audio signal. The spectral defects may include peaks or troughs in the frequency response diagram of the super-mid component, residual mid component, overside component, or residual side component exceeding a predetermined threshold (e.g., 10 dB) as an artifact occurring as part of the crosstalk processing. The spectral defects may be estimated spectral defects.
[0226] In some embodiments, the filtering of the spectral quadrature components for sub-band spatial processing in step 2250 and crosstalk compensation in step 2260 can be integrated into a single filtering operation for each spectral quadrature component selected for filtering.
[0227] In some embodiments, filtering of the super / residual middle / side components for subband spatial processing or crosstalk compensation may be performed in combination with filtering for other purposes, such as gain application, amplitude- or delay-based translation, stereo processing, reverberation, dynamic range processing (such as compression and limiting), self-chorus and / or flange effects, machine learning-based methods to vocal or instrumental style shifting, conversion or resynthesis of linear or nonlinear audio processing techniques and effects, or other types of processing using any of the super-middle, residual middle, super-side, and residual side components.
[0228] Filtering can be performed in the frequency domain or the time domain. In some embodiments, the intermediate and side components are transformed from the time domain to the frequency domain, super- and / or residual components are generated in the frequency domain, filtering is performed in the frequency domain, and the filtered components are transformed to the time domain. In other embodiments, the super- and / or residual components are transformed to the time domain, and filtering is performed on these components in the time domain.
[0229] The audio processing system generates a left output channel and a right output channel 2270 from the filtered super-intermediate component. In some embodiments, the left and right output channels are further based on at least one of the filtered residual intermediate component, the filtered super-intermediate component, and the filtered residual intermediate component. Example computer
[0230] Figure 23 is a block diagram of a computer 2300 according to one of some embodiments. The computer 2300 is an example of a computing device including a circuit system implementing an audio system (such as audio system 100, 202, or 1200). At least one processor 2302 coupled to a chipset 2304 is illustrated. The chipset 2304 includes a memory controller hub 2320 and an input / output (I / O) controller hub 2322. A memory 2306 and a graphics adapter 2312 are coupled to the memory controller hub 2320, and a display device 2318 is coupled to the graphics adapter 2312. A storage device 2308, a keyboard 2310, a pointer device 2314, and a network adapter 2316 are coupled to the I / O controller hub 2322. The computer system 2300 may include various types of input or output devices. Other embodiments of the computer 2300 have different architectures. For example, in some embodiments, memory 2306 is directly coupled to processor 2302.
[0231] Storage device 2308 includes one or more non-transitory computer-readable storage media, such as a hard disk, optical disc read-only memory (CD-ROM), DVD, or a solid-state memory device. Memory 2306 stores program code (including one or more instructions) and data used by processor 2302. The program code may correspond to the processing patterns described with reference to Figures 1 to 3.
[0232] The pointer device 2314 is used in conjunction with the keyboard 2310 to input data into the computer system 2300. The graphics adapter 2312 displays images and other information on the display device 2318. In some embodiments, the display device 2318 includes a touchscreen capability for receiving user input and selection. The network adapter 2316 couples the computer system 2300 to a network. Some embodiments of the computer 2300 have components that differ from those shown in FIG. 23 and / or other components.
[0233] The circuit system may include one or more processors that execute program code stored in a non-transitory computer-readable medium, the program code configuring one or more processors to implement an audio system or an audio system module when executed by the one or more processors. Other examples of a circuit system implementing an audio system or an audio system module may include an integrated circuit, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other types of computer circuitry.
[0234] Additional considerations
[0235] The exemplary benefits and advantages of the disclosed configuration include attribution to the dynamic audio enhancement of an enhanced audio system adapting to a device and associated audio presentation system, as well as other relevant information available to the device OS, such as use case information (e.g., indicating that the audio signal is used for music playback rather than for gaming). The enhanced audio system may be integrated into a device (e.g., using a software development kit) or stored on a remote server that can be accessed on demand. In this way, a device does not need to devote storage or processing resources to the maintenance of an audio enhancement system specific to its audio presentation system or audio presentation configuration. In some embodiments, the enhanced audio system enables queries for different levels of presentation system information, allowing effective audio enhancement to be applied across available device-specific presentation information at different levels.
[0236] In this specification, a plurality of instances may be implemented as components, operations, or structures of a single instance. Although individual operations of one or more methods are drawn and described as separate operations, one or more individual operations may be performed simultaneously, and not necessarily in the order shown. Structures and functions that appear as separate components in an instance configuration may be implemented as a combined structure or component. Similarly, structures and functions that appear as a single component may be implemented as a single component. Such and other changes, modifications, additions, and improvements fall within the scope of this document.
[0237] Certain embodiments herein are described as comprising logic or a number of components, modules, or mechanisms. Modules may constitute software modules (e.g., code embodied on a machine-readable medium or in a transmitted signal) or hardware modules. A hardware module is a tangible unit capable of performing certain operations and can be configured or configured in a particular manner. In exemplary embodiments, one or more computer systems (e.g., a standalone client or server computer system) or one or more hardware modules (e.g., a processor or a group of processors) of a computer system may be configured by software (e.g., an application or a portion of an application) to operate as one of the hardware modules performing some of the operations described herein.
[0238] Various operations of the exemplary methods described herein may be performed at least in part by one or more processors, which may be temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, these processors may constitute a processor implementation module that performs one or more operations or functions. In some exemplary embodiments, the modules involved herein include processor implementation modules.
[0239] Similarly, the methods described herein can be implemented at least in part by a processor. For example, at least some operations of a method can be performed by one or more processors or a hardware module implemented by a processor. The performance of certain operations can be distributed among one or more processors, rather than residing in a single machine, but deployed across several machines. In some exemplary embodiments, one or more processors may be located in a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed across several locations.
[0240] Unless otherwise expressly stated, the discussion herein using terms such as “processing,” “operation,” “calculation,” “determination,” “presentation,” “display,” or similar terms may refer to the actions or programs of a machine (e.g., a computer) that manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memory units (e.g., volatile memory, non-volatile memory, or combinations thereof), registers, or other machine components that receive, store, transmit, or display information.
[0241] As used herein, any reference to "an embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The phrase "in an embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.
[0242] Some embodiments may use the expressions "coupled" and "connected," as well as their derivatives, for description. It should be understood that these terms are not intended to be synonymous with each other. For example, some embodiments may use the term "connected" to indicate that two or more elements are in direct physical or electrical contact with each other. In another instance, some embodiments may use the term "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. Embodiments are not limited to this background.
[0243] As used herein, the terms "comprising," "including," "having," or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a procedure, method, article, or apparatus that includes one of a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to the procedure, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" means an inclusive "or" rather than an exclusive "or." For example, a condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).
[0244] Additionally, the use of "a" is for describing elements and components of the embodiments herein. This is for convenience only and to give the general meaning of the invention. This description should be interpreted as including one or at least one, and the singular includes the plural, unless it clearly has a different meaning.
[0245] Some parts of this description describe embodiments in terms of the algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are typically used by those skilled in data processing techniques to effectively communicate their working essence to others skilled in the art. These operations, when described in a functional, computational, or logical manner, should be understood as being implemented by computer programs or equivalent circuits, microcode, or the like. Furthermore, it has been shown that it is sometimes convenient to configure these operations as modules without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.
[0246] Any of the steps, operations, or procedures described herein may be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented using a computer program product including a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or procedures.
[0247] The embodiments may also relate to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for the desired purpose, and / or may include a general-purpose computing device that can be selectively started or reconfigured by a computer program stored in a computer. This computer program may be stored on a non-transitory tangible computer-readable storage medium or any type of media suitable for storing electronic instructions, and may be coupled to a computer system bus. Furthermore, any computing system mentioned in this specification may include a single processor or may employ an architecture designed to increase computing power using multiple processors.
[0248] The embodiment may also relate to a product generated by one of the computing programs described herein. Such a product may include information derived from a computing program, wherein the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of a computer program product or other combination of data described herein.
[0249] Upon reading this invention, those skilled in the art will understand additional alternative structures and functional designs for a system and a program related to the audio content of the principles disclosed herein. Therefore, although specific embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the precise constructions and components disclosed herein. Various modifications, alterations, and variations, as understood by those skilled in the art, can be made to the configuration, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
[0250] Finally, the language used in this specification has been chosen primarily for readability and pedagogical purposes, and is not intended to define or limit the patent rights. Therefore, the scope of the patent rights is not intended to be limited in detail, but rather is limited by the scope of any application filed based on one of these applications. Thus, the disclosure of embodiments is intended to illustrate, and not limit, the scope of the patent rights set forth in the following application scopes. [Simplified Explanation of the Diagram]
[0009] Figure 1 is a block diagram of an audio processing system according to one of some embodiments.
[0010] Figure 2 is a block diagram of a computing system environment according to one of some embodiments.
[0011] Figure 3 illustrates a curve of HRTF measured at a 60-degree elevation angle according to some embodiments.
[0012] Figure 4 illustrates a graph showing an example of a perceptual cue characterized by a target magnitude function corresponding to an infinitely attenuating narrow region at 11 kHz, according to some embodiments.
[0013] Figure 5 illustrates a frequency diagram generated according to some embodiments by driving a second-order all-pass filter segment with coefficients shown in Table 1 using white noise.
[0014] Figure 6 is a block diagram of a PSM module implemented using Hilbert transform according to one or more embodiments.
[0015] Figure 7 is a block diagram of one of the Hilbert transform modules according to one or more embodiments.
[0016] Figure 8 illustrates a frequency response generated by driving the HPSM module of Figure 6 with white noise according to some embodiments, showing the difference between one of the multiple channels (center) and one of the multiple channels (side).
[0017] Figure 9 is a block diagram of a PSM module implemented using an FNORD filter network according to some embodiments.
[0018] Figure 10A is a detailed block diagram of a PSM module 900 according to one of some embodiments.
[0019] Figure 10B is a block diagram of a broadband phase rotator implemented in the all-pass filter module of a PSM module according to some embodiments.
[0020] Figure 11 illustrates a frequency response diagram of the output frequency response of an FNORD filter network configured to achieve an amplitude response of a vertical filament of 60 degrees according to some embodiments.
[0021] Figure 12 is a block diagram of an audio processing system 1000 according to one or more embodiments.
[0022] Figure 13A is a block diagram of one of the orthogonal component generators according to one or more embodiments.
[0023] Figure 13B is a block diagram of one of the orthogonal component generators according to one or more embodiments.
[0024] Figure 13C is a block diagram of an orthogonal component generator according to one or more embodiments.
[0025] Figure 14A is a block diagram of one of the orthogonal component processor modules according to one or more embodiments.
[0026] Figure 14B illustrates a block diagram of one of the orthogonal component processor modules according to one or more embodiments.
[0027] Figure 15 is a block diagram of one of the sub-band space processor modules according to one or more embodiments.
[0028] Figure 16 is a block diagram of one of the crosstalk compensation processor modules according to one or more embodiments.
[0029] Figure 17 is a block diagram of one of the crosstalk analog processor modules according to one or more embodiments.
[0030] Figure 18 is a block diagram of one of the crosstalk cancellation processor modules according to one or more embodiments.
[0031] Figure 19 is a flowchart of a procedure for PSM processing using a Hilbert transform perceptual sound field modification (HPSM) module according to one or more embodiments.
[0032] Figure 20 is a flowchart of another procedure for PSM processing using a first-order nonorthogonal rotatable decorrelation (FNORD) filter network, according to some embodiments.
[0033] Figure 21 is a flowchart of a procedure for spatial processing using at least one of a super-center, residual center, super-side, or residual side component according to one or more embodiments.
[0034] Figure 22 is a flowchart of a procedure for subband spatial processing and crosstalk compensation processing using at least one of a super-center, residual center, super-side or residual side component, according to one or more embodiments.
[0035] Figure 23 is a block diagram of a computer according to one of some embodiments.
[0036] The figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.
Claims
1. A method for encoding spatial cues along a sagittal plane into a monophonic signal to generate a plurality of resulting channels, comprising a processing circuit system: determining a target amplitude response of one of the mid- or side-components of the plurality of resulting channels based on a spatial cues associated with a frequency-dependent phase shift; converting the target amplitude response of the mid- or side-component into a transfer function of a single-input multiple-output (SMILE) all-pass filter; and processing the monophonic signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
2. The method of claim 1, wherein the target amplitude response of the middle or side component of the obtained channel is determined according to a compensated null.
3. The method of claim 2, wherein the compensation zero point is in the range of 8 kHz to 16 kHz for the purpose of encoding vertical spatial cues.
4. As in request item 1, where: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and further includes using an inverse discrete Fourier transform (idFT) to convert the target amplitude response into the coefficients of the single-input multiple-output all-pass filter.
5. As in request item 1, where: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and further includes using a phase-vocoder to convert the target amplitude response into coefficients of the single-input multiple-output all-pass filter.
6. The method of claim 1, wherein the target amplitude response defines one or more parameter spatial cues, including one or more of a target broadband attenuation, a critical point, a filter characteristic, and a sound field location.
7. The method of claim 6, wherein the filter characteristic comprises one of the following: a high-pass filter characteristic; a low-pass filter characteristic; a band-pass filter characteristic; or a band-reject filter characteristic.
8. A system for encoding a spatial cue along a sagittal plane into a mono signal to generate a plurality of resulting channels, comprising: One or more computing devices are configured to: determine a target amplitude response of one of the middle and side components of the plurality of obtained channels based on a spatial cue associated with a frequency-dependent phase shift; convert the target amplitude response of the middle or side component into a transfer function of a single-input multiple-output (SMILE) all-pass filter; and process the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
9. The system of claim 8, wherein the target amplitude response of the middle or side component of the obtained channel is determined based on a compensated zero point.
10. The system of claim 9, wherein the compensation zero point is in the range of 8 kHz to 12 kHz for the purpose of encoding vertical spatial cues.
11. As in request item 8, the system wherein: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and the one or more computing devices are further configured to use an inverse discrete Fourier transform (idFT) to convert the target amplitude response into the coefficients of the single-input multiple-output all-pass filter.
12. The system as described in request item 8, wherein: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and the one or more computing devices are further configured to use a phase vocoder to convert the target amplitude response into the coefficients of the single-input multiple-output all-pass filter.
13. The system of claim 8, wherein the target amplitude response defines one or more parameter spatial cues, including one or more of a target broadband attenuation, a critical point, a filter characteristic, and a sound field location.
14. The system of claim 13, wherein the filter characteristic includes one of the following: a high-pass filter characteristic; a low-pass filter characteristic; a band-pass filter characteristic; or a band-reject filter characteristic.
15. A non-transitory computer-readable medium comprising stored instructions for encoding a spatial cue along a sagittal plane into a mono signal to generate a plurality of resulting channels, the instructions, when executed by at least one processor, configuring the at least one processor to: determine a target amplitude response of a central or lateral component of the plurality of resulting channels based on a spatial cue associated with a frequency-dependent phase shift; convert the target amplitude response of the central or lateral component into a transfer function of a single-input multiple-output (SMILE) all-pass filter; and process the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.
16. The non-transitory computer-readable medium of claim 15, wherein the target amplitude response of the middle or side component of the obtained channel is determined according to a compensated zero point.
17. The non-transitory computer-readable medium of claim 16, wherein the compensation zero point is in the range of 8 kHz to 12 kHz, for the purpose of encoding vertical spatial cues.
18. A non-transitory computer-readable medium as described in claim 15, wherein: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and the one or more processors are further configured to use an inverse discrete Fourier transform (idFT) to convert the target amplitude response into the coefficients of the single-input multiple-output all-pass filter.
19. A non-transitory computer-readable medium as described in claim 15, wherein: The target amplitude response of the middle or side component of the plurality of obtained channels is determined based on the amplitude within the frequency range; and the one or more processors are further configured to use a phase vocoder to convert the target amplitude response into the coefficients of the single-input multiple-output all-pass filter.
20. The non-transitory computer-readable medium of claim 15, wherein the target amplitude response defines constraints on one or more of the plurality of resulting channels, including one or more of a target broadband attenuation, a critical point, and a filter characteristic.
Citation Information
Patent Citations
Auto-focus in low-profile folded optics multi-camera system
CN106164732A
Optical systems, metrology apparatus and associated methods
TW201921145A
Optical systems, metrology apparatus and associated methods
TW202001444A
Steering of monaural sources of sound using head related transfer functions
US6611603B1