Colorless generation of elevation angle perceptual suggestions using all-pass filter networks

By encoding spatial cues into mono audio signals using all-pass filter networks, the method improves spatial perception and immersion in audio content, addressing the limitations of existing technologies.

JP7797594B2Active Publication Date: 2026-01-13BOOMCLOUD 360 INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024166285
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-01
Filing Date
2024-09-25
Publication Date
2026-01-13
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

Existing audio encoding technologies struggle to effectively incorporate spatial cues into mono audio signals, limiting the perception of space and immersion in audio content.

Method used

The method involves encoding spatial sensitivity along a sagittal plane into a monaural signal using a single-input, multiple-output all-pass filter network, which processes the signal to generate multiple channels with spatial cues, and employs Hilbert Transforms to separate and combine audio components for enhanced spatial perception.

Benefits of technology

This approach enhances the perceived spatial quality of audio, allowing for improved immersion and clarity by distinguishing sound sources in space, reducing the need for multiple speakers, and providing a more immersive audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797594000048
    Figure 0007797594000048
  • Figure 0007797594000049
    Figure 0007797594000049
  • Figure 0007797594000050
    Figure 0007797594000050
Patent Text Reader

Abstract

To provide a system that includes one or more computing devices that encode spatial perceptual cues into a monaural channel to generate a plurality of output channels.SOLUTION: A computing device determines a target amplitude response for the mid and side channels of the plurality of output channels, defining a spatial perception associated with one or more frequency-dependent phase shifts. The computing device determines a transfer function of a single-input, multi-output all-pass filter based on the target amplitude response, determines coefficients of the all-pass filter based on the transfer function, and processes the monaural channel with the coefficients of the all-pass filter to generate the plurality of channels having the encoded spatial perceptual cues. The all-pass filter is configured to be colorless with respect to the individual output channels, allowing the placement of spatial cues into the audio stream to be decoupled from the overall coloration of the audio.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to audio processing, and more particularly to encoding spatial cues into audio content. [Background technology]

[0002] Audio content can be encoded to include spatial characteristics of a sound field, allowing a user to perceive a sense of space in the sound field. For example, audio from a particular source (e.g., a voice or an instrument) may be mixed with the audio content in a way that creates a sense of space associated with the audio, such as the perception that the audio is coming to the user from a particular direction of arrival or is located in a particular type of location (e.g., a small room, a large auditorium, etc.). Summary of the Invention

[0003] Some embodiments include a method for encoding spatial sensitivity along a sagittal plane into a monaural signal to generate multiple resultant channels, the method including: determining, by a processing circuit, a target amplitude response for a mid or side component of the resulting multiple channels based on the spatial sensitivity associated with a frequency-dependent phase shift; converting the target amplitude response for either the mid or side component into a transfer function for a single-input, multiple-output all-pass filter; and processing the monaural signal using the all-pass filter, the all-pass filter configured based on the transfer function.

[0004] Some embodiments include a system for generating multiple channels from a mono channel, where the multiple channels are encoded with one or more spatial cues. The system includes one or more computing devices configured to determine a target amplitude response for a mid or side component of the multiple channels based on the spatial cues associated with a frequency-dependent phase shift. The one or more computers are further configured to convert the target amplitude response for either the mid or side component to a transfer function for a single-input, multiple-output all-pass filter and process the mono signal using the all-pass filter, where the all-pass filter is configured based on the transfer function.

[0005] Some embodiments include a non-transitory computer-readable medium including instructions stored thereon for generating multiple channels from a mono channel, wherein the multiple channels are encoded with one or more spatial implications, the instructions, when executed by at least one processor, configure the at least one processor to: determine a target amplitude response for a mid or side component of the resulting multiple channels based on the spatial implications associated with a frequency-dependent phase shift, transform the target amplitude response for the mid or side component to a transfer function of a single-input multiple-output all-pass filter, and process the mono signal using the all-pass filter, wherein the all-pass filter is configured based on the transfer function.

[0006] Some embodiments relate to spatially shifting portions of audio content (e.g., vocalizations) using a series of Hilbert Transforms. Some embodiments include one or more processors and a non-transitory computer-readable medium. The computer-readable medium includes stored program code that, when executed by the one or more processors, configures the one or more processors to: separate an audio channel into low-frequency and high-frequency components; apply a first Hilbert Transform to the high-frequency components to generate a first left leg component and a first right leg component, where the first left leg component is 90 degrees out of phase with the first right leg component; apply a second Hilbert Transform to the first right leg component to generate a second left leg component and a second right leg component, where the second left leg component is 90 degrees out of phase with the second right leg component; combine the first left leg component with the low-frequency components to generate a left channel; and combine the second right leg component with the low-frequency components to generate a right channel.

[0007] Some embodiments include a non-transitory computer-readable medium containing stored program code that, when executed by one or more processors, configures the one or more processors to: separate an audio channel into low and high frequency components, apply a first Hilbert transform to the high frequency components to generate a first left leg component and a first right leg component, where the first left leg component is 90 degrees out of phase with the first right leg component, apply a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, where the second left leg component is 90 degrees out of phase with the second right leg component, combine the first left leg component with the low frequency component to generate a left channel, and combine the second right leg component with the low frequency component to generate a right channel.

[0008] Some embodiments include a method executed by one or more processors, comprising: separating an audio channel into low-frequency and high-frequency components; applying a first Hilbert transform to the high-frequency components to generate a first left leg component and a first right leg component, the first left leg component being 90 degrees out of phase with the first right leg component; applying a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, the second left leg component being 90 degrees out of phase with the second right leg component; combining the first left leg component with the low-frequency components to generate a left channel; and combining the second right leg component with the low-frequency components to generate a right channel. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram illustrating an audio processing system, according to some embodiments. [Figure 2] FIG. 1 is a block diagram illustrating a computing system environment according to some embodiments. [Figure 3] 10A-10C illustrate graphs showing sampled HRTFs measured at an elevation angle of 60 degrees, in accordance with some embodiments. [Figure 4] FIG. 10 illustrates a graph showing an example of a perceptual cue characterized by a target amplitude function corresponding to a narrow region of infinite attenuation at 11 kHz, according to some embodiments. [Figure 5] 10A-10C illustrate frequency plots produced by driving a second-order all-pass filter section having coefficients shown in Table 1 with white noise, according to some embodiments. [Figure 6] FIG. 1 is a block diagram illustrating a PSM module implemented using a Hilbert transform in accordance with one or more embodiments. [Figure 7] FIG. 2 is a block diagram illustrating a Hilbert transformer module according to one or more embodiments. [Figure 8] FIG. 7 illustrates a frequency plot produced by driving the HPSM module in FIG. 6 with white noise, showing the output frequency response of the sum of multiple channels (mid) and the difference of multiple channels (side), in accordance with some embodiments. [Figure 9] FIG. 1 is a block diagram illustrating a PSM module implemented using an FNORD filter network, according to some embodiments. [Figure 10A] FIG. 9 is a detailed block diagram illustrating a PSM module 900, according to some embodiments. [Figure 10B] FIG. 2 is a block diagram illustrating a Broadband Phase Rotator implemented within an all-pass filter module of a PSM module, according to some embodiments. [Figure 11] FIG. 10 illustrates a frequency response graph showing the output frequency response of an FNORD filter network configured to achieve an amplitude response for a 60 degree vertical cue, according to some embodiments. [Figure 12] FIG. 10 is a block diagram illustrating an audio processing system 1000 according to one or more embodiments. [Figure 13A] FIG. 2 is a block diagram illustrating an orthogonal component generator according to one or more embodiments. [Figure 13B] FIG. 2 is a block diagram illustrating a quadrature component generator according to one or more embodiments. [Figure 13C] FIG. 2 is a block diagram illustrating a quadrature component generator according to one or more embodiments. [Figure 14A] FIG. 2 is a block diagram illustrating a quadrature component processor module according to one or more embodiments. [Figure 14B] FIG. 1 shows a block diagram illustrating a quadrature component processor module according to one or more embodiments. [Figure 15]FIG. 2 is a block diagram illustrating a subband spatial processor module according to one or more embodiments. [Figure 16] FIG. 2 is a block diagram illustrating a crosstalk compensation processor module according to one or more embodiments. [Figure 17] FIG. 2 is a block diagram illustrating a crosstalk simulation processor module according to one or more embodiments. [Figure 18] FIG. 2 is a block diagram illustrating a crosstalk cancellation processor module according to one or more embodiments. [Figure 19] 1 is a flowchart illustrating a process for PSM processing using a Hilbert Transform Perceptual Soundstage Modification (HPSM) module in accordance with one or more embodiments. [Figure 20] 10 is a flowchart illustrating another process for PSM processing using a First Order Non-Orthogonal Rotation-Based Decorrelation (FNORD) filter network, according to some embodiments. [Figure 21] 1 is a flowchart illustrating a process for spatial processing using at least one of a hyper mid component, a residual mid component, a hyper side component, or a residual side component, according to one or more embodiments. [Figure 22]FIG. 10 is a flowchart illustrating a process for subband spatial processing and compensation for crosstalk processing using at least one of a hypermid component, a residual mid component, a hyperside component, or a residual side component, in accordance with one or more embodiments. [Figure 23] FIG. 1 is a block diagram illustrating a computer according to some embodiments.

[0010] The figures depict various embodiments for purposes of illustration only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein. DETAILED DESCRIPTION OF THE INVENTION

[0011] [Detailed explanation] The drawings (FIG.) and the following description relate to preferred embodiments for purposes of illustration only. It should be noted from the following description that alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be used without departing from the principles of the claimed subject matter.

[0012] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying figures. It should be noted that, wherever possible, like or similar reference numbers may be used in the figures and may indicate like or similar functionality. The figures depict embodiments of the disclosed system (or method) for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein.

[0013] Encoding spatial perceptual cues into mono audio sources may be desirable in a variety of applications involving the presentation of multiple simultaneous streams of audible content. Examples of such applications include: Conferencing use case: The addition of spatial cues applied to one or more distant talkers can improve overall voice intelligibility and contribute to a greater overall sense of immersion for the listener. Video and music playback / streaming use cases: One or more audio channels, or signal components of one or more audio channels, can be enhanced by adding spatial cues to improve the intelligibility or sense of space of the audio or other elements of the mix. Collaborative Viewing Entertainment Use Case: When the streams are separate content channels, such as one or more distant talkers and entertainment program material, that need to be mixed to form an immersive experience, applying spatial perceptual cues to one or more elements can enhance the sense of perceptual distinction between the elements of the mix and widen the listener's perceptual bandwidth.

[0014] Embodiments relate to audio systems that modify the perceived spatial quality (e.g., sound stage and overall position relative to the head of a target listener) of one or more audio channels. In some embodiments, modifying the perceived spatial quality of an audio channel can be used to separate the coloration of a particular source from its perceived position in space and / or to reduce the required number of amplifiers and speakers needed to encode such effects.

[0015] The audio signal processing performed by the audio system is referred to as perceptual sound stage modification (PSM) processing. The perceptual result of PSM processing is referred to herein as spatial shifting. The psychoacoustic effect is typically experienced by the user as an overall shift of the sound source above, around, or toward the head, perceptually distinguishing the sound source from the rest of the audio content. This psychoacoustic effect results from the phase and time relationship between the left and right channels, as accentuated by a network of all-pass filters and delays. In some embodiments, this filter and delay network may be implemented as one or more second-order allpass sections, such as a series of Hilbert transforms, or using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network, each of which will be described in more detail below. The perceptual result of PSM processing may vary depending on various listening configurations (e.g., headphones, speakers, etc.). For some content and algorithm configurations, the result may be the impression that the perceived signal is spread (e.g., diffuse) around the listener's head. For mono input signals (e.g., non-spatial audio signals), the diffuse effect of PSM processing can be used for mono-to-stereo upmixing.

[0016] In some embodiments, an audio system may separate a target portion of an audio signal from the remaining portion of the audio signal, apply various configurations of PSM processing to perceptually shift the target portion, and then mix the processing results (e.g., unprocessed or differently processed) back with the remaining portion. Such a system may be perceived as clarifying, elevating, or otherwise differentiating the target portion from the overall audio mix. In some embodiments, PSM processing is used to perceptually shift a portion of an audio signal, including a sung or spoken voice. By convention, voices in television, movie, or music audio streams are often located in the center of the soundstage and are therefore part of the mid component (also called the non-spatial or non-correlated component) of a stereo or multi-channel audio signal. Therefore, PSM processing may be applied to the mid component of an audio signal or to the hyper-mid component, which includes the spectral energy of side components (also called the spatial or non-correlated component) that are removed from the spectral energy of the mid component.

[0017] PSM processing may be combined with other types of processing. For example, an audio system may apply processing to shifted portions of an audio signal to perceptually transform the shifted portions and distinguish them from other components in the mix. These types of additional processing may include one or more of single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., limiting, compression, expansion, etc.), single-band or multi-band gain or delay, crosstalk processing (e.g., crosstalk cancellation processing and / or crosstalk simulation processing), or crosstalk processing compensation. In some embodiments, PSM processing may be performed in conjunction with mid / side processing, such as subband spatial processing, in which subbands of mid and side components of the audio signal generated via PSM processing are gain-adjusted to enhance the sense of spatiality of the sound field.

[0018] Separation of audio channels for PSM processing can be achieved in various ways. In some embodiments, PSM processing may be performed on spectrally orthogonal audio components, such as the hyper-mid components of an audio signal. In other embodiments, PSM processing is performed on an audio channel associated with a sound source (e.g., vocalization), and the processed channel is then mixed with other audio content (e.g., background music).

[0019] The following description focuses primarily on upmixing a mono signal to stereo (i.e., two output channels) because the majority of audio presentation equipment is stereo, but it will be understood that the techniques described can be easily generalized to include a greater number of channels. Stereo embodiments may be described in terms of mid / side processing, where the phase difference between the left and right channels results in complementary regions of amplification and attenuation in the mid / side space.

[0020] [Example of a voice processing system] 1 is a block diagram illustrating an audio processing system 100 according to one or more embodiments. System 100 uses PSM processing to spatially shift audio signals and apply other types of spatial (e.g., mid / side) processing. Some embodiments of system 100 have different components than those described herein. Similarly, in some cases, functionality may be distributed among components in a different manner than that described herein.

[0021] System 100 includes a PSM module 102, an L / RM / S converter module 104, a component processor module 106, an M / SL / R converter module 108, and a crosstalk processor module 108. PSM module 102 receives input audio 120 and generates spatially shifted left and right channels 122 and 124. Operation of PSM 102, according to various embodiments, is described in more detail below in connection with Figures 6-11.

[0022] The L / RM / S transformer module 104 receives a left channel 122 and a right channel 124 and generates a mid component 126 (e.g., a non-spatial component) and a side component 128 (e.g., a spatial component) from the channels 122 and 124. In some embodiments, the mid component 126 is generated based on the addition of the left channel 122 and the right channel 122, and the side component 128 is generated based on the difference of the left channel 122 and the right channel 124. In some embodiments, the transformation of a point in L / R space to a point in M / S space may be expressed according to equation (1) as follows:

[0023]

number

[0024] Meanwhile, the inverse transformation can be expressed according to equation (2) as follows:

[0025]

number

[0026] It is understood that in other embodiments, other L / RM / S type transforms may be used to generate the mid and side components 126 and 128. In some embodiments, the transforms shown in equations (1) and (2) may be used instead of a true orthonormal form in which both the forward and inverse transforms are scaled by √2 due to reduced computational complexity. For ease of explanation, regardless of the specific transform used, the convention of transforming the coordinates of row vectors by multiplication on the right and the notation of the transformed coordinates having a basis as a label on them will be used, as shown in equation (3) below:

[0027]

number

[0028] The component processor module 106 processes the mid component 126 to generate a processed mid component 130 and processes the side component 128 to generate a processed side component 314. The processing on each of the components 126 and 128 may include various types of filter extraction, such as spatially sensitive processing (e.g., amplitude or delay-based panning, binaural processing, etc.), single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., compression, expansion, limiting, etc.), single-band or multi-band gain or delay stages, addition of audio effects, or other types of processing. In some embodiments, the component processor module 106 uses the mid component 126 and the side component 128 to perform subband spatial processing and / or crosstalk compensation processing. Subband spatial processing is processing performed on frequency subbands of the mid component and the side component to spatially enhance the audio signal. Crosstalk compensation processing is processing that adjusts for spectral artifacts caused by crosstalk processing, such as crosstalk correction for speakers or crosstalk simulation for headphones. The various components that may be included in the component processor module 106 are further described with reference to Figures 12A-13.

[0029] The M / SL / R transformer module 108 receives the processed mid component 130 and the processed side component 132 and generates a processed left component 134 and a processed right component 136. In some embodiments, the M / SL / R transformer module 108 transforms the processed mid component 130 and the side component 132 based on the inverse of the transform performed by the L / RM / S transformer module 104, e.g., the processed left component 134 is generated based on the addition of the processed mid component 130 and the processed side component 132, and the processed right component 136 is generated based on the difference of the processed mid component 130 and the processed side component 132. Other M / SL / R transform types may be used to generate the processed left component 134 and the processed right component 136.

[0030] The crosstalk processor module 110 receives the processed left component 134 and the processed right component 136 and performs crosstalk processing. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is processing performed on an audio signal (e.g., output through headphones) to simulate the effect of loudspeakers. Crosstalk cancellation is processing performed on an audio signal (e.g., output through speakers) to reduce crosstalk caused by the loudspeakers. The crosstalk processor module 110 outputs a left channel 138 and a right output channel 140. In some embodiments, crosstalk processing (e.g., simulation or cancellation) may be performed before component processing, such as before converting the left channel 122 and the right channel 124 into mid and side components. Various components that may be included in the crosstalk processor module 110 are further described with reference to FIGS. 15 and 16 .

[0031] In some embodiments, the PSM module 100 is incorporated into the component processor module 106. The L / RM / S converter module 104 receives a left channel and a right channel, which may represent (e.g., stereo) inputs to the audio processing system 100. The L / RM / S converter module 104 uses the left and right input channels to generate a mid component and a side component. The PSM module 100 of the component processor module 106 processes the mid component and / or the side component as input to generate the left and right channels, as discussed herein for the input audio 102. The component processor module 106 may also perform other types of processing on the mid and side components, and the M / SL / R converter module 108 generates the left and right channels from the processed mid and side components. The left channel generated by the HPSM module 100 is combined with the left channel generated by the M / SL / R converter module 108 to generate the processed left component. The right channel produced by the PSM module 100 is combined with the right channel produced by the M / SL / R converter module 108 to produce a processed right component.

[0032] System 100 provides left channel 138 to left speaker 112 and right channel 140 to right speaker 114. Speakers 112 and 114 may be components of a smartphone, tablet, smart speaker, laptop, desktop, exercise machine, etc. Speakers 112 and 114 may be part of an appliance that includes system 100, or may be separate from system 100 such that they are connected to system 100 via a network. The network may include wired and / or wireless connections. The network may include a local area network, a wide area network (including, for example, the Internet), or a combination thereof.

[0033] 2 is a block diagram illustrating a computing system environment 200 according to some embodiments. Computing system 200 may include an audio system 202, which may include one or more computing devices (e.g., servers, etc.) connected to user devices 210a and 210b via a network 208. Audio system 202 provides audio content to user devices 210a and 210b (individually referred to as user devices 210) via network 208. Network 208 facilitates communication between system 202 and user devices 210. Network 106 may include various types of networks, including the Internet.

[0034] The audio system 202 includes one or more processors 204 and a computer-readable medium 206. The one or more processors 204 execute program modules that cause the one or more processors 204 to perform functions, such as generating multiple output channels from a mono channel. The processor 204 may include a central processing unit (CPU), a graphics processing unit (GPU), a controller, a state machine, other types of processing circuitry, or a combination of one or more of these. The processor 204 may further include local memory for storing, among other things, program modules, operating system data.

[0035] The computer-readable medium 206 is a non-transitory storage medium that stores program code for the PSM module 102, the component processor module 106, the crosstalk processor module 110, the L / R converter module 104 and the M / S converter module 108, and the channel summation module 212. The PSM module 102 generates multiple output channels from the mono channel, which may be further processed using the component processor module 106, the crosstalk processor module 110, and / or the L / R converter module 104 and the M / S converter module 108. The system 202 provides the multiple output channels to a user device 210a, which includes multiple speakers 214 for rendering each of the output channels.

[0036] The channel summation module 212 generates a mono output channel by adding together multiple output channels generated by the PSM module 102 and / or other modules. The system 202 provides the mono output channel to the user equipment 210b, which includes a single speaker 216 for rendering the mono output channel. In some embodiments, the channel summation module 212 is located in the user equipment 210b. The audio system 202 provides multiple output channels to the user equipment 210b, which converts the multiple channels into a mono output channel for the speaker 216. The user equipment 210 presents audio content to the user. The user equipment 210 may be a user's computing device, such as a music player, a smart speaker, a smartphone, a wearable device, a tablet, a laptop, or a desktop.

[0037] [Mid / Side Space Coloration] In some embodiments, spatial implications are encoded into the audio signal by creating coloration effects in the mid / side spaces while avoiding coloration effects in the left / right spaces. In some embodiments, this is achieved by applying all-pass filters with specifically selected characteristics to the left / right spaces to produce the desired coloration in the mid / side spaces. For example, in a two-channel system, the relationship between the left / right phase angle and the mid / side gain can be expressed using the following equation (4):

[0038]

number

[0039] where:

[0040]

number

[0041] is a two-dimensional row vector composed of the mid and side target gain factors, each in decibels, at a particular frequency ω,

[0042]

number

[0043] is the objective function for the phase relationship between the left and right channels.

[0044]

number

[0045] Solving equation (4) for σ provides the frequency-dependent phase differential required to apply to left / right space according to equations (5) and (6) below:

[0046]

number

[0047] Note that if the constraint that the system be colorless in left-right space is applied, then only the transfer functions for the mid or side components can be specified. Therefore, the system of equations (5) and (6) becomes over-determined, and only one of the above equations can be solved without breaking the required symmetry. In some embodiments, control over either the mid or side components can be obtained by selecting a particular equation. If the constraint that the system be colorless in left-right space were removed, further degrees of freedom could be achieved. In systems with more than two channels, various techniques, such as pairwise transforms or hierarchical sum-and-difference transforms, could be used instead of mid and side.

[0048] [Example implementation of an all-pass filter for encoding elevation cues] In some embodiments, spatial perceptual cues may be encoded into audio signals by embedding frequency-dependent amplitude cues (i.e., coloration) into mid-space / side-space while constraining the left and right signals to be colorless. For example, elevation cues (e.g., spatial perceptual cues located along the sagittal plane) may be encoded using this framework because left and right cues with respect to elevation are theoretically symmetric in coloration.

[0049] In some embodiments, a notable feature in the Head-Related Transfer Function (HRTF)-based elevation angle suggestion is a notch starting at approximately 8 kHz and rising monotonically as a function of elevation angle to approximately 16 kHz, which is used to derive the appropriate coloration of the mid channel for encoding the elevation angle. Using this mid-encoded cue, a corresponding frequency-dependent phase shift can be derived, which can be further used to derive a function implemented via a filter network (e.g., PSM module 100) such as the one described below. In some embodiments, the HRTF-based elevation angle suggestion can be characterized as a notch starting at approximately 8 kHz and rising monotonically as a function of elevation angle to approximately 12 kHz.

[0050] For ease of explanation, the following example filter framework will be described in connection with encoding the same perceptual suggestion according to some embodiments, where the target elevation angle is 60 degrees (e.g., spatially shifting audio content 60 degrees above horizontal). However, it will be understood that in other embodiments, similar techniques are used to encode perceptual suggestions having different elevation angles. FIG. 3 illustrates a graph showing sampled HRTFs measured at a 60-degree elevation angle, according to some embodiments. FIG. 4 illustrates a graph showing an example of a perceptual suggestion characterized by a target amplitude function corresponding to a narrow region of infinite attenuation at approximately 11 kHz, according to some embodiments. Such suggestions could be used to create an elevation perception in most people across a wide variety of presentation scenarios. While the graph in FIG. 4 illustrates a simplified sampled HRTF, it will be understood that more complex suggestions can also be derived based on the framework described herein.

[0051] [Design using second-order all-pass section] In some embodiments, PSM module 100 is implemented using two independent cascades of second-order all-pass filters and delay elements to achieve the desired phase shift in left / right space to encode perceptual suggestions as described above in connection with FIG. 4. In some embodiments, the second-order sections are implemented as biquad sections, whose coefficients are applied to feedback and feedforward taps of up to two delayed samples. As discussed herein, the convention of naming feedback coefficients A1 and A2 of one sample and two samples, respectively, and feedforward coefficients B0, B1, and B2 of zero sample, one sample, and two samples, respectively, is used.

[0052] In some embodiments, the PSM module 100 is implemented using a second-order all-pass filter configured to perform pole and zero cancellation, allowing the amplitude component of the transfer function to remain flat while the phase response is modified. Using all-pass filter sections for both channels in left and right space ensures a specific phase shift across the spectrum. This has the added benefit of allowing for a given phase offset between the left and right sides, resulting in an increased sense of spaciousness in addition to the desired nulls in the mid / side space.

[0053] Table 1 below illustrates an example set of biquad coefficients that may be used in the framework of a second-order all-pass filter with an additional two-sample delay on the right channel, according to some embodiments. The biquad coefficients illustrated in Table 1 may be designed for a sampling rate of 44.1 kHz, but may also be used for systems with other sampling rates (e.g., 48 kHz, etc.).

[0054] [Table 1]

[0055] A network of filters with the coefficients shown in Table 1 can produce an appropriate phase response in left / right space, resulting in a significant null / amplification in mid / side space at 11 kHz. Figure 5 illustrates a frequency plot produced by driving a second-order all-pass filter section with the coefficients shown in Table 1 with white noise, showing the output frequency response of the sum 502 of multiple channels (mid) and the difference 504 of multiple channels (side), according to some embodiments.

[0056] In some embodiments, PSM module 100 implemented using second-order all-pass filter sections may be further enhanced using crossover networks to filter out processing for frequency regions that do not require it. The use of crossover networks may increase the flexibility of the implementation by allowing further processing for perceptually significant cues to filter out unnecessary auditory data.

[0057] In some embodiments, the PSM module 100 implemented using a second-order all-pass filter section may be implemented using a network of serially chained Hilbert Transforms, as will be described in more detail below.

[0058] [Example of a Hilbert Transform Perceptual Sound Stage Modification (HPSM) module] 6 is a block diagram illustrating a PSM module implemented using a Hilbert transform, according to one or more embodiments. The PSM module 600, also known as a Hilbert transform-based perceptual sound stage modification (HPSM) module, applies a network of serially concatenated Hilbert transforms to an input audio 602 (which may correspond to the input audio 120 shown in FIG. 1) to perceptually shift the input audio 602.

[0059] Module 600 includes a crossover network module 604, a gain unit 610, a gain unit 612, a Hilbert transformer module 614, a Hilbert transformer module 620, a delay unit 626, a gain unit 628, a delay unit 630, a gain unit 630, a gain unit 632, a summing unit 634, and a summing unit 636. Some embodiments of module 600 have different components than those described herein. Likewise, in some cases, functionality may be distributed among components in a different manner than described herein.

[0060] The crossover network module 604 receives the input audio 602 and generates a low frequency component 606 and a high frequency component 608. The low frequency component includes subbands of the input audio 602 that have lower frequencies than the subbands of the high frequency component 608. In some embodiments, the low frequency component 606 includes a first portion of the input audio that includes low frequencies, and the high frequency component 608 includes the remaining portion of the input audio that includes high frequencies.

[0061] As described in more detail below, the high frequency components 608 are processed using a series of Hilbert Transforms, while the low frequency components 606 bypass the series of Hilbert Transforms, after which the low frequency components and the processed high frequency components 608 are recombined. The crossover frequency between the frequency components 606 and the high frequency components 608 may be adjustable. For example, more frequencies may be included in the high frequency components 608 to increase the perceived strength of the spatial shift caused by the HPSM module 600, while more frequencies may be included in the low frequency components 606 to reduce the perceived strength of the shift. In another example, the crossover frequency is set so that frequencies for a sound source of interest (e.g., speech, etc.) are included in the high frequency components 608.

[0062] The input audio 602 may include a mono channel or may be a mixdown of a stereo or other multi-channel signal (e.g., surround sound, ambisonics, etc.). In some embodiments, the input audio 602 is audio content associated with a sound source that is to be incorporated into an audio mix. For example, the input audio 602 may be a vocalization that is processed by the module 600, and the processing results are combined with other audio content (e.g., background music, etc.) to generate the audio mix.

[0063] Gain unit 610 applies a gain to low frequency components 606, and gain unit 612 applies a gain to high frequency components 608. Gain units 610 and 612 may be used to adjust the overall levels of low frequency components 606 and high frequency components 608 relative to one another. In some embodiments, gain unit 610 or gain unit 612 may be omitted from module 600.

[0064] Hilbert transformer modules 614 and 620 apply a series of Hilbert transforms to the high frequency components 608. Hilbert transformer module 614 applies a Hilbert transform to the high frequency components 608 to produce left leg components 616 and right leg components 618. The left leg components 616 and right leg components 618 are audio components that are 90 degrees out of phase with each other. In some embodiments, the left leg components 616 and right leg components 618 are out of phase with each other by an angle other than 90 degrees, such as, for example, between 20 and 160 degrees.

[0065] The Hilbert transformer module 620 applies a Hilbert transform to the right leg component 618 generated by the Hilbert transformer module 614 to generate the left leg component 122 and the right leg component 624. The left leg component 622 and the right leg component 624 are audio components that are 90 degrees out of phase with each other. In some embodiments, the Hilbert transformer module 620 generates the right leg component 624 without generating the left leg component 122. In some embodiments, the left leg component 622 and the right leg component 624 are out of phase with each other by an angle other than 90 degrees, such as, for example, between 20 and 160 degrees.

[0066] In some embodiments, each of the Hilbert transformer modules 614 and 620 is implemented in the time domain and includes a cascaded all-pass filter and a delay, as described in more detail below in connection with Figure 7. In other embodiments, the Hilbert transformer modules 614 and 620 are implemented in the frequency domain.

[0067] Delay unit 626, gain unit 628, delay unit 630, and gain unit 632 provide adjustment controls for manipulating the perceptual results of processing by module 600. Delay unit 626 applies a time delay to left leg component 616 generated by Hilbert transformer module 614. Gain unit 628 applies a gain to left leg component 616. In some embodiments, delay unit 626 or gain unit 628 may be omitted from module 600.

[0068] The delay unit 630 applies a time delay to the right leg component 624 generated by the Hilbert transformer module 620. The gain unit 632 applies a gain to the right leg component 624. In some embodiments, the delay unit 630 or the gain unit 632 may be omitted from the module 600.

[0069] A summing unit 634 combines the low frequency component 606 with a left leg component 616 to generate a left channel 642. The left leg component 616 is the output from the first Hilbert transformer module 614 in the series. The left leg component 616 may include a delay applied by a delay unit 626 and a gain applied by a gain unit 628.

[0070] A summing unit 636 combines the low frequency component 606 with a right leg component 624 to generate a right channel 644. The right leg component 624 is an output from the second Hilbert transformer module 620 in the series. The right leg component 624 may include a delay applied by a delay unit 626 and a gain applied by a gain unit 628.

[0071] 7 is a block diagram illustrating a Hilbert transformer module 700 according to one or more embodiments. The Hilbert transformer module 700 is an example of the Hilbert transformer module 614 or the Hilbert transformer module 620. The Hilbert transformer module 700 receives an input component 702 and uses the input component 702 to generate a left leg component 712 and a right leg component 724. Some embodiments of the Hilbert transformer module 700 have different components than those described herein. Similarly, in some cases, functionality may be distributed among components in a different manner than described herein.

[0072] The Hilbert transformer module 700 includes an all-pass filter cascade module 740 for generating a left leg component 712, a delay unit 714, and an all-pass filter cascade module 742 for generating a right leg component 724. The all-pass filter cascade module 714 includes a series of all-pass filters 704, 706, 708, and 710. The delay unit 714 applies a time delay to the input component 702. The all-pass filter cascade module 742 includes a series of all-pass filters 716, 718, 720, and 722. Each of the all-pass filters 704-710 and 716-722 passes frequencies with equal gain while varying the phase relationship between different frequencies. In some embodiments, each all-pass filter 704-710 and 716-722 is a biquad filter as defined by equation (7).

[0073]

number

[0074] where z is a complex variable and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. Different biquadratic filters may contain different coefficients to apply different phase changes.

[0075] The all-pass filter cascade modules 740 and 742 may each include a different number of all-pass filters. The Hilbert transformer module 700 is an eighth-order filter with eight all-pass filters, four for each of the left leg component 712 and the right leg component 724. In other embodiments, the Hilbert transformer module 700 is an eighth-order filter (e.g., four all-pass filters for each of the all-pass filter cascade modules 740 and 742) or a sixth-order filter (e.g., three all-pass filters for each of the all-pass filter cascade modules 740 and 742).

[0076] As discussed above in connection with FIG. 6 , module 600 includes a series of Hilbert transformer modules 614 and 620. Using a Hilbert transformer module 700 for each of Hilbert transformer modules 614 and 620, left leg component 616 is generated by one all-pass filter cascade module 740 applied to high frequency component 608. Right leg component 624 is generated by two passes through Hilbert transformer module 700 with two delay units 714 and two all-pass filter cascade modules 742. In some embodiments, Hilbert transformer modules 614 and 620 may be different. For example, Hilbert transformer modules 614 and 620 may include filters of different orders, such as an eighth-order filter for one of the Hilbert transformer modules and a sixth-order filter for another of the Hilbert transformer modules.

[0077] When the Hilbert transformer module 700 is used for the Hilbert transformer modules 614 and 620, the right leg component 624 has a phase and delay relationship to the right leg component 618, which is created by the all-pass filter and delay of the Hilbert transformer module 620. The right leg component 624 also has a phase and delay relationship to the high frequency component 608, which is created by the all-pass filter and delay in the Hilbert transformer modules 614 and 620. In some embodiments, the Hilbert transformer module 620 uses the left leg component 616 rather than the right leg component 618 to generate the left leg component 622 and the right leg component 624. This results in the right leg component 624 having a phase and delay relationship to the high frequency component 608, which is created by the all-pass filter (e.g., and no delay) of the Hilbert transformer 614 and the delay and all-pass filter of the Hilbert transformer module 620.

[0078] FIG. 8 illustrates a frequency plot produced by driving an HPSM module (as described in FIG. 6) with white noise, according to some embodiments, showing the output frequency response of the sum 802 of multiple channels (mid) and the difference 804 of multiple channels (side).

[0079] As shown in Figure 8, this filter does produce the desired perceptual suggestion in the region around 11 kHz, while also imparting additional mid- and side coloration at lower frequencies. In some embodiments, this can be corrected by applying a crossover network (such as the crossover network module 604 shown in Figure 6) to the input audio so that the HPSM module processes only audio data within the desired frequency range (e.g., high frequency content), or by directly removing the pole / zero pairs that correspond to that region of the spectral transform.

[0080] [Design using First Order Non-Orthogonal Rotation-Based Decorrelation (FNORD)] In some embodiments, a similar perceptual effect may be achieved using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network. Figure 9 is a block diagram illustrating a PSM module 900 implemented using an FNORD filter network, according to some embodiments. The PSM module 900 may correspond to the PSM module 102 illustrated in Figure 1, but provides for decorrelation of a mono channel into multiple channels and includes a magnitude response module 902, an all-pass filter configuration module 904, and an all-pass filter module 906. The PSM module 900 processes a mono input channel x(t) 912 to generate a channel y(t) that is provided to a speaker 910a. a (t), and channel y provided to speaker 910b (which may correspond to left speaker 112 and right speaker 114 illustrated in FIG. 1). b9 illustrates the PSM module 900 as including a magnitude response module 902 and a filter configuration module 904 in addition to an all-pass filter module 906, in some embodiments, the PSM module 900 may include the all-pass filter module 906 with the magnitude response module 902 and / or the filter configuration module 904 implemented separately from the PSM module 900.

[0081] The amplitude response module 902 determines a target amplitude response that defines one or more spatial cues to be encoded into output channel y(t) (e.g., the mid and side components of output channel y(t)). The target amplitude response is defined by the relationship between the amplitude and frequency values ​​of the channel (e.g., the mid and side components of the channel), such as amplitude as a function of frequency. In some embodiments, the target amplitude response defines one cu or spatial cues on the channel, which may include a target broadband attenuation, a target subband attenuation, a critical point, a filter characteristic, or a sound stage position. The amplitude response module 902 may receive data 914 and mono channel x(t) 912 and use these inputs to determine the target amplitude response. The data 914 may include information such as characteristics of the spatial cues to be encoded, characteristics of the presentation equipment (e.g., one or more speakers), the expected content of the audio data, or the listener's perceptual capabilities in a scene. In some embodiments, mono channel x(t) 912 may correspond to audio input 120 illustrated in Figure 1, or a portion of the audio input (e.g., high-frequency components of the input audio, such as high-frequency components 608 of input audio 602 illustrated in Figure 6). In embodiments in which mono channel x(t) 912 corresponds to a portion of the audio input, output channel y(t) may be combined with a channel corresponding to the remaining portion of the audio input (e.g., having low-frequency components, as illustrated in Figure 1) to generate a combined output channel.

[0082] Target broadband attenuation is a specification of attenuation across all frequencies. Target subband attenuation is a specification of amplitude for a frequency range defined by a subband. A target amplitude response may include one or more target subband attenuation values, each for a different subband.

[0083] A critical point is a specification of the curvature of a filter's target magnitude response and is described as the frequency value at which the gain for one of the output channels (e.g., the side component of an output channel) reaches a predefined value, such as -3 dB or -∞ dB. The location of this point can have an overall effect on the curvature of the target magnitude response. One example of a critical point corresponds to the frequency at which the target magnitude response reaches -∞ dB. This critical point is a null point because the behavior of the target magnitude response is to null signals at frequencies close to this point. Another example of a critical point corresponds to the frequency at which the target magnitude response reaches -3 dB. This critical point is a crossover point because the behavior of the target magnitude responses for the sum and difference channels (e.g., the mid and side components of a channel) intersect at this point.

[0084] Filter characteristics are parameters that specify how to filter out the mid and side components of a channel. Examples of filter characteristics include high-pass, low-pass, band-pass, or band-stop characteristics. The filter characteristics describe the shape of the resulting sum as if it were the result of equalization filtering. Equalization filtering can be described in terms of which frequencies can pass the filter or which frequencies are rejected. Thus, a low-pass characteristic allows frequencies below an inflection point to pass and attenuates frequencies above the inflection point. A high-pass characteristic does the opposite by passing frequencies above the inflection point and attenuating frequencies below the inflection point. A band-pass characteristic allows frequencies in a band surrounding the inflection point to pass and attenuates other frequencies. A band-stop characteristic rejects frequencies in a band surrounding the inflection point and allows other frequencies to pass.

[0085] The target magnitude response may define multiple spatial implications to be encoded into the output channel y(t). For example, the target magnitude response may specify spatial implications specified by critical points and the filter characteristics of the mid- or side-components of an all-pass filter. In another example, the target magnitude response may specify spatial implications specified by target broadband attenuation, critical points, and filter characteristics. Although described as independent specifications, the specifications may be interdependent across most regions of the parameter space. This result may be due to phase nonlinearities in the system. To address this, additional, higher-level descriptors of the target magnitude response may be devised that are nonlinear functions of the target magnitude response parameters.

[0086] The filter configuration module 904 determines the characteristics of a single-input, multiple-output all-pass filter based on the target magnitude response received from the magnitude response module 902. In particular, the filter configuration module determines a transfer function of the all-pass filter based on the target magnitude response, and determines coefficients of the all-pass filter based on the transfer function. The all-pass filter is a decorrelation filter that encodes the spatial implications described in terms of the target magnitude response and is applied to the mono input channel x(t) to generate the output channel y a (t) and y b Generate (t).

[0087] The all-pass filter may include various configurations and parameters based on the spatial implication and / or constraints specified by the target amplitude response. A filter having a target amplitude response for the encoded spatial implication may be colorless, preserving, for example, the spectral content (e.g., overall) of the individual output channels (e.g., left / right output channels). This filter is then used to encode the elevation implication by embedding coloration in the mid / side space in the form of frequency-dependent amplitude implication while preserving the spectral content of the left and right signals. Because the filter is colorless, monaural content can be placed at a specific location in the sound stage (e.g., as specified by the target elevation angle), and the spatial placement of the sound is decoupled from its overall coloration.

[0088] 10A and 10B are block diagrams illustrating an example of a PSM module based on a first-order non-orthogonal rotation-based decorrelation (FNORD) technique, according to some embodiments. FIG. 10A shows a detailed view of the PSM module 900, according to some embodiments, while FIG. 10B provides a more detailed view of the wideband phase rotator 1004 within the all-pass filter module 906 of the PSM module 900, according to some embodiments.

[0089] As in FIG. 10A, the all-pass filter module 906 receives a monophonic input audio signal x(t) 912, a rotation control parameter θ bf 1048, and the linear coefficient β bf 1050 receives the input speech signal x(t) 912 and the rotation control parameters θ bf 1048 is utilized by a broadband phase rotator 1004, which has a rotation control parameter θ bf1048 to process the input audio signal 912 to generate a left wideband rotated component 1020 and a right wideband rotated component 1022. The left wideband rotated component 1020 is then provided to a narrow-band phase rotator 1024 for further processing, while the right wideband rotated component 1022 is provided to output channel y of the PSM module 900 according to some embodiments. b (t) (e.g., as the right output channel). Narrowband phase rotator 1024 receives left wideband rotation component 1020 from wideband phase rotator 1004 and first order coefficient β bf 1050 to generate a narrowband rotation component 1028, which is then transmitted to an output channel y (e.g., the left output channel) of the PSM module 900. a (t) is provided.

[0090] According to some embodiments, the control data 914 for configuring the magnitude response module 902 includes a critical point f c 1038, filter characteristics θ bf 1036, and sound stage position Γ 1040. This data is provided to the PSM module 900 via the magnitude response module 902, which determines the critical points coc 1044 (in radians), the filter characteristic θ bf 1042, and a secondary term φ 1046. In some embodiments, the magnitude response module 902 determines a parameter of the control data 914 (e.g., a critical point f c 1038, filter characteristics θ bf 1036, and / or sound stage position Γ 1040). bf 1042 is the filter characteristic θ bf1036. These mid-representations 1042, 1044, and 1046 are provided to the filter construction module 904, which then constructs at least the first-order coefficients β bf 1050 and rotation control parameters 1048. bf 1050 is provided to the all-pass filter module 906 via a first order all-pass filter 1026. In some embodiments, the rotation control parameter θ bf 1048 is the filter characteristic θ bf 1036 and 1042, while in other embodiments this parameter may be scaled for convenience. For example, in some embodiments, the filter characteristic is associated with a parameter range (e.g., 0 to 0.5) with a meaningful center point, and the rotation control parameter is scaled relative to the filter characteristic to change the parameter range, e.g., to 0 to 1. In some embodiments, the filter characteristic is scaled linearly (e.g., to maintain increased resolution at the poles compared to the center point), while in other embodiments, a non-linear mapping may be used (e.g., to increase numerical resolution relative to the center point). In the following equation, the rotation control parameter θ bf is treated as unscaled, but it is understood that the same principles can be applied when the rotation control parameters are scaled. bf 1048 is provided to the all-pass filter module 906 via a wideband phase rotator 1004 .

[0091] 10B details an example implementation of the wideband phase rotator 1004, according to some embodiments. The wideband phase rotator 1004 receives the monophonic input audio signal x(t) 912 and the rotation control parameter θ bf1048. The input speech signal x(t) 912 is first processed by a Hilbert transformer module 1006 to generate a left leg component 1008 and a right leg component 1010. The Hilbert transformer module 1006 may be implemented using the configuration shown in FIG. 7, although it will be understood that other implementations of the Hilbert transformer module 1006 may be used in other embodiments. The left leg component 1008 and the right leg component 1010 are provided to a 2D orthogonal rotation module 1012. The left leg component 1008 is also provided to the output of the wideband phase rotator 1004 as a right wideband rotation component 1022. Because the wideband phase rotator 1004 is configured to rotate the left and right leg signals relative to each other, one way to achieve this in some embodiments is to hold the left leg component 1008 constant as the right wideband rotation component 1022 and rotate the left and right leg components to form the left wideband rotation component 1020.

[0092] In addition to the left leg component 1008 and the right leg component 1010, the 2D orthogonal rotation module 1012 receives a rotation control parameter θ from the filter configuration module 904 according to some embodiments, as shown in FIG. 10A. bf 1048. The 2D orthogonal rotation module 1012 uses this data to generate a left rotation component 1014 and a right rotation component 1016. The projection module 1018 then receives the left rotation component 1014 and the right rotation component 1016, which are combined (e.g., added) to form a left wideband rotation component 1020. As shown in FIG. 10A, the wideband phase rotator 1004 rotates the left output channel y of the PSM module. a (t) is the narrowband rotation component 1028, which is then fed to the right output channel y b10A and 10B , narrowband phase rotator 1028 outputs left wideband rotation component 1020 to narrowband phase rotator 1024 to generate right wideband rotation component 1022 as (t) (which either bypasses narrowband phase rotator 1024 or passes through it unchanged). In other embodiments, narrowband rotation component 1028 and left leg component 1008 (which serves as right wideband rotation component 1022 in the embodiment shown in FIGS. 10A and 10B ) are instead fed to right output channel y b (t) and the left output channel y a (t) are each mapped to (t).

[0093] In some embodiments, the PSM module 900 may be formally described by the following equation (8):

[0094]

number

[0095] In some embodiments, this single-input multiple-output all-pass filter is composed of several parts, each of which will be described in turn. According to some embodiments, these parts include: A f , A b , and H2.

[0096] According to some embodiments, A f may correspond to narrowband phase rotator 1024 in FIG. 10A. f is a first order all-pass filter with one channel output assuming the form of equation (9).

[0097]

number

[0098] Here, β f are the coefficients of the filter, which range from -1 to +1. The second output of the filter may simply pass the input unchanged. Therefore, according to some embodiments, filter A fThe implementation of can be defined by equation (10):

[0099]

number

[0100] A f The transfer function of is the differential phase shift from one output to the other.

[0101]

number

[0102] This differential phase shift is a function of the center frequency (radian frequency) ω, as defined by equation (11).

[0103]

number

[0104] where the target amplitude response is determined by changing θ in either (5) or (6) depending on whether the response should be placed in the mid (5) or side (6).

[0105]

number

[0106] The frequency f at which the total gain αf=3 dB can be used as the critical point for adjustment. c is defined by the following equation:

[0107]

number

[0108] and

[0109]

number

[0110] By normalizing the target amplitude response to 0 dB, this critical point is determined by the parameter f c This corresponds to the -3 dB point. f The outputs of are subscripted to identify only the output of the first channel that is used, according to some embodiments.

[0111] In equation (8), A b is a single-input multiple-output all-pass filter, which may correspond to wideband phase rotator 1004 in FIG. 10A. b can be formally defined as in equation (14).

[0112]

number

[0113] where H2(x(t)) is the discrete form of the filter, implemented using a pair of orthogonal all-pass filters, defined using a continuous-time prototype according to equation (15), i.e.,

[0114]

number

[0115] In some embodiments, an all-pass filter

[0116]

number

[0117] constrains a 90 degree phase relationship between the two output signals, and a unity magnitude relationship between the input signal and both output signals, but does not necessarily guarantee a particular phase relationship between the input (mono) signal and either of the two (stereo) output signals.

[0118]

number

[0119] The discrete form of

[0120]

number

[0121] and is defined by its operation on the mono signal x(t). The result is a two-dimensional vector as defined by equation (16):

[0122]

number

[0123] Discrete single-input multiple-output all-pass filters

[0124]

number

[0125] corresponds to Hilbert transformer module 1006 in Figure 10B, and may also correspond to Hilbert transformer module 700 in Figure 7, according to some embodiments. In equation (14), θ determines the angle of rotation of the first output relative to the second output of Ab, according to some embodiments.

[0126] Finally, the parameter A supplied to the complete system in Eq. (8) bfmay be determined according to some embodiments as follows: These parameters include the rotation control parameter θ in FIG. bf 1048 and the first coefficient β bf β bf and θ bf In some embodiments, β bf is the center radian frequency ω c can be determined from

[0127]

number

[0128] where ω c is calculated using equation (12) to find the desired center frequency f c In FIG. 10A, ω c is the critical point ω c Corresponding to 1044, f c is the critical point f c 1038, the operation of equation (17) is partially performed within the filter construction module 904, resulting in the first order coefficient β bf In some embodiments, the quadratic term φ 1046 is calculated via equation (18) as bf and can be derived from the Boolean sound stage position parameter Γ, i.e.,

[0129]

number

[0130] This quadratic term φ 1046 is provided to the filter configuration module 904 by the magnitude response module 902 in FIG. 10A.

[0131] In some embodiments, a high-level parameter f c , θ bf , and Γ may be sufficient to intuitively and conveniently tune this system. According to such an embodiment, the center frequency fc determines the inflection point in Hz where the target magnitude response asymptotically approaches -∞ dB. The parameter θ bf Therefore, the inflection point f c θ bf < 1 / 4, the characteristic is low-pass with a null at fc, and the spectral slope of the target amplitude function is θ bf As θ increases, it smoothly interpolates from favorable low frequencies to flat. bf If <1 / 2, then θ bf As increases, the characteristic c Smoothly interpolate from flat to high-pass with a null at θ bf At the point f = 1 / 4, the target amplitude function is purely band-reject, f c The parameter Γ is null at f c and θ bf is a Boolean value that places the target amplitude function determined by either the mid-channel (i.e., L+R) or the side-channel (i.e., LR) Due to the global constraints on both outputs to the filter network, the effect of Γ is to switch between complementary target amplitude responses.

[0132] In some embodiments, to achieve an amplitude response for a 60 degree vertical angle, the FNORD filter network described above is configured with a parameter f c =11kHz, θ bf = 0.13, and Γ = 1. Figure 11 illustrates a frequency response graph showing the output frequency response of an FNORD filter network configured to achieve an amplitude response for a 60 degree vertical angle, according to some embodiments. Figure 11 illustrates the output frequency response for a mid component 1110 and a side component 1120, where the FNORD filter network is driven by white noise. In some embodiments, the filter parameters f c , θ bf , and / or Γ are selected based on an analysis of the HRTF-based elevation angle suggestion at the desired angle.

[0133] In some embodiments, the PSM module 900 uses frequency-domain specifications for the all-pass filters. For example, in some cases, more complex spatial cues, such as those sampled from anthropometric datasets, may be required. Within certain limitations, the techniques described above may be used to embed arbitrary cues into the phase differences of the audio stream based on a magnitude-frequency-domain representation of the cues. For example, the filter configuration module 904 may derive K phase angles from a vectorized target magnitude response of K narrow-band attenuation coefficients in the mid or side using an equation in the form of Equation (5) or (6).

[0134]

number

[0135] The vectorized transfer function of

[0136]

number

[0137] The phase angle vector θ generates a Finite Impulse Response filter defined by equation (19):

[0138]

number

[0139] where DFT -1 is the inverse discrete Fourier transform (idft) and

[0140]

number

[0141] Then, a vector of 2(K-1) IFR filter coefficients Bn(θ) is applied to x(t) as defined by equation (20), i.e.,

[0142]

number

[0143] Here,

[0144]

number

[0145] represents a convolution operation.

[0146] To replicate the effect from the previous example and achieve the target amplitude response corresponding to a 60 degree height cue, the observed HRIR

[0147]

number

[0148] is sampled and applied to a DFT of length 2(K-1), resulting in

[0149]

number

[0150] , which can be used to generate the target magnitude response vector

[0151]

number

[0152] can be used to determine

[0153]

number

[0154] where:

[0155]

number

[0156] and

[0157]

number

[0158] are operations that return the real and imaginary components of a complex number, respectively, and all operations are applied vector component-wise. This target magnitude response, inserted into either the mid or side, is applied to one of equations (5) or (6) to obtain a vector of K phase angles:

[0159]

number

[0160] can be determined, from which an FIR filter B can be derived. This filter is then inserted into equation (19) to derive a single-input multiple-output all-pass filter.

[0161] While equations (19) and (20) provide an effective means for constraining the target magnitude response, their implementation often relies on relatively high-order FR filters resulting from an inverse DFT operation, which may not be suitable for resource-constrained systems. In such cases, a lower-order infinite impulse response (IIR) implementation may be used, as described in connection with equation (8).

[0162] The all-pass filter module 906, as configured by the filter configuration module 904, applies an all-pass filter to the mono channel x(t) to produce the output channel y a (t) and y b The application of the all-pass filter to the channel x(t) may be performed as defined by equations (8), (20), or as depicted in FIG. 9 or FIG. 10A. The all-pass filter module 906 applies the all-pass filter to the channel y(t). a (t) to speaker 910a, channel y b Each output channel is provided to a respective speaker, such as speaker 910b, and so on. Although not shown in FIG. 9, output channel y a (t) and y b It will be understood that (t) may be provided to speakers 910a and 910b via one or more intervening components (e.g., component processor module 106, crosstalk processor module 110, and / or L / RM / S converter module 104 and M / SL / R converter module 108, as shown in FIG. 1).

[0163] [Hypermid Processing] In some embodiments, PSM processing may be performed on a target portion of a received audio signal, such as a mid component of the audio signal or a hyper-mid component of the audio signal. FIG. 12 is a block diagram illustrating an audio processing system 1200 according to one or more embodiments. System 1200 generates a hyper-mid component to isolate a target portion of the audio signal (e.g., voicing) and performs PSM processing on the hyper-mid component to spatially shift the target portion. Some embodiments of system 1200 have different components than those described herein. Similarly, in some cases, functionality may be distributed among components in a manner different from that described herein.

[0164] The system 1200 includes an L / RM / S converter module 1206 , a quadrature component generator module 1212 , a quadrature component processor module 1214 that includes the PSM module 102 , and a crosstalk processor module 1224 .

[0165] The L / RM / S converter module 1206 receives the left channel 1202 and the right channel 1204 and generates a mid component 1208 and a side component 1210 from the channels 1202 and 1204. The description regarding the L / RM / S converter module 104 may be applicable to the L / RM / S converter module 1206.

[0166] The quadrature component generator module 1212 processes the mid component 1208 and the side component 1210 to generate at least one of a hypermid component M1, a hyperside component S1, a residual mid component M2, and a residual side component S2. The hypermid component M1 is the spectral energy of the mid component 1208 with the spectral energy of the side component 1210 removed. The hyperside component S1 is the spectral energy of the mid component 1208 with the spectral energy of the side component 1210 removed. The residual mid component M2 is the spectral energy of the hypermid component M1 with the spectral energy of the mid component 1208 removed. The mid component M2 is the spectral energy of the hypermid component M1 with the spectral energy of the mid component 1208 removed. The residual side component S2 is the spectral energy of the hypermid component 1210 with the spectral energy of the hyperside component S1 removed. The system 1200 processes at least one of the hyper-mid component M1, the hyper-side component S1, the residual mid component M2, and the residual side component S2 to generate a left channel 1242 and a right output channel 1244. The quadrature component generator module 1212 is further described with reference to Figures 13A, 13B, and 13C.

[0167] The quadrature component processor module 1214 processes one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and / or residual side component S2, converting the processed components into a processed left component 1220 and a processed right component 1222. The description of the component processor module 106 may be applicable to the quadrature component processor module 1214, except that the processing is performed on the hypermid component M1, hyperside component S1, residual mid component M2, and / or residual side component rather than on the mid and side components. For example, the processing on the components M1, M2, S1, and S2 may include various types of spatially sensitive processing (e.g., amplitude or delay-based panning, binaural processing, etc.), single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., compression, expansion, limiting, etc.), single-band or multi-band gain or delay stages, addition of audio effects, or other types of processing. In some embodiments, the quadrature component processor module 1214 performs subband spatial processing and / or crosstalk compensation processing using the hyper-mid component M1, the hyper-side component S1, the residual mid component M2, and / or the residual side component S2. The quadrature component processor module 1214 may further include an L / RM / S converter for converting the components M1, S2, S1, and S2 into a processed left component 1220 and a processed right component 1222.

[0168] The quadrature component processor module 1214 further includes a PSM module 102, which may operate on one or more of the hyper-mid component M1, the hyper-side component S1, the residual mid component M2, and / or the residual side component S2. For example, the PSM module 102 may receive the hyper-mid component M1 as an input and generate spatially shifted left and right channels. The hyper-mid component M1 may include, for example, an isolated portion of the audio signal representing voicing and therefore may be selected for HPSM processing. The left channel generated by the PSM module 102 is used to generate the processed left component 1020, and the right channel generated by the PSM module 102 is used to generate the processed right component 1222. The quadrature component processor module 1214 is further described with reference to FIG. 12.

[0169] The crosstalk processor module 1224 receives the processed left component 1220 and the processed right component 1222 and performs crosstalk processing thereon. The crosstalk processor module 1224 outputs a left channel 1242 and a right channel 1244. Descriptions related to the crosstalk processor module 1224 may be applicable to the crosstalk processor module 1224. In some embodiments, crosstalk processing (e.g., simulation or cancellation, etc.) may be performed before quadrature component processing, such as before conversion of the left channel 1202 and the right channel 1204 into mid and side components. The left channel 1242 may be provided to the left speaker 112, and the right channel 1244 may be provided to the right speaker 114.

[0170] [Example of quadrature component generator] 13A through 13C are block diagrams illustrating quadrature component generator modules 1313, 1323, and 1343, respectively, according to one or more embodiments. Quadrature component generator modules 1313, 1323, and 1343 are examples of quadrature component generator module 1212. Some embodiments of modules 1313, 1323, and 1343 have different components than described herein. Similarly, in some cases, functionality may be distributed among components in a different manner than described herein.

[0171] 13A, the quadrature component generator module 1313 includes a subtraction unit 1305, a subtraction unit 1309, a subtraction unit 1315, and a subtraction unit 1319. As described above, the quadrature component generator module 1313 receives the mid component 1208 and the side component 1210 and outputs one or more of a hyper-mid component M1, a hyper-side component S1, a residual mid component M2, and a residual side component S2.

[0172] The subtraction unit 1305 removes the spectral energy of the side component 1210 from the spectral energy of the mid component 1208 to generate the hyper-mid component M1. For example, the subtraction unit 1305 subtracts the amplitude of the side component 1210 in the frequency domain from the amplitude of the mid component 1208 in the frequency domain, leaving only the phase, to generate the hyper-mid component M1. The subtraction in the frequency domain may be performed using a Fourier transform on the time-domain signals to generate signals in the frequency domain, and then subtracting the signals in the frequency domain. In another example, the subtraction in the frequency domain may be performed in other ways, such as using a wavelet transform instead of a Fourier transform. The subtraction unit 1309 removes the spectral energy of the hyper-mid component M1 from the spectral energy of the mid component 1208 to generate the residual mid component M2. For example, subtraction unit 1309 subtracts the amplitude of hyper-mid component M1 in the frequency domain from the amplitude of mid component 1208 in the frequency domain, while retaining only the phase, to generate residual mid component M2. While subtracting side from mid in the time domain yields the right channel of the original signal, the above operation in the frequency domain separates and distinguishes between the portion of the spectral energy of the mid component (called M1 or hyper-mid) that differs from the spectral energy of the mid component and the portion of the spectral energy of the side component (called M2 or residual mid) that is the same as the spectral energy of the side component.

[0173] In some embodiments, if subtracting the spectral energy of the side component 1210 from the spectral energy of the mid component 1006 results in a negative value for the hyper-mid component M1 (e.g., for one or more bins in the frequency domain), additional processing may be used. In some embodiments, if subtracting the spectral energy of the side component 1210 from the spectral energy of the mid component 1208 results in a negative value, the hyper-mid component M1 is clamped to a value of zero. In some embodiments, the hyper-mid component M1 is wrapped around to its minimum value by taking the absolute value of the negative value as the value of the hyper-mid component M1. Other types of processing may be used if subtracting the spectral energy of the side component 1210 from the spectral energy of the mid component 1208 results in a negative value for M1. Similar additional processing may be used if the result of the subtraction to generate the hyper-side component S1, residual side component S2, or residual mid component M2 is negative, such as clamping at zero, wrapping around to a minimum value, or other processing. Fixing the hypermid component M1 to 0 provides spectral orthogonality between M1 and both side components when the subtraction results in a negative value. Similarly, fixing the hyperside component S1 to 0 provides spectral orthogonality between S1 and both mid components when the subtraction results in a negative value. By creating orthogonality between the hypermid and side components and their corresponding appropriate mid / side components (i.e., side components relative to hypermid, mid components relative to hyperside, etc.), the resulting residual mid component M2 and residual side component S2 contain spectral energy that is not orthogonal to (i.e., common to) their corresponding appropriate mid / side components. That is, applying a 0 fixation to the hypermid and deriving the residual mid using the M1 component results in a hypermid component that has no spectral energy in common with the side component and a residual mid component that has completely common spectral energy with the side component.If the hyperside is fixed to 0, the same relationship applies to the hyperside and residual side. When applying frequency-domain processing, there is typically a trade-off in resolution between frequency and time information. As frequency resolution increases (i.e., as the FFT window size and the number of frequency bins increase), time resolution decreases, and vice versa. Because the spectral subtraction described above occurs on a frequency bin-by-frequency basis, in certain situations, such as when removing voicing energy from hypermid components, it may be more preferable to have a large FFT window size (e.g., 8192 samples, resulting in 4096 frequency bins given a real-valued input signal). In other situations, higher time resolution may be required, resulting in lower overall latency and lower frequency resolution (e.g., 512 sample FFT window size, resulting in 256 frequency bins given a real-valued input signal). In the latter case, the low frequency resolution of the mid and side components may produce audible spectral artifacts when they are subtracted from each other to derive the hypermid and hyperside components S1, because the spectral energy in each frequency bin is an average representation of the energy across too wide a frequency range. In this case, taking the absolute value of the difference between mid and side when deriving hypermid M1 or hyperside S1 can help reduce perceptual artifacts by allowing for deviations from true orthogonality in the components for each frequency bin. In addition to or instead of wrapping around zero, a coefficient can be applied to the subtracted value to scale it between 0 and 1, providing a way to interpolate between perfect orthogonality of the hypermid and residual mid / side components at one extreme (i.e., when the value is 1) and identical hypermid M1 and hyperside S1 to their corresponding original mid and side components at the other extreme (i.e., when the value is 0).

[0174] The subtraction unit 1315 removes the spectral energy of the mid component 1208 in the frequency domain from the spectral energy of the side component 1210 in the frequency domain, while retaining only the phase, to generate the hyper side component S1. For example, the subtraction unit 1315 subtracts the amplitude of the mid component 1208 in the frequency domain from the amplitude of the side component 1210 in the frequency domain, while retaining only the phase, to generate the hyper side component S1. The subtraction unit 1319 removes the spectral energy of the hyper side component S1 from the spectral energy of the side component 1210, to generate the residual side component S2. For example, the subtraction unit 1319 subtracts the amplitude of the hyper side component S1 in the frequency domain from the amplitude of the side component 1210 in the frequency domain, while retaining the phase, to generate the residual side component S2.

[0175] 5B, quadrature component generator module 1323 is similar to quadrature component generator module 1313 in that quadrature component generator module 1323 receives mid component 1006 and side component 1210 and generates hyper-mid component M1, residual mid component M2, hyper-side component S1, and residual side component S2. Quadrature component generator module 1323 differs from quadrature generator module 1313 by generating hyper-mid component M1 and hyper-side component S1 in the frequency domain and then transforming these components back to the time domain to generate residual mid component M2 and residual side component S2. The quadrature component generator module 1323 includes a forward FFT unit 1320, a band pass unit 1322, a subtraction unit 1324, a hyperside processor 1325, an inverse FFT unit 1326, a time delay unit 1328, a subtraction unit 1330, a forward FFT unit 1332, a band pass unit 1334, a subtraction unit 1336, a hyperside processor 1337, an inverse FFT unit 1340, a time delay unit 1342, and a subtraction unit 1344.

[0176] A forward fast Fourier transform (FFT) unit 1320 applies a forward FFT to the mid component 1208, converting it to the frequency domain. The converted mid component 1208 includes amplitude and phase. A bandpass unit 1322 applies a bandpass filter to the frequency-domain mid component 1208, where the bandpass filter specifies frequencies in the hyper-mid component M1. For example, to isolate the typical human vocal range, the bandpass filter may specify frequencies between 300 Hz and 8000 Hz. In another example, to remove audio content associated with the typical human vocal range, the bandpass filter may preserve low frequencies (e.g., generated by a bass guitar or drums) and high frequencies (e.g., generated by a cymbal) in the hyper-mid component M1. In other embodiments, the quadrature component generator module 1323 applies various other filters to the frequency-domain mid component 1208 in addition to and / or instead of the bandpass filter applied by the bandpass unit 1322. In some embodiments, the quadrature component generator module 1323 does not include a bandpass unit 1322 and does not apply any filter to the frequency-domain mid component 1208. In the frequency domain, a subtraction unit 1324 subtracts the side component 1210 from the filtered mid component to generate the hyper-mid component M1. In other embodiments, in addition to and / or instead of subsequent processing applied to the hyper-mid component M1, such as that performed by a quadrature component processor module (e.g., the quadrature component processor module in FIG. 12), the quadrature component generator module 1323 applies various audio enhancements to the frequency-domain hyper-mid component M1. The hyper-mid processor 1325 performs processing on the hyper-mid component M1 in the frequency domain before converting it to the time domain. This processing may include subband spatial processing and / or crosstalk compensation processing.In some embodiments, the hypermid processor 1325 performs processing on the hypermid component M1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1326 applies an inverse FFT to the hypermid component M1 and transforms it back into the time domain. The hypermid component M1 in the frequency domain includes the amplitude of M1 and the phase of the mid component 1208, which the inverse FFT unit 1326 transforms into the time domain. The time delay unit 1328 applies a time delay to the mid component 1208 so that the mid component 1208 and the hypermid component M1 arrive at the subtraction unit 1330 simultaneously. The subtraction unit 1330 subtracts the hypermid component M1 in the time domain from the time-delayed mid component 1208 in the time domain to generate a residual mid component M2. In this example, the spectral energy of the hypermid component M1 is removed from the spectral energy of the mid component 1208 using processing in the time domain.

[0177] The forward FFT unit 1332 applies a forward FFT to the side component 1210, transforming it into the frequency domain. The frequency-domain transformed side component 1210 includes amplitude and phase. The bandpass unit 1334 applies a bandpass filter to the frequency-domain side component 1210. The bandpass filter specifies the frequency in the hyper-side component S1. In other embodiments, the quadrature component generator module 1323 applies various other filters to the frequency-domain side component 1210 in addition to and / or instead of a bandpass filter. In the frequency domain, the subtraction unit 1336 subtracts the mid component 1208 from the filtered side component 1210 to generate the hyper-side component S1. In other embodiments, the quadrature component generator module 1323 applies various audio enhancements to the frequency-domain hyper-side component S1 in addition to and / or instead of subsequent processing applied to the hyper-side component S1, such as performed by a quadrature component processor (e.g., quadrature component processor module 1214). The hyperside processor 1337 performs processing on the hyperside component S1 in the frequency domain prior to its transformation to the time domain. This processing may include subband spatial processing and / or crosstalk compensation processing. In some embodiments, the hyperside processor 1337 performs processing on the hyperside component S1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1340 applies an inverse FFT to the hyperside component S1 in the frequency domain to generate the hyperside component S1 in the time domain. The hyperside component S1 in the frequency domain includes the amplitude S1 and phase of the side component 1210, and the inverse FFT unit 1326 transforms the side component 1210 to the time domain. The time delay unit 1342 time delays the side component 1210 so that the side component 1210 arrives at the subtraction unit 1344 at the same time as the hyperside component S1. Subsequently, a subtraction unit 1344 subtracts the hyper side component in the time domain S1 from the time-delayed side component in the time domain 1210 to generate a residual side component S2.In this example, the spectral energy of the hyper side component S1 is removed from the spectral energy of the side component 1210 using processing in the time domain.

[0178] In some embodiments, the hypermid processor 1325 and the hyperside processor 1337 may be omitted if the processing performed by these components is performed by the orthogonal component processor module 1214.

[0179] In FIG. 13C, the quadrature component generator module 1343 is similar to the quadrature component generator module 1323 in that it receives the mid component 1208 and the side component 1210 and generates a hyper-mid component M1, a residual mid component M2, a hyper-side component S1, and a residual side component S2, except that the quadrature component generator module 1343 generates components M1, M2, S1, and S2 in the frequency domain and then converts these components to the time domain. The quadrature component generator module 1343 includes a forward FFT unit 1347, a band pass unit 1349, a subtraction unit 1351, a hyper-mid processor 1352, a subtraction unit 1353, a residual mid processor 1354, an inverse FFT unit 1355, an inverse FFT unit 1355, an inverse FFT unit 1357, a forward FFT unit 1361, a band pass unit 1363, a subtraction unit 1365, a hyper-side processor 1366, a subtraction unit 1367, a residual side processor 1368, an inverse FFT unit 1369, and an inverse FFT unit 1371.

[0180] The forward FFT unit 1347 applies a forward FFT to the mid component 1208, transforming it into the frequency domain. The mid component 1208 transformed into the frequency domain includes amplitude and phase. The forward FFT unit 1361 applies a forward FFT to the side component 1210, transforming it into the frequency domain. The side component 1210 transformed into the frequency domain includes amplitude and phase. The bandpass unit 1349 applies a bandpass filter to the frequency domain mid component 1208, which specifies the frequency of the hyper-mid component M1. In some embodiments, the quadrature component generator module 1343 applies various other filters to the frequency domain mid component 1208 in addition to and / or instead of the bandpass filter. The subtraction unit 1351 subtracts the frequency domain side component 1210 from the frequency domain mid component 1208 to generate the hyper-mid component M1 in the frequency domain. The hyper-mid processor 1352 performs processing on the hyper-mid component M1 in the frequency domain before transforming it to the time domain. In some embodiments, the hyper-mid processor 1352 performs subband spatial processing and / or crosstalk compensation processing. In some embodiments, the hyper-mid processor 1352 performs processing on the hyper-mid component M1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1357 applies an inverse FFT to the hyper-mid component M1 and transforms it back to the time domain. The hyper-mid component M1 in the frequency domain includes the amplitude M1 and phase of the mid component 1208, which the inverse FFT unit 1357 transforms back to the time domain. The subtraction unit 1353 subtracts the hyper-mid component M1 from the mid component 1208 in the frequency domain to generate a residual mid component M2. The residual mid processor 1354 performs processing on the residual mid component M2 in the frequency domain before transforming it to the time domain. In some embodiments, the residual mid processor 1354 performs subband spatial processing and / or crosstalk compensation processing on the residual mid component M2.In some embodiments, the residual mid processor 1354 performs processing on the residual mid component M2 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1355 applies an inverse FFT to transform the residual mid component M2 into the time domain. The residual mid component M2 in the frequency domain includes the amplitude M2 ​​and phase of the mid component 1208, which the inverse FFT unit 1355 transforms into the time domain.

[0181] The bandpass unit 1363 applies a bandpass filter to the frequency-domain side component 1210. The bandpass filter specifies the frequencies in the hyperside component S1. In other embodiments, the quadrature component generator module 1343 applies various other filters to the frequency-domain side component 1210 in addition to and / or instead of a bandpass filter. In the frequency domain, the subtraction unit 1365 subtracts the mid component 1208 from the filtered side component 1210 to generate the hyperside component S1. The hyperside processor 1366 performs processing on the hyperside component S1 in the frequency domain prior to transformation to the time domain. In some embodiments, the hyperside processor 1366 performs subband spatial processing and / or crosstalk compensation processing on the hyperside component S1. In some embodiments, the hyperside processor 1366 performs processing on the hyperside component S1 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1371 applies an inverse FFT to transform the hyper-side component S1 back into the time domain. The hyper-side component S1 in the frequency domain includes the amplitude S1 and phase of the side component 1210, which the inverse FFT unit 1371 transforms into the time domain. The subtraction unit 1367 subtracts the hyper-side component S1 from the side component 1210 in the frequency domain to generate a residual side component S2. The residual side processor 1368 performs processing of the residual side component S2 in the frequency domain prior to transforming it to the time domain. In some embodiments, the residual side processor 1368 performs subband spatial processing and / or crosstalk compensation processing on the residual side component S2. In some embodiments, the residual side processor 1368 performs processing on the residual side component S2 instead of and / or in addition to processing that may be performed by the quadrature component processor module 1214. An inverse FFT unit 1369 applies an inverse FFT to the residual side component S2 to transform it into the time domain.The residual side component S2 in the frequency domain contains the amplitude S2 and phase of the side component 1210, which the inverse FFT unit 1369 transforms into the time domain.

[0182] In some embodiments, the hyper-mid processor 1352, hyper-side processor 1366, residual mid processor 1354, or residual side processor 1368 may be omitted if the processing performed by these components is performed by the orthogonal component processor module 1214.

[0183] [Example of quadrature component processor] 14A is a block diagram illustrating a quadrature component processor module 1417 according to one or more embodiments. The quadrature component processor module 1417 is an example of the quadrature component processor module 1412. Some embodiments of the module 1417 have different components than those described herein. Similarly, in some cases, functionality may be distributed among the components in a different manner than described herein.

[0184] The quadrature component processor module 1417 includes a component processor module 1420 , a PSM module 102 , a summing unit 1422 , an M / SL / R converter module 1424 , a summing unit 1426 , and a summation 1428 .

[0185] The component processor module 1420 performs similar processing to the component processor module 106, except that it uses the hypermid component M1, the hyperside component S1, the residual mid component M2, and / or the residual side component S2 instead of the mid component and the side component. For example, the component processor module 1420 performs subband spatial processing and / or crosstalk compensation processing on at least one of the hypermid component M1, the residual mid component M2, the hyperside component S1, and the residual side component S2. As a result of the subband spatial processing and / or crosstalk compensation by the component processor module 1420, the quadrature component processor module 1417 outputs at least one of processed M1, processed M2, processed S1, and processed S2. In some embodiments, one or more of the components M1, M2, S1, or S2 may bypass the component processor module 1420.

[0186] In some embodiments, the quadrature component processor module 1417 performs subband spatial processing and / or crosstalk compensation processing on at least one of the hyper-mid component M1, the residual mid component M2, the hyper-side component S1, and the residual side component S2 in the frequency domain. The quadrature component generator module 410 may provide the frequency-domain components M1, M2, S1, or S2 to the quadrature component processor module 1417 without performing an inverse FFT. After generating the processed M1, processed M2, and processed side components 1442, the quadrature component processor module 1417 may perform an inverse FFT to transform these components back to the time domain. In some embodiments, the quadrature component processor module 1417 performs an inverse FFT on the processed M1, processed M2, processed S1, and processed S1 to generate the processed side components 1446 in the time domain.

[0187] 15 and 16 show example components of the quadrature component processor module 1417. In some embodiments, the quadrature component processor module 1417 performs both subband spatial processing and crosstalk compensation processing. The processing performed by the quadrature component processor module 1417 is not limited to subband spatial processing or crosstalk compensation processing. Any type of spatial processing using mid-space / side-space may be performed by the quadrature component processor module 1417, such as by using hyper-mid components instead of mid components or hyper-side components instead of side components. Other types of processing may include gain application, amplitude or delay-based panning, binaural processing, reverberation, dynamic range processing such as compression and limiting, and other linear or nonlinear audio processing techniques and effects, ranging from chorus or flanging to machine learning-based approaches to vocal or instrument style transfer, transformation, or resynthesis, etc.

[0188] The PSM module 102 receives the processed M1 and applies PSM processing to spatially shift the processed M1, resulting in a left channel 1432 and a right channel 1434. Although the PSM module 102 is shown as being applied to the hypermid component M1, the PSM module may be applied to one or more of the components M1, M2, S1, or S2. In some embodiments, the components processed by the PSM module 102 bypass processing by the component processor module 1420. For example, the PSM module 102 may process the hypermid component M1 rather than the processed M1.

[0189] The summing unit 1422 adds the processed S1 to the processed S2 to generate the processed side component 1442. The M / SL / R transformer module 1424 uses the processed M2 and the processed side component 1442 to generate a processed left component 1444 and a processed right component 1446. In some embodiments, the processed left side component 1444 is generated based on the addition of the processed M2 and the processed side component 1442, and the processed right component 1446 is generated based on the difference of the processed M2 and the processed side component 1442. Other M / SL / R type transforms may be used to generate the processed left component 1444 and the processed right component 1446.

[0190] A summing unit 1426 adds a left channel 1432 from the PSM module 102 to the processed left component 1444 to generate a left channel 1452. A summing unit 1428 adds a right channel 1434 from the PSM module 102 to the processed right component 1446 to generate a right channel 1454. More generally, one or more left channels from the PSM module 102 may be added to a left component (e.g., generated using a hyper / residual component not processed by the PSM module 102) from the M / SL / R converter module 1424 to generate the left channel 1452, and one or more right channels from the PSM module 102 may be added to a right component (e.g., generated using a hyper / residual component not processed by the PSM module 102) from the M / SL / R converter module 1424 to generate the right channel 1454.

[0191] As such, the quadrature component processor module 1417 applies PSM processing to the hyper-mid component M1 of the audio signal as separated by the L / RM / S transformer module 1206 and the quadrature component generator module 1212. The PSM-enhanced stereo signal (including the left channel 1432 and the right channel 1434) can then be added to the residual left signal / residual right signal (e.g., the processed left component 1444 and the processed right component 1446 generated without the hyper-mid component). Other approaches for separating the components of the input signal used for PSM processing may be used in addition to or instead of this example, including machine learning-based source separation.

[0192] In some embodiments, the quadrature component processor module 1417 applies PSM processing to the mid component M of the audio signal instead of the hyper-mid component ML. FIG. 14B illustrates a block diagram showing the quadrature component processor module 1419 according to one or more embodiments. In some embodiments, the quadrature component processor module 1419 in FIG. 14B may be implemented as part of an audio processing system similar to the system 1200 shown in FIG. 12, but without the quadrature component generator module 1212, whereby the quadrature component processor module 1419 receives mid and side component signals (e.g., mid component 1208 and side component 1210) instead of the hyper-mid component, hyper-side component, residual mid component, and residual side component. In some embodiments, the quadrature component processor module 1410 includes a component processor module similar to component processor module 106 to generate processed mid and processed side components from the received mid and side components (not shown). The PSM module 102 receives a mid component M (or processed mid) and applies PSM processing to spatially shift the received mid signal to generate a PSM-processed left channel 1432 and a PSM-processed right channel 1434 in the mid signal, which are combined with a side component S (or processed side) by an M / SL / R converter module 1424 to generate a left channel 1452 and a right channel 1454. For example, as shown in FIG. 14B , the M / SL / R converter module 1424 generates the left channel 1452 as a sum of the PSM-processed left channel 1432 and the side component S using a summation unit 1460, and generates the right channel 1452 as a sum of the PSM-processed right channel 1434 and the side component S using a subtraction unit 1462. In other words, the M / SL / R converter module 1424 serves to mix the side signal (which, in a left-right context, is in the subspace defined by the left component, being the inverse of the right component) into a PSM-processed stereo signal in left-right space by combining the signal for the left channel and the signal inverse for the right channel.

[0193] [Example of a subband spatial processor] 15 is a block diagram illustrating a subband spatial processor module 1510 according to one or more embodiments. The subband spatial processor module 1510 is an example of a component of a component processor module 106 or 1520. The subband spatial processor module 1510 includes a mid EQ filter 1504(1), a mid EQ filter 1504(2), a mid EQ filter 1504(3), a mid EQ filter 1504(4), a side EQ filter 1506(1), a side EQ filter 1506(2), a side EQ filter 1506(3), and a side EQ filter 1506(4). Some embodiments of the subband spatial processor module 1510 have different components than those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.

[0194] The subband spatial processor module 1510 receives the non-spatial components Ym and the spatial components Ys and gain adjusts one or more subbands of these components to provide spatial enhancement. If the subband spatial processor module 1510 is part of the component processor module 1420, the non-spatial components Ym may be the hyper-mid components M1 or the residual mid components M2. The spatial components Ys may be the hyper-side components S1 or the residual side components S2. If the subband spatial processor module 1510 is part of the component processor module 106, the non-spatial components Ym may be the mid components 126 and the spatial components Ys may be the side components 128.

[0195] The subband spatial processor module 1510 receives the non-spatial component Ym and applies mid EQ filters 1504(1) through 1504(4) to different subbands of Ym to generate emphasized non-spatial component Em. The subband spatial processor module 1510 also receives the spatial component Ys and applies side EQ filters 1506(1) through 1506(4) to different subbands of Ys to generate emphasized spatial component Es. The subband filters may include various combinations of peaking filters, notch filters, low-pass filters, high-pass filters, low-emphasis filters, high-emphasis filters, band-pass filters, band-stop filters, and / or all-pass filters. The subband filters may also apply gain to each subband. More specifically, the subband spatial processor module 1510 includes a subband filter for each of the n frequency subbands in the non-spatial component Ym and a filter subband for each of the n subbands in the spatial component Ys. For n=4 subbands, for example, the subband spatial processor module 1510 includes a series of subband filters for the non-spatial component Ym, including a mid equalization (EQ) filter 1504(1) for subband(1), a mid equalization (EQ) filter 1504(2) for subband(2), a mid EQ filter 1504(3) for subband(3), and a mid EQ filter 1504(4) for subband(4). Each mid EQ filter 1504 applies filter extraction to a frequency subband portion of the non-spatial component Ym to produce an emphasized non-spatial component Em.

[0196] The subband spatial processor module 1510 further includes a series of subband filters for the frequency subbands of the spatial component Ys, including a side equalization (EQ) filter 1506(1) for subband(1), a side EQ filter 1506(2) for subband(2), a side EQ filter 1506(3) for subband(3), and a side EQ filter 1506(4) for subband(4). Each side EQ filter 1506 applies a filter extraction to a frequency subband portion of the spatial component Ys to produce an enhanced spatial component Es.

[0197] Each of the n frequency subbands in the non-spatial component Ym and the spatial component Ys may correspond to a frequency range. For example, frequency subband (1) may correspond to 0 to 300 Hz, frequency subband (2) may correspond to 300 to 510 Hz, frequency subband (3) may correspond to 510 to 2700 Hz, and frequency subband (4) may correspond to 2700 Hz to the Nyquist frequency. In some embodiments, each of the n frequency subbands is a set of integrated critical bands. The critical bands may be determined using a corpus of audio samples from a wide variety of musical genres. The long-term average energy ratio of the mid- to side-components across the 24 Bark scale critical bands is determined from the samples. Contiguous frequency bands with similar long-term average ratios are then grouped together to form a set of critical bands. The range of the frequency subbands and the number of frequency subbands may be adjustable.

[0198] In some embodiments, the subband spatial processor module 1510 processes the residual mid component M2 as the non-spatial component Ym and uses one of the side component, hyperside component S1, or residual side component S2 as the spatial component Ys.

[0199] In some embodiments, the subband spatial processor module 1510 processes one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and residual side component S2. The filters applied to each subband of these components may be different. The hypermid component M1 and residual mid component M2 may each be processed as described for the non-spatial component Ym. The hyperside component S1 and residual side component S2 may each be processed as described for the spatial component Ys.

[0200] [Example of a crosstalk compensation processor] 16 is a block diagram illustrating a crosstalk compensation processor module 1610 according to one or more embodiments. The crosstalk compensation processor module 1610 is an example of a component of the component processor module 106 or 1420. Some embodiments of the crosstalk compensation processor module 1610 have different components than those described herein. Similarly, in some cases, functionality may be distributed among the component components in a different manner than described herein.

[0201] The crosstalk compensation processor module 1610 includes a mid component processor 1620 and a side component processor 1630. The crosstalk compensation processor module 1610 receives the non-spatial component Ym and the spatial component Ys and applies filter extraction to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or subsequent) crosstalk processing. If the crosstalk compensation processor module 1610 is part of the component processor module 1420, the non-spatial component Ym may be the hyper-mid component M1 or the residual mid component M2. The spatial component Ys may be the hyper-side component S1 or the residual side component S2. If the crosstalk compensation processor module 1610 is part of the component processor module 106, the non-spatial component Ym may be the mid component 126, and the spatial component Ys may be the side component 128.

[0202] The crosstalk compensation processor module 1610 receives the non-spatial component Ym, and the mid component processor 1620 applies a set of filters to generate an enhanced non-spatial crosstalk compensation component Zm. The crosstalk compensation processor module 1610 also receives the spatial subband component Ys, and applies a set of filters in the side component processor 1630 to generate an enhanced spatial subband component Es. The mid component processor 1620 includes a plurality of filters 1640, such as m mid filters 1640(a), 1640(b) through 1640(m), where each of the m mid filters 1640 filters a non-spatial component Xm. m Therefore, the mid-component processor 1620 processes one of the m frequency bands in the non-spatial component X m In some embodiments, the mid filter 1640 generates the mid-crosstalk compensation channel Zm by processing the non-spatial component X with simulated crosstalk processing. mThe frequency response plot is constructed using the frequency response plot of the crosstalk signal. Furthermore, by analyzing the frequency response plot, any spectral imperfections, such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB), that occur as artifacts of crosstalk processing can be estimated. These artifacts are primarily caused by the addition of delayed and possibly inverted contralateral signals to corresponding ipsilateral signals in the crosstalk processing, thereby effectively introducing a comb filter-like frequency response to the final rendering result. A mid-crosstalk compensation channel Zm can be generated by the mid-component processor 1620 to compensate for the estimated peaks or troughs, with each of m frequency bands corresponding to a peak or trough. Specifically, based on the specific delay, filter extraction frequency, and gain applied to the crosstalk processing, the peaks or troughs in the frequency response will shift up or down, causing variable amplification and / or variable attenuation of energy in specific regions of the spectrum. Each of the mid-filters 1640 may be configured to adjust one or more of the peaks and troughs.

[0203] The side component processor 1630 includes multiple filters 1650, such as m side filters 1650(a), 1650(b), through 1650(m). The side component processor 1630 processes the spatial component Xs to generate the side crosstalk compensation channel Zs. In some embodiments, a frequency response plot of the spatial component X with crosstalk processing can be obtained through simulation. By analyzing the frequency response plot, any spectral imperfections, such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB), that occur as artifacts of the crosstalk processing can be estimated. The side crosstalk compensation channel Zs can be generated by the side component processor 1630 to compensate for the estimated peaks or troughs. Specifically, based on the specific delays, filter extraction frequencies, and gains applied to the crosstalk processing, the peaks or troughs in the frequency response shift up or down, causing variable amplification and / or variable attenuation of energy in specific regions of the spectrum. Each of the side filters 1650 can be configured to adjust one or more peaks or troughs. In some embodiments, the mid component processor 1620 and the side component processor 1630 may include different numbers of filters.

[0204] In some embodiments, the mid filter 1640 and the side filter 1650 may include biquadratic filters having a transfer function defined by equation (7). One way to implement such a filter is with a direct form I topology as defined by equation (22), i.e.,

[0205]

number

[0206] where X is the input vector and Y is the output. Other topologies can be used depending on their maximum word length and saturated summation operation. Biquads can then be used to implement second-order filters with real-valued inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and then transformed to discrete time by a bilinear transform. Furthermore, the resulting shift in center frequency and bandwidth can be compensated for using frequency warping.

[0207] For example, the peaking filter may have an S-plane transfer function defined by equation (23):

[0208]

number

[0209] where s is a complex variable, A is the peak amplitude, Q is the filter "quality", and the digital filter coefficients are defined by the following equation (24):

[0210]

number

[0211] where ω is the filter center frequency in radians,

[0212]

number

[0213] Furthermore, the filter quality Q can be defined by equation (25), i.e.,

[0214]

number

[0215] where Δf is the bandwidth and f c is the center frequency. The mid filter 1640 is shown as being in series and the side filter 1650 is shown as being in series. In some embodiments, the mid filter 1640 filters the mid component X m and the side filter is applied in parallel to the side component Xs.

[0216] In some embodiments, the crosstalk compensation processor module 1610 processes each of the hyper-mid component M1, hyper-side component S1, residual mid component M2, and residual side component S2. The filters applied to each of these components may be different.

[0217] [Example of a crosstalk processor] 17 is a block diagram illustrating a crosstalk simulation processor module 1700 according to one or more embodiments. The crosstalk simulation processor module 1700 is an example of the crosstalk processor module 110 or the crosstalk processor module 1224. Some embodiments of the crosstalk simulation processor module 1700 have different components than those described herein. Similarly, in some cases, functionality may be distributed among the components in a different manner than that described herein.

[0218] The crosstalk simulation processor module 1700 generates contralateral sound components for output to stereo headphones, thereby providing a loudspeaker-like listening experience on the headphones. L may be the processed left component 134 / 1220 and the right input channel X R may be the processed right component 136 / 1222.

[0219] The crosstalk simulation processor module 1700 is L , the left head shadow low pass filter 1702, the left head shadow high pass filter 1724, the left crosstalk delay 1704, and the left head shadow gain 1710. The crosstalk simulation processor module 1700 processes the right input channel X R The crosstalk simulation processor module 1500 further includes a right head shadow low pass filter 1706, a right head shadow high pass filter 1726, a right crosstalk delay 1708, and a right head shadow gain 1712 to process the right head shadow low pass filter 1706, the right head shadow high pass filter 1726, the right crosstalk delay 1708, and the right head shadow gain 1712. The crosstalk simulation processor module 1500 further includes a summing unit 1714 and a summing unit 1716.

[0220] The left head shadow low-pass filter 1702 and the left head shadow high-pass filter 1724 model the frequency response of the signal after passing through the listener's head, which is the left input channel X L The output of the left head shadow high pass filter 1724 is fed to the left crosstalk delay 1704, which applies a time delay that represents the transaural distance passed by the contralateral sound component relative to the ipsilateral sound component. The left head shadow gain 1710 applies a gain to the output of the left crosstalk delay 1704 to modulate the right-left simulation channels W L Generate.

[0221] Similarly, the right input channel X R For right head shadow low pass filter 1706 and right head shadow high pass filter 1726 apply modulation to the right input channel XR that models the frequency response in the listener's head. The output of right head shadow high pass filter 1726 is fed to right crosstalk delay 1708, which applies a time delay. Right head shadow gain 1712 applies a gain to the output of right crosstalk delay 1708 to generate the right crosstalk simulation channel W. RGenerate.

[0222] The application of the head shadow low pass filter, head shadow high pass filter, crosstalk delay, and head shadow gain to each of the left and right channels may be performed in different orders.

[0223] The summing unit 1714 outputs the right crosstalk simulation channel W R and left input channel X L to produce the left output channel O L The summing unit 1716 generates the left crosstalk simulation channel W L and right input channel X R to produce the left output channel O R Generate.

[0224] 18 is a block diagram illustrating a crosstalk cancellation processor module 1800 according to one or more embodiments. The crosstalk cancellation processor module 1800 is an example of the crosstalk processor module 110 or the crosstalk processor module 1224. Some embodiments of the cancellation processor module 1800 have different components than those described herein. Similarly, in some cases, functionality may be distributed among the components in a different manner than described herein.

[0225] The crosstalk cancellation processor module 1800 is L and right input channel X R Receives channel X L , X R Perform crosstalk cancellation on the left output channel O L and right output channel O R Generates the left input channel X L may be the processed left component 134 / 1220 and the right input channel X R may be the processed right component 136 / 1222.

[0226] The crosstalk cancellation processor module 1800 includes an in-band / out-of-band divider 1810, inverters 1820 and 1822, contralateral estimators 1830 and 1840, combiners 1850 and 1852, and an in-band / out-of-band combiner 1860. These components work in conjunction to combine the input channel T L , T R into in-band and out-of-band components, perform crosstalk cancellation on the in-band component, and output channel O L , O R Generate.

[0227] By dividing the input audio signal T into different frequency band components and performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed on specific frequency bands while avoiding degradations in other frequency bands. If crosstalk cancellation is performed without dividing the input audio signal T into different frequency bands, the audio signal after such crosstalk cancellation may exhibit significant attenuation or amplification of non-spatial and spatial components at low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12,000 Hz), or both. By selectively performing crosstalk cancellation on in-band components (e.g., between 250 Hz and 14,000 Hz), where most of the influential spatial implications reside, the overall energy in the mix can be maintained balanced across the spectrum, particularly in the non-spatial components.

[0228] The in-band / out-of-band divider 1810 divides the input channel T L , T R In-band channel T L,In , T R,In , and the out-of-band channel T L,Out , T R,Out In particular, the in-band / out-of-band splitter 1810 splits the left enhancement compensation channel T L the left in-band channel T L,In and the left out-of-band channel T L,OutSimilarly, the in-band / out-of-band splitter 1810 splits the right enhancement compensation channel T R Right in-band channel T R,In , right out-of-band channel T R,Out Each in-band channel may include a portion of the respective input channel corresponding to a frequency range including, for example, 250 Hz to 14 kHz. The range of the frequency band may be adjustable, for example, according to speaker parameters.

[0229] The inverter 1820 and the contralateral estimator 1830 work together to generate the left contralateral cancellation component S L and generate the left in-band channel T L,In Similarly, the inverter 1822 and the contralateral estimator 1840 work together to compensate for the right contralateral cancellation component S R and generate the right in-band channel T R,In Compensation is performed for the contralateral acoustic component caused by

[0230] In one approach, the inverter 1820 is connected to the in-band channel T L,In Receive the received in-band channel T L,In Invert the polarity of the inverted in-band channel T L,In’ The contralateral estimator 1830 generates the inverted in-band channel T L,In’ and receives the inverted in-band channel T corresponding to the contralateral acoustic component through filter extraction. L,In’ The filter extraction extracts a portion of the inverted in-band channel T L,In’ , the portion extracted by the contralateral estimator 1830 is the in-band channel T L,In Therefore, the part extracted by the contralateral estimator 1830 is the left contralateral cancellation component S L and this is the corresponding in-band channel T R,In In addition to the in-band channel T L,InIn some embodiments, the inverter 1820 and the contralateral estimator 1830 are implemented in different sequences.

[0231] The inverter 1822 and the contralateral estimator 1840 calculate the in-band channel T R,In Perform the same operation for the right contralateral cancellation component S R Therefore, for the sake of brevity, a detailed description thereof will be omitted herein.

[0232] In one implementation, the contralateral estimator 1830 includes a filter 1832, an amplifier 1834, and a delay unit 1836. The filter 1832 has an inverting input channel T L,In’ Receive the inverted in-band channel T corresponding to the contralateral acoustic component through a filter extraction function. L,In’ An example implementation of the filter is a notch filter or a high-pass filter with a center frequency chosen between 5000 Hz and 10000 Hz and a Q chosen between 0.5 and 1 / 0. The gain (G dB ) can be derived from equation (26).

[0233]

number

[0234] where D is the delay by the delay unit 1836 in samples, for example at a sampling rate of 48 KHz. An alternative implementation is a low pass filter with a corner frequency selected between 5000 Hz and 10000 Hz and a Q selected between 0.5 and 1.0. Additionally, the amplifier 1834 filters the extracted portion with a corresponding gain factor G L,In and a delay unit 1836 delays the amplified output from the amplifier 1834 according to a delay function D to generate the left contralateral cancellation component S L The contralateral estimator 1840 generates a filter 1842, an amplifier 1844, and an inverted in-band channel TR,In’Perform a similar operation on the right contralateral cancellation component S R In one example, the contralateral estimators 1830, 1840 generate the left and right contralateral cancellation components S according to the following equations: L , S R That is,

[0235]

number

[0236] where F[] is the filter function and D[] is the delay function.

[0237] The crosstalk cancellation configuration can be determined by speaker parameters. In one example, the filter center frequency, delay, amplifier gain, and filter gain can be determined according to the angle formed between the two speakers relative to the listener. In some embodiments, values ​​between speaker angles are used to interpolate other values.

[0238] The combiner 1850 cancels the right contralateral component S R the left in-band channel T L,In and combine it into the left in-band crosstalk channel U L and combiner 1852 generates the left contralateral cancellation component S L Right in-band channel T R,In to form the right in-band crosstalk channel U R The in-band / out-band combiner 1860 generates the left in-band crosstalk channel U L out-of-band channel T L,Out Combined with the left output channel O L and generates the right in-band crosstalk channel U R out-of-band channel T R,Out and combine it with the right output channel O R Generate.

[0239] Therefore, the left output channel O L is the in-band channel T attributed to the contralateral sound. R,InThe right contralateral cancellation component S corresponds to the inverse of a part of R and the right output channel O R is the in-band channel T attributed to the contralateral sound. L,In The left contralateral cancellation component S corresponds to the inverse of a part of L In this configuration, the right output channel O R When the wavefront of the ipsilateral acoustic component output by the right loudspeaker according to L Similarly, the left output channel O L When the wavefront of the ipsilateral acoustic component output by the left loudspeaker according to R Therefore, the contralateral sound component can be reduced to enhance spatial detectability.

[0240] [Example of PSM process flow] 19 is a flow chart illustrating a process 1900 for PSM processing according to one or more embodiments. Process 1900 may include fewer or additional steps, and the steps may be performed in a different order. In some embodiments, PSM processing may be performed using a Hilbert Transform Perceptual Sound Stage Modification (HPSM) module.

[0241] In step 1905, an audio processing system (e.g., PSM module 102 of audio processing system 100 or 1200) separates the input channels into low-frequency and high-frequency components. A crossover frequency defining the boundary between the low-frequency and high-frequency components may be adjustable so that frequencies subject to PSM processing are included in the high-frequency components. In some embodiments, the audio processing system applies gain to the low-frequency and / or high-frequency components.

[0242] An input channel may be a particular portion of an audio signal that is extracted for PSM processing. In some embodiments, an input channel is a mid or side component of an audio signal (e.g., stereo or multi-channel). In some embodiments, an input channel is a hyper-mid, hyper-side, residual mid, or residual side component of an audio signal. In some embodiments, an input channel is associated with a sound source, such as a voice or instrument, that is to be combined with other sounds into an audio mix.

[0243] In step 1910, the audio processing system applies a first Hilbert transform to the high frequency components to generate a first left leg component and a first right leg component, the first left leg component being 90 degrees out of phase with the first right leg component.

[0244] In step 1915, the speech processing system applies a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, the first left leg component being 90 degrees out of phase with the first right leg component.

[0245] In some embodiments, the audio processing system may apply a delay and / or gain to the first left leg component. The audio processing system may apply a delay and / or gain to the second right leg component. These gains and delays may be used to manipulate the perceptual results of the PSM processing.

[0246] In step 1920, the audio processing system combines the first left leg component with the low frequency component to generate a left channel. In step 1925, the audio processing system combines the second right leg component with the low frequency component to generate a right channel. The left channel can be provided to a left speaker and the right channel can be provided to a right speaker.

[0247] 20 is a flowchart illustrating another process 2000 for PSM processing using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network, according to some embodiments. The process shown in FIG. 20 may be performed by components of an audio system (e.g., systems 100, 202, or 1200). Other entities may perform some or all of the steps in FIG. 2 in other embodiments. Embodiments may include different and / or additional steps, or may perform steps in a different order.

[0248] In step 2005, the audio system determines a target amplitude response that defines one or more spatial cues to be encoded into the mono audio signal to generate the resulting multiple channels. The one or more spatial cues are associated with one or more frequency-dependent amplitude cues that are encoded into the mid-space / side-space of the resulting channels, and the mid-space / side-space of these channels does not change the overall coloration of the resulting channels. The one or more spatial cues may include at least one elevation angle cues associated with a target elevation angle. Each elevation angle cues may correspond to one or more frequency-dependent amplitude cues that are encoded into the mid-space / side-space of the audio signal, such as a target amplitude function corresponding to a narrow region of infinite attenuation at one or more specific frequencies. On the other hand, left and right cues relative to elevation angle are typically symmetric in coloration, so that the left and right signals can be constrained to be colorless. In some embodiments, the spatial cues may be based on sampled HRTFs.

[0249] In some embodiments, the target amplitude response may further define one or more parametric spatial suggestions, which may include a target broadband attenuation, a target subband attenuation, a critical point, a filter characteristic, and / or a sound stage location where the suggestion should be embedded. The critical point may be a 3 dB inflection point. The filter characteristic may include one of a high-pass filter characteristic, a low-pass characteristic, a band-pass characteristic, or a band-stop characteristic. The sound stage location may include a mid-channel or a side-channel, or, if the number of output channels is greater than two, other subspaces in the output space, such as those determined via pairwise and / or hierarchical addition and / or subtraction. The one or more spatial suggestions may be determined based on the characteristics of the presentation equipment (e.g., the speaker's frequency response, the speaker's location, etc.), the expected content of the audio data, the perceptual capabilities of listeners in the scene, or the minimum quality expected of the audio presentation system involved. For example, if a speaker is unable to adequately reproduce frequencies below 200 Hz, spatial suggestions embedded in this range should be avoided. Similarly, if the expected audio content is speech, the audio system may select a target amplitude response that affects only those frequencies within the expected bandwidth of speech to which the ear is most sensitive. If the listener will derive audible cues from other sources in the situation, such as an array of speakers in the area, the audio system may determine a target amplitude response that complements those simultaneous cues.

[0250] In step 2010, the audio system determines a transfer function for a single-input, multiple-output all-pass filter based on the target magnitude response. This transfer function defines the relative rotations in phase angle of the output channels. This transfer function represents, for each output, the effect that the filter network has on its input in terms of phase angle rotations as a function of frequency.

[0251] In step 2015, the audio system determines coefficients of an all-pass filter based on the transfer function. These coefficients are selected in a manner that best suits the type of suggestion and / or constraint and are to be applied to the incoming audio stream. Some example coefficient sets are defined in equations (12), (13), (17), and (19). In some embodiments, determining the coefficients of the all-pass filter based on the transfer function includes using an inverse discrete Fourier transform (IDFT). In this case, the coefficient set may be determined as defined by equation (19). In some embodiments, determining the coefficients of the all-pass filter based on the transfer function includes using a phase-vocoder. In this case, the coefficient set may be determined as defined by equation (19), except that the coefficient set is to be applied to the frequency domain before resynthesizing the time-domain data. In some embodiments, these coefficients include at least a rotation control parameter and a primary coefficient, which are determined based on the received critical point parameters, filter characteristic parameters, and sound stage position parameters.

[0252] The audio system 2020 processes the mono channel with coefficients of an all-pass filter to generate multiple channels. For example, in some embodiments, the all-pass filter module receives the mono audio channel, performs wideband phase rotation on the mono audio channel to generate multiple wideband rotated component channels (e.g., left and right wideband rotated components) based on the rotation control parameter, and performs narrowband phase rotation on at least one of the multiple wideband rotated component channels based on the linear coefficient to determine a narrowband rotated component channel, which, together with one or more remaining wideband rotated component channels, form the multiple channels output by the audio system.

[0253] In some embodiments, if the system is operating in the time domain using an IIR implementation, as in equation (8), these coefficients may adjust the appropriate feedback and feedforward delays. If an FIR implementation is used, as in equation (19), only a feedforward delay may be used. If the coefficients are determined and applied in the spectral domain, the coefficients may be applied to the spectral data as complex multiplications before resynthesis. An audio system may provide multiple output channels to presentation equipment, such as user equipment connected to the audio system via a network.

[0254] The example PSM processing flows described above each utilize a network of all-pass filters to encode spatial suggestion by perceptually positioning mono content at a specific location in the sound stage (e.g., a location associated with a target elevation angle). Because the networks of all-pass filters described herein are colorless, these filters allow the user to decouple the spatial placement of audio from its overall coloration.

[0255] [Spatial processing of orthogonal components] 21 is a flowchart illustrating a process 2100 for spatial processing using at least one of the hyper-mid, residual mid, hyper-side, or residual side components, according to one or more embodiments. Spatial processing may include application of gain, amplitude or delay-based panning, binaural processing, reverberation processing, dynamic range processing such as compression and limiting, linear or nonlinear audio processing techniques and effects, chorus effects, flanging effects, machine learning-based approaches to vocal or instrument style transfer, transformation, or resynthesis, among other techniques. This process may be performed to provide spatially enhanced audio to a user's device. This process may include fewer or additional steps, and the steps may be performed in a different order.

[0256] In step 2110, an audio processing system (such as, for example, audio processing system 1200) receives an input audio signal (such as, for example, left channel 1202 and right input channel 1204). In some embodiments, the input audio signal may be a multi-channel audio signal including multiple left and right channel pairs. Each left and right channel pair may be processed as described herein for the left and right input channels.

[0257] In step 2120, the audio processing system generates non-spatial mid components (e.g., mid components 1208, etc.) and spatial side components (e.g., side components 1210, etc.) from the input audio signal. In some embodiments, an L / RM / S converter (e.g., L / RM / S converter module 1206, etc.) performs the conversion of the input audio signal into mid and side components.

[0258] In step 2130, the audio processing system generates at least one of a hyper-mid component (e.g., hyper-mid component M1), a hyper-side component (e.g., hyper-side component S1), a residual mid component (e.g., residual mid component M2), and a residual side component (e.g., residual side component S2). The audio processing system may generate at least one and / or all of the components listed above. The hyper-mid component includes the spectral energy of the side component removed from the spectral energy of the mid component. The residual mid component includes the spectral energy of the hyper-mid component removed from the spectral energy of the side component. The hyper-side component includes the spectral energy of the mid component removed from the spectral energy of the side component. The residual side component includes the spectral energy of the hyper-side component removed from the spectral energy of the side component. The processing used to generate M1, M2, S1, or S2 may be performed in the frequency domain or the time domain.

[0259] In step 2140, the audio processing system filters at least one of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component to enhance the audio signal. The filtering may include HPSM processing, in which a series of Hilbert transforms are applied to the high frequency components of the hyper-mid component, the residual mid component, the hyper-side component, or the residual side component. In one example, the hyper-mid component undergoes HPSM processing, while one or more of the residual mid component, the hyper-side component, or the residual side component undergoes other types of filtering.

[0260] The filter extraction may include PSM processing, where the spatial implications are colorlessly encoded either through parametric specification of the spatial implications, as described in more detail above in connection with Figures 10A and 10B, or via anthropometric sample extraction of the HRTF data, as described above in connection with equation (20). In one example, the hyper-mid component undergoes PSM processing, while one or more of the residual mid, hyper-side, or residual side components undergo no or other types of filter extraction.

[0261] The filter extraction may include other types of filter extraction, such as spatially-indicated processing. The spatially-indicated processing may include adjusting the frequency-dependent amplitude or frequency-dependent delay of the hyper-mid component, the residual mid component, the hyper-side component, or the residual side component. Examples of spatially-indicated processing include amplitude- or delay-based panning, or binaural processing.

[0262] The filter extraction may include dynamic range processing, such as compression or limiting. For example, the hypermid, residual mid, hyperside, or residual side components may be compressed according to a compression ratio when exceeding a threshold level for compression. In another example, the hypermid, residual mid, hyperside, or residual side components may be limited to a maximum level when exceeding a threshold level for limiting.

[0263] Filter extraction may include machine learning-based modifications to the hyper-mid, residual mid, hyper-side, or residual side components. Some examples include machine learning-based vocal or instrument style transfer, transformation, or resynthesis.

[0264] The filter extraction of the hyper-mid, residual mid, hyper-side, or residual side components may include application of gain, reverberation processing, and other linear or non-linear audio processing techniques and effects, ranging from chorus and / or flanging, or other types of processing. In some embodiments, the filter extraction may include filter extraction for subband spatial processing and crosstalk compensation, as described in more detail below in connection with FIG. 22.

[0265] The filter extraction may be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are transformed from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain, and filter extraction is performed in the frequency domain, and the filter extracted components are transformed to the time domain, and filter extraction is performed on these components in the time domain.

[0266] In step 2150, the audio processing system generates a left output channel (e.g., left output channel 1242) and a right output channel (e.g., right output channel 1244) using one or more of the filter-extracted hyper / residual components. For example, the M / S to L / R conversion may be performed using a mid component or a side component generated from at least one of the filter-extracted hyper-mid component, the filter-extracted residual mid component, the filter-extracted hyper-side component, or the filter-extracted residual side component. In another example, the filter-extracted hyper-mid component or the filter-extracted residual mid component may be used as the mid component for the M / S to L / R conversion, or the filter-extracted hyper-side component or the residual side component may be used as the side component for the M / S to L / R conversion.

[0267] [Subband spatial processing of quadrature components and crosstalk processing] FIG. 22 is a flowchart illustrating a process 2200 for subband spatial processing and crosstalk compensation using at least one of hyper-mid, residual mid, hyper-side, or residual side components, according to one or more embodiments. Crosstalk processing may include crosstalk cancellation or crosstalk simulation. Subband spatial processing may be performed to provide audio content with enhanced spatial detectability (e.g., sound stage enhancement), such as by creating the perception that sound is directed to the listener from a wide area rather than a specific point in space corresponding to the loudspeaker's location, thereby providing a more immersive listening experience for the listener. Crosstalk simulation may be used on audio output to headphones to simulate the experience of a loudspeaker with contralateral crosstalk. Crosstalk cancellation may be used on audio output to loudspeakers to remove the effects of crosstalk interference. Crosstalk compensation compensates for spectral imperfections caused by crosstalk cancellation or crosstalk simulation. The process may include fewer or additional steps, and the steps may be performed in a different order. The hyper and residual mid / side components can be manipulated in various ways for various purposes. For example, in the case of crosstalk compensation, a targeted subband filter extraction may be applied only to the hyper-mid component M1 (where the majority of vocal dialog energy in most movie content occurs) in an effort to remove spectral artifacts resulting from crosstalk processing in only that component. For soundstage enhancement, with or without crosstalk processing, a targeted subband gain may be applied to the residual mid component M2 and residual side component S2.For example, the residual mid component M2 may be attenuated and the residual side component S2 may be de-amplified to increase the distance between these components in terms of gain without causing a dramatic overall change in the perceived sound of the final L / R signal, while also avoiding attenuation in the hyper-mid M1 component (e.g., which is often that part of the signal that contains the majority of the vocalization energy) (which, if done tastefully, can increase spatial detectability).

[0268] In step 2210, the audio processing system receives an input audio signal including a left channel and a right channel. In some embodiments, the input audio signal may be a multi-channel audio signal including multiple left and right channel pairs. Each left and right channel pair may be processed as described herein for left and right input channels.

[0269] In step 2220, the audio processing system applies crosstalk processing to the received input audio signal, the crosstalk processing including at least one of crosstalk simulation and crosstalk cancellation.

[0270] In steps 2230 through 2260, the audio processing system performs subband spatial processing and crosstalk compensation for crosstalk processing using one or more of the hypermid, hyperside, residual mid, or residual side components. In some embodiments, crosstalk processing may be performed after the processing of steps 2230 through 2260.

[0271] In step 2230, the audio processing system generates a mid component and a side component from the (eg, crosstalk processed) audio signal.

[0272] In step 2240, the audio processing system generates at least one of a hyper-mid component, a residual mid component, a hyper-side component, and a residual side component. The audio processing system may generate at least one and / or all of the components listed above.

[0273] In step 2250, the audio processing system filter-extracts at least one subband of the hyper-mid, residual mid, hyper-side, and residual side components and applies subband spatial processing to the audio signal. Each subband may include a frequency range that may be defined by a set of critical bands. In some embodiments, the subband spatial processing further includes time-delaying at least one subband of the hyper-mid, residual mid, hyper-side, and residual side components. In some embodiments, the filter-extracting includes applying HPSM processing.

[0274] In step 2260, the audio processing system filters out at least one of the hyper-mid component, the residual mid component, the hyper-side component, and the residual side component to compensate for spectral imperfections from crosstalk processing of the input audio signal. The spectral imperfections may include peaks or troughs in a frequency response plot of the hyper-mid component, the residual mid component, the hyper-side component, or the residual side component that exceed a predetermined threshold (e.g., 10 dB) and occur as artifacts of crosstalk processing. The spectral imperfections may be estimated spectral imperfections.

[0275] In some embodiments, the filter extraction of spectral orthogonal components for subband spatial processing in step 2250 and the crosstalk compensation in step 2260 may be combined into a single filter extraction operation for each spectral orthogonal component selected for filter extraction.

[0276] In some embodiments, filter extraction on the hyper / residual mid / side components for subband spatial processing or crosstalk compensation may be performed in conjunction with filter extraction for other purposes, such as applying gain, amplitude or delay-based panning, binaural processing, reverberation processing, dynamic range processing such as compression and limiting, linear or non-linear audio processing techniques and effects ranging from chorus and / or flanging, machine learning-based approaches to vocal or instrument style transfer, transformation, or resynthesis, or other types of processing using any of the hyper-mid, residual mid, hyper-side, or residual side components.

[0277] The filter extraction may be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are transformed from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain, the filter extraction is performed in the frequency domain, and the filter-extracted components are transformed to the time domain. In other embodiments, the hyper and / or residual components are transformed to the time domain, and the filter extraction is performed on these components in the time domain.

[0278] In step 2270, the audio processing system generates left and right output channels from the filter-extracted hyper-mid component. In some embodiments, the left and right output channels are additionally based on at least one of the filter-extracted residual mid component, the filter-extracted hyper-side component, and the filter-extracted residual side component.

[0279] [Computer example] FIG. 23 is a block diagram illustrating a computer 2300, according to some embodiments. The computer 2300 is an example of a computing device that includes circuitry implementing an audio system, such as audio system 100, 202, or 1200. Illustrated is at least one processor 2302 coupled to a chipset 2304. The chipset 2304 includes a memory controller hub 2320 and an input / output (I / O) controller hub 2322. The memory 2306 and graphics adapter 2312 are coupled to the memory controller hub 2320, and the display 2318 is coupled to the graphics adapter 2312. The storage device 2308, keyboard 2310, pointing device 2314, and network adapter 2316 are coupled to the I / O controller hub 2322. The computer 2300 may include various types of input or output devices. Other embodiments of the computer 2300 have different architectures. For example, in some embodiments, the memory 2306 is directly connected to the processor 2302.

[0280] The storage device 2308 includes one or more non-transitory computer-readable storage media, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 2306 stores program code (comprised of one or more instructions) and data used by the processor 2302. This program code may correspond to the processing aspects described with reference to Figures 1-3.

[0281] A pointing device 2314 is used in combination with the keyboard 2310 to input data into the computer system 2300. A graphics adapter 2312 displays images and other information on a display 2318. In some embodiments, the display 2318 includes touch screen functionality for receiving user inputs and user selections. A network adapter 2316 connects the computer system 2300 to a network. Some embodiments of the computer 2300 have different and / or additional components than those shown in FIG. 23 .

[0282] The circuitry may include one or more processors executing program code stored on a non-transitory computer-readable medium that, when executed by the one or more processors, configures the one or more processors to implement an audio system or a module of an audio system. Other examples of circuitry implementing an audio system or a module of an audio system may include integrated circuits, such as an application specific integrated circuit (AS1C), a field programmable gate array (FPGA), or other types of computer circuitry.

[0283] [Additional Considerations] Example benefits and advantages of the disclosed configuration include dynamic audio enhancements resulting from an audio system that is adaptively enhanced to the device and associated audio rendering system, as well as other relevant information made available by the device OS, such as use case information (e.g., indicating that an audio signal is used for music playback rather than gaming). The enhanced audio system can either be integrated into the device (using a software development kit) or stored on a remote server for on-demand access. In this way, the device need not allocate storage or processing resources to maintaining an audio enhancement system that is specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system allows for various levels of querying of rendering system information, such that effective audio enhancements can be applied across various levels of available device-specific rendering information.

[0284] Throughout this specification, a component, operation, or structure described as a single instance may be implemented by multiple instances. Although individual operations in one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously and the operations need not be performed in the order illustrated. Structures and functions presented as separate components in example configurations may be implemented as combined structures or components. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are included within the scope of the subject matter herein.

[0285] Certain embodiments are described herein as including logic, or multiple components, modules, or mechanisms. A module may constitute either a software module (e.g., code embodied on a machine-readable medium or in a transmission signal) or a hardware module. A hardware module is a tangible unit capable of performing specific operations, and may be configured or arranged in a particular way. In example embodiments, one or more computer systems (e.g., standalone, client, or server computer systems), or one or more hardware modules of a computer system (e.g., a processor or group of processors), are configured as hardware modules that are operated by software (e.g., an application or portion of an application) to perform specific operations described herein.

[0286] Various operations in the example methods described herein may be performed, at least in part, by one or more processors that are configured temporarily (e.g., by software) or permanently configured to perform the associated operations. Whether configured temporarily or permanently, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. Modules referred to herein may, in some example embodiments, include processor-implemented modules.

[0287] Similarly, the methods described herein may be at least partially processor-implemented. For example, at least some of the operations in the methods may be performed by one or more processors or processor-implemented hardware modules. Performance for a particular operation may reside within a single machine or may be distributed among one or more processors deployed across multiple machines. In some example embodiments, a processor or processors may be located in a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed across multiple locations.

[0288] Unless otherwise specified, descriptions herein using terms such as "processing," "computing," "calculating," "determining," "presenting," or "displaying" may refer to machine (e.g., computer) operations or processes that manipulate or transform data represented as physical (e.g., electronic, magnetic, or optical) quantities, for example, in one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, send, transmit, or display information.

[0289] As used herein, any reference to "one embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" in various places in this specification do not necessarily all refer to the same embodiment.

[0290] Some embodiments may be described using the terms "coupled" and "connected," along with their derivatives. It should be understood that these terms are not intended as synonyms for each other. For example, some embodiments may be described using the term "connected" to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments may be described using the term "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other. The embodiments are not limited in this context.

[0291] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements and may include other elements not expressly listed or that are inherent in such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or, not an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), or both A and B are true (or absent).

[0292] Furthermore, the use of "a" or "an" is used to describe elements and components of embodiments herein. This is done merely for convenience and to understand the ordinary meaning of the present invention. This description should be interpreted to include one or at least one, and the singular also includes the plural unless it is clear that it is meant otherwise.

[0293] In some portions of this description, embodiments are described in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. These operations, while described functionally, computationally, or logically, will be understood to be implemented by computer programs or equivalent electrical circuits, or microcode, or the like. Further, it has proven convenient at times to refer to arrangements of these operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0294] Any of the steps, operations, or processes described herein may be performed or implemented by one or more hardware or software modules, alone or in combination with other devices. In some embodiments, the software modules are implemented by a computer program product that includes a computer-readable medium containing computer program code, which can be executed by a computer processor to perform any or all of the steps, operations, or processes described.

[0295] Embodiments may also relate to apparatus for performing the operations herein. This apparatus may be specifically constructed for the required purposes and / or may include general-purpose computing equipment selectively activated or reconfigured by a computer program stored in the computer. Such computer programs may be stored on a non-transitory, tangible, computer-readable storage medium or any type of medium suitable for storing electronic instructions that may be coupled to a computer system bus. Furthermore, all computing systems referred to herein may include a single processor or may be architectures employing multiple processor designs to increase computing power.

[0296] Embodiments may also relate to articles of manufacture produced by the computing processes described herein. Such articles of manufacture may include information obtained from the computing processes, which information is stored on a non-transitory, tangible computer-readable storage medium, and may include any embodiment of a computer program product or other data combination described herein.

[0297] Upon reading this disclosure, those skilled in the art will recognize still additional alternative structural and functional designs of systems and processes for audio content decorrelation according to the principles disclosed herein. Thus, while particular embodiments and applications have been illustrated and described, it should be understood that the disclosed embodiments are not limited to the precise structure and components disclosed herein. Various modifications, changes, and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope, as defined in the appended claims.

[0298] Finally, the language used herein has been chosen primarily for ease of reading and educational purposes, and may not be chosen to delineate or limit patent rights. Accordingly, it is intended that the scope of patent rights be limited not by this detailed description, but rather by any claims that issue on an application based on this specification. Accordingly, the disclosure of the embodiments is intended to illustrate, but not limit, the scope of patent rights, which are set forth in the following claims.

Claims

1. 1. A system comprising: one or more processors; A non-transitory computer-readable medium that, when executed by the one or more processors, causes the one or more processors to: Separating an audio channel into low frequency and high frequency components; applying a first Hilbert transform to the high frequency components to generate a first left leg component and a first right leg component, the first left leg component being 90 degrees out of phase with the first right leg component; applying a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, the second left leg component being 90 degrees out of phase with the second right leg component; combining the first left leg component with the low frequency component to generate a left channel; combining the second right leg component with the low frequency component to generate a right channel; a stored non-transitory computer readable medium including stored program code configured to perform the A system comprising:

2. 2. The system of claim 1, wherein the program code further configures the one or more processors to apply a first gain to the low frequency components and a second gain to the high frequency components, the first gain and the second gain being different.

3. 2. The system of claim 1, wherein the program code further configures the one or more processors to apply a first delay to the first left leg component and a second delay to the second right leg component, the first delay and the second delay being different.

4. 2. The system of claim 1, wherein the program code further configures the one or more processors to apply a first gain to the first left leg component and a second gain to the second right leg component, the first gain and second gain being different.

5. The program code for configuring the one or more processors to apply the first Hilbert transform to the high frequency components may further include: applying a first series of all-pass filters to the high frequency component to generate the first left leg component; applying a first delay and a second series of all-pass filters to the high frequency component to generate the first right leg component; and The program code configuring the one or more processors to apply the second Hilbert transform to the first right leg component may further include: applying a third series of all-pass filters to the first right leg component to generate the second left leg component; applying a second delay and a fourth series of all-pass filters to the first right leg component to generate the second right leg component; Configure it to do 2. The system of claim 1 .

6. The program code causes the one or more processors to: generating mid and side components from left and right input channels of an audio signal; generating a hyper-mid component that includes the spectral energy of the side component removed from the spectral energy of the mid component; 2. The system of claim 1, further configured to generate the audio channel by:

7. The program code causes the one or more processors to: generating a mid component and a side component from the left channel and the right channel; applying filter extraction to the mid and side components; generating a left output channel and a right output channel from the filtered mid component and the filtered side component; 10. The system of claim 1, further configured to:

8. 10. The system of claim 1, wherein the program code further configures the one or more processors to generate the audio channels by combining channels of a multi-channel audio signal.

9. 10. The system of claim 1, wherein the program code further configures the one or more processors to generate the audio channels by isolating portions of an audio signal.

10. 10. The system of claim 1, wherein the high frequency components include sounds for vocalization.

11. A non-transitory computer-readable medium having stored thereon program code that, when executed by one or more processors, causes the one or more processors to: Separating an audio channel into low frequency and high frequency components; applying a first Hilbert transform to the high frequency components to generate a first left leg component and a first right leg component, the first left leg component being 90 degrees out of phase with the first right leg component; applying a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, the second left leg component being 90 degrees out of phase with the second right leg component; combining the first left leg component with the low frequency component to generate a left channel; combining the second right leg component with the low frequency component to generate a right channel; 1. A non-transitory computer-readable medium configured to:

12. 12. The computer-readable medium of claim 11, wherein the program code further configures the one or more processors to apply a first gain to the low frequency components and a second gain to the high frequency components, the first gain and the second gain being different.

13. 12. The computer-readable medium of claim 11, wherein the program code further configures the one or more processors to apply a first delay to the first left leg component and a second delay to the second right leg component, wherein the first delay and the second delay are different.

14. 12. The computer-readable medium of claim 11, wherein the program code further configures the one or more processors to apply a first gain to the first left leg component and a second gain to the second right leg component, wherein the first gain and the second gain are different.

15. The program code for configuring the one or more processors to apply the first Hilbert transform to the high frequency components may further include: applying a first series of all-pass filters to the high frequency component to generate the first left leg component; applying a first delay and a second series of all-pass filters to the high frequency component to generate the first right leg component; and The program code configuring the one or more processors to apply the second Hilbert transform to the first right leg component may further include: applying a third series of all-pass filters to the first right leg component to generate the second left leg component; applying a second delay and a fourth series of all-pass filters to the first right leg component to generate the second right leg component; Configure it to do 12. The computer-readable medium of claim 11.

16. The program code causes the one or more processors to: generating mid and side components from left and right input channels of an audio signal; generating a hyper-mid component that includes the spectral energy of the side component removed from the spectral energy of the mid component; 12. The computer-readable medium of claim 11, further configured to generate the audio channel by:

17. The program code causes the one or more processors to: generating a mid component and a side component from the left channel and the right channel; applying filter extraction to the mid and side components; generating a left output channel and a right output channel from the filtered mid component and the filtered side component; 12. The computer-readable medium of claim 11, further configured to:

18. 12. The computer-readable medium of claim 11, wherein the program code further configures the one or more processors to generate the audio channels by combining channels of a multi-channel audio signal.

19. 12. The computer-readable medium of claim 11, wherein the program code further configures the one or more processors to generate the audio channels by separating portions of an audio signal.

20. 12. The computer-readable medium of claim 11, wherein the high frequency components include sounds for vocalizations.

21. 1. A method, comprising: Separating an audio channel into low frequency and high frequency components; applying a first Hilbert transform to the high frequency components to generate a first left leg component and a first right leg component, the first left leg component being 90 degrees out of phase with the first right leg component; applying a second Hilbert transform to the first right leg component to generate a second left leg component and a second right leg component, the second left leg component being 90 degrees out of phase with the second right leg component; combining the first left leg component with the low frequency component to generate a left channel; combining the second right leg component with the low frequency component to generate a right channel; A method comprising:

22. applying, by the one or more processors, a first gain to the low frequency components and a second gain to the high frequency components, the first gain and the second gain being different.

22. The method of claim 21 further comprising:

23. applying, by the one or more processors, a first delay to the first left leg component and a second delay to the second right leg component, wherein the first delay and the second delay are different.

22. The method of claim 21 further comprising:

24. applying, by the one or more processors, a first gain to the first left leg component and a second gain to the second right leg component, the first gain and the second gain being different.

22. The method of claim 21 further comprising:

25. applying the first Hilbert transform to the high frequency components applying a first series of all-pass filters to the high frequency component to generate the first left leg component; applying a first delay and a second series of all-pass filters to the high frequency component to generate the first right leg component; Including, applying the second Hilbert transform to the first right leg component applying a third series of all-pass filters to the first right leg component to generate the second left leg component; applying a second delay and a fourth series of all-pass filters to the first right leg component to generate the second right leg component; 22. The method of claim 21 further comprising:

26. by said one or more processors; generating mid and side components from left and right input channels of an audio signal; generating a hyper-mid component that includes the spectral energy of the side component removed from the spectral energy of the mid component; 22. The method of claim 21, further comprising generating the audio channel by:

27. by said one or more processors; generating a mid component and a side component from the left channel and the right channel; applying filter extraction to the mid and side components; generating a left output channel and a right output channel from the filtered mid component and the filtered side component; 22. The method of claim 21 further comprising:

28. 22. The method of claim 21, further comprising generating, by the one or more processors, the audio channels by combining channels of a multi-channel audio signal.

29. 22. The method of claim 21, further comprising generating, by the one or more processors, the audio channels by isolating portions of an audio signal.

30. 22. The method of claim 21, wherein the high frequency components include sounds for vocalization.

Citation Information

Patent Citations

  • Signal processing method and signal processing apparatus

    JP2009276268A

  • Modulating device, demodulating device, information transmission system, modulating method and demodulating method

    JP2010016625A

  • Hybrid derivation of surround sound audio channels by controllably combining ambient signal components and matrix-decoded signal components.

    JP2010529780A

  • Audio spatial environment up-mixer

    US20060093152A1

  • Audio Spatial Environment Engine

    US20090060204A1