Colorless generation of elevation-perceptual suggestions using a pass-through filter network
By encoding spatial cues into monaural audio signals using allpass and pass-through filters, the method generates multiple channels, enhancing spatial perception and immersion in audio experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-17
AI Technical Summary
Existing audio encoding technologies fail to effectively encode spatial cues in monaural audio signals, limiting the ability to create immersive and spatially distinct audio experiences.
A method and system that utilize a single-input, multiple-output allpass filter and pass-through filters to encode spatial cues into monaural signals, generating multiple channels with perceptual shifts and enhancements, including Hilbert transforms and non-orthogonal rotation-based decorrelation techniques.
Enhances spatial perception in audio content, allowing for immersive experiences and improved clarity by creating distinct sound sources and reducing the need for additional amplifiers and speakers.
Smart Images

Figure 2026048986000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure generally relates to audio processing, and more specifically to the encoding of spatial cues to audio content. [Background technology]
[0002] Audio content can be encoded to include the spatial characteristics of a sound field, enabling users to perceive a sense of space within that sound field. For example, the sound of a specific sound source (e.g., voice or instrument) may be mixed into audio content in a way that creates a sense of space associated with the sound, such as the perception that the sound is coming from a specific direction or is located in a specific type of place (e.g., a small room, a large auditorium). [Overview of the project]
[0003] Some embodiments include a method for encoding spatial implications along a sagittal plane into a monaural signal to generate a resulting multiple channels. The method includes the steps of: determining a target amplitude response for the mid-channel or side-channel of the resulting multiple channels based on spatial implications associated with a frequency-dependent phase shift using a processing circuit; converting the target amplitude response for either the mid-channel or side-channel into a transfer function for a single-input, multiple-output allpass filter; and processing the monaural signal using the allpass filter, wherein the allpass filter is configured based on the transfer function.
[0004] Some embodiments include a system for generating multiple channels from a mono channel, wherein the multiple channels are encoded by one or more spatial suggestions. The system includes one or more computing devices configured to determine a target amplitude response for the mid- or side-components of the multiple channels based on spatial suggestions associated with frequency-dependent phase shifts. One or more computers are further configured to convert the target amplitude response for either the mid- or side-components into a transfer function for a single-input, multiple-output pass-through filter, and to process the mono signal using the pass-through filter, wherein the pass-through filter is configured based on the transfer function.
[0005] Some embodiments include a non-temporary computer-readable medium containing instructions stored for generating multiple channels from a mono channel, wherein the multiple channels are encoded using one or more spatial suggestions, and the instructions, when executed by at least one processor, configure at least one processor to: determine a target amplitude response for the mid or side components of the resulting multiple channels based on a spatial suggestion associated with a frequency-dependent phase shift; convert the target amplitude response for the mid or side components into a transfer function of a single-input, multiple-output pass-through filter; and process the mono signal using the pass-through filter, the pass-through filter being configured based on the transfer function.
[0006] Several embodiments relate to spatially shifting a portion of audio content (e.g., speech) using a series of Hilbert transforms. Some embodiments include one or more processors and a non-temporal computer-readable medium. The computer-readable medium includes stored program code that, when executed by one or more processors, configures one or more processors to: separate an audio channel into low-frequency and high-frequency components; apply a first Hilbert transform to the high-frequency component to generate a first left-foot component and a first right-foot component, wherein the first left-foot component is 90 degrees out of phase with respect to the first right-foot component; apply a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component; combine the first left-foot component with the low-frequency component to generate a left channel; and combine the second right-foot component with the low-frequency component to generate a right channel.
[0007] Some embodiments include a non-temporary computer-readable medium containing stored program code. When executed by one or more processors, this program code configures one or more processors to: separate an audio channel into low-frequency and high-frequency components; apply a first Hilbert transform to the high-frequency component to generate a first left-foot component and a first right-foot component, wherein the first left-foot component is 90 degrees out of phase with respect to the first right-foot component; apply a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component; combine the first left-foot component with the low-frequency component to generate a left channel; and combine the second right-foot component with the low-frequency component to generate a right channel.
[0008] Some embodiments include a method performed by one or more processors. The method includes the steps of: separating an audio channel into low-frequency and high-frequency components; applying a first Hilbert transform to the high-frequency components to generate a first left-foot component and a first right-foot component, wherein the first left-foot component is 90 degrees out of phase with respect to the first right-foot component; applying a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component; coupling the first left-foot component with the low-frequency components to generate a left channel; and coupling the second right-foot component with the low-frequency components to generate a right channel. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing several embodiments of a speech processing system. [Figure 2] This block diagram shows a computing system environment in several embodiments. [Figure 3] This figure illustrates graphs showing the HRTF of sample extraction, measured at an elevation angle of 60 degrees, according to several embodiments. [Figure 4] This figure illustrates a graph showing an example of a perceptual cue, characterized by a target amplitude function corresponding to a narrow region of infinite attenuation at 11 kHz, according to several embodiments. [Figure 5] This figure illustrates frequency plots generated by driving a second-order pass-through filter section with coefficients shown in Table 1 using white noise, according to several embodiments. [Figure 6] This is a block diagram showing a PSM module implemented using a Hilbert transform, according to one or more embodiments. [Figure 7] This is a block diagram showing one or more embodiments of a Hilbert transducer module. [Figure 8] Figure 6 illustrates frequency plots generated by driving an HPSM module, including white noise, in several embodiments, showing the output frequency response of the sum of multiple channels (mids) and the difference of multiple channels (sides). [Figure 9] This block diagram shows a PSM module implemented using an FNORD filter network, according to several embodiments. [Figure 10A] This is a detailed block diagram showing the PSM module 900 in several embodiments. [Figure 10B] This block diagram shows a broadband phase rotator implemented within a full-pass filter module of a PSM module, according to several embodiments. [Figure 11] This figure illustrates frequency response graphs showing the output frequency response of FNORD filter networks configured to achieve an amplitude response to a 60-degree vertical cue, according to several embodiments. [Figure 12] This is a block diagram showing a voice processing system 1000 according to one or more embodiments. [Figure 13A] This is a block diagram showing an orthogonal component generator according to one or more embodiments. [Figure 13B] Block diagram of an orthogonal component generator according to one or more embodiments. [Figure 13C] Block diagram of an orthogonal component generator according to one or more embodiments. [Figure 14A] This is a block diagram showing one or more embodiments of an orthogonal component processor module. [Figure 14B] A block diagram showing an orthogonal component processor module according to one or more embodiments is shown. [Figure 15]This is a block diagram showing one or more embodiments of a subband spatial processor module. [Figure 16] This is a block diagram showing a crosstalk compensation processor module according to one or more embodiments. [Figure 17] This is a block diagram showing a crosstalk simulation processor module according to one or more embodiments. [Figure 18] This is a block diagram showing a crosstalk cancellation processor module according to one or more embodiments. [Figure 19] This flowchart shows the process for PSM processing using a Hilbert Transform Perceptual Soundstage Modification (HPSM) module, according to one or more embodiments. [Figure 20] This flowchart shows another process for PSM processing using a First Order Non-Orthogonal Rotation-Based Decorrelation (FNORD) filter network, according to several embodiments. [Figure 21] This is a flowchart illustrating a process for spatial processing using at least one of the following components according to one or more embodiments: a hyper-mid component, a residual-mid component, a hyper-side component, or a residual-side component. [Figure 22]This is a flowchart showing a process for subband spatial processing and compensation for crosstalk processing using at least one of the following components according to one or more embodiments: a hypermid component, a residual mid component, a hyperside component, or a residual side component. [Figure 23] A computer block in several embodiments.
[0010] The figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following considerations that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein. [Modes for carrying out the invention]
[0011] [Detailed explanation] The drawings (FIG.) and the following description relate to preferred embodiments for illustrative purposes only. It should be noted that from the following description, alternative embodiments of the structures and methods disclosed herein will be readily recognizable as viable alternatives that can be used without departing from the principles of the asserted subject matter.
[0012] This specification will provide detailed references to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that, wherever possible, similar or identical reference numerals may be used in the drawings to indicate similar or identical functions. These drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein.
[0013] Encoding spatial perceptual cues into monaural audio sources can be desirable in various applications involving the presentation of multiple simultaneous streams of audible content. Examples of such applications include: • Example of use in meetings: Adding spatial perceptual suggestions applied to one or more remote speakers can improve overall voice intelligibility and contribute to an increased sense of immersion for the listener. • Examples of use for video and music playback / streaming: One or more audio channels, or signal components of one or more audio channels, can be enhanced by adding spatial perceptual suggestions to improve the clarity or spatial sense of the audio or other elements in the mix. • Example of use for collaborative viewing entertainment: The stream consists of individual content channels, such as one or more remote speakers and entertainment program material, which need to be mixed to form an immersive experience. Applying spatial perceptual suggestions to one or more elements can enhance the sense of perceptual distinction between elements in the mix and broaden the listener's perceptual bandwidth.
[0014] The embodiments relate to a speech system that modifies the perceptual spatial quality (e.g., sound stage and overall position relative to the head of the target listener) in one or more speech channels. In some embodiments, the modification of the perceptual spatial quality of a speech channel may be used to isolate the coloration of a particular source from its perceptual position in space and / or to reduce the number of amplifiers and speakers required to encode such an effect.
[0015] The audio signal processing performed by the audio system is called Perceptual Soundstage Modification (PSM) processing. The perceived result of PSM processing is referred to herein as spatial shifting. Psychoacoustic effects are typically experienced by the user as a perceptual distinction of the sound source from other parts of the audio content, and as the sound source being shifted overall above, around, or towards the head. This psychoacoustic effect is derived from the phase and time relationship between the left and right channels, as enhanced by a network of all-pass filters and delays. In some embodiments, this network of filters and delays may be implemented as one or more second-order all-pass sections, such as a series of Hilbert transforms, or using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network, each of which will be described in more detail later. The perceived result of PSM processing may vary depending on the listening configuration (e.g., headphones or speakers). With certain content and algorithm configurations, the result can give the impression that the perceived signal is spreading (e.g., diffused) around the listener's head. For monaural input signals (e.g., non-spatial audio signals), the diffusion effect of PSM processing can be used for upmixing from monaural to stereo.
[0016] In some embodiments, an audio system may separate a target portion of an audio signal from the residual portion, perceptually shift the target portion by applying various configurations in PSM processing, and then mix the processed result back with the residual portion (e.g., unprocessed or processed differently). Such a system may be perceived as performing clarification, elevation, or otherwise differential differentiation of the target portion within the overall audio mix. In some embodiments, PSM processing is used to perceptually shift a portion of an audio signal, including sung voice or spoken voice. By convention, vocalizations in television, film, or music audio streams are often placed at the center of the soundstage and are therefore part of the mid-range component (also called the non-spatial component or correlated component) of a stereo or multi-channel audio signal. Thus, PSM processing may be applied to the mid-range component of an audio signal, or to the hyper-mid-range component, which includes the spectral energy of the side components (also called the spatial component or non-correlated component) that are removed from the spectral energy of the mid-range component.
[0017] PSM processing can be combined with other types of processing. For example, an audio system may apply processing to the shifted portion of an audio signal to perceptually transform the shifted portion and distinguish it from other components in the mix. These types of additional processing may include one or more of the following: single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., limiting, compression, expansion, etc.), single-band or multi-band gain or delay, crosstalk processing (e.g., crosstalk cancellation processing and / or crosstalk simulation processing, etc.), or compensation for crosstalk processing. In some embodiments, PSM processing may be performed with mid / side processing, such as subband spatial processing, in which the mid and side subbands of the audio signal generated via PSM processing are gain-adjusted to enhance the sense of spatiality of the sound field.
[0018] The separation of audio channels for PSM processing can be achieved in various ways. In some embodiments, PSM processing may be performed on spectrally orthogonal audio components, such as the hypermid component of an audio signal. In other embodiments, PSM processing is performed on audio channels associated with a sound source (e.g., utterance), and the processed channels are then mixed with other audio content (e.g., background music).
[0019] The following explanation primarily focuses on upmixing a mono signal to stereo (i.e., two output channels), given that most audio presentation equipment is stereo; however, it is understood that the techniques described can be easily generalized to include a greater number of channels. The stereo embodiment can be described in terms of mid-processing / side-processing, where the phase difference between the left and right channels constitutes the complementary regions of amplification and attenuation in the mid-space / side-space.
[0020] [Example of a speech processing system] Figure 1 is a block diagram showing one or more embodiments of an audio processing system 100. The system 100 uses PSM processing to spatially shift the audio signal and apply other types of spatial processing (e.g., mid / side processing). Some embodiments of the system 100 have components different from those described herein. Similarly, in some cases, functions can be distributed among the components in ways different from those described herein.
[0021] System 100 includes a PSM module 102, an L / RM / S converter module 104, a component processor module 106, an M / SL / R converter module 108, and a crosstalk processor module 108. The PSM module 102 receives the input audio 120 and generates spatially shifted left channel 122 and right channel 124. The operation of the PSM 102 in various embodiments will be described in more detail later with reference to Figures 6 to 11.
[0022] The L / RM / S converter module 104 receives the left channel 122 and the right channel 124 and generates a mid component 126 (e.g., a non-spatial component) and a side component 128 (e.g., a spatial component) from these channels 122 and 124. In some embodiments, the mid component 126 is generated based on the sum of the left channel 122 and the right channel 122, and the side component 128 is generated based on the difference between the left channel 122 and the right channel 124. In some embodiments, the transformation of a point in L / R space to a point in M / S space can be expressed according to equation (1) as follows:
[0023]
number
[0024] On the other hand, the inverse transform can be expressed as follows according to equation (2): that is,
[0025]
number
[0026] In other embodiments, it is understood that other L / RM / S type transformations may be used to generate the mid component 126 and the side component 128. In some embodiments, the transformations shown in equations (1) and (2) may be used instead of the true orthonormal form, where both the forward and reverse transformations are scaled by √2, due to the reduction in computational complexity. For ease of explanation, regardless of the specific transformation used, the convention of transforming the coordinates of row vectors by multiplication on the right and the notation on which the transformed coordinates are superimposed as labels will be used, as shown in equation (3) below. That is,
[0027]
number
[0028] The component processor module 106 processes the mid component 126 to produce a processed mid component 130 and processes the side component 128 to produce a processed side component 314. Processing of each component 126 and 128 may include various types of filter extraction, such as spatial suggestion processing (e.g., amplitude or delay-based panning, binaural processing), single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., compression, expansion, limiting), single-band or multi-band gain or delay steps, addition of speech effects, or other types of processing. In some embodiments, the component processor module 106 uses the mid component 126 and side component 128 to perform subband spatial processing and / or crosstalk compensation processing. Subband spatial processing is processing performed on the frequency subbands of the mid and side components to spatially enhance the speech signal. Crosstalk compensation is a process that adjusts for spectral artifacts caused by crosstalk processing, such as crosstalk correction for speakers or crosstalk simulation for headphones. Various components that may be included in the component processor module 106 are further described with reference to Figures 12A to 13.
[0029] The M / SL / R converter module 108 receives the processed mid component 130 and the processed side component 132 and generates the processed left component 134 and the processed right component 136. In some embodiments, the M / SL / R converter module 108 converts the processed mid component 130 and the side component 132 based on the inverse of the conversion performed by the L / RM / S converter module 104, for example, the processed left component 134 is generated based on the addition of the processed mid component 130 and the processed side component 132, and the processed right component 136 is generated based on the difference between the processed mid component 130 and the processed side component 132. Other M / SL / R conversion types may be used to generate the processed left component 134 and the processed right component 136.
[0030] The crosstalk processor module 110 receives the processed left component 134 and the processed right component 136 and performs crosstalk processing. Crosstalk processing includes, for example, crosstalk simulation or crosstalk cancellation. Crosstalk simulation is a process performed on an audio signal (e.g., output via headphones) to simulate the effect of a loudspeaker. Crosstalk cancellation is a process performed on an audio signal (e.g., output via speakers) to reduce crosstalk caused by a loudspeaker. The crosstalk processor module 110 outputs a left channel 138 and a right output channel 140. In some embodiments, crosstalk processing (e.g., simulation or cancellation) may be performed before component processing, such as before converting the left channel 122 and the right channel 124 into mid and side components. Various components that may be included in the crosstalk processor module 110 will be further described with reference to Figures 15 and 16.
[0031] In some embodiments, the PSM module 100 is integrated into the component processor module 106. The L / RM / S converter module 104 receives the left and right channels, which may represent (e.g., stereo) inputs to the audio processing system 100. The L / RM / S converter module 104 uses the left and right input channels to generate the mid and side components. The PSM module 100 of the component processor module 106 processes the mid and / or side components as inputs, as discussed herein with respect to the input audio 102, to generate the left and right channels. The component processor module 106 may also perform other types of processing on the mid and side components, and the M / SL / R converter module 108 generates the left and right channels from the processed mid and side components. The left channel generated by the HPSM module 100 is combined with the left channel generated by the M / SL / R converter module 108 to generate the processed left component. The right channel generated by the PSM module 100 is coupled with the right channel generated by the M / SL / R converter module 108 to produce the processed right component.
[0032] System 100 provides the left channel 138 to the left speaker 112 and the right channel 140 to the right speaker 114. Speakers 112 and 114 may be components of a smartphone, tablet, smart speaker, laptop, desktop, exercise machine, etc. Speakers 112 and 114 may be part of a device including System 100, or they may be separate from System 100 so as to be connected to System 100 via a network. The network may include wired and / or wireless connections. The network may include a local area network, a wide area network (e.g., the Internet), or a combination thereof.
[0033] Figure 2 is a block diagram showing a computing system environment 200 according to several embodiments. The computing system 200 may include a voice system 202, which may include one or more computing devices (e.g., servers) connected to user devices 210a and 210b via a network 208. The voice system 202 provides voice content to user devices 210a and 210b (also individually referred to as user devices 210) via the network 208. The network 208 facilitates communication between the system 202 and the user devices 210. The network 106 may include various types of networks, including the Internet.
[0034] The audio system 202 includes one or more processors 204 and a computer-readable medium 206. The one or more processors 204 execute program modules that cause the one or more processors 204 to perform functions such as generating multiple output channels from a mono channel. The processors 204 may include a central processing unit (CPU), a graphics processing unit (GPU), a controller, a state machine, other types of processing circuits, or one or more combinations thereof. The processors 204 may further include, among other things, program modules and local memory for storing operating system data.
[0035] The computer-readable medium 206 is a non-temporary storage medium that stores program code for the PSM module 102, the component processor module 106, the crosstalk processor module 110, the L / R converter module 104 and the M / S converter module 108, and the channel summing module 212. The PSM module 102 generates multiple output channels from a mono channel, which can be further processed using the component processor module 106, the crosstalk processor module 110, and / or the L / R converter module 104 and the M / S converter module 108. The system 202 provides the multiple output channels to a user device 210a, which includes multiple speakers 214 for rendering each of the output channels.
[0036] The channel summing module 212 generates a mono output channel by adding together multiple output channels generated by the PSM module 102 and / or other modules. The system 202 provides the mono output channel to the user device 210b, which includes a single speaker 216 for rendering the mono output channel. In some embodiments, the channel summing module 212 is located in the user device 210b. The audio system 202 provides multiple output channels to the user device 210b, which converts the multiple channels into a mono output channel for the speaker 216. The user device 210 presents the audio content to the user. The user device 210 may be the user's computing device, such as a music player, smart speaker, smartphone, wearable device, tablet, laptop, or desktop.
[0037] [Colorization of mid-space / side-space] In some embodiments, spatial implications are encoded in the audio signal by creating a coloration effect in the mid-space / side-space while avoiding a coloration effect in the left-space / right-space. In some embodiments, this is achieved by applying a pass-through filter with specifically selected characteristics to the left-space / right-space to produce the desired coloration in the mid-space / side-space. For example, in a two-channel system, the relationship between the left-phase angle / right-phase angle and the mid-gain / side-gain can be expressed using equation (4) below. That is,
[0038]
number
[0039] Here,
[0040]
number
[0041] This is a two-dimensional row vector composed of mid and side target gain factors in decibels at a specific frequency ω.
[0042]
number
[0043] This is the target function for the phase relationship between the left and right channels.
[0044]
number
[0045] Solving equation (4) for , we obtain the frequency-dependent phase difference required for application to the left space / right space according to equations (5) and (6) below. That is,
[0046]
number
[0047] Note that when the constraint that the system is colorless in left-right space is applied, only one of the transfer functions of the mid-component or the side-component can be specified. Therefore, the system of equations (5) and (6) becomes over-determined, and only one of the above equations can be solved without breaking the required symmetry. In some embodiments, control over either the mid-component or the side-component can be obtained by selecting a specific equation. If the constraint that the system is colorless in left-right space is removed, further degrees of freedom may be achieved. In systems with more than two channels, various techniques such as pairwise transforms or hierarchical addition and difference transforms can be used instead of mid and side transforms.
[0048] [Example implementation of a pass-through filter for encoding elevation cues] In some embodiments, spatial perceptual suggestions can be encoded into the audio signal by embedding frequency-dependent amplitude suggestions (i.e., coloration) in the mid-space / side-space, while constraining the left-right signals to be colorless. For example, elevation suggestions (e.g., spatial perceptual suggestions located along the sagittal plane) can be encoded using this framework, since the left-right suggestions for elevation are theoretically symmetric in coloration.
[0049] In some embodiments, a notable feature of the head-related transfer function (HRTF) based elevation cue is a notch that starts at approximately 8 kHz and rises monotonically as a function of elevation to approximately 16 kHz, which is used to derive a suitable coloration of the mid-channel for encoding the elevation. Using this mid-encoded cue, a corresponding frequency-dependent phase shift can be derived, which can be further used to derive a function implemented via a filter network (e.g., PSM module 100) as described later. In some embodiments, the HRTF-based elevation cue may be characterized as a notch that starts at approximately 8 kHz and rises monotonically as a function of elevation to approximately 12 kHz.
[0050] For the sake of clarity, the following examples of filter frameworks are described in relation to the encoding of the same perceptual suggestion, according to several embodiments, where the target elevation angle is 60 degrees (e.g., spatially shifting audio content 60 degrees above the horizontal). However, in other embodiments, similar techniques are understood to be used to encode perceptual suggestions with different elevation angles. Figure 3 illustrates a graph showing the HRTF of sample extraction measured at an elevation angle of 60 degrees, according to several embodiments. Figure 4 illustrates a graph showing an example of perceptual suggestion, according to several embodiments, characterized by a target amplitude function corresponding to a narrow region of infinite attenuation at approximately 11 kHz. Such suggestions could be used to create elevation angle perception for most people across a wide range of presentation scenarios. While the graph in Figure 4 illustrates a simplified HRTF of sample extraction, more complex suggestions are understood to be deriveable based on the framework described herein.
[0051] [Design using a secondary full-area pass-through section] In some embodiments, the PSM module 100 is implemented using two independent cascades and delay elements in a second-order pass-through filter to achieve a desired phase shift in left space / right space in order to encode perceptual implications as described above in relation to Figure 4. In some embodiments, the second-order sections are implemented as biquad sections, and their coefficients are applied to feedback taps and feedforward taps of up to two delayed samples. As considered herein, the convention is used to name the feedback coefficients A1 and A2 for one sample and two samples, respectively, and the feedforward coefficients B0, B1, and B2 for zero sample, one sample, and two samples, respectively.
[0052] In some embodiments, the PSM module 100 is implemented using a second-order full-pass filter configured to perform pole and zero cancellation, allowing the amplitude components of the transfer function to remain flat while the phase response is modified. By using full-pass filter sections for both channels in left and right space, a specific phase shift across the entire spectrum can be guaranteed. This has the additional advantage of allowing a given phase offset between the left and right sides, resulting in an increased sense of spatial breadth, in addition to the desired null in mid-space / side space.
[0053] Table 1 below illustrates an example set of biquad coefficients that may be used in a second-order pass-through filter framework with an additional two-sample delay on the right channel, according to several embodiments. The biquad coefficients illustrated in Table 1 may be designed for a 44.1 kHz sampling rate, but may be used in systems with other sampling rates (e.g., 48 kHz).
[0054] [Table 1]
[0055] A network of filters with the coefficients shown in Table 1 can produce a suitable phase response in the left / right space, resulting in a significant null / amplification in the mid / side space at 11 kHz. Figure 5 illustrates frequency plots generated by driving a second-order pass-through filter section with the coefficients shown in Table 1 with white noise, according to several embodiments, showing the output frequency response of the sum of multiple channels (mid) 502 and the difference of multiple channels (side) 504.
[0056] In some embodiments, the PSM module 100, implemented using a second-order full-pass filter section, may be further extended using a crossover network to exclude processing in frequency domains where it is not needed. The use of a crossover network can increase the flexibility of the embodiment by allowing further processing of perceptually important implications in order to eliminate unnecessary auditory data.
[0057] In some embodiments, the PSM module 100, which is implemented using a second-order pass-through filter section, may be implemented using a network of sequentially chained Hilbert transforms, as will be described in more detail later.
[0058] [Example of a Hilbert Transformation Perceptual Soundstage Modification (HPSM) module] Figure 6 is a block diagram showing a PSM module implemented using a Hilbert transform, according to one or more embodiments. The PSM module 600, also known as a Hilbert transform perceptual soundstage modification (HPSM) module, applies a network of serially chained Hilbert transforms to an input sound 602 (which may correspond to the input sound 120 shown in Figure 1) to perceptually shift the input sound 602.
[0059] Module 600 includes a crossover network module 604, a gain unit 610, a gain unit 612, a Hilbert transducer module 614, a Hilbert transducer module 620, a delay unit 626, a gain unit 628, a delay unit 630, a gain unit 630, a gain unit 632, an adder unit 634, and an adder unit 636. Some embodiments of Module 600 have components different from those described herein. Similarly, in some cases, functions may be distributed among the components in a manner different from that described herein.
[0060] The crossover network module 604 receives the input audio 602 and generates a low-frequency component 606 and a high-frequency component 608. The low-frequency component includes a subband of the input audio 602 having frequencies lower than the subband of the high-frequency component 608. In some embodiments, the low-frequency component 606 includes a first portion of the input audio that includes low frequencies, and the high-frequency component 608 includes the remainder of the input audio that includes high frequencies.
[0061] As will be described in more detail later, the high-frequency component 608 is processed using a series of Hilbert transforms, while the low-frequency component 606 bypasses the series of Hilbert transforms, after which the low-frequency component and the processed high-frequency component 608 are recombined. The crossover frequency between frequency component 606 and high-frequency component 608 may be adjustable. For example, more frequencies may be included in the high-frequency component 608 to increase the perceived intensity of the spatial shift by the HPSM module 600, while more frequencies may be included in the low-frequency component 606 to decrease the perceived intensity of the shift. In another example, the crossover frequency is set so that frequencies corresponding to the sound source of interest (e.g., vocalizations) are included in the high-frequency component 608.
[0062] The input audio 602 may include a mono channel, or it may be a mixdown of a stereo signal or other multi-channel signal (e.g., surround sound, ambisonics, etc.). In some embodiments, the input audio 602 is audio content associated with a sound source that should be incorporated into the audio mix. For example, the input audio 602 may be a speech that is processed by module 600, and the processing result is combined with other audio content (e.g., background music) to generate the audio mix.
[0063] Gain unit 610 applies gain to the low-frequency component 606, and gain unit 612 applies gain to the high-frequency component 608. Gain units 610 and 612 may be used to adjust the overall levels of the low-frequency component 606 and the high-frequency component 608 relative to each other. In some embodiments, gain unit 610 or gain unit 612 may be omitted from module 600.
[0064] Hilbert transformer modules 614 and 620 apply a series of Hilbert transforms to the high-frequency component 608. Hilbert transformer module 614 applies the Hilbert transform to the high-frequency component 608 to generate the left-foot component 616 and the right-foot component 618. The left-foot component 616 and the right-foot component 618 are speech components that are 90 degrees out of phase with respect to each other. In some embodiments, the left-foot component 616 and the right-foot component 618 are out of phase with respect to each other at an angle other than 90 degrees, such as between 20 and 160 degrees.
[0065] The Hilbert transform module 620 applies a Hilbert transform to the right leg component 618 generated by the Hilbert transform module 614 to generate the left leg component 122 and the right leg component 624. The left leg component 622 and the right leg component 624 are speech components that are 90 degrees out of phase with respect to each other. In some embodiments, the Hilbert transform module 620 generates the right leg component 624 without generating the left leg component 122. In some embodiments, the left leg component 622 and the right leg component 624 are out of phase with respect to each other at an angle other than 90 degrees, such as between 20 and 160 degrees.
[0066] In some embodiments, each of the Hilbert transducer modules 614 and 620 is implemented in the time domain and includes cascaded pass-through filters and delays, as will be described in more detail later in relation to Figure 7. In other embodiments, the Hilbert transducer modules 614 and 620 are implemented in the frequency domain.
[0067] The delay unit 626, gain unit 628, delay unit 630, and gain unit 632 provide adjustment controls for manipulating the perceived results of the process by module 600. Delay unit 626 applies a time delay to the left leg component 616 generated by the Hilbert transducer module 614. Gain unit 628 applies a gain to the left leg component 616. In some embodiments, delay unit 626 or gain unit 628 may be omitted from module 600.
[0068] The delay unit 630 applies a time delay to the right-foot component 624 generated by the Hilbert transducer module 620. The gain unit 632 applies gain to the right-foot component 624. In some embodiments, the delay unit 630 or the gain unit 632 may be omitted from module 600.
[0069] The summing unit 634 combines the low-frequency component 606 with the left-leg component 616 to generate the left channel 642. The left-leg component 616 is the output from the first Hilbert converter module 614 in the series. The left-leg component 616 may include a delay applied by the delay unit 626 and a gain applied by the gain unit 628.
[0070] The summing unit 636 combines the low-frequency component 606 with the right-foot component 624 to generate the right channel 644. The right-foot component 624 is, in this sequence, the output from the second Hilbert transducer module 620. The right-foot component 624 may include the delay applied by the delay unit 626 and the gain applied by the gain unit 628.
[0071] Figure 7 is a block diagram showing one or more embodiments of a Hilbert transducer module 700. The Hilbert transducer module 700 is an example of a Hilbert transducer module 614 or a Hilbert transducer module 620. The Hilbert transducer module 700 receives an input component 702 and uses the input component 702 to generate a left-foot component 712 and a right-foot component 724. Some embodiments of the Hilbert transducer module 700 have components different from those described herein. Similarly, in some cases, functions may be distributed among the components in a manner different from that described herein.
[0072] The Hilbert converter module 700 includes a cascade module 740 of pass-through filters for generating the left leg component 712, a delay unit 714, and a cascade module 742 of pass-through filters for generating the right leg component 724. The cascade module 714 of pass-through filters includes a series of pass-through filters 704, 706, 708, and 710. The delay unit 714 applies a time delay to the input component 702. The cascade module 742 of pass-through filters includes a series of pass-through filters 716, 718, 720, and 722. Each of the pass-through filters 704 to 710 and 716 to 722 passes frequencies with equal gain while varying the phase relationship between different frequencies. In some embodiments, each of the pass-through filters 704 to 710 and 716 to 722 is a biquad filter as defined by equation (7).
[0073]
number
[0074] Here, z is a complex variable, and a0, a1, a2, b0, b1, and b2 are digital filter coefficients. Different biquadratic filters may contain different coefficients to apply different phase shifts.
[0075] The pass-through filter cascade modules 740 and 742 may each contain a different number of pass-through filters. The Hilbert transducer module 700 is an eighth-order filter with a total of eight pass-through filters, four for each of the left leg component 712 and the right leg component 724. In other embodiments, the Hilbert transducer module 700 is an eighth-order filter (e.g., four pass-through filters for each of the pass-through filter cascade modules 740 and 742) or a sixth-order filter (e.g., three pass-through filters for each of the pass-through filter cascade modules 740 and 742).
[0076] As discussed above in relation to Figure 6, module 600 includes a series of Hilbert converter modules 614 and 620. Using Hilbert converter module 700 for each of Hilbert converter modules 614 and 620, the left leg component 616 is generated by a single pass-through filter cascade module 740 applied to the high-frequency component 608. The right leg component 624 is generated by two passes through Hilbert converter module 700 by a cascade module 742 of two delay units 714 and two pass-through filters. In some embodiments, Hilbert converter modules 614 and 620 may be different. For example, Hilbert converter modules 614 and 620 may include filters of different orders, such as an 8th-order filter for one of the Hilbert converter modules and a 6th-order filter for the other Hilbert converter module.
[0077] When Hilbert transformer module 700 is used with Hilbert transformer modules 614 and 620, the right-foot component 624 includes a phase and delay relationship with the right-foot component 618, which is generated by the pass-through filter and the delay of Hilbert transformer module 620. The right-foot component 624 also includes a phase and delay relationship with the high-frequency component 608, which is generated by the pass-through filter and delay in Hilbert transformer modules 614 and 620. In some embodiments, Hilbert transformer module 620 uses the left-foot component 616 instead of the right-foot component 618 to generate the left-foot component 622 and the right-foot component 624. This results in the right-foot component 624 having a phase and delay relationship with the pass-through filter (e.g., and without delay) of Hilbert transformer 614 and the high-frequency component 608, which is generated by the delay and pass-through filter of Hilbert transformer module 620.
[0078] Figure 8 illustrates frequency plots generated by driving an HPSM module (as described in Figure 6) with white noise according to several embodiments, showing the output frequency responses of the sum 802 of multiple channels (mids) and the difference 804 of multiple channels (sides).
[0079] As shown in Figure 8, this filter certainly produces the desired perceptual suggestion in the approximately 11 kHz range, while also adding additional coloration to the mids and sides at lower frequencies. In some embodiments, this can be corrected by applying a crossover network (such as the crossover network module 604 shown in Figure 6) to the input audio so that the HPSM module processes only the audio data within the desired frequency range (e.g., high-frequency components), or by directly removing the pole / zero pairs corresponding to that region of the spectral transform.
[0080] [Design using First-Order Non-Orthogonal Rotation-Based Correlation Removal (FNORD)] In some embodiments, similar perceptual effects can be achieved using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network. Figure 9 is a block diagram showing a PSM module 900 implemented using an FNORD filter network in some embodiments. The PSM module 900 may correspond to the PSM module 102 illustrated in Figure 1, but is configured to decorrelate a mono channel into multiple channels and includes an amplitude response module 902, a full-pass filter configuration module 904, and a full-pass filter module 906. The PSM module 900 processes the mono input channel x(t) 912 to provide channel y to the speaker 910a. a (t), and Channel y provided to speaker 910b (which may correspond to the left speaker 112 and right speaker 114 illustrated in Figure 1) bIt generates multiple output channels, such as (t). Although two output channels are illustrated, the PSM module 900 can generate any number of output channels (each referred to as channel y(t)). The PSM module 900 can be implemented as part of a computing device such as a music player, speaker, smart speaker, smartphone, wearable device, tablet, laptop, or desktop. Figure 9 illustrates the PSM module 900 as including an amplitude response module 902 and a filter configuration module 904 in addition to the pass-through filter module 906, but in some embodiments, the PSM module 900 may include a pass-through filter module 906 with an amplitude response module 902 and / or a filter configuration module 904 that are implemented separately from the PSM module 900.
[0081] The amplitude response module 902 determines a target amplitude response that defines one or more spatial implications to be encoded in the output channel y(t) (e.g., the mid and side components of the output channel y(t)). The target amplitude response is defined by the relationship between the amplitude and frequency values of the channel (e.g., the mid and side components of the channel), such as amplitude as a function of frequency. In some embodiments, the target amplitude response defines one or more spatial implications on the channel, which may include a target broadband attenuation, target subband attenuation, critical point, filter characteristics, or soundstage position. The amplitude response module 902 may receive data 914 and a mono channel x(t) 912 and use these inputs to determine the target amplitude response. Data 914 may include information such as the characteristics of the spatial implications to be encoded, the characteristics of the presentation equipment (e.g., one or more speakers), the expected content of the audio data, or the listener's perceptual ability in a scene. In some embodiments, the mono channel x(t)912 may correspond to the audio input 120 illustrated in Figure 1, or a portion of the audio input (for example, a high-frequency component of the input audio, such as the high-frequency component 608 of the input audio 602 illustrated in Figure 6). In embodiments where the mono channel x(t)912 corresponds to a portion of the audio input, the output channel y(t) may be coupled with a channel corresponding to the remainder of the audio input (for example, having low-frequency components, as illustrated in Figure 1) to produce a combined output channel.
[0082] The target broadband attenuation is the specification for attenuation across all frequencies. The target subband attenuation is the specification for amplitude over the frequency range defined by the subband. The target amplitude response may include one or more target subband attenuation values for different subbands.
[0083] A critical point is a specification of the curvature of the filter's target amplitude response, described as a frequency value where the gain for one of the output channels (e.g., the side component of the output channel) is a predefined value, such as -3dB or -∞dB. The location of this point can have an overall impact on the curvature of the target amplitude response. One example of a critical point is the frequency at which the target amplitude response becomes -∞dB. This critical point is a null point because the behavior of the target amplitude response is to nullify the signal at frequencies close to this point. Another example of a critical point is the frequency at which the target amplitude response becomes -3dB. This critical point is a crossover point because the behavior of the target amplitude response for the summing and differenceping channels (e.g., the mid- and side components of the channels) intersect at this point.
[0084] Filter characteristics are parameters that specify how the mid and side components of a channel are filtered out. Examples of filter characteristics include high-pass, low-pass, band-pass, or band-reject characteristics. Filter characteristics describe the shape of the resulting sum, as if it were the result of equalization filtering. Equalization filtering can be described in terms of which frequencies can pass through the filter and which frequencies are rejected. Thus, a low-pass characteristic allows frequencies below the inflection point to pass through and frequencies above the inflection point to be attenuated. A high-pass characteristic does the opposite by allowing frequencies above the inflection point to pass through and frequencies below the inflection point to be attenuated. A band-pass characteristic allows frequencies in the band around the inflection point to pass through and other frequencies to be attenuated. A band-reject characteristic rejects frequencies in the band around the inflection point and allows other frequencies to pass through.
[0085] The target amplitude response can define multiple spatial implications encoded in the output channel y(t). For example, the target amplitude response may specify spatial implications designated by the critical point and the filter characteristics of the mid- or side-pass components of a full-pass filter. In another example, the target amplitude response may specify spatial implications designated by the target broadband attenuation, critical point, and filter characteristics. Although described as independent specifications, these specifications may be interdependent over most regions of the parameter space. This result may be due to the system being nonlinear with respect to phase. To address this, additional, higher-level descriptors of the target amplitude response can be devised, which are nonlinear functions of the target amplitude response parameters.
[0086] The filter configuration module 904 determines the characteristics of a single-input, multiple-output whole-pass filter based on the target amplitude response received from the amplitude response module 902. Specifically, the filter configuration module determines the transfer function of the whole-pass filter based on the target amplitude response, and determines the coefficients of the whole-pass filter based on that transfer function. The whole-pass filter is an uncorrelated filter that encodes spatial implications described with respect to the target amplitude response, applied to a monaural input channel x(t) and output channel y a (t) and y b Generate (t).
[0087] A full-pass filter can have various configurations and parameters based on the spatial implications and / or constraints defined by the target amplitude response. A filter with a target amplitude response of the spatial implications to be encoded can be colorless, preserving the spectral content (e.g., the whole) of individual output channels (e.g., left / right output channels). Therefore, this filter is used to encode elevation indications by embedding coloration in the mid-space / side-space in the form of frequency-dependent amplitude indications while maintaining the spectral content of the left and right signals. Because the filter is colorless, monaural content can be placed at a specific position in the soundstage (e.g., as specified by the target elevation angle), and the spatial placement of the sound is decoupled from its overall coloration.
[0088] Figures 10A and 10B are block diagrams showing examples of PSM modules based on first-order non-orthogonal rotation-based decorrelation (FNORD) technique according to several embodiments. Figure 10A shows a detailed view of the PSM module 900 according to several embodiments, while Figure 10B provides a more detailed view of the broadband phase rotator 1004 within the full-pass filter module 906 of the PSM module 900 according to several embodiments.
[0089] As shown in Figure 10A, the full-pass filter module 906 receives a monaural input audio signal x(t) 912 and a rotation control parameter θ. bf 1048, and the linear coefficient β bf Information is received in the format 1050. Input audio signal x(t)912 and rotation control parameter θ bf 1048 is utilized by a broadband phase rotator 1004, which controls the rotation parameter θ bfThe input audio signal 912 is processed using 1048 to generate a left wideband rotation component 1020 and a right wideband rotation component 1022. Next, the left wideband rotation component 1020 is provided to a narrow-band phase rotator 1024 for further processing, while the right wideband rotation component 1022 is output as the output channel y b (t) of the PSM module 900 (e.g., as the right output channel). The narrow-band phase rotator 1024 receives the left wideband rotation component 1020 from the wide-band phase rotator 1004 and receives the primary coefficient β bf 1050 from the filter configuration module 904 to generate a narrow-band rotation component 1028, and the narrow-band rotation component 1028 is provided as the output channel y a (t) of the PSM module 900 (e.g., such as the left output channel).
[0090] According to some embodiments, the control data 914 for configuring the amplitude response module 902 may include a critical point f c 1038, a filter characteristic θ bf 1036, and a sound stage position Γ 1040. This data is provided to the PSM module 900 via the amplitude response module 902, and the amplitude response module 902 determines a mid representation of the data in the form of a critical point coc 1044 (in radians), a filter characteristic θ bf 1042, and a secondary term φ 1046. In some embodiments, the amplitude response module 902 changes one or more of the parameters of the control data 914 (e.g., the critical point f c 1038, the filter characteristic θ bf 1036, and / or the sound stage position Γ 104) based on one or more parameters of the input audio signal x(t) 912. In some embodiments as shown in FIG. 10A, the filter characteristic θ bf 1042 is the filter characteristic θ bfIt is equivalent to 1036. These mid-representations 1042, 1044, and 1046 are provided to the filter configuration module 904, which provides at least the first-order coefficient β bf Generate filter configuration data that may include 1050 and rotation control parameter 1048. (First-order coefficient β) bf 1050 is provided to the pass-through filter module 906 via the primary pass-through filter 1026. In some embodiments, the rotation control parameter θ bf 1048 is a filter characteristic θ bf While equivalent to 1036 and 1042, in other embodiments, this parameter may be scaled for convenience. For example, in some embodiments, the filter characteristics are associated with a parameter range (e.g., 0 to 0.5) having a meaningful center point, and the rotation control parameter is scaled relative to the filter characteristics to change the parameter range, for example, from 0 to 1. In some embodiments, the filter characteristics are scaled linearly (e.g., to maintain an increase in resolution at the poles compared to the center point), but in other embodiments, a nonlinear mapping may be used (e.g., to increase the numerical resolution with respect to the center point). In the following equation, the rotation control parameter θ bf Although it is treated as not being scaled, it is understood that the same principle can be applied even when the rotation control parameter is scaled. Rotation control parameter θ bf 1048 is supplied to the whole-pass filter module 906 via the broadband phase rotater 1004.
[0091] Figure 10B illustrates in detail several implementation examples of the broadband phase rotator 1004 according to various embodiments. The broadband phase rotator 1004 uses a monaural input audio signal x(t)912 and a rotation control parameter θ bfInformation is received in the format 1048. The input audio signal x(t)912 is first processed by the Hilbert converter module 1006 to generate the left leg component 1008 and the right leg component 1010. The Hilbert converter module 1006 may be implemented using the configuration shown in Figure 7, but it is understood that other implementations of the Hilbert converter module 1006 may be used in other embodiments. The left leg component 1008 and the right leg component 1010 are provided to the 2D orthogonal rotation module 1012. The left leg component 1008 is also provided to the output of the broadband phase rotator 1004 as the right broadband rotation component 1022. Since the broadband phase rotator 1004 is configured to rotate the left leg and right leg signals relative to each other, one way to achieve this in some embodiments is to keep the left leg component 1008 constant as the right broadband rotation component 1022 and rotate the left leg component and the right leg component to form the left broadband rotation component 1020.
[0092] In addition to the left leg component 1008 and the right leg component 1010, the 2D orthogonal rotation module 1012 controls the rotation parameter θ from the filter configuration module 904 according to several embodiments, as shown in Figure 10A. bf 1048 can also be received. The 2D orthogonal rotation module 1012 uses this data to generate a left rotation component 1014 and a right rotation component 1016. The projection module 1018 then receives the left rotation component 1014 and the right rotation component 1016, which are combined (e.g., added) to form a left broadband rotation component 1020. As shown in Figure 10A, the broadband phase rotator 1004 receives the left output channel y of the PSM module. a (t) is the narrowband rotation component 1028, and the right output channel y of the PSM module bThe left broadband rotation component 1020 is output to the narrowband phase rotator 1024 in order to generate a right broadband rotation component 1022 as (t) (which bypasses the narrowband phase rotator 1024 or passes through it unchanged). In other embodiments, the narrowband rotation component 1028 and the left leg component 1008 (which functions as the right broadband rotation component 1022 in the embodiments shown in Figures 10A and 10B) are instead output to the right output channel y b (t) and left output channel y a Each receives a mapping to (t).
[0093] In some embodiments, the PSM module 900 can be formally described by the following equation (8):
[0094]
number
[0095] In some embodiments, this single-input, multiple-output pass-through filter consists of several parts, each of which will be described in turn. According to some embodiments, these components include A f , A b , and H2 may be included.
[0096] According to some embodiments, A f This can correspond to the narrowband phase rotator 1024 in Figure 10A. f This is a first-order pass filter having a single-channel output that assumes the form of equation (9).
[0097]
number
[0098] Here, β f is the filter coefficient, which is in the range of -1 to +1. The second output of the filter can simply pass through without changing the input. Therefore, according to some embodiments, filter A fThe implementation of can be defined by equation (10). That is,
[0099]
number
[0100] A f The transfer function is the differential phase shift from one output to the other.
[0101]
number
[0102] This is expressed as follows. This difference phase shift is a function of the center frequency (radian frequency) ω, as defined by equation (11).
[0103]
number
[0104] Here, the target amplitude response is determined by using either equation (5) or (6) depending on whether the response should be positioned in the mid (equation (5)) or the side (equation (6)).
[0105]
number
[0106] A total gain αf = 3dB can be used as a critical point for adjustment at the frequency f c It is defined by the following formula: That is,
[0107]
number
[0108] and
[0109]
number
[0110] By normalizing the target amplitude response to 0 dB, this critical point is obtained by parameter f c This corresponds to a point of -3dB. In equation (8), A f The output is subscripted to show only the output of the first channel used according to some embodiments.
[0111] In equation (8), A b This is a single-input, multiple-output full-pass filter, which can correspond to the broadband phase rotator 1004 in Figure 10A. b It can be formally defined as shown in equation (14).
[0112]
number
[0113] Here, H2(x(t)) is the discrete form of the filter, implemented using a pair of orthogonal total-pass filters, and defined using a continuous-time prototype according to equation (15). That is,
[0114]
number
[0115] In some embodiments, a full-pass filter
[0116]
number
[0117] This imposes constraints on a 90-degree phase relationship between the two output signals, as well as a unity-magnitude relationship between the input signal and both output signals, but does not necessarily guarantee a specific phase relationship between the input (mono) signal and either of the two (stereo) output signals.
[0118]
number
[0119] The discrete form is
[0120]
number
[0121] It is denoted as and defined by its action on a monaural signal x(t). The result is a two-dimensional vector, as defined by equation (16). That is,
[0122]
number
[0123] A discrete single-input, multiple-output full-pass filter.
[0124]
number
[0125] According to several embodiments, this corresponds to the Hilbert transducer module 1006 in Figure 10B and can also correspond to the Hilbert transducer module 700 in Figure 7. In equation (14), θ determines the rotation angle of the first output relative to the second output of Ab, according to several embodiments.
[0126] Finally, parameter A supplied to the complete system in equation (8) bfThese parameters can be determined according to several embodiments as follows: These parameters include the rotation control parameter θ in Figure 10A. bf 1048 and the linear coefficient β bf β bf and θ bf It may include β. In some embodiments, bf This is the center radian frequency (ω). c Therefore, it can be determined as follows:
[0127]
number
[0128] Here, ω c The desired center frequency f is calculated using equation (12). c It can be calculated from ω. In Figure 10A, c ω is the critical point c Compatible with 1044, f c The critical point f c Corresponding to 1038, the operation of equation (17) is partially performed within the filter configuration module 904, resulting in the linear coefficient β bf This results in the quadratic term φ1046, via equation (18), θ bf And it can be derived from the Boolean soundstage position parameter Γ. That is,
[0129]
number
[0130] This quadratic term φ1046 is provided to the filter configuration module 904 by the amplitude response module 902 in Figure 10A.
[0131] In some embodiments, the high-level parameter f c , θ bf , and Γ may be sufficient to intuitively and conveniently adjust this system. According to such embodiments, the center frequency fc This determines the inflection point in Hz where the target amplitude response asymptotically approaches -∞dB. Parameter θ bf Therefore, the inflection point f c This enables control over the filter characteristics related to 0 < θ. bf When <1 / 4, the characteristic is low-pass, null at fc, and the spectral gradient in the target amplitude function is θ bf As the coefficient increases, it interpolates smoothly from favorable low frequencies to a flat response. 1 / 4 < θ bf If <1 / 2, θ bf As it increases, its properties change, f c It smoothly interpolates from a flat frequency with a null in θ to a high-pass frequency range. bf At the point where =1 / 4, the target amplitude function is purely band-rejected, and f c It becomes null in this case. The parameter Γ is f c and θ bf This is a Boolean value that places the target amplitude function, determined by Γ, in either the mid-channel (i.e., L+R) or the side channel (i.e., LR). Due to the full-range constraint on both outputs to the filter network, the action of Γ is to switch between complementary target amplitude responses.
[0132] In some embodiments, to achieve an amplitude response to a 60-degree vertical suggestion, the FNORD filter network described above uses parameter f c =11kHz, θ bf It can be configured using =0.13 and Γ=1. Figure 11 illustrates a frequency response graph showing the output frequency response of an FNORD filter network configured to achieve an amplitude response to a 60-degree vertical suggestion according to several embodiments. Figure 11 illustrates the output frequency response at the mid component 1110 and the side component 1120, and the FNORD filter network is driven by white noise. In some embodiments, the filter parameter f c , θ bf , and / or Γ are selected based on an analysis of HRTF-based elevation angle suggestions at the desired angle.
[0133] In some embodiments, the PSM module 900 uses a frequency-domain specification for the pass-through filter. For example, in some cases, more complex spatial implications, such as those sampled from anthropometric datasets, may be required. Within certain limitations, the techniques described above are used to embed arbitrary implications into the phase difference of the audio stream based on the magnitude frequency-domain representation of the implications. For example, the filter configuration module 904 uses an equation in the form of equation (5) or (6) to vectorize K phase angles from the target amplitude response of K narrow-band attenuation coefficients at the mid or side.
[0134]
number
[0135] The vectorized transfer function can be determined. That is,
[0136]
number
[0137] The phase angle vector θ generates a finite impulse response filter defined by equation (19). That is,
[0138]
number
[0139] Here, DFT -1 This is the inverse Discrete Fourier Transform (IDFT) and
[0140]
number
[0141] This represents the following. Then, the vector of 2(K-1) IFR filter coefficients Bn(θ) is applied to x(t) as defined by equation (20). That is,
[0142]
number
[0143] Here,
[0144]
number
[0145] This represents a convolution operation.
[0146] To reproduce the effect from the previous example and achieve the target amplitude response corresponding to a 60-degree height cue, the observed HRIR
[0147]
number
[0148] A sample is extracted and applied to a DFT of length 2(K-1), and as a result
[0149]
number
[0150] This can generate the target amplitude response vector using the following calculation.
[0151]
number
[0152] It can be used to determine, that is,
[0153]
number
[0154] Here,
[0155]
number
[0156] and
[0157]
number
[0158] These operations return the real and imaginary components of a complex number, and all operations are applied to each vector component. This target amplitude response, inserted into either the mid or side, is applied to one of equations (5) or (6) to form a vector of K phase angles.
[0159]
number
[0160] This can be determined, and from this, FIR filter B can be derived. Next, this filter is inserted into equation (19) to derive a single-input, multiple-output, full-pass filter.
[0161] Equations (19) and (20) provide effective means for constraining the target amplitude response, but their implementation often relies on relatively high-order FR filters obtained as a result of inverse DFT operations. This may not be suitable for resource-constrained systems. In such cases, a low-order infinite impulse response (IIR) implementation may be used, as described in relation to equation (8).
[0162] The whole-pass filter module 906, composed of the filter configuration module 904, applies the whole-pass filter to the mono channel x(t) and outputs to the output channel y a (t) and y b (t) is generated. The application of a pass filter to channel x(t) can be performed as defined by equations (8), (20) or as depicted in Figure 9 or Figure 10A. The pass filter module 906 processes channel y a (t) to speaker 910a, channel y b Each output channel is provided to its respective speaker, such as (t) to speaker 910b. Although not shown in Figure 9, output channel y a (t) and y b (t) is understood to be provided to speakers 910a and 910b via one or more intervening components (for example, component processor module 106, crosstalk processor module 110, and / or L / RM / S converter module 104 and M / SL / R converter module 108, as shown in Figure 1).
[0163] [Hypermid Processing] In some embodiments, the PSM process may be performed on a target portion of the received audio signal, such as the mid component of the audio signal or the hyper-mid component of the audio signal. FIG. 12 is a block diagram showing an audio processing system 1200 according to one or more embodiments. The system 1200 generates a hyper-mid component, separates a target portion (e.g., voicing, etc.) of the audio signal, and performs a PSM process on the hyper-mid component to spatially shift the target portion. Some embodiments related to the system 1200 have components different from those described herein. Similarly, in some cases, functions can be distributed among components in a manner different from that described herein.
[0164] The system 1200 includes an L / R-M / S converter module 1206, an orthogonal component generator module 1212, an orthogonal component processor module 1214 including a PSM module 102, and a crosstalk processor module 1224.
[0165] The L / R-M / S converter module 1206 receives the left channel 1202 and the right channel 1204 and generates a mid component 1208 and a side component 1210 from the channels 1202 and 1204. The description of the L / R-M / S converter module 104 may be applicable to the L / R-M / S converter module 1206.
[0166] The orthogonal component generator module 1212 processes the mid component 1208 and the side component 1210 to generate at least one of a hyper mid component M1, a hyper side component S1, a residual mid component M2, and a residual side component S2. The hyper mid component M1 is the spectral energy of the spectral energy mid component 1208 from which the spectral energy of the side component 1210 has been removed. The hyper side component S1 is the spectral energy of the mid component 1208 from which the spectral energy of the side component 1210 has been removed. The residual mid component M2 is the spectral energy of the hyper mid component M1 from which the spectral energy of the mid component 1208 has been removed. The mid component M2 is the spectral energy of the hyper mid component M1 from which the spectral energy of the mid component 1208 has been removed. The residual side component S2 is the spectral energy of the hyper side component 1210 from which the spectral energy of the hyper side component S1 has been removed. The system 1200 generates a left channel 1242 and a right output channel 1244 by processing at least one of the hyper mid component M1, the hyper side component S1, the residual mid component M2, and the residual side component S2. The orthogonal component generator module 1212 will be further described with reference to FIGS. 13A, 13B, and 13C.
[0167] The orthogonal component processor module 1214 processes one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and / or residual side component S2, and converts the processed components into a processed left component 1220 and a processed right component 1222. The description of the component processor module 106 may be applicable to the orthogonal component processor module 1214, except that the processing is performed on the hypermid component M1, hyperside component S1, residual mid component M2, and / or residual side, rather than the mid and side components. For example, processing on components M1, M2, S1, and S2 may include various types of processing, such as processing of spatial implications (e.g., amplitude or delay-based panning, binaural processing, etc.), single-band or multi-band equalization, single-band or multi-band dynamics processing (e.g., compression, expansion, limiting, etc.), single-band or multi-band gain or delay steps, addition of sound effects, or other types of processing. In some embodiments, the orthogonal component processor module 1214 performs subband spatial processing and / or crosstalk compensation processing using the hypermid component M1, the hyperside component S1, the residual mid component M2, and / or the residual side component S2. The orthogonal component processor module 1214 may further include an L / RM / S converter for converting components M1, S2, S1, and S2 into a processed left component 1220 and a processed right component 1222.
[0168] The orthogonal component processor module 1214 further includes a PSM module 102 which can operate on one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and / or residual side component S2. For example, the PSM module 102 may receive the hypermid component M1 as input and generate spatially shifted left and right channels. The hypermid component M1 may include, for example, an isolated portion of a speech signal representing a utterance and can therefore be selected for HPSM processing. The left channel generated by the PSM module 102 is used to generate the processed left component 1020, and the right channel generated by the PSM module 102 is used to generate the processed right component 1222. The orthogonal component processor module 1214 will be further described with reference to Figure 12.
[0169] The crosstalk processor module 1224 receives the processed left component 1220 and the processed right component 1222 and performs crosstalk processing on them. The crosstalk processor module 1224 outputs the left channel 1242 and the right channel 1244. The description of the crosstalk processor module 1224 may be applicable to the crosstalk processor module 1224. In some embodiments, crosstalk processing (e.g., simulation or cancellation) may be performed before orthogonal component processing, such as before the conversion of the left channel 1202 and the right channel 1204 to the mid and side components. The left channel 1242 may be provided to the left speaker 112, and the right channel 1244 may be provided to the right speaker 114.
[0170] [Example of an orthogonal component generator] Figures 13A to 13C are block diagrams showing orthogonal component generator modules 1313, 1323, and 1343 according to one or more embodiments, respectively. Orthogonal component generator modules 1313, 1323, and 1343 are examples of orthogonal component generator module 1212. Some embodiments of modules 1313, 1323, and 1343 have components different from those described herein. Similarly, in some cases, functions can be distributed among the components in a manner different from that described herein.
[0171] Referring to Figure 13A, the orthogonal component generator module 1313 includes subtraction units 1305, 1309, 1315, and 1319. As described above, the orthogonal component generator module 1313 receives the mid component 1208 and the side component 1210 and outputs one or more of the hyper-mid component M1, hyper-side component S1, residual mid component M2, and residual side component S2.
[0172] The subtraction unit 1305 generates a hypermid component M1 by removing the spectral energy of the side component 1210 from the spectral energy of the mid component 1208. For example, the subtraction unit 1305 generates a hypermid component M1 by subtracting the amplitude of the side component 1210 in the frequency domain from the amplitude of the mid component 1208 in the frequency domain, while retaining only the phase. Subtraction in the frequency domain can be performed using a Fourier transform on a time-domain signal to generate a signal in the frequency domain, and then subtracting that signal in the frequency domain. In other examples, subtraction in the frequency domain can also be performed in other ways, such as using a wavelet transform instead of a Fourier transform. The subtraction unit 1309 generates a residual mid component M2 by removing the spectral energy of the hypermid component M1 from the spectral energy of the mid component 1208. For example, the subtraction unit 1309 subtracts the amplitude of the hypermid component M1 in the frequency domain from the amplitude of the mid component 1208 in the frequency domain, while retaining only the phase, to generate the residual mid component M2. While subtracting the side from the mid in the time domain yields the right channel of the original signal, the above operation in the frequency domain separates and distinguishes the portion of the mid component with a spectral energy different from that of the mid component (called M1 or hypermid) from the portion of the side component with a spectral energy equal to that of the side component (called M2 or residual mid).
[0173] In some embodiments, if subtracting the spectral energy of side component 1210 from the spectral energy of mid component 1006 results in a negative value for hypermid component M1 (for example, for one or more bins in the frequency domain), additional processing may be used. In some embodiments, if subtracting the spectral energy of side component 1210 from the spectral energy of mid component 1208 results in a negative value, hypermid component M1 is fixed to 0. In some embodiments, hypermid component M1 is returned to its minimum value by taking the absolute value of the negative value as the value of hypermid component M1. If subtracting the spectral energy of side component 1210 from the spectral energy of mid component 1208 results in a negative value for M1, other types of processing may be used. Similar additional processing may be used if the result of the subtraction that generates hyperside component S1, residual side component S2, or residual mid component M2 is negative, such as being fixed at 0, returning to the minimum value (wrap around), or other processing. Fixing the hypermid component M1 to 0 provides spectral orthogonality between M1 and both side components when the result of subtraction is negative. Similarly, fixing the hyperside component S1 to 0 provides spectral orthogonality between S1 and both mid components when the result of subtraction is negative. By creating orthogonality between the hypermid component and the side components, and their corresponding appropriate mid / side components (i.e., side components relative to hypermid, mid components relative to hyperside, etc.), the introduced residual mid component M2 and residual side component S2 contain spectral energies that are not orthogonal to (i.e., common with) their corresponding appropriate mid / side components. That is, when fixing the hypermid to 0 and using its M1 component to derive the residual mid, a hypermid component is produced that does not have spectral energies common with the side components, and a residual mid component that has spectral energies completely common with the side components.If the hyperside is fixed at 0, the same relationship applies to the hyperside and residual side. When applying frequency domain processing, there is typically a trade-off in resolution between frequency information and time information. As frequency resolution increases (i.e., as the FFT window size and the number of frequency bins increase), time resolution decreases, and vice versa. Because the spectral subtraction described above occurs on a frequency bin basis, in certain situations, such as when removing vocal energy from the hypermid component, a larger FFT window size (e.g., 8192 samples, giving a real-valued input signal, resulting in 4096 frequency bins) may be preferable. In other situations, higher time resolution may be required, thus necessitating lower overall latency and lower frequency resolution (e.g., a 512-sample FFT window size, giving a real-valued input signal, resulting in 256 frequency bins). In the latter case, the low-frequency resolution of the mid and side components, when subtracted from each other to derive the hypermid component M1 and hyperside component S1, can generate audible spectral artifacts because the spectral energy of each frequency bin is an average representation of energy over a frequency range that is too broad. In this case, obtaining the absolute value of the difference between the mid and side components when deriving the hypermid M1 or hyperside S1 can help mitigate perceptual artifacts by allowing deviations from true orthogonality in the components for each frequency bin. In addition to, or instead of, folding zero, a coefficient can be applied to the subtracted value to scale it between 0 and 1, so that at one extreme (i.e., when the value is 1), there is perfect orthogonality between the hyper and residual mid / side components, and at the other extreme (i.e., when the value is 0), there is a method for interpolating between their corresponding original mid and side components and the identical hypermid M1 and hyperside S1.
[0174] The subtraction unit 1315 generates a hyperside component S1 by subtracting the spectral energy of the mid-range component 1208 in the frequency domain from the spectral energy of the side component 1210 in the frequency domain, while retaining only the phase. For example, the subtraction unit 1315 generates a hyperside component S1 by subtracting the amplitude of the mid-range component 1208 in the frequency domain from the amplitude of the side component 1210 in the frequency domain, while retaining only the phase. The subtraction unit 1319 generates a residual side component S2 by subtracting the spectral energy of the hyperside component S1 from the spectral energy of the side component 1210. For example, the subtraction unit 1319 generates a residual side component S2 by subtracting the amplitude of the hyperside component S1 in the frequency domain from the amplitude of the side component 1210 in the frequency domain, while retaining the phase.
[0175] In Figure 5B, the orthogonal component generator module 1323 is similar to the orthogonal component generator module 1313 in that it receives the mid component 1006 and the side component 1210 and generates the hypermid component M1, the residual mid component M2, the hyperside component S1, and the residual side component S2. The orthogonal component generator module 1323 differs from the orthogonal generator module 1313 in that it generates the hypermid component M1 and the hyperside component S1 in the frequency domain, and then converts these components back into the time domain to generate the residual mid component M2 and the residual side component S2. The orthogonal component generator module 1323 includes a forward FFT unit 1320, a bandpass unit 1322, a subtraction unit 1324, a hypermid processor 1325, an inverse FFT unit 1326, a time delay unit 1328, a subtraction unit 1330, a forward FFT unit 1332, a bandpass unit 1334, a subtraction unit 1336, a hyperside processor 1337, an inverse FFT unit 1340, a time delay unit 1342, and a subtraction unit 1344.
[0176] The forward fast Fourier transform (FFT) unit 1320 applies a forward FFT to the mid-range component 1208, transforming it into the frequency domain. The frequency-domain mid-range component 1208 includes amplitude and phase. The band-pass unit 1322 applies a band-pass filter to the frequency-domain mid-range component 1208, which specifies the frequencies in the hyper-mid-range component M1. For example, to isolate the typical human vocal range, the band-pass filter may specify frequencies between 300 Hz and 8000 Hz. In another example, to remove speech content associated with the typical human vocal range, the band-pass filter may retain low frequencies (e.g., generated by a bass guitar or drums) and high frequencies (e.g., generated by a cymbal) in the hyper-mid-range component M1. In other embodiments, the orthogonal component generator module 1323 applies various other filters to the frequency-domain mid-range component 1208 in addition to and / or instead of the band-pass filter applied by the band-pass unit 1322. In some embodiments, the orthogonal component generator module 1323 does not include a bandpass unit 1322 and does not apply any filter to the frequency-domain mid component 1208. In the frequency domain, the subtraction unit 1324 subtracts the side component 1210 from the filtered mid component to generate the hypermid component M1. In other embodiments, in addition to and / or instead of subsequent processing applied to the hypermid component M1, such as that performed by an orthogonal component processor module (e.g., the orthogonal component processor module in Figure 12), the orthogonal component generator module 1323 applies various audio enhancements to the frequency-domain hypermid component M1. The hypermid processor 1325 performs processing on the hypermid component M1 in the frequency domain before converting it to the time domain. This processing may include subband spatial processing and / or crosstalk compensation processing.In some embodiments, the hypermid processor 1325 performs processing on the hypermid component M1 instead of, and / or in addition to, the processing that could be performed by the orthogonal component processor module 1214. The inverse FFT unit 1326 applies an inverse FFT to the hypermid component M1 and converts the hypermid component M1 back to the time domain. The hypermid component M1 in the frequency domain includes the amplitude of M1 and the phase of the mid component 1208, and the inverse FFT unit 1326 converts these to the time domain. The time delay unit 1328 applies a time delay to the mid component 1208 so that the mid component 1208 and the hypermid component M1 arrive at the subtraction unit 1330 simultaneously. The subtraction unit 1330 subtracts the hypermid component M1 in the time domain from the time-delayed mid component 1208 in the time domain to produce the residual mid component M2. In this example, the spectral energy of the hypermid component M1 is removed from the spectral energy of the mid component 1208 using processing in the time domain.
[0177] The forward FFT unit 1332 applies a forward FFT to the side component 1210, converting it to the frequency domain. The frequency domain side component 1210 includes amplitude and phase. The bandpass unit 1334 applies a bandpass filter to the frequency domain side component 1210. The bandpass filter specifies the frequency in the hyperside component S1. In other embodiments, the quadrature component generator module 1323 applies various other filters to the frequency domain side component 1210 in addition to and / or instead of the bandpass filter. In the frequency domain, the subtraction unit 1336 subtracts the mid component 1208 from the filtered side component 1210 to generate the hyperside component S1. In other embodiments, in addition to and / or instead of subsequent processing applied to the hyperside component S1, such as that performed by a quadrature component processor (e.g., quadrature component processor module 1214), the quadrature component generator module 1323 applies various speech enhancements to the frequency domain hyperside component S1. The hyperside processor 1337 performs processing on the hyperside component S1 in the frequency domain prior to its conversion to the time domain. This processing may include subband space processing and / or crosstalk compensation processing. In some embodiments, the hyperside processor 1337 performs processing on the hyperside component S1 in place of, and / or in addition to, the processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1340 applies an inverse FFT to the hyperside component S1 in the frequency domain to generate the hyperside component S1 in the time domain. The hyperside component S1 in the frequency domain includes the amplitude S1 and phase of the side component 1210, and the inverse FFT unit 1326 converts the side component 1210 to the time domain. The time delay unit 1342 time delays the side component 1210 so that it arrives at the subtraction unit 1344 at the same time as the hyperside component S1. Next, the subtraction unit 1344 subtracts the hyper-side component S1 in the time domain from the time-delayed side component 1210 in the time domain to generate the residual side component S2.In this example, using a time-domain process, the spectral energy of the hyperside component S1 is removed from the spectral energy of the side component 1210.
[0178] In some embodiments, the hypermid processor 1325 and the hyperside processor 1337 may be omitted if the processing performed by these components is performed by the orthogonal component processor module 1214.
[0179] In Figure 13C, the orthogonal component generator module 1343 is similar to the orthogonal component generator module 1323 in that it receives the mid component 1208 and the side component 1210 and generates the hypermid component M1, the residual mid component M2, the hyperside component S1, and the residual side component S2, except that the orthogonal component generator module 1343 generates components M1, M2, S1, and S2 in the frequency domain and then converts these components into the time domain. The orthogonal component generator module 1343 includes a forward FFT unit 1347, a bandpass unit 1349, a subtraction unit 1351, a hypermid processor 1352, a subtraction unit 1353, a residual mid processor 1354, an inverse FFT unit 1355, an inverse FFT unit 1355, an inverse FFT unit 1357, a forward FFT unit 1361, a bandpass unit 1363, a subtraction unit 1365, a hyperside processor 1366, a subtraction unit 1367, a residual side processor 1368, an inverse FFT unit 1369, and an inverse FFT unit 1371.
[0180] The forward FFT unit 1347 applies a forward FFT to the mid component 1208, converting it to the frequency domain. The mid component 1208 converted to the frequency domain includes amplitude and phase. The forward FFT unit 1361 applies a forward FFT to the side component 1210, converting it to the frequency domain. The side component 1210 converted to the frequency domain includes amplitude and phase. The band-pass unit 1349 applies a band-pass filter to the frequency domain mid component 1208, and the band-pass filter specifies the frequency of the hypermid component M1. In some embodiments, the orthogonal component generator module 1343 applies various other filters to the frequency domain mid component 1208 in addition to and / or instead of the band-pass filter. The subtraction unit 1351 subtracts the frequency domain side component 1210 from the frequency domain mid component 1208 to generate the hypermid component M1 in the frequency domain. The hypermid processor 1352 performs processing on the hypermid component M1 in the frequency domain before converting it to the time domain. In some embodiments, the hypermid processor 1352 performs subband spatial processing and / or crosstalk compensation processing. In some embodiments, the hypermid processor 1352 performs processing on the hypermid component M1 in place of and / or in addition to processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1357 applies an inverse FFT to the hypermid component M1 and converts it back to the time domain. The hypermid component M1 in the frequency domain includes the amplitude M1 and phase of the mid component 1208, and the inverse FFT unit 1357 converts this to the time domain. The subtraction unit 1353 subtracts the hypermid component M1 from the mid component 1208 in the frequency domain to generate the residual mid component M2. The residual mid processor 1354 performs processing on the residual mid component M2 in the frequency domain before converting the residual mid component M2 to the time domain. In some embodiments, the residual mid-processor 1354 performs subband spatial processing and / or crosstalk compensation processing on the residual mid-component M2.In some embodiments, the residual mid-processor 1354 performs processing on the residual mid-component M2 in place of and / or in addition to the processing that can be performed by the quadrature component processor module 1214. The inverse FFT unit 1355 applies an inverse FFT to convert the residual mid-component M2 into the time domain. The residual mid-component M2 in the frequency domain includes the amplitude M2 and phase of the mid-component 1208, which the inverse FFT unit 1355 converts into the time domain.
[0181] The bandpass unit 1363 applies a bandpass filter to the frequency-domain side component 1210. The bandpass filter specifies the frequency in the hyper-side component S1. In other embodiments, the quadrature component generator module 1343 applies various other filters to the frequency-domain side component 1210 in addition to and / or instead of the bandpass filter. In the frequency domain, the subtraction unit 1365 subtracts the mid component 1208 from the filtered side component 1210 to generate the hyper-side component S1. The hyper-side processor 1366 performs processing on the hyper-side component S1 in the frequency domain prior to conversion to the time domain. In some embodiments, the hyper-side processor 1366 performs subband space processing and / or crosstalk compensation processing on the hyper-side component S1. In some embodiments, the hyper-side processor 1366 performs processing on the hyper-side component S1 in addition to and instead of processing that may be performed by the quadrature component processor module 1214. The inverse FFT unit 1371 applies an inverse FFT to convert the hyperside component S1 back to the time domain. The hyperside component S1 in the frequency domain includes the amplitude S1 and phase of the side component 1210, which the inverse FFT unit 1371 converts to the time domain. The subtraction unit 1367 subtracts the hyperside component S1 from the side component 1210 in the frequency domain to generate the residual side component S2. The residual side processor 1368 performs processing on the residual side component S2 in the frequency domain prior to conversion to the time domain. In some embodiments, the residual side processor 1368 performs subband space processing and / or crosstalk compensation processing on the residual side component S2. In some embodiments, the residual side processor 1368 performs processing on the residual side component S2 instead of and / or in addition to the processing that can be performed by the quadrature component processor module 1214. The inverse FFT unit 1369 applies the inverse FFT to the residual side component S2 and converts it to the time domain.The residual side component S2 in the frequency domain includes the amplitude S2 and phase of the side component 1210, and the inverse FFT unit 1369 converts this to the time domain.
[0182] In some embodiments, the hypermid processor 1352, hyperside processor 1366, residual mid processor 1354, or residual side processor 1368 may be omitted if the processing performed by these components is performed by the orthogonal component processor module 1214.
[0183] [Example of an orthogonal component processor] Figure 14A is a block diagram showing one or more embodiments of the orthogonal component processor module 1417. The orthogonal component processor module 1417 is an example of the orthogonal component processor module 1412. Some embodiments of module 1417 have components different from those described herein. Similarly, in some cases, functions can be distributed among the components in a manner different from that described herein.
[0184] The orthogonal component processor module 1417 includes a component processor module 1420, a PSM module 102, an adder unit 1422, an M / SL / R converter module 1424, an adder unit 1426, and an adder 1428.
[0185] The component processor module 1420 performs the same processing as the component processor module 106, except that it uses a hyper-mid component M1, a hyper-side component S1, a residual mid component M2, and / or a residual side component S2, instead of a mid component and a side component. For example, the component processor module 1420 performs subband spatial processing and / or crosstalk compensation processing on at least one of the hyper-mid component M1, residual mid component M2, hyper-side component S1, and residual side component S2. As a result of the subband spatial processing and / or crosstalk compensation by the component processor module 1420, the orthogonal component processor module 1417 outputs at least one of the processed M1, processed M2, processed S1, and processed S2. In some embodiments, one or more of the components M1, M2, S1, or S2 may bypass the component processor module 1420.
[0186] In some embodiments, the quadrature component processor module 1417 performs subband spatial processing and / or crosstalk compensation processing on at least one of the hypermid component M1, residual mid component M2, hyperside component S1, and residual side component S2 in the frequency domain. The quadrature component generator module 410 may provide the quadrature component processor module 1417 with components M1, M2, S1, or S2 in the frequency domain without performing an inverse FFT. After generating the processed M1, processed M2, and processed side component 1442, the quadrature component processor module 1417 may perform an inverse FFT to convert these components back to the time domain. In some embodiments, the quadrature component processor module 1417 performs an inverse FFT on the processed M1, processed M2, processed S1, and processed S1 to generate the processed side component 1446 in the time domain.
[0187] Examples of components of the orthogonal component processor module 1417 are shown in FIGS. 15 and 16. In some embodiments, the orthogonal component processor module 1417 performs both subband spatial processing and crosstalk compensation processing. The processing performed by the orthogonal component processor module 1417 is not limited to subband spatial processing or crosstalk compensation processing. Any type of spatial processing that uses midspace / side space can be performed by the orthogonal component processor module 1417, such as by using hypermid components instead of mid components or hyperside components instead of side components. Other types of processing may include the application of gain, amplitude- or delay-based panning, binaural processing, reverberation processing, dynamic range processing such as compression and limiting, and machine learning-based approaches to transfer, conversion, or resynthesis from chorus or flanging to vocal style or instrument style, among other linear or non-linear audio processing techniques and effects.
[0188] The PSM module 102 receives the processed M1, applies PSM processing to spatially shift the processed M1, resulting in the left channel 1432 and the right channel 1434. The PSM module 102 is shown as being applied to the hypermid component M1, but the PSM module can be applied to one or more of the components M1, M2, S1, or S2. In some embodiments, the component processed by the PSM module 102 bypasses the processing by the component processor module 1420. For example, the PSM module 102 can process the hypermid component M1 instead of the processed M1.
[0189] The addition unit 1422 adds the processed S1 to the processed S2 to generate the processed side component 1442. The M / SL / R converter module 1424 uses the processed M2 and the processed side component 1442 to generate the processed left component 1444 and the processed right component 1446. In some embodiments, the processed left side component 1444 is generated based on the addition of the processed M2 and the processed side component 1442, and the processed right component 1446 is generated based on the difference between the processed M2 and the processed side component 1442. Other M / SL / R type conversions may be used to generate the processed left component 1444 and the processed right component 1446.
[0190] Adding unit 1426 adds the left channel 1432 from PSM module 102 to the processed left component 1444 to generate the left channel 1452. Adding unit 1428 adds the right channel 1434 from PSM module 102 to the processed right component 1446 to generate the right channel 1454. More generally, one or more left channels from PSM module 102 may be added to the left component from M / SL / R converter module 1424 (e.g., generated using hyper / residual components not processed by PSM module 102) to generate the left channel 1452, and one or more right channels from PSM module 102 may be added to the right component from M / SL / R converter module 1424 (e.g., generated using hyper / residual components not processed by PSM module 102) to generate the right channel 1454.
[0191] Therefore, the orthogonal component processor module 1417 applies PSM processing to the hypermid component M1 of the audio signal so that it is separated by the L / RM / S converter module 1206 and the orthogonal component generator module 1212. The PSM-enhanced stereo signal (including the left channel 1432 and the right channel 1434) can then be added to the residual left / residual right signals (e.g., the processed left component 1444 and the processed right component 1446, which are generated without the hypermid component). In addition to or instead of this example, other approaches may be used to separate the components of the input signal used for PSM processing, including machine learning-based source separation.
[0192] In some embodiments, the orthogonal component processor module 1417 applies PSM processing to the mid component M of the audio signal instead of the hypermid component ML. Figure 14B illustrates a block diagram showing the orthogonal component processor module 1419 in one or more embodiments. In some embodiments, the orthogonal component processor module 1419 in Figure 14B may be implemented as part of an audio processing system similar to the system 1200 shown in Figure 12, but without the orthogonal component generator module 1212, thereby allowing the orthogonal component processor module 1419 to receive mid and side component signals (e.g., mid component 1208 and side component 1210, etc.) instead of the hypermid component, hyperside component, residual mid component, and residual side component. In some embodiments, the orthogonal component processor module 1410 includes a component processor module similar to the component processor module 106 to generate processed mid and processed side components from the received mid and side components (not shown). The PSM module 102 receives the mid component M (or processed mid), applies PSM processing to spatially shift the received mid signal, and generates PSM-processed left channel 1432 and PSM-processed right channel 1434 in the mid signal. These are then combined with the side component S (or processed side) by the M / SL / R converter module 1424 to generate left channel 1452 and right channel 1454. For example, as shown in Figure 14B, the M / SL / R converter module 1424 uses an adder unit 1460 to generate left channel 1452 as the sum of the PSM-processed left channel 1432 and side component S, and uses a subtractor unit 1462 to generate right channel 1452 as the sum of the PSM-processed right channel 1434 and side component S. In other words, the M / SL / R converter module 1424 combines the signal for the left channel with the signal inversion for the right channel to perform the service of mixing the side signal (which exists in a subspace defined by the left component, and is the inverse of the right component in the left-right reference) into the PSM-processed stereo signal in the left-right space.
[0193] [Example of a subband spatial processor] Figure 15 is a block diagram showing a subband spatial processor module 1510 according to one or more embodiments. The subband spatial processor module 1510 is an example of the components of the component processor module 106 or 1520. The subband spatial processor module 1510 includes mid-EQ filters 1504(1), 1504(2), 1504(3), 1504(4), side-EQ filters 1506(1), 1506(2), 1506(3), and 1506(4). Some embodiments of the subband spatial processor module 1510 have components different from those described herein. Similarly, in some cases, functions can be distributed among the components in ways different from those described herein.
[0194] The subband spatial processor module 1510 receives the non-spatial component Ym and the spatial component Ys, and adjusts the gain of one or more subbands of these components to provide spatial enhancement. If the subband spatial processor module 1510 is part of the component processor module 1420, the non-spatial component Ym may be the hyper-mid component M1 or the residual mid component M2. The spatial component Ys may be the hyper-side component S1 or the residual side component S2. If the subband spatial processor module 1510 is part of the component processor module 106, the non-spatial component Ym may be the mid component 126, and the spatial component Ys may be the side component 128.
[0195] The subband spatial processor module 1510 receives the non-spatial component Ym and applies mid-EQ filters 1504(1) to 1504(4) to different subbands of Ym to generate an enhanced non-spatial component Em. The subband spatial processor module 1510 also receives the spatial component Ys and applies side-EQ filters 1506(1) to 1506(4) to different subbands of Ys to generate an enhanced spatial component Es. The subband filters can include various combinations of peaking filters, notch filters, low-pass filters, high-pass filters, low-pass filters, high-pass filters, band-pass filters, band-stop filters, and / or full-pass filters. Furthermore, gain can be applied to each subband of the subband filters. More specifically, the subband spatial processor module 1510 includes subband filters for each of the n frequency subbands in the non-spatial component Ym and filter subbands for each of the n subbands in the spatial component Ys. For n=4 subbands, for example, the subband spatial processor module 1510 includes a series of subband filters for the non-spatial component Ym, including a mid-equalization (EQ) filter 1504(1) for subband (1), a mid-equalization (EQ) filter 1504(2) for subband (2), a mid-EQ filter 1504(3) for subband (3), and a mid-EQ filter 1504(4) for subband (4). Each mid-EQ filter 1504 applies filter extraction to the frequency subband portion of the non-spatial component Ym to generate an enhanced non-spatial component Em.
[0196] The subband spatial processor module 1510 further includes a series of subband filters for the frequency subbands of the spatial component Ys, including a side equalization (EQ) filter 1506(1) for subband (1), a side EQ filter 1506(2) for subband (2), a side EQ filter 1506(3) for subband (3), and a side EQ filter 1506(4) for subband (4). Each side EQ filter 1506 applies filter extraction to the frequency subband portion of the spatial component Ys to generate an enhanced spatial component Es.
[0197] Each of the n frequency subbands in the non-spatial component Ym and spatial component Ys can correspond to a certain frequency range. For example, frequency subband (1) may correspond to 0–300 Hz, frequency subband (2) to 300–510 Hz, frequency subband (3) to 510–2700 Hz, and frequency subband (4) to 2700 Hz up to the Nyquist frequency. In some embodiments, each of the n frequency subbands is a set of integrated critical bands. The critical bands can be determined using a corpus of speech samples from a wide range of musical genres. The long-term average energy ratio of the mid- or side-components across the critical bands on a 24-Bark scale is determined from the samples. Then, consecutive frequency bands having similar long-term average ratios are grouped together to form a set of critical bands. The range and number of frequency subbands may be adjustable.
[0198] In some embodiments, the subband spatial processing module 1510 processes the residual mid component M2 as a non-spatial component Ym, and uses one of the side component, hyperside component S1, or residual side component S2 as the spatial component Ys.
[0199] In some embodiments, the subband spatial processing module 1510 processes one or more of the hypermid component M1, hyperside component S1, residual mid component M2, and residual side component S2. The filters applied to each subband of these components may differ. The hypermid component M1 and residual mid component M2 may be processed as described for the non-spatial component Ym, respectively. The hyperside component S1 and residual side component S2 may be processed as described for the spatial component Ys, respectively.
[0200] [Example of a crosstalk compensation device] Figure 16 is a block diagram showing a crosstalk compensation module 1610 according to one or more embodiments. The crosstalk compensation module 1610 is an example of a component of the component processor module 106 or 1420. Some embodiments of the crosstalk compensation module 1610 have components different from those described herein. Similarly, in some cases, functions can be distributed among the component components in a manner different from that described herein.
[0201] The crosstalk compensation processor module 1610 includes a mid-component processor 1620 and a side-component processor 1630. The crosstalk compensation processor module 1610 receives a non-spatial component Ym and a spatial component Ys, and applies filter extraction to one or more of these components to compensate for spectral defects caused by (e.g., subsequent or preceding) crosstalk processing. If the crosstalk compensation processor module 1610 is part of the component processor module 1420, the non-spatial component Ym may be a hyper-mid component M1 or a residual mid component M2. The spatial component Ys may be a hyper-side component S1 or a residual side component S2. If the crosstalk compensation processor module 1610 is part of the component processor module 106, the non-spatial component Ym may be a mid component 126, and the spatial component Ys may be a side component 128.
[0202] The crosstalk compensation processor module 1610 receives the non-spatial component Ym, and the mid-component processor 1620 applies a set of filters to generate an enhanced non-spatial crosstalk compensation component Zm. The crosstalk compensation processor module 1610 also receives the spatial subband component Ys, and the side-component processor 1630 applies a set of filters to generate an enhanced spatial subband component Es. The mid-component processor 1620 includes multiple filters 1640, such as m mid-filters 1640(a), 1640(b) through 1640(m), etc., where each of the m mid-filters 1640 processes the non-spatial component X m It processes one of the m frequency bands in the . Therefore, the mid-component processor 1620 processes the non-spatial component X m By processing this, a mid-crosstalk compensation channel Zm is generated. In some embodiments, the mid-filter 1640 performs crosstalk processing by simulation on the non-spatial component X mThe frequency response plot is constructed using the following. Furthermore, by analyzing the frequency response plot, arbitrary spectral defects such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB) can be estimated as artifacts of the crosstalk processing. These artifacts are mainly caused by the addition of delayed and possibly inverted contralateral signals to the corresponding ipsilateral signals during crosstalk processing, thereby effectively introducing a comb filter-like frequency response to the final rendering result. The mid-crosstalk compensation channel Zm can be generated by the mid-component processor 1620 to compensate for the estimated peaks or troughs, with each of the m frequency bands corresponding to a peak or trough. Specifically, based on the specific delay, filter extraction frequency, and gain applied to the crosstalk processing, the peaks or troughs in the frequency response are shifted up or down, causing variable amplification and / or variable attenuation of energy in specific regions of the spectrum. Each of the mid-filters 1640 can be configured to adjust one or more of the peaks and troughs.
[0203] The side component processor 1630 includes multiple filters 1650, such as m side filters 1650(a), 1650(b) through 1650(m). The side component processor 1630 generates a side crosstalk compensation channel Zs by processing the spatial component Xs. In some embodiments, the frequency response plot of spatial X with crosstalk processing can be obtained through simulation. By analyzing the frequency response plot, arbitrary spectral defects, such as peaks or troughs in the frequency response plot that exceed a predetermined threshold (e.g., 10 dB), can be estimated as artifacts of the crosstalk processing. The side crosstalk compensation channel Zs can be generated by the side component processor 1630 to compensate for the estimated peaks or troughs. Specifically, based on the specific delay, filter extraction frequency, and gain applied to the crosstalk processing, the peaks or troughs in the frequency response are shifted up or down, causing variable amplification and / or variable attenuation of energy in a specific region of the spectrum. Each of the side filters 1650 may be configured to adjust one or more peaks or troughs. In some embodiments, the mid-component processor 1620 and the side-component processor 1630 may include a different number of filters.
[0204] In some embodiments, the mid-filter 1640 and the side-filter 1650 may include biquadratic filters having a transfer function defined by equation (7). One way to implement such filters is a direct form I topology as defined by equation (22). That is,
[0205]
number
[0206] Here, X is the input vector and Y is the output. Other topologies can be used depending on their maximum word length and summation operation. A bilinear topology can then be used to implement a quadratic filter with real-valued inputs and outputs. To design a discrete-time filter, a continuous-time filter is designed and then converted to discrete-time by a bilinear transform. Furthermore, the resulting shifts in center frequency and bandwidth can be compensated for using frequency warping.
[0207] For example, a peaking filter may have an S-plane transfer function defined by equation (23). That is,
[0208]
number
[0209] Here, s is a complex variable, A is the peak amplitude, Q is the filter "quality", and the digital filter coefficients are defined by the following equation (24).
[0210]
number
[0211] Here, ω0 is the filter center frequency in radians,
[0212]
number
[0213] Furthermore, the filter quality Q can be defined by equation (25). That is,
[0214]
number
[0215] Here, Δf is the bandwidth, and f c is the center frequency. The mid-filter 1640 is shown as being in series, and the side filter 1650 is shown as being in series. In some embodiments, the mid-filter 1640 is the mid-component X m The side filter is applied in parallel to the first component, and the side filter is applied in parallel to the side component Xs.
[0216] In some embodiments, the crosstalk compensation module 1610 processes each of the hypermid component M1, hyperside component S1, residual mid component M2, and residual side component S2. The filters applied to each of these components may differ.
[0217] [Example of a crosstalk processor] Figure 17 is a block diagram showing a crosstalk simulation processor module 1700 according to one or more embodiments. The crosstalk simulation processor module 1700 is an example of the crosstalk processor module 110 or the crosstalk processor module 1224. Some embodiments of the crosstalk simulation processor module 1700 have components different from those described herein. Similarly, in some cases, functions can be distributed among the components in a manner different from that described herein.
[0218] The crosstalk simulation processor module 1700 generates contralateral sound components for output to stereo headphones, thereby providing a loudspeaker-like listening experience on the headphones. Left input channel X L This could be the processed left component 134 / 1220, and the right input channel X R This could be the processed right component 136 / 1222.
[0219] The crosstalk simulation processor module 1700 includes a left head shadow low-pass filter 1702, a left head shadow high-pass filter 1724, a left crosstalk delay 1704, and a left head shadow gain 1710 to process the left input channel X L The crosstalk simulation processor module 1700 further includes a right head shadow low-pass filter 1706, a right head shadow high-pass filter 1726, a right crosstalk delay 1708, and a right head shadow gain 1712 to process the right input channel X R The crosstalk simulation processor module 1500 further includes an addition unit 1714 and an addition unit 1716.
[0220] The left head shadow low-pass filter 1702 and the left head shadow high-pass filter 1724 apply modulation to the left input channel X L which models the frequency response of the signal after passing through the listener's head. The output of the left head shadow high-pass filter 1724 is supplied to the left crosstalk delay 1704, which applies a time delay. This time delay represents the transaural distance through which the contralateral acoustic component passes with respect to the ipsilateral sound component. The left head shadow gain 1710 applies a gain to the output of the left crosstalk delay 1704 to generate the right-to-left simulation channel W L
[0221] Similarly, for the right input channel X R the right head shadow low-pass filter 1706 and the right head shadow high-pass filter 1726 apply modulation to the right input channel XR which models the frequency response in the listener's head. The output of the right head shadow high-pass filter 1726 is supplied to the right crosstalk delay 1708, which applies a time delay. The right head shadow gain 1712 applies a gain to the output of the right crosstalk delay 1708 to generate the right crosstalk simulation channel W R Generate it.
[0222] The application of the head shadow low-pass filter, the head shadow high-pass filter, the crosstalk delay, and the head shadow gain for each of the left and right channels can be performed in different orders.
[0223] The addition unit 1714 adds the right crosstalk simulation channel W R and the left input channel X L to generate the left output channel O L The addition unit 1716 adds the left crosstalk simulation channel W L and the right input channel X R to generate the left output channel O R Generate it.
[0224] FIG. 18 is a block diagram showing a crosstalk cancellation processor module 1800 according to one or more embodiments. The crosstalk cancellation processor module 1800 is an example of the crosstalk processor module 110 or the crosstalk processor module 1224. Some embodiments regarding the cancellation processor module 1800 have components different from those described herein. Similarly, in some cases, functions can be distributed among components in a manner different from that described herein.
[0225] The crosstalk cancellation processor module 1800 receives the left input channel X <00001\02> and the right input channel X R performs crosstalk cancellation on channel X L X R to generate the left output channel O L and the right output channel O R Generate it. The left input channel X L can be the processed left component 134 / 1220, and the right input channel X R can be the processed right component 136 / 1222.
[0226] The crosstalk cancellation processor module 1800 includes an in-band / out-of-band divider 1810, inverters 1820 and 1822, contra-side estimators 1830 and 1840, couplers 1850 and 1852, and an in-band / out-of-band coupler 1860. These components work together to cancel input channel T L , T R The signal is split into in-band and out-of-band components, and crosstalk cancellation is performed on the in-band component to output channel O L , O R Generates.
[0227] By splitting an input audio signal T into different frequency band components and performing crosstalk cancellation on selective components (e.g., in-band components), crosstalk cancellation can be performed on specific frequency bands while avoiding degradation in other frequency bands. If crosstalk cancellation is performed without splitting the input audio signal T into different frequency bands, the resulting audio signal may show significant attenuation or amplification of non-spatial and spatial components at low frequencies (e.g., below 350 Hz), high frequencies (e.g., above 12000 Hz), or both. By selectively performing crosstalk cancellation on in-band components (e.g., between 250 Hz and 14000 Hz), where the majority of influential spatial implications reside, it is possible to preserve the overall energy that maintains balance across the spectrum of the mix, particularly in the non-spatial components.
[0228] The in-band / out-of-band divider 1810 controls the input channel T L , T R In-band channel T L,In , T R,In , and out-of-band channel T L,Out , T R,Out They are separated into two parts. In particular, the in-band / out-of-band divider 1810 is the left enhanced compensation channel T L left in-band channel T L,In and left out-of-band channel T L,OutIt is separated into. Similarly, the in-band / out-of-band divider 1810 is the right enhanced compensation channel T R right-band channel T R,In Right out-of-band channel T R,Out The signal is separated into these bands. Each in-band channel may encompass a portion of the respective input channels corresponding to a frequency range including, for example, 250 Hz to 14 kHz. The frequency band range may be adjustable, for example, according to the speaker parameters.
[0229] The inverter 1820 and the contra-side estimator 1830 work together to generate the left contra-side cancellation component S L Generates the left intraband channel T L,In Compensation is performed for the contra-side acoustic component caused by the following. Similarly, the inverter 1822 and the contra-side estimator 1840 work together to compensate for the right contra-side cancellation component S R Generates the right intraband channel T R,In Compensation is performed for contralateral acoustic components caused by this.
[0230] One approach is that inverter 1820 has an in-band channel T L,In Received, received in-band channel T L,In The polarity is reversed, and the inverted inband channel T L,In’ The contra-side estimator 1830 generates the inverted in-band channel T. L,In’ Receives and, through filter extraction, extracts the inverted in-band channel T corresponding to the contralateral acoustic component. L,In’ A portion is extracted. Filter extraction is performed on the inverted inband channel T L,In’ Since it is performed on, the portion extracted by the contraside estimator 1830 is the in-band channel T that belongs to the contraside acoustic component. L,In This results in a partial inversion. Therefore, the portion extracted by the contralateral estimator 1830 is the left contralateral cancellation component S L This corresponds to the in-band channel T R,In In addition, the in-band channel T L,InThe contra-side acoustic components resulting from this can be reduced. In some embodiments, the inverter 1820 and the contra-side estimator 1830 are implemented in different sequences.
[0231] The inverter 1822 and the opposite-side estimator 1840 are used for in-band channel T R,In Perform a similar operation with respect to the right-contrast cancellation component S R This generates [the specified output]. Therefore, for the sake of brevity, a detailed explanation is omitted in this specification.
[0232] In one implementation example, the contra-side estimator 1830 includes a filter 1832, an amplifier 1834, and a delay unit 1836. The filter 1832 inverts the input channel T L,In’ It receives the signal and, through the filter extraction function, inverts the intraband channel T corresponding to the contralateral acoustic component. L,In’ A portion of this is extracted. An example of a filter implementation is a notch filter or high-pass filter with a center frequency selected between 5000Hz and 10000Hz, and a Q selected between 0.5 and 1 / 0. Gain in decibels (G dB ) can be derived from equation (26).
[0233]
number
[0234] Here, D is the amount of delay by the sample-level delay unit 1836, for example, at a sampling rate of 48 kHz. An alternative implementation is a low-pass filter having a corner frequency selected between 5000 Hz and 10000 Hz, and a Q selected between 0.5 and 1.0. Furthermore, the amplifier 1834 controls the extracted portion with the corresponding gain coefficient G L,In The output is amplified by the delay unit 1836, which delays the amplified output from amplifier 1834 according to the delay function D, thereby creating the left-contrast cancellation component S. L The contra-side estimator 1840 generates the filter 1842, the amplifier 1844, and the inverted in-band channel. TR,In’Perform a similar operation on the right-contrast cancellation component S R This generates the left and right contra-side cancellation components S, for example, the contra-side estimators 1830 and 1840 are calculated according to the following formula. L S R This generates...
[0235]
number
[0236] Here, F[] is the filter function and D[] is the delay function.
[0237] The crosstalk cancellation configuration can be determined by the speaker parameters. In one example, the filter center frequency, delay, amplifier gain, and filter gain may be determined according to the angle formed between the two speakers relative to the listener. In some embodiments, the values between the speaker angles are used to interpolate other values.
[0238] Coupler 1850 has a right-opposite cancellation component S R left in-band channel T L,In Combined with the left in-band crosstalk channel U L The coupler 1852 generates the left contralateral cancellation component S L right-band channel T R,In Combined with the right-band crosstalk channel U R The in-band crosstalk coupler 1860 generates the left in-band crosstalk channel U L out-of-band channel T L,Out Combined with, left output channel O L Generates the right intraband crosstalk channel U R out-of-band channel T R,Out Combined with, the right output channel O R Generates.
[0239] Therefore, left output channel O L This is an intraband channel T belonging to the opposite sound. R,InRight-opposite cancellation component S, which corresponds to the inverse of a portion of it. R Includes right output channel O R This is an intraband channel T belonging to the opposite sound. L,In Left-contrast cancellation component S, which corresponds to the inverse of a portion of it. L This includes the right output channel O R Accordingly, when the wavefront of the ipsilateral acoustic component output by the right loudspeaker reaches the right ear, the left output channel O L Accordingly, the wavefront of the contralateral acoustic component output by the left loudspeaker can be canceled out. Similarly, the left output channel O L Accordingly, when the wavefront of the ipsilateral acoustic component output by the left loudspeaker reaches the left ear, the right output channel O R Accordingly, the wavefront of the contralateral acoustic component output by the right speaker can be canceled out. Therefore, the contralateral acoustic component can be reduced to improve spatial detectability.
[0240] [Example of PSM process flow] Figure 19 is a flowchart of process 1900 of PSM processing according to one or more embodiments. Process 1900 may include fewer or additional steps, and the steps may be performed in a different order. In some embodiments, PSM processing may be performed using a Hilbert Transform Perceptual Sound Stage Modification (HPSM) module.
[0241] In step 1905, the audio processing system (e.g., the PSM module 102 of the audio processing system 100 or 1200) separates the input channel into low-frequency and high-frequency components. The crossover frequency defining the boundary between the low-frequency and high-frequency components may be adjustable so that the frequencies subject to PSM processing are included in the high-frequency components. In some embodiments, the audio processing system applies gain to the low-frequency and / or high-frequency components.
[0242] An input channel may be a specific portion of an audio signal extracted for PSM processing. In some embodiments, the input channel is the mid-range or side component of an audio signal (e.g., stereo or multi-channel). In some embodiments, the input channel is the hyper-mid-range, hyper-side-range, residual-mid-range, or residual-side-range component of an audio signal. In some embodiments, the input channel is associated with a sound source, such as a voice or instrument, which should be coupled with other sounds to the audio mix.
[0243] In step 1910, the audio processing system applies a first Hilbert transform to the high-frequency components to generate a first left-foot component and a first right-foot component, the first left-foot component being 90 degrees out of phase with respect to the first right-foot component.
[0244] In step 1915, the speech processing system applies a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, the first left-foot component being 90 degrees out of phase with respect to the first right-foot component.
[0245] In some embodiments, the speech processing system applies a delay and / or gain to the first left branch component. The speech processing system may also apply a delay and / or gain to the second right branch component. These gains and delays may be used to manipulate the perceptual results of PSM processing.
[0246] In step 1920, the audio processing system combines the first left foot component with the low-frequency component to generate the left channel. In step 1925, the audio processing system combines the second right foot component with the low-frequency component to generate the right channel. The left channel may be supplied to the left speaker, and the right channel may be supplied to the right speaker.
[0247] Figure 20 is a flowchart of another process 2000 for PSM processing using a first-order non-orthogonal rotation-based decorrelation (FNORD) filter network, according to several embodiments. The processing shown in Figure 20 may be performed by a component of a voice system (e.g., system 100, 202, or 1200). Other entities may perform some or all of the steps in Figure 2, in other embodiments. Embodiments may include different steps and / or additional steps, or the steps may be performed in a different order.
[0248] In step 2005, the audio system determines a target amplitude response that defines one or more spatial suggestions encoded in a monaural audio signal to generate the resulting multiple channels, one or more of which are associated with one or more frequency-dependent amplitude suggestions encoded in the mid-space / side-space of the resulting channels, and these mid-space / side-spaces do not alter the overall coloration of the resulting channels. One or more spatial suggestions may include at least one elevation angle suggestion associated with a target elevation angle. Each elevation angle suggestion may correspond to one or more frequency-dependent amplitude suggestions encoded in the mid-space / side-space of the audio signal, such as a target amplitude function corresponding to a narrow region of infinite attenuation at one or more specific frequencies. On the other hand, the left and right suggestions for elevation angle are typically symmetrical to the coloration, so the left and right signals may be constrained to be colorless. In some embodiments, the spatial suggestions may be based on a sampled HRTF.
[0249] In some embodiments, the target amplitude response may further define one or more parametric spatial implications, which may include a target broadband attenuation, a target subband attenuation, a critical point, filter characteristics, and / or the location of the soundstage to which the implications should be embedded. The critical point may be an inflection point at 3 dB. The filter characteristics may include one of the following: high-pass filter characteristics, low-pass filter characteristics, band-pass filter characteristics, or band-reject filter characteristics. The location of the soundstage may include the mid-channel or side channels, or, if there are more than two output channels, other subspaces within the output space, such as per pair and / or determined via hierarchical addition and / or difference. One or more spatial implications may be determined based on the characteristics of the presentation equipment (e.g., speaker frequency response, speaker location), the expected content of the audio data, the listener's perceptual ability in the context, or the minimum quality expected of the audio presentation system in question. For example, if the speaker is unable to adequately reproduce frequencies below 200 Hz, spatial implications embedded in this range should be avoided. Similarly, if the expected audio content is speech, the speech system may select a target amplitude response that affects only the frequencies in which the ear is most sensitive, within the expected bandwidth of the speech. If the listener will derive audible cues from other sources in the situation, such as an array of speakers in the vicinity, the speech system may determine a target amplitude response that complements those simultaneous cues.
[0250] In step 2010, the audio system determines a transfer function for a single-input, multiple-output pass-through filter based on the target amplitude response. This transfer function defines the relative rotation of the output channels in terms of phase angle. This transfer function represents the effect the filter network has on its input for each output in terms of phase angle rotations as a function of frequency.
[0251] In step 2015, the audio system determines the coefficients of a pass-through filter based on the transfer function. These coefficients are selected in a manner best suited to the type of implication and / or constraints and applied to the incoming audio stream. Several examples of coefficient sets are defined in equations (12), (13), (17), and (19). In some embodiments, determining the coefficients of a pass-through filter based on the transfer function involves using the inverse discrete Fourier transform (idft). In this case, the coefficient set may be determined as defined by equation (19). In some embodiments, determining the coefficients of a pass-through filter based on the transfer function involves using a phase vocoder. In this case, the coefficient set may be determined as defined by equation (19), except that it is applied to the frequency domain before resynthesizing the time-domain data. In some embodiments, these coefficients include at least rotational control parameters and first-order coefficients, which are determined based on the received critical point parameters, filter characteristic parameters, and soundstage position parameters.
[0252] The audio system 2020 processes a monaural channel using coefficients of a pass-through filter to generate multiple channels. For example, in some embodiments, a pass-through filter module receives a monaural audio channel, performs broadband phase rotation on the monaural audio channel to generate multiple broadband rotation component channels (e.g., left and right broadband rotation components) based on rotation control parameters, performs narrowband phase rotation on at least one of the multiple broadband rotation component channels based on a linear coefficient to determine a narrowband rotation component channel, which, together with one or more of the remaining broadband rotation component channels, forms multiple channels output by the audio system.
[0253] In some embodiments, as in equation (8), when the system is operating in the time domain using an IIR implementation, these coefficients can be used to adjust appropriate feedback and feedforward delays. When an FIR implementation is used, as in equation (19), only the feedforward delay may be used. When the coefficients are determined and applied in the spectral domain, the coefficients may be applied to the spectral data as a complex multiplication before resynthesis. The audio system may provide multiple output channels to presentation equipment, such as user equipment connected to the audio system via a network.
[0254] The PSM processing flow examples described above encode spatial implications by perceptually positioning monaural content at specific locations within the soundstage (e.g., locations associated with a target elevation angle) using a network of full-pass filters. Because the network of full-pass filters described herein is colorless, these filters allow the user to decouple the spatial placement of the audio from its overall coloration.
[0255] [Spatial processing of orthogonal components] Figure 21 is a flowchart of a process 2100 for spatial processing using at least one of the hypermid, residualmid, hyperside, or residualside components in one or more embodiments. Spatial processing may include dynamic range processing such as gain application, amplitude or delay-based panning, binaural processing, reverberation processing, compression and limiting, linear or nonlinear speech processing techniques and effects, chorus effects, flanging effects, and, among other techniques, machine learning-based approaches to the transfer, transformation, or resynthesis of vocal style or instrument style. This process may be performed to provide spatially enhanced speech to the user's device. This process may include fewer or additional steps, and the steps may be performed in different orders.
[0256] In step 2110, an audio processing system (e.g., audio processing system 1200) receives an input audio signal (e.g., left channel 1202 and right input channel 1204). In some embodiments, the input audio signal may be a multi-channel audio signal comprising multiple left-right channel pairs. Each left-right channel pair may be processed with respect to the left and right input channels as described herein.
[0257] In step 2120, the audio processing system generates a non-spatial mid component (e.g., mid component 1208) and a spatial side component (e.g., side component 1210) from the input audio signal. In some embodiments, an L / RM / S converter (e.g., an L / RM / S converter module 1206) performs the conversion of the input audio signal into mid and side components.
[0258] In step 2130, the audio processing system generates at least one of the following: a hypermid component (e.g., hypermid component M1), a hyperside component (e.g., hyperside component S1), a residual mid component (e.g., residual mid component M2), and a residual side component (e.g., residual side component S2). The audio processing system may generate at least one and / or all of the components listed above. The hypermid component includes the spectral energy of the side component removed from the spectral energy of the mid component. The residual mid component includes the spectral energy of the hypermid component removed from the spectral energy of the mid component. The hyperside component includes the spectral energy of the mid component removed from the spectral energy of the side component. The residual side component includes the spectral energy of the hyperside component removed from the spectral energy of the side component. The processing used to generate M1, M2, S1, or S2 may be performed in the frequency domain or the time domain.
[0259] In step 2140, the audio processing system enhances the audio signal by filtering out at least one of the hypermid, residualmid, hyperside, and residualside components. The filtering may include HPSM processing, in which a series of Hilbert transforms are applied to the high-frequency components of the hypermid, residualmid, hyperside, or residualside components. In one example, the hypermid component undergoes HPSM processing, while one or more of the residualmid, hyperside, or residualside components undergo other types of filtering.
[0260] Filter extraction may include PSM processing, and the spatial implications are colorless encoded either through a parametric specification of the spatial implications, as described above in more detail in relation to Figures 10A and 10B, or through sampling of anthropometric measurements of HRTF data, as described above in relation to equation (20). In one example, the hypermid component undergoes PSM processing, while one or more of the residual mid component, hyperside component, or residual side component are not subjected to filter extraction or undergo other types of filter extraction.
[0261] Filter extraction may include other types of filtering, such as processing of spatial implications. Processing of spatial implications may include adjusting the frequency-dependent amplitude or frequency-dependent delay of hypermid, residualmid, hyperside, or residualside components. Examples of processing of spatial implications include amplitude or delay-based panning, or binaural processing.
[0262] Filter extraction may include dynamic range processing, such as compression or limiting. For example, hypermid, residualmid, hyperside, or residualside components may be compressed according to the compression ratio at which they exceed a threshold level for compression. In another example, hypermid, residualmid, hyperside, or residualside components may be limited to the maximum level at which they exceed a threshold level for limiting.
[0263] Filter extraction may include machine learning-based modifications to hypermid, residual mid, hyperside, or residual side components. Some examples include machine learning-based transfer, transformation, or resynthesis of vocal or instrumental styles.
[0264] Filter extraction of hypermid, residualmid, hyperside, or residualside components may include other linear or nonlinear audio processing techniques and effects, ranging from gain application, reverberation processing, and chorus and / or flanging, or other types of processing. In some embodiments, filter extraction may include filter extraction for subband spatial processing and crosstalk compensation, as will be described in more detail later in relation to Figure 22.
[0265] Filter extraction may be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are converted from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain and filtered in the frequency domain, the filtered components are converted to the time domain and filtered in the time domain for these components.
[0266] In step 2150, the audio processing system uses one or more of the filtered hyper / residual components to generate the left output channel (e.g., left output channel 1242) and the right output channel (e.g., right output channel 1244). For example, M / S to L / R conversion may be performed using a mid or side component generated from at least one of the filtered hypermid component, filtered residual mid component, filtered hyperside component, or filtered residual side component. In another example, a filtered hypermid component or filtered residual mid component may be used as the mid component for M / S to L / R conversion, or a filtered hyperside component or residual side component may be used as the side component for M / S to L / R conversion.
[0267] [Subband space processing and crosstalk processing of orthogonal components] Figure 22 is a flowchart of process 2200 for subband spatial processing and crosstalk compensation using at least one of the hypermid, residual mid, hyperside, or residual side components, according to one or more embodiments. Crosstalk processing may include crosstalk cancellation or crosstalk simulation. Subband spatial processing may be performed to provide audio content with enhanced spatial discoverability (e.g., enhanced soundstage) by creating the perception that sound is directed to the listener from a wide area rather than a specific point in space corresponding to the loudspeaker's location, thereby providing the listener with a more immersive listening experience. Crosstalk simulation may be used for audio output to headphones to simulate a loudspeaker experience with contra-side crosstalk. Crosstalk cancellation may be used for audio output to loudspeakers to eliminate the effects of crosstalk interference. Crosstalk compensation compensates for spectral defects caused by crosstalk cancellation or crosstalk simulation. The process may include fewer or additional steps, and the steps may be performed in a different order. Hyper and residual mid / side components can be manipulated in various ways depending on the purpose. For example, in the case of crosstalk compensation, the target subband filter extraction may be applied only to the hyper-mid component M1 (where the majority of vocal dialogue energy in many film contents originates) in an effort to remove spectral artifacts resulting from crosstalk processing in that component only. When enhancing the soundstage with or without crosstalk processing, the target subband gain may be applied to the residual mid component M2 and residual side component S2.For example, the residual mid-range component M2 may be attenuated, and the residual side-range component S2 may be inversely amplified, thereby increasing the distance between these components in terms of gain without causing a dramatic overall change in the perceived acoustics of the final L / R signal, while also avoiding attenuation in the hyper-mid-range component M1 (which is often the portion of the signal containing the majority of the vocal energy). (If done well, this can increase spatial detectability.)
[0268] In step 2210, the audio processing system receives an input audio signal, including a left channel and a right channel. In some embodiments, the input audio signal may be a multi-channel audio signal, including multiple left-right channel pairs. Each left-right channel pair may be processed with respect to the left and right input channels as described herein.
[0269] In step 2220, the audio processing system applies crosstalk processing to the received input audio signal. Crosstalk processing includes at least one of crosstalk simulation and crosstalk cancellation.
[0270] In steps 2230 to 2260, the audio processing system performs subband spatial processing and crosstalk compensation for crosstalk processing using one or more of the hypermid, hyperside, residualmid, or residualside components. In some embodiments, crosstalk processing may be performed after the processing in steps 2230 to 2260.
[0271] In step 2230, the audio processing system generates mid-range and side-range components from the audio signal (e.g., after crosstalk processing).
[0272] In step 2240, the audio processing system generates at least one of the hypermid component, residual mid component, hyperside component, and residual side component. The audio processing system may generate at least one and / or all of the components listed above.
[0273] In step 2250, the audio processing system filters and extracts at least one subband from the hypermid, residual mid, hyperside, and residual side components and applies subband spatial processing to the audio signal. Each subband may include a frequency range that may be defined by a set of critical bands. In some embodiments, the subband spatial processing further includes time-delaying at least one subband from the hypermid, residual mid, hyperside, and residual side components. In some embodiments, the filtering includes applying HPSM processing.
[0274] In step 2260, the audio processing system filters out at least one of the hypermid, residualmid, hyperside, and residualside components to compensate for spectral defects from crosstalk processing of the input audio signal. Spectral defects may include peaks or troughs in the frequency response plot of the hypermid, residualmid, hyperside, or residualside components that exceed a predetermined threshold (e.g., 10 dB) as artifacts of crosstalk processing. Spectral defects may be estimated spectral defects.
[0275] In some embodiments, the filtering of orthogonal components for subband spatial processing in step 2250 and the crosstalk compensation in step 2260 can be integrated into a single filtering operation for each orthogonal component selected for filtering.
[0276] In some embodiments, filter extraction in hyper / residual mid / side components for subband spatial processing or crosstalk compensation may be performed in connection with filter extraction for other purposes, such as dynamic range processing including gain application, amplitude or delay-based panning, binaural processing, reverberation processing, compression and limiting, linear or nonlinear speech processing techniques and effects ranging from chorus and / or flanging, machine learning-based approaches to transferring, transforming, or resynthesizing vocal or instrumental styles, or other types of processing using any of the hypermid, residual mid, hyperside, or residual side components.
[0277] Filter extraction can be performed in the frequency domain or the time domain. In some embodiments, the mid and side components are converted from the time domain to the frequency domain, the hyper and / or residual components are generated in the frequency domain, filter extraction is performed in the frequency domain, and the filtered components are converted to the time domain. In other embodiments, the hyper and / or residual components are converted to the time domain, and filter extraction is performed on these components in the time domain.
[0278] In step 2270, the audio processing system generates left and right output channels from the filtered hypermid component. In some embodiments, the left and right output channels are additionally based on at least one of the filtered residual mid component, the filtered hyperside component, and the filtered residual side component.
[0279] [Computer example] Figure 23 is a block diagram showing a computer 2300 according to several embodiments. The computer 2300 is an example of a computing device that includes circuitry implementing a voice system, such as voice system 100, 202, or 1200. At least one processor 2302 coupled to a chipset 2304 is illustrated. The chipset 2304 includes a memory controller hub 2320 and an input / output (I / O) controller hub 2322. Memory 2306 and a graphics adapter 2312 are coupled to the memory controller hub 2320, and a display 2318 is coupled to the graphics adapter 2312. A storage device 2308, a keyboard 2310, a pointing device 2314, and a network adapter 2316 are coupled to the I / O controller hub 2322. The computer 2300 may include various types of input or output devices. Other embodiments of the computer 2300 have different architectures. For example, in some embodiments, memory 2306 is directly connected to the processor 2302.
[0280] The storage device 2308 includes one or more non-temporary computer-readable storage media, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. Memory 2306 holds program code (consisting of one or more instructions) and data used by the processor 2302. This program code may correspond to the processing modes described with reference to Figures 1 to 3.
[0281] The pointing device 2314 is used in conjunction with the keyboard 2310 to input data into the computer system 2300. The graphics adapter 2312 displays images and other information on the display 2318. In some embodiments, the display 2318 includes touchscreen functionality for receiving user input and user selection. The network adapter 2316 connects the computer system 2300 to a network. Some embodiments of the computer 2300 have components and / or other components different from those shown in Figure 23.
[0282] The circuit may include one or more processors that execute program code stored in a non-temporary computer-readable medium, and when executed by one or more processors, this program code configures one or more processors to implement a voice system or a module of a voice system. Other examples of circuits that implement a voice system or a module of a voice system may include integrated circuits, such as application-specific integrated circuits (AS1Cs), field-programmable gate arrays (FPGAs), or other types of computer circuits.
[0283] [Additional considerations] Examples of the advantages and merits of the disclosed configuration include dynamic audio enhancement resulting from an audio system that is enhanced in accordance with the device and associated audio rendering system, as well as other relevant information made available by the device OS, such as use case information (e.g., indicating that the audio signal is used for music playback rather than gaming). The enhanced audio system may be integrated into the device (using a software development kit) or stored on a remote server for on-demand access. In this way, the device does not need to allocate storage or processing resources for the maintenance of the audio enhancement system, which is specific to its audio rendering system or audio rendering configuration. In some embodiments, the enhanced audio system enables various levels of querying of rendering system information so that effective audio enhancement can be applied across various levels of available device-specific rendering information.
[0284] Throughout this specification, a component, operation, or structure described as a single instance may be implemented by multiple instances. While individual actions in one or more ways are illustrated and described as separate actions, one or more separate actions may be performed simultaneously, and it is not necessary for them to be performed in the illustrated order. Structures and functions presented as separate components in the examples may be implemented as combined structures or components. Similarly, structures and functions presented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements are included within the scope of this specification.
[0285] Certain embodiments described herein include logic, or a number of components, modules, or mechanisms. A module may constitute either a software module (e.g., code embodied on a machine-readable medium or in a transmitted signal) or a hardware module. A hardware module is a tangible unit capable of performing a particular operation and may be configured or arranged in a particular manner. In the examples of embodiments, one or more computer systems (e.g., a standalone, client, or server computer system), or one or more hardware modules of a computer system (e.g., a processor or a group of processors), are configured by software (e.g., an application or a part of an application) as hardware modules that operate to perform a particular operation described herein.
[0286] The various operations in the method examples described herein may be performed, at least partially, by one or more processors that are configured either temporarily (e.g., by software) or permanently to perform the relevant operations. Whether configured temporarily or permanently, such processors may constitute a processor implementation module that operates to perform one or more operations or functions. The modules referred to herein may include processor implementation modules in some embodiments.
[0287] Similarly, the methods described herein can be implemented at least partially by a processor. For example, at least some of the operations in the method may be performed by one or more processors or processor implementation hardware modules. Performance for a particular operation may be distributed among one or more processors, not only residing within a single machine but also deployed across multiple machines. In some embodiments, the processor or group of processors may be located in a single location (e.g., in a home environment, an office environment, or as a server farm), while in other embodiments, the processors may be distributed across numerous locations.
[0288] Unless otherwise specified, the descriptions herein using terms such as “processing,” “computing,” “calculating,” “determining,” “presenting,” or “displaying” may refer to the operation or process of a machine (e.g., a computer) that manipulates or transforms data expressed as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, transmit, or display information.
[0289] As used herein, any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in relation to that embodiment is included in at least one embodiment. The phrase “in one embodiment” appearing in various places herein does not necessarily refer to the same embodiment.
[0290] Some embodiments may be described using the terms “coupled” and “connected,” along with their derivatives. These terms should be understood not to be intended as synonyms. For example, some embodiments may be described using the term “connected” to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments may be described using the term “coupled” to indicate that two or more elements are in direct physical or electrical contact. However, the term “coupled” can also mean that two or more elements are not in direct contact with each other but are still coordinating or interacting with each other. Embodiments are not limited to this situation.
[0291] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus including a list of elements is not necessarily limited to those elements alone, but may include other elements not expressly enumerated or that are specific to such process, method, article, or apparatus. Furthermore, unless expressly stated otherwise, “or” refers to an inclusive or not an exclusive or. For example, condition A or B is satisfied by any one of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or do not exist).
[0292] Furthermore, the use of "a" or "an" is used to describe elements and components of embodiments herein. This is done simply for convenience and to understand the usual meaning in the invention. This description should be interpreted as including one or at least one, and unless it becomes clear that it is otherwise, the singular also includes the plural.
[0293] In some parts of this description, embodiments are described in terms of algorithms and symbolic representations of operations on information. Descriptions and representations of these algorithms are commonly used by those skilled in data processing technology to effectively communicate the nature of the work to others skilled in the art. These operations are described functionally, computationally, or logically, but are understood to be implemented by computer programs, or equivalent electrical circuits, or microcode. Furthermore, without loss of generality, it has sometimes proven convenient to refer to arrangements of these operations as modules. The operations described and the modules associated with them may be embodied in software, firmware, hardware, or any combination thereof.
[0294] Any of the steps, operations, or processes described herein may be performed or implemented by one or more hardware or software modules alone or in combination with other devices. In some embodiments, a software module is implemented by a computer program product including a computer-readable medium containing computer program code, and the computer program code may be executed by a computer processor to perform any or all of the steps, operations, or processes described herein.
[0295] Embodiments may also relate to apparatus for performing the operations described herein. Such apparatus may include general-purpose computing devices that can be specifically constructed for the required purpose and / or selectively invoked or reconfigured by computer programs stored in a computer. Such computer programs may be stored in non-temporary, tangible computer-readable storage media or any type of media suitable for storing electronic instructions that can be coupled to a computer system bus. Furthermore, all computing systems referred to herein may include a single processor or may be architectures employing multiple processor designs to enhance computing power.
[0296] Embodiments may also relate to products generated by computing processes described herein. Such products may include information obtained from computing processes, which is stored in a non-temporary, tangible, computer-readable storage medium, and may include any embodiment of a computer program product or any combination of other data described herein.
[0297] A person skilled in the art will understand, upon reading this disclosure, further additional and alternative structural and functional designs of systems and processes for decorrelation of audio content based on the principles disclosed herein. Therefore, while specific embodiments and applications are illustrated and described, it should be understood that the disclosed embodiments are not limited to the exact structures and components disclosed herein. Various modifications, changes, and variations will be apparent to a person skilled in the art, without departing from the spirit and scope defined in the appended claims, in the arrangement, operation, and details of the methods and apparatus disclosed herein.
[0298] Finally, the language used herein has been selected primarily for readability and educational purposes, and not to describe or limit the patent rights. Therefore, the scope of the patent rights is intended to be limited not by this detailed description, but rather by any claims issued in connection with an application based herein. Accordingly, the disclosure of embodiments is intended to be illustrative, not limiting, the scope of the patent rights described in the following claims.
Claims
1. It is a system, One or more processors, A non-temporary computer-readable medium that, when executed by one or more of the aforementioned processors, Separating the audio channel into low-frequency and high-frequency components, Determining a target amplitude response that defines one or more spatial implications, The aforementioned target amplitude response is converted into a transfer function for constructing a single-input, multiple-output full-pass filter, The high-frequency components are processed using the full-pass filter configured based on the transfer function to generate a plurality of processed high-frequency components. The low-frequency component is coupled to at least one of the plurality of processed high-frequency components to generate one or more processed channels. To perform this, a non-temporary computer-readable medium containing stored program code comprising one or more of the aforementioned processors and A system characterized by comprising the following features.
2. The plurality of processed high-frequency components include a first processed high-frequency component and a second processed high-frequency component, and the program code is The first processed high-frequency component is coupled to the low-frequency component to generate the left channel, The second processed high-frequency component is coupled to the low-frequency component to generate the right channel. The system according to claim 1, further characterized in that one or more processors are configured to perform the following.
3. The system according to claim 1, characterized in that the transfer function corresponds to a frequency-dependent phase shift.
4. The system according to claim 1, characterized in that the target amplitude response defines a spatial indication for the mid-frequency or side-frequency components of the plurality of processed high-frequency components.
5. The system according to claim 1, wherein the single-input, multiple-output full-pass filter is configured to apply a Hilbert transform to the high-frequency components to generate a left-foot and a right-foot component.
6. The aforementioned single-input, multiple-output full-pass filter is, Applying a first series of full-pass filters to the high-frequency components generates the left-foot component, The right-foot component is generated by applying a first delay and a second series of full-pass filters. The system according to claim 5, characterized in that it is configured to perform the following.
7. The aforementioned single-input, multiple-output full-pass filter is, The left leg component is output as a rotated right leg component, Performing a 2D rotation on the left leg component and the right leg component, To output a rotated left leg component that corresponds to the combined output of the 2D rotation performed on the left leg component and the right leg component. The system according to claim 5, characterized in that it is configured to perform the following.
8. The aforementioned single-input, multiple-output full-pass filter is, The first Hilbert transform is applied to the high-frequency component to generate a first left-foot component and a first right-foot component, wherein the first left-foot component is 90 degrees out of phase with respect to the first right-foot component. The method involves applying a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component, and the plurality of processed high-frequency components include the first left-foot component and the second right-foot component. The system according to claim 1, characterized in that it is configured to perform the following.
9. The aforementioned program code is: To generate mid-range and side-range components from the received audio signal, The method involves generating a hypermid component, which includes the spectral energy of the side component removed from the spectral energy of the mid component, wherein the audio channel corresponds to at least the hypermid component. The system according to claim 1, further characterized in that one or more processors are configured to perform the following.
10. The system according to claim 1, characterized in that the target amplitude response defines a compensation null for encoding vertical spatial implications.
11. The system according to claim 1, wherein the program code further configures one or more processors to convert the target amplitude response into coefficients for the single-input, multiple-output, full-pass filter using the inverse discrete Fourier transform (idft).
12. The system according to claim 1, further comprising configuring one or more processors to use a phase vocoder to convert the target amplitude response into coefficients for a single-input, multiple-output, full-pass filter.
13. The system according to claim 1, characterized in that the target amplitude response defines one or more parametric spatial suggestions, including one or more of the target broadband attenuation, critical point, filter characteristics, and soundstage position.
14. A non-temporary computer-readable medium containing stored program code, wherein the program code is executed by one or more processors. Separating the audio channel into low-frequency and high-frequency components, Determining a target amplitude response that defines one or more spatial implications, The aforementioned target amplitude response is converted into a transfer function for constructing a single-input, multiple-output full-pass filter, The high-frequency components are processed using the full-pass filter configured based on the transfer function to generate a plurality of processed high-frequency components. The low-frequency component is combined with at least one of the plurality of processed high-frequency components to generate one or more processed channels. A non-temporary computer-readable medium characterized by configuring one or more processors to perform the following.
15. The plurality of processed high-frequency components include a first processed high-frequency component and a second processed high-frequency component, and the program code is The first processed high-frequency component is coupled to the low-frequency component to generate the left channel, The second processed high-frequency component is coupled to the low-frequency component to generate the right channel. The non-temporary computer-readable medium according to claim 14, further comprising configuring one or more processors to perform the following:
16. The non-transient computer-readable medium according to claim 14, characterized in that the transfer function corresponds to a frequency-dependent phase shift.
17. The non-transient computer-readable medium according to claim 14, characterized in that the target amplitude response defines the spatial indication of the mid- or side components of the plurality of processed high-frequency components.
18. The non-temporary computer-readable medium according to claim 14, characterized in that the single-input, multiple-output, wide-range currency filter is configured to apply a Hilbert transform to the high-frequency components to generate left-foot and right-foot components.
19. The aforementioned single-input, multiple-output full-pass filter is, Applying a first series of full-pass filters to the high-frequency components generates the left-foot component, Applying a first delay and a second series of full-pass filters to the high-frequency components generates the right-foot component. A non-temporary computer-readable medium according to claim 18, characterized in that it is configured to perform the following.
20. The aforementioned single-input, multiple-output full-pass filter is, Outputting the left leg component as the rotated right leg component, Performing a 2D rotation on the left leg component and the right leg component, To output a rotated left leg component that corresponds to the combined output of the 2D rotation performed on the left leg component and the right leg component. A non-temporary computer-readable medium according to claim 18, characterized in that it is configured to perform the following.
21. The aforementioned single-input, multiple-output full-pass filter is, The first Hilbert transform is applied to the high-frequency component to generate a first left-legged component and a first right-legged component, wherein the first left-legged component is 90 degrees out of phase with respect to the first right-legged component. The method involves applying a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component, and the plurality of processed high-frequency components include the first left-foot component and the second right-foot component. A non-temporary computer-readable medium according to claim 14, characterized in that it is configured to perform the following.
22. The aforementioned program code is: To generate mid-range and side-range components from the received audio signal, The method involves generating a hypermid component, which includes the spectral energy of the side component removed from the spectral energy of the mid component, wherein the audio channel corresponds to at least the hypermid component. The non-temporary computer-readable medium according to claim 14, further comprising configuring one or more processors to perform the following:
23. The non-temporary computer-readable medium according to claim 14, characterized in that the target amplitude response defines a compensation null for encoding vertical spatial implications.
24. The non-temporary computer-readable medium according to claim 14, further comprising configuring one or more processors to convert the target amplitude response into coefficients for the single-input, multiple-output, full-pass filter using an inverse discrete Fourier transform (idft).
25. The non-temporary computer-readable medium according to claim 14, further comprising configuring one or more processors to use a phase vocoder to convert the target amplitude response into coefficients for the single-input, multiple-output, full-pass filter, wherein the program code further comprises a phase vocoder.
26. The non-temporary computer-readable medium according to claim 14, characterized in that the target amplitude response defines one or more parametric spatial suggestions, including one or more of a target broadband attenuation, a critical point, filter characteristics, and soundstage position.
27. By one or more processors, The steps include separating the audio channel into low-frequency and high-frequency components, A step of determining a target amplitude response that defines one or more spatial implications, The steps include converting the target amplitude response into a transfer function for constructing a single-input, multiple-output full-pass filter, The steps include: processing the high-frequency components using the full-pass filter configured based on the transfer function to generate a plurality of processed high-frequency components; The steps include: combining the low-frequency component with at least one of the plurality of processed high-frequency components to generate one or more processed channels; A method characterized by comprising:
28. The plurality of processed high-frequency components include a first processed high-frequency component and a second processed high-frequency component, The steps include: coupling the first processed high-frequency component with the low-frequency component to generate the left channel; The steps include: coupling the second processed high-frequency component with the low-frequency component to generate the right channel; The method according to claim 27, further comprising:
29. The method according to claim 27, characterized in that the transfer function corresponds to a frequency-dependent phase shift.
30. The method according to claim 27, characterized in that the target amplitude response defines the spatial indication of the mid-frequency or side-frequency components of the plurality of processed high-frequency components.
31. The method according to claim 27, further comprising the step of applying a Hilbert transform to the high-frequency components using the single-input, multiple-output full-pass filter to generate a left-bundle component and a right-bundle component.
32. The aforementioned single-input, multiple-output full-pass filter, The steps include applying a first series of full-pass filters to the high-frequency components to generate the left-foot component, The steps include: applying a first delay and a second series of full-pass filters to the high-frequency components to generate the right-foot component; The method according to claim 31, further comprising:
33. The aforementioned single-input, multiple-output full-pass filter, The steps include outputting the left leg component as the rotated right leg component, The steps include performing a 2D rotation on the left leg component and the right leg component, A step of outputting a rotated left leg component that corresponds to the combined output of the 2D rotation performed on the left leg component and the right leg component. The method according to claim 31, further comprising:
34. The aforementioned single-input, multiple-output full-pass filter, A step of applying a first Hilbert transform to the high-frequency component to generate a first left-foot component and a first right-foot component, wherein the first left-foot component is 90 degrees out of phase with respect to the first right-foot component. A step of applying a second Hilbert transform to the first right-foot component to generate a second left-foot component and a second right-foot component, wherein the second left-foot component is 90 degrees out of phase with respect to the second right-foot component, and the plurality of processed high-frequency components include the first left-foot component and the second right-foot component. The method according to claim 27, further comprising:
35. A step of generating mid-range and side-range components from the received audio signal, A step of generating a hypermid component, which includes the spectral energy of the side component removed from the spectral energy of the mid component, wherein the audio channel corresponds to at least the hypermid component. The method according to claim 27, further comprising:
36. The method according to claim 27, characterized in that the target amplitude response defines a compensation null for encoding vertical spatial implications.
37. The method according to claim 27, further comprising the step of using the inverse discrete Fourier transform (idft) to convert the target amplitude response into coefficients for the single-input multiple-output full-pass filter.
38. The method according to claim 27, further comprising the step of using a phase vocoder to convert the target amplitude response into coefficients for the single-input multiple-output full-pass filter.
39. The method according to claim 27, characterized in that the target amplitude response defines one or more parametric spatial suggestions, including one or more of the target broadband attenuation, critical point, filter characteristics, and soundstage position.