Customized Binaural Rendering of Audio Content

The method and system for customizable binaural rendering of audio content address the challenge of maintaining scene stability in spatial audio by separating audio signals and modifying them based on the listening context, resulting in an enhanced spatial audio experience.

JP2025516333APending Publication Date: 2025-05-27DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024565097
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2023-05-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing technologies face challenges in rendering spatial audio content binaurally through headphones or earphones, particularly in maintaining scene stability when the listener moves their head.

Method used

A method and system for customizable binaural rendering of audio content, which involves separating a stereo audio signal into steer and diffuse signals, determining diffusion signal modification parameters based on the listening context, and generating an output multi-channel signal to redistribute or attenuate the diffuse signal, thereby enhancing scene stability or immersion.

Benefits of technology

The solution allows for a balanced customization of the listening experience, prioritizing either scene stability or immersion based on the type of audio content, thereby improving the overall spatial audio experience when using headphones or earphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516333000001_ABST
    Figure 2025516333000001_ABST
Patent Text Reader

Abstract

A method, system, and medium for processing audio are provided. In some embodiments, a method for processing audio includes receiving a stereo audio signal. The method may include separating the stereo audio signal into a steer signal and a diffusion signal. The method may include determining one or more diffusion signal modification parameters based on a current listening context, where the one or more diffusion signal modification parameters indicate a proportion of the diffusion signal to be redistributed to one or more output channels in an output multi-channel signal, or a degree of attenuation to be applied to the diffusion signal. The method may include generating the output multi-channel signal based on the steer signal, the diffusion signal, and the one or more diffusion signal modification parameters. The method may include providing the output multi-channel signal to a virtualizer for rendering as a binaural audio signal for playback on a wearable device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - References to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 63 / 497,025, filed Apr. 19, 2023, and PCT Application No. PCT / CN2022 / 0901993, filed May 5, 2022, and incorporates the entire disclosures of each of these applications by reference.

[0002] Technical Field The present disclosure relates to systems, methods, and media for customized binaural rendering of audio content.

Background Art

[0003] Background Viewers of media content are increasingly interested in spatial audio that can create an immersive experience. For example, when listening to immersive audio content, a listener may feel as if the audio content is surrounding them. However, rendering spatial audio content can be difficult, especially when the sound is rendered binaurally through headphones or earphones.

[0004] Notation and Nomenclature Throughout the present disclosure, including the claims, "speaker" and "loudspeaker", and "audio playback transducer" are used interchangeably to represent any acoustic radiation transducer (or set of transducers) driven by a single speaker feed. A typical headset includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter) that are driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuit branches connected to different transducers.

[0005] Throughout the present disclosure, including the claims, the expression performing an operation (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) on a signal or data is used in a broad sense to mean performing the operation directly on the signal or data or performing the operation on a processed version of the signal or data (e.g., a pre-filtered or pre-processed version of the signal before receiving the execution of the operation).

[0006] Throughout the present disclosure, including the claims, the expression "system" is used in a broad sense to mean a device, system, or subsystem. For example, a subsystem implementing a decoder may sometimes be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.

[0007] Throughout the present disclosure, including the claims, the term "processor" is used in a broad sense to mean a system or device that is programmable or otherwise configurable (e.g., by software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors configured by programming and / or otherwise to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chip sets, among others.

SUMMARY OF THE INVENTION

[0008] Overview A method, system, and medium are provided for customizable binaural rendering of audio content. In some embodiments, the method includes receiving a stereo audio signal. The method may further include separating the stereo audio signal into a steer signal and a diffuse signal, where the steer signal corresponds to the directional content in the stereo audio signal and the diffuse signal corresponds to the background content in the stereo audio signal. The method may further include determining one or more diffuse signal modification parameters based on a current listening context, where the one or more diffuse signal modification parameters indicate a proportion of the diffuse signal to be redistributed to one or more output channels in an output multi-channel signal, or a degree of attenuation to be applied to the diffuse signal. The method may further include generating the output multi-channel signal based on the steer signal, the diffuse signal, and the one or more diffuse signal modification parameters. The method may further include providing the output multi-channel signal to a virtualizer for rendering as a binaural audio signal for playback on a wearable device.

[0009] In some examples, the one or more output channels include at least one of a left channel, a right channel, or a center channel.

[0010] In some examples, generating the output multi-channel signal includes obtaining a dithering matrix, generating a modified dithering matrix using the dithering matrix and the one or more spread signal correction parameters, generating a spread multi-channel signal using the modified dithering matrix, and generating the output multi-channel signal based on the spread multi-channel signal and the steer signal. In some examples, the one or more spread signal correction parameters cause the ratio of the spread signals to be redistributed to the one or more output channels, and generating the modified dithering matrix includes determining a matrix dot product of a norm associated with the dithering matrix, a matrix representing the spread signal correction parameters, a matrix associated with the one or more spread signal correction parameters, and the dithering matrix. In some examples, the one or more spread signal correction parameters include one spread signal redistribution correction parameter indicating redistribution of the spread signals in the multi-channel output, and the norm normalizes the energy of the spread signals. In some examples, the one or more spread signal correction parameters cause the degree of attenuation applied to the spread signals, and generating the output multi-channel signal includes performing energy normalization configured such that the energy of the output multi-channel signal is the same as the energy of the stereo audio signal. In some examples, the energy normalization is performed by either an upmixer that generates the output multi-channel signal or the virtualizer.

[0011] In some examples, the current listening context includes one of a movie content viewing mode, a music listening mode, or a game play mode. In some examples, the current listening context is the movie content viewing mode, and the one or more diffusion signal correction parameters are in the range of about 0.8 to 1. In some examples, the current listening context is the music listening mode, and the one or more diffusion signal correction parameters are in the range of about 0 to 0.2. In some examples, the current listening context is the game play mode, and the one or more diffusion signal correction parameters have a value smaller than the value associated with the movie content viewing mode.

[0012] In some examples, the one or more diffusion signal correction parameters are received from a user of the wearable device. In some examples, the one or more diffusion signal correction parameters are received via a user interface.

[0013] In some examples, the generation of the output multi-channel signal is performed on a companion user device associated with the wearable device, and the virtualizer includes one or more components executed on the wearable device. In some examples, it further includes transmitting data from the companion user device to the wearable device via a BLUETOOTH communication protocol.

[0014] In some examples, the wearable device includes one of earphones or headphones.

[0015] In some examples, the wearable device has one or more sensors that collect sensor data that can be used to generate head tracking information related to the wearer of the wearable device

[0016] In some examples, the virtualizer is configured to render the binaural audio signal based on the output multi-channel signal and the head tracking information.

[0017] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein. Memory devices may include, but are not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having stored software.

[0018] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially executing the methods disclosed herein. In some aspects, the apparatus may be or may include an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.

[0019] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions of the following figures may not be drawn to exact scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0020]

Figure 1

[0021]

Figure 2

[0022]

Figure 3

[0023]

Figure 4

[0024]

Figure 5

[0025]

Figure 6

[0026]

Figure 7

[0027] Like reference numerals and designations in the various drawings indicate like elements.

DETAILED DESCRIPTION OF THE INVENTION

[0028] Detailed Description of Embodiments Viewers of media content are showing an increasing interest in spatial audio that creates a sense of immersion. For example, when listening to immersive audio content, a listener may feel as if the audio content is surrounding them. However, rendering spatial audio content can be difficult, especially when the sound is rendered binaurally via headphones or earphones. For example, spatial audio rendered binaurally via headphones or earphones may feel unstable when the user moves their head. As an example, if a listener is listening to music and desires the perception that the vocalist and / or instrumentalist is in front of them, and the spatial audio is rendered binaurally and the listener moves their head (e.g., to look around), the listener may perceive a jump or discontinuity. Listeners may have different preferences regarding whether to prioritize immersion or scene stability, which may further depend on the type of audio content being listened to, so rendering spatial audio via headphones and / or earphones that perform head orientation determination can be particularly difficult. For example, a listener may prioritize scene stability, in which case, when listening to music, a direct signal or a steered signal (such as vocals or instrumental performance) is perceived as fixed in front of the listener. In other words, such rendering can make the listener recognize that they are listening to music in front of a front speaker. Conversely, when listening to audio content related to a movie, a listener may prioritize immersion. Customizing the listening experience of spatial audio presented via headphones or earphones can be particularly challenging. As used herein, "immersiveness" refers to rendering audio data in a manner that is perceived as three-dimensional and surrounding the user.Immersive audio content may involve rendering an audio object as having a predetermined spatial position with respect to a listener having a specific azimuth and / or elevation. For example, immersive audio data can provide a listening experience where the audio sound is rendered to be perceived not only in front of the listener but also surrounding the listener. As a specific example, immersive audio content may include the sound of an airplane or helicopter rendered such that the listener perceives the sound above their head. The techniques disclosed herein enable a user to adjust the audio scene, for example, by enabling the fixation of an audio object at a specific perceived position (e.g., the screen on which the content is rendered), or by enabling the audio object to be perceived as enveloping or surrounding the user. For example, the techniques described herein enable a balance between scene stability (generally referring to a listening experience where the position of the audio object does not change even when the listener moves their head (e.g., moves left or right, looks around)) and immersion. By enabling the listener to balance scene stability and immersion, it may be possible for the listener to customize the listening experience based on the type of content they are listening to.

[0029] Disclosed herein is a technique for generating a customized multi-channel output signal based on a current listening context. The customized multi-channel output signal can be generated by considering a diffusion signal modification parameter that attenuates the diffusion signal (thereby making the direct signal or the steer signal more perceptible, and as a result, the perception of scene stability can be increased), or by redistributing at least a portion of the diffusion signal to one or more output channels such as the left, right, or center channel. When at least a portion of the diffusion signal is redistributed, the degree to which the diffusion signal is distributed can depend on the listening context. For example, when the listening context indicates that the listener is watching a movie, none of the diffusion signal may be distributed, or a relatively small percentage of the diffusion signal may be distributed to the left, right, and center channels, thereby maintaining the immersive feeling. Conversely, when the listening context indicates that the listener is listening to music or playing a game, a larger percentage of the diffusion signal may be distributed to one or more output channels, thereby increasing the perception of scene stability when the user moves their head.

[0030] In some aspects, the customized multi-channel output signal can be generated by an upmixer component of the device. The device can be, for example, a user device (such as a mobile phone, tablet computer, laptop computer, desktop computer, gaming console, television, etc.) that presents audio content via paired or connected headphones or earphones. The customized multi-channel output signal can then be rendered as a binaural audio signal by a virtualizer component. In some embodiments, the virtualizer component can be part of the headphones or earphones such that the rendering as a binaural audio signal depends on the orientation of the user's head.

[0031] FIG. 1 is a block diagram of a system configured to generate and utilize customized binaural audio rendering in some embodiments. As shown, an upmixer 102 receives a stereo audio signal including a left audio signal and a right audio signal. The upmixer 102 further receives diffusion signal modification information. The diffusion signal modification information may indicate how the diffusion signal is attenuated or reassigned to one or more channels, such as one or more of the left, right, and center channels of the customized multi-channel output signal generated by the upmixer 102. The diffusion signal modification information may correspond to a current listening context. Examples of listening contexts include the user listening to music, watching a movie, playing a game (e.g., a computer game), etc. The diffusion signal modification information may include one or more parameters, and each set of one or more parameters may be associated with a given listening context. The set of one or more parameters may be stored (e.g., in a memory) in a device used to present the audio content, and the device may be configured to search for the set of diffusion signal modification parameters corresponding to the current listening context. The upmixer 102 may then generate a multi-channel audio signal. An example of an upmixer system is shown in FIG. 3 and will be described later in connection with this figure. Techniques for generating a customized multi-channel audio signal based on diffusion signal attenuation information are shown in FIGS. 4, 5, and 6 and will be described later in connection with these figures.

[0032] The upmixer 102 may provide the multichannel audio signal to the virtualizer 104. Note that in some embodiments, the upmixer 102 may be implemented on a companion device (e.g., a mobile phone, a tablet computer, a laptop computer, etc.) that provides an audio signal for playback by a paired set of headphones or earphones, and the virtualizer 104 may be implemented on the paired headphones or earphones. In such an embodiment, the upmixer 102 may transmit the multichannel audio signal to the virtualizer 104 via BLUETOOTH, or another wireless communication protocol. Note that in some embodiments, the upmixer 102 and the virtualizer 104 may be implemented on the same device.

[0033] The virtualizer 104 may receive a multi-channel audio signal and render the multi-channel audio signal as a binaural audio signal suitable for reproduction via, for example, headphones or earphones. Note that the virtualizer 104 may render the multi-channel audio signal based on head tracking information obtained using one or more sensors (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers). The one or more sensors can be disposed on the headphones or earphones. Since the multi-channel audio signal is generated based on the current listening context (e.g., indicating the type of audio content or video content being consumed by the user), in a manner dependent on the listening context, the diffusion signal may be attenuated (e.g., boosting the audio signal rendered in front of the user relative to the diffusion signal), or reallocated to one or more channels such as one or more of the left, right, and center channels of the upmixed multi-channel audio signal. Thereafter, the virtualizer can render the multi-channel audio signal as a binaural audio signal in a manner dependent on the orientation of the listener's head. As a result, the binaural audio signal can be presented in a manner dependent on both the orientation of the listener's head and the listening context such that the diffusion signal is attenuated or redistributed in a customized manner along the user's listening preferences for various types of audio content. As an example, when the listening context corresponds to watching a movie, the binaural audio signal can be rendered to have a substantial sense of immersion regardless of the orientation of the user's head. As another example, when the listening context corresponds to listening to music, the binaural audio signal can be rendered such that the vocals and instruments are perceived to be in front of the user regardless of the orientation of the user's head, thereby improving the stability of the scene.

[0034] Figures 2A, 2B, and 2C show the effects when the value of the diffusion signal correction parameter, generally represented as β in this specification, is changed. Generally, β indicates the degree to which the diffusion signal is spread or re-allocated to other channels of the multi-channel mix (e.g., left, right, and / or center channels). Note that in Figures 2A, 2B, and 2C, L s and R s represent the left and right surround signals respectively, and L, R, C represent the left, right, and center channels of the 5.1 upmix signal respectively. Referring to Figure 2A, when β is 1, none of the diffusion signals from the left and right surround channels are spread or re-allocated to the left, right, and center channels. Referring to Figure 2B, when β is between 0 and 1, some of the diffusion signals from the left and right surround channels are spread or re-allocated to the left, right, and center channels, but some of the diffusion signals remain in the left and right surround channels. Referring to Figure 2C, when β is 0, all of the diffusion signals from the left and right surround channels are spread or re-allocated to the left, right, and center channels. It should be noted that the distribution of the diffusion signal may also be performed for other upmix formats such as 7.1 upmix.

[0035] In some embodiments, the diffusion signal modification may be performed by an upmixer. In some embodiments, the upmixer may be a component or module of a user device that provides audio content for playback. For example, the user device may be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a gaming console, a television, or the like. The upmixer may be configured to receive a stereo audio signal (e.g., a left signal and a right signal) and generate a customized multi-channel output signal based on diffusion signal modification parameters. The diffusion signal modification parameters may be received (e.g., acquired and / or identified) by the upmixer based on the current listening context of the user of the user device. The upmixer may be configured to separate a direct signal from a diffusion signal, where the direct signal corresponds to, for example, vocal and / or instrument sounds, and the diffusion signal generally corresponds to ambient and / or environmental sounds. The upmixer may be configured to pan the direct signal such that the direct signal is rendered as if it were located at a single point. The upmixer may be configured to decorrelate and spread the diffusion signal such that the diffusion signal is attenuated or reassigned to output channels, thereby affecting the stability and / or immersion of the audio content scene experienced by the listener. The upmixer may be configured to combine the panned direct signal with the spread and / or attenuated diffusion signal into a customized multi-channel output signal. Note that the upmixer may be configured to receive the stereo signal in the time domain and generate the multi-channel output signal in the time domain. However, in some embodiments, the upmixer may generally be configured to perform processing in the frequency domain. In other words, in some embodiments, the upmixer may convert the received stereo signal to the frequency domain before separating the direct signal from the diffusion signal, spreading and / or attenuating the diffusion signal to other channels, generating the combined multi-channel output signal, etc.Subsequently, the upmixer may convert the multi-channel output signal from the frequency domain to the time domain before providing the multi-channel output signal to the virtualizer for rendering.

[0036] FIG. 3 is a block diagram of an exemplary upmixer 300 in some embodiments. In some embodiments, the upmixer 300 may be implemented using one or more processors or controllers of, for example, a user device (such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a game console, etc.). An example of such a controller is the control system 710 of FIG. 7.

[0037] As shown, the upmixer 300 can receive stereo audio signals generally represented herein as L T (n) (e.g., left stereo signal) and R T (n) (e.g., right stereo signal), where n represents the current audio frame. The stereo signal may be converted from the time domain to the frequency domain using time-frequency conversion blocks 302a and 302b. For example, the time-frequency conversion block 302a can output a frequency domain representation of the left stereo signal generally represented herein as L T (m,k), and the time-frequency conversion block 302b can output a frequency domain representation of the right stereo signal generally represented herein as R T (m,k). Here, m represents the time block index and k represents the frequency index.

[0038] The frequency domain representation of the stereo audio signal may be provided to the statistical estimation block 304. The statistical estimation block 304 may generate estimated parameters X(m,b), Y(m,b), and T(m,b), which may be provided to the separation block 306. The frequency domain representation of the stereo audio signal may also be provided to the separation block 306 as shown in FIG. 3.

[0039] The separation block 306 may be configured to separate the direct signal and the spread signal. The direct signal may be represented as S L (m,k) and S R (m,k) for the left and right direct signals respectively. The spread signal can be represented as d L (m,k) and d R (m,k). In some embodiments, the spread signal can be obtained by subtracting the steer signal estimated from the input stereo signal, where the subtraction is performed in the frequency domain. For example, in some embodiments, the left and right spread signals can be determined as follows: [Number]

[0040] In the above equation, W(m,b) represents the steer signal separation matrix.

[0041] The direct signals S L (m,k) and S R (m,k) may be supplied to the panning block 308. The panning block 308 may be configured to place the direct signal at a specific position. Note that the panning block 308 may utilize the parameters generated by the statistical estimation block 304. Further, note that the panning block 308 is configured to output a multi-channel output of the direct signal. For example, the multi-channel direct signal may have N channels, where N is the total number of channels. As an example, N is 5 for a 5.1 upmix and N is 7 for a 7.1 upmix.

[0042] The decorrelation and spreading block 310 receives the spread signals d L (m,k) and d R(m,k) may be received. The decorrelation and spreading block 310 may be further configured to receive, obtain, or determine a spreading signal correction parameter that may be specified by a diffusion energy adjustment matrix B. The spreading signal correction parameter may be received or determined based on the current listening context. The decorrelation and spreading block 310 may be configured to modify the spreading signal such that the spreading signal is attenuated (thereby making the direct signal more prominent), or such that the spreading signal is reassigned to other output channels (e.g., the left, right, and center channels in a 5.1 upmix). The decorrelation and spreading block 310 may modify the spreading signal by generating a modified spreading matrix that controls the degree to which the spreading signal is present in various output channels. The spreading matrix, generally represented herein as O, may be obtained based on parameters generated by the statistical estimation block 304 and may be modified based on the spreading signal correction parameter. Techniques for generating the modified spreading matrix are shown in FIGS. 5 and 6 and are described below in connection with these figures. The modified spreading signal is generally represented herein as Z d1 (m,k),…Z dN (m,k). Here, N is the number of channels of the multi-channel output signal.

[0043] The panned multi-channel direct signal may be combined with the modified spreading signal generally represented herein as Z d1 (m,k),…Z dN (m,k) to generate an output multi-channel signal in the frequency domain. Here, N is the number of channels of the multi-channel output signal. Next, the output signal may be converted to the time domain to generate a multi-channel output signal in the time domain generally represented herein as Z 1 (n),…Z N (n). Here, n is the audio frame number and N is the number of output channels. The conversion to the time domain may be implemented via a set of frequency-time conversion blocks such as blocks 312a and 312b.

[0044] As described above, the upmixer can generate a customized multi-channel output based on the current listening context. The current listening context may indicate the type of content presented by the user device (e.g., whether the content is music content, video content such as a movie or a TV show, game content, etc.), whether the audio content is presented via paired headphones or earphones, and so on. For example, the diffusion signal correction parameter may be applied when the audio content is presented via paired headphones or earphones, and may not be applied when the audio content is presented directly by the user device (e.g., via the speakers of the user device). In some aspects, the diffusion signal correction parameter may be obtained, retrieved, or otherwise determined based on the type of content presented. For example, the user device may store different sets of diffusion signal correction parameters applicable to different types of content. As an example, the diffusion signal correction parameter may include a diffusion signal attenuation parameter that attenuates the diffusion signal (thereby rendering the vocals and instrument sounds more prominent) in response to determining that the current listening context corresponds to the playback of music content. As another example, the diffusion signal correction parameter may include a first set of diffusion signal correction parameters that prevent the diffusion signal from being reallocated to other output channels at all in response to determining that the current listening context corresponds to the playback of a movie, thereby rendering the movie audio content more immersively. As yet another example, the diffusion signal correction parameter may include a second set of diffusion signal correction parameters that reallocate or redistribute a portion of the diffusion signal to other output channels in response to determining that the current listening context corresponds to playing a video game, thereby increasing the stability of the scene when the user moves their head at the expense of immersive perception.

[0045] Regardless of whether the spread signal is attenuated or reassigned to another output channel, the spread signal correction parameter can be applied by modifying the spread matrix to generate a modified spread matrix. Thus, the modified spread matrix can indicate the degree to which the spread signal is attenuated or redistributed. The modified spread signal is determined based on the modified spread matrix. Then, by combining the multi-channel modified spread signal with the multi-channel direct signal and converting the combined signal into the time domain, a customized multi-channel output signal can be determined. Note that when the spread signal is attenuated (e.g., when the spread signal correction parameter corresponds to a parameter that renders direct signals such as vocals and / or instruments more prominent), the multi-channel output signal may have less energy than the input stereo signal. Thus, in such cases, in some embodiments, the multi-channel output signal may be normalized so that the output signal has the same energy as the input stereo signal.

[0046] FIG. 4 is a flowchart of an example process 400 for generating a customized multi-channel output signal in some embodiments. In some aspects, the blocks of process 400 may be executed on a user device. For example, the user device may be one that plays audio content (or audio content associated with video content) via paired headphones or earphones. Examples of such user devices include mobile phones, tablet computers, laptop computers, desktop computers, gaming consoles, televisions, and the like. The blocks of process 400 may be executed by one or more processors or controllers of the user device. An example of such a controller is control system 710 shown in FIG. 7 and described later in connection with this figure. In some aspects, the blocks of process 400 may be executed in an order other than that shown in FIG. 4. In some aspects, two or more blocks of process 400 may be executed substantially in parallel. In some aspects, one or more blocks of process 400 may be omitted.

[0047] Process 400 may begin by receiving a stereo audio signal at 402. The stereo audio signal may be received by an upmixer. As described above, herein, a stereo audio signal is generally represented as L T (n) (e.g., left stereo signal) and R T (n) (e.g., right stereo signal), where n represents the current audio frame. Note that after receiving the stereo audio signal, process 400 may convert the stereo audio signal from the time domain to the frequency domain. For example, process 400 may utilize a short-time Fourier transform (STFT).

[0048] At 404, process 400 can separate a stereo audio signal into a direct signal and a diffuse signal. Process 400, as shown in FIG. 3 and described above in connection with this figure, can separate the stereo signal into a direct signal and a diffuse signal based on a statistical estimate performed using the frequency domain representation of the stereo audio signal. Generally herein, the direct signals are, respectively, for the left and right direct signals, S L (m,k) and S R (m,k), and generally herein, the diffuse signals are, respectively, for the left and right diffuse signals, d L (m,k) and d R (m,k). In some embodiments, the diffuse signal can be obtained by subtracting the steering signal estimated from the input stereo signal using, for example, the steering signal separation matrix as described above in connection with FIG. 3.

[0049] At 406, process 400 can determine one or more diffuse signal modification parameters based on the current listening context. As described above, the current listening context can indicate the type of audio content being presented (e.g., whether the audio content is music content, audio content related to a movie or television show, audio content related to a video game, etc.), and / or whether the audio content is being played via headphones and / or earphones.

[0050] In some embodiments, the diffuse signal modification parameter(s) can include a parameter configured to attenuate the diffuse signal. Attenuation can be performed when increasing the stability of the scene while the user's head is moving is prioritized. For example, attenuation of the diffuse signal may be performed when the current listening context indicates that the audio content is music being listened to using paired headphones or earphones.

[0051] In some embodiments, the diffusion signal modification parameter(s) may include a parameter indicating the degree to which the diffusion signal is redistributed to one or more output channels (e.g., one or more of the left, right, and / or center channels), thereby causing a change in the degree of immersion perceived by the user. For example, if the diffusion signal modification parameter(s) indicates that the diffusion signal is not redistributed to one or more output channels (e.g., if the current listening context indicates that the audio content is related to a movie or television show and thus a high perception of immersion is desired), the diffusion signal modification parameter(s) may not redistribute any portion of the diffusion signal. Conversely, if the diffusion signal modification parameter(s) indicates that the diffusion signal is at least partially redistributed to other output channels (e.g., if the current listening context indicates that the audio content is related to content other than a movie or television show, such as music, video games, or podcasts), at least a portion of the diffusion signal may be redistributed, for example, to the left, right, and center channels, thereby increasing the perception of scene stability at the expense of the perception of immersion.

[0052] Note that the diffusion signal correction parameter(s) may be specified by the user of the user device or may be programmed into the user device, for example, by the manufacturer of the user device. The diffusion signal correction parameter(s) may include different sets of diffusion signal correction parameter(s) each applicable to a different listening context. Process 400 may then obtain the diffusion signal correction parameter applicable to the current listening context. If the diffusion signal correction parameter(s) are specified and / or modified by the user of the user device, the parameter(s) may be specified or modified via, for example, a user interface presented on the user device. For example, the user interface may include a slider control or other user interface control that enables the user to adjust the diffusion signal correction parameter for different listening contexts. The settings are saved, for example, in the memory of the user device for use during future listening sessions. Note that in some embodiments, the diffusion signal correction parameter may be received from a wearable device (e.g., a pair of earphones or headphones) and / or from a companion user device (e.g., a paired mobile device).

[0053] At 408, process 400 may generate an output multi-channel signal based on a direct signal, a spread signal, and one or more spread signal modification parameters. For example, process 400 may apply the spread signal modification parameters by applying the spread signal modification parameters to a dithering matrix to generate a modified dithering matrix. The modified dithering matrix may then be used to generate a modified spread signal that is attenuated or redistributed with respect to the original spread signal. The modified spread signal may then be combined with the direct signal to generate an output multi-channel signal. In some embodiments, the output multi-channel signal may then be converted to the time domain. An example of a process for generating an attenuated spread signal by generating a modified dithering matrix is shown in FIG. 5 and will be described later in connection with this figure. An example of a process for distributing a spread signal to other output channels by generating a modified dithering matrix is shown in FIG. 6 and will be described later in connection with this figure.

[0054] In some embodiments, at 410, process 400 can optionally perform energy normalization. For example, energy normalization can be performed in consideration of a decrease in the total energy of the output multi-channel signal when using the attenuated spread signal if the spread signal is attenuated at block 408. By energy normalization, the normalized multi-channel signal can have a total energy similar to that of the stereo audio signal received at block 402. Note that if the spread signal is redistributed to other output channels instead of being attenuated, energy normalization need not be performed and block 410 may be omitted. Further note that energy normalization may be performed by either an upmixer or a virtualizer. When energy normalization is performed by a virtualizer, block 410 may be performed after block 412, which will be described later. Energy normalization may be performed either in the time domain or the frequency domain.

[0055] At 412, process 400 can provide an output multi-channel signal to a virtualizer for rendering as a binaural audio signal to be played on a wearable device. The wearable device can include headphones or earphones. The virtualizer may be implemented either on a user device that provides the audio content or on the wearable device (e.g., headphones or earphones). In some aspects, the wearable device can include one or more sensors that can be used to determine the orientation of the user's head. Next, the virtualizer can render the binaural audio signal based on the head orientation. Thus, in the case where the diffusion signal correction parameter is a parameter that prioritizes scene stability, the rendering of the binaural audio signal may be such that the direct signal (e.g., vocals, musical instruments, etc.) is perceived in front of the user regardless of the user's head orientation. Conversely, in the case where the diffusion signal correction parameter is a parameter that prioritizes immersion, the binaural audio signal may be rendered such that the diffusion signal is distributed to give the perception of immersion even when the user moves their head. Note that since the diffusion signal correction parameter is specific to the listening context (which may indicate the type of audio content being presented), scene stability may be prioritized for some types of audio content while immersion may be prioritized for other types of audio content. Further, in some aspects, the diffusion signal correction parameter can be set or adjusted by the end user of the user device, so the listener can control which of scene stability or immersion is prioritized for different types of audio content and the degree to which each is prioritized.

[0056] In some embodiments, the diffuse signal modification parameter may include a diffuse signal attenuation parameter that attenuates the diffuse signal. Thereby, the steer signal or the direct signal can be rendered so as to be fixed and perceived in front of the listener even when the listener moves their head while wearing headphones or earphones. In other words, the steer signal or the direct signal can be rendered to be more prominent and to be perceived as being fixed in front, thereby increasing the stability of the listener's scene even while the listener is moving their head. The attenuation of the diffuse signal may be performed when the current listening context is listening to music content. This is because the steer signal or the direct signal may include a vocal or instrumental performance that is advantageously rendered to be more prominent. In some embodiments, the attenuation of the diffuse signal may include applying the diffuse signal attenuation parameter to the spreading matrix to generate a modified spreading matrix. The modified spreading matrix can be used to generate an attenuated modified diffuse signal for the original diffuse signal. Note that since the diffuse signal is attenuated, the multi-channel output signal including the attenuated diffuse signal and the multi-channel direct signal may have lower energy than the original stereo signal. Therefore, normalization may be performed to normalize the energy of the multi-channel output signal to the energy of the original stereo signal so that the rendered binaural signal does not become lower than the desired volume.

[0057] FIG. 5 is a flowchart of an example process 500 for attenuating a spread signal in some embodiments. In some embodiments, the blocks of process 500 may be executed on a user device. For example, the user device may be one that plays audio content (or audio content associated with video content) via paired headphones or earphones. Examples of such user devices include mobile phones, tablet computers, laptop computers, desktop computers, gaming consoles, televisions, and the like. The blocks of process 500 may be executed by one or more processors or controllers of the user device. An example of such a controller is control system 710 shown in FIG. 7 and described later in connection with this figure. In some embodiments, the blocks of process 500 may be executed in an order other than that shown in FIG. 5. In some embodiments, two or more blocks of process 500 may be executed substantially in parallel. In some embodiments, one or more blocks of process 500 may be omitted.

[0058] Process 500 can begin by obtaining, at 502, one or more spread signal attenuation parameters, a dither matrix, and a spread signal. Note that the spread signal may be provided in the frequency domain. As described above in connection with FIG. 3, the spread signal is d 1 (m,k),…d jIt can be represented as (m, k). Here, j is the number of spread signals, m is the time block index, and k is the frequency band. The spreading matrix can be represented as O in this specification and can have an N×j dimension. N is the number of channels. For example, in the case of a 5.1 channel upmix, N may be 5 and j may be 2. In other examples, N may be 7, 9, etc., and j may be 2, 3, 4, etc. In some embodiments, the spreading matrix may be generated based on statistical estimation parameters estimated by the statistical estimation block 304, as shown in FIG. 3 and described above in connection with this figure. The spread signal attenuation parameter can generally be represented as a matrix B having an N×j dimension similar to the matrix O in this specification. The spread signal attenuation parameter may be obtained from the memory of the user device and may be set based on the user's preference regarding the degree of attenuation of the spread signal for various listening contexts (e.g., for various types of audio content). It should be noted that in some aspects, each element of the matrix β may be between 0 and 1, a value of 0 indicates complete attenuation of the spread signal, and a value of 1 indicates no attenuation of the spread signal.

[0059] In 504, process 500 may generate a modified spreading matrix using the spreading matrix and one or more spread signal attenuation parameters. For example, in some aspects, the modified spreading matrix O' may be determined by:

Equation

[0060] In 506, process 500 can generate an attenuated spread multi-channel signal using the modified spreading matrix and the spread multi-channel signal. Generally denoted as Z in this specification d1 ,… Z dNNote that the spread multi-channel signal represented as is conventionally generated by multiplying the vector formed by the spread signal by the spreading matrix. However, to generate a attenuated spread multi-channel signal, process 500 may multiply the vector formed by the spread signal by a modified spreading matrix incorporating a spread signal attenuation parameter. For example, in some embodiments, the spread multi-channel signal may be determined by:

Number

[0061] The spread multi-channel signal is shown in FIGS. 3 and 4 and, as described above in connection with these figures, can be combined with the multi-channel direct signal to generate a multi-channel output signal in the frequency domain. The signal in the frequency domain may then be converted to the time domain as described above in connection with FIGS. 3 and 4.

[0062] Note that due to the attenuation of the spread signal, the total energy in the output multi-channel signal is attenuated, so the energy may be normalized before rendering the binaural audio signal. For example, the energy may be normalized based on the stereo signal. Such normalization may be performed by an upmixer, virtualizer, or any other component.

[0063] In some embodiments, the diffusion signal modification parameter may include a parameter that redistributes at least a portion of the diffusion signal to one or more output channels, such as the left, right, and / or center channels. Such redistribution may help effectively balance the scene stability and immersion in the presentation of immersive audio content. For example, when spatial audio is rendered from a device such as a tablet computer or a mobile phone and presented via the user's headphones or earphones, distributing at least a portion of the diffusion signal to other output channels can increase the perception of scene stability, especially when the user's head orientation changes. In some cases, the degree of redistribution of the diffusion signal may depend on the listening context. For example, in a listening context where the user is watching a movie, none of the diffusion signal may be redistributed, or a relatively low percentage (e.g., 10%, 15%, etc.) of the diffusion signal may be redistributed to prioritize immersion. As another example, in a listening context where the user is watching or playing a game, a relatively large percentage of the diffusion signal may be redistributed (e.g., larger compared to the percentage distributed in the movie-watching context). As an example, when watching or playing a game, 60%, 70%, 80%, etc. of the diffusion signal may be redistributed.

[0064] Similar to what was described above in connection with FIG. 5, in some embodiments, the modified diffusion signal may be determined by generating a modified spreading matrix using the diffusion signal modification parameter. However, unlike what was described above in connection with FIG. 5, since the total energy of the multi-channel output signal should not be attenuated, the modified spreading matrix may be generated using a normalization parameter that helps maintain the norm (e.g., Frobenius norm) of the spreading matrix. Further, in some embodiments, the diffusion signal modification parameter may include a single diffusion signal modification parameter. The use of a single diffusion signal modification parameter may make the user configuration easier.

[0065] FIG. 6 is a flowchart of process example 600 for distributing a spread signal in some embodiments. In some embodiments, the blocks of process 600 may be executed on a user device. For example, the user device may play audio content (or audio content associated with video content) via paired headphones or earphones. Examples of such user devices include mobile phones, tablet computers, laptop computers, desktop computers, gaming consoles, televisions, and the like. The blocks of process 600 may be executed by one or more processors or controllers of the user device. An example of such a controller is control system 710 shown in FIG. 7 and described hereinafter in connection with this figure. In some embodiments, the blocks of process 600 may be executed in an order other than the order shown in FIG. 6. In some embodiments, two or more blocks of process 600 may be executed substantially in parallel. In some embodiments, one or more blocks of process 600 may be omitted.

[0066] Process 600 can begin at 602 by obtaining one or more spread signal modification parameters, a spreading matrix, and a spread stereo signal. Note that the spread stereo signal may be provided in the frequency domain. As described above in connection with FIG. 3, the spread stereo signal can be represented as d 1 (m,k),…d j (m,k). Here, j is the number of spread signals, m is the time block index, and k is the frequency band. The spreading matrix may be represented as O herein and may have an N×j dimension. N is the number of channels. For example, in the case of a 5.1 channel upmix, N may be 5 and j may be 2. In other examples, N may be 7, 9, etc., and j may be 2, 3, 4, etc. In some embodiments, the spreading matrix may be generated based on statistical estimation parameters estimated by statistical estimation block 304, as shown in FIG. 3 and described above in connection with this figure.

[0067] In some aspects, the diffusion signal modification parameter(s) may be represented as a matrix B that can generally have an N×j dimension, similar to matrix O herein. Or, in some aspects, the diffusion signal modification parameter(s) may be a single parameter generally represented as β herein. The use of a single parameter may alleviate the complexity in configuring the user for the diffusion signal modification parameter(s) for various listening contexts. The diffusion signal modification parameter(s) may be obtained from the memory of the user device and may be set based on the user's preference regarding the degree to which the diffusion signal is distributed to other output channels for various listening contexts (e.g., for various types of audio content). In some aspects, each element of matrix B may be between 0 and 1, where a value of 0 indicates that the diffusion signal in a specified channel (e.g., left surround and right surround) is completely redistributed to other output channels (e.g., left, right, and / or center channels), and a value of 1 indicates that the diffusion signal is not redistributed at all.

[0068] In 604, process 600 can generate a modified dithering matrix using the dithering matrix and one or more diffusion signal modification parameters. Since the multi-channel output signal should not be attenuated when the diffusion signal is redistributed to other output channels, the modified dithering matrix can be determined using one or more normalization parameters. For example, the modified dithering matrix may be determined by the dot product of a normalization matrix, a matrix representing the diffusion signal modification parameter(s), and the dithering matrix. As an example, when the diffusion signal modification parameter(s) are composed of a plurality of diffusion signal modification parameters each corresponding to a specific output channel and a diffusion signal and are arranged in matrix B, the modified dithering matrix can be determined as follows:

Equation

[0069] In the above formula, the elements of the NORM matrix may be determined as values such that the Frobenius norm of O’ is the same as the Frobenius norm of O.

[0070] As another example, when the spread signal correction parameter consists of a single spread signal correction parameter β, the modified dithering matrix can be determined as follows:

Equation

[0071] In the above exemplary formula, the normalization matrix, the spread signal correction matrix, and the dithering matrix may each have an N×j dimension, where N represents the number of output channels in the upmix and j represents the number of spread signals. In some embodiments, the normalization matrix may be used so as not to change the Frobenius norm O of the dithering matrix. In the above formula, for the case of an example of five output channels (N) and two spread signals (j), the value of NORM used in the normalization matrix can be determined as follows:

Equation

[0072] In 606, process 600 can generate a spread multi-channel signal using a modified dithering matrix in which at least a portion of the spread stereo signal is distributed to the output channels. The modified spread multi-channel signal is generally represented herein as Z d1 ,… Z dN and these are conventionally generated by multiplying the vector formed by the spread signal by the dithering matrix. However, to generate the modified spread multi-channel signal, process 500 may multiply the vector formed by the spread signal by a modified dithering matrix incorporating the spread signal correction parameter. For example, in some embodiments, the spread multi-channel signal can be determined as follows:

Equation

[0073] As an example, when there are five output channels (N) and two spread signals (j), the spread multi-channel signal can be determined as follows:

Number

[0074] The spread multi-channel signal is then shown in FIGS. 3 and 4 and, as described above in connection with these figures, can be combined with the multi-channel direct signal to generate a multi-channel output signal in the frequency domain. The signal in the frequency domain may be converted to the time domain as described above in connection with FIGS. 3 and 4.

[0075] FIG. 7 is a block diagram showing an example of components of an apparatus capable of implementing various aspects of the present disclosure. Similar to the other figures provided herein, the types and numbers of elements shown in FIG. 7 are merely illustrative. Other aspects may include more, fewer, and / or different types and numbers of elements. According to some embodiments, apparatus 700 may be configured to perform at least some of the methods disclosed herein. In some aspects, apparatus 700 may be or include one or more components of a television, an audio system, a mobile device (such as a mobile phone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0076] According to some alternative aspects, apparatus 700 may be or include a server. In some such examples, apparatus 700 may be or include an encoder. Thus, in some cases, apparatus 700 may be a device configured to be used in an audio environment such as a home audio environment, and in other cases, apparatus 700 may be a device configured to be used in the "cloud," such as a server.

[0077] In this example, device 700 includes interface system 705 and control system 710. Interface system 705 may be configured to communicate with one or more other devices in an audio environment in some aspects. The audio environment may be a home audio environment in some examples. In other examples, the audio environment may be another type of environment such as an office environment, an automotive environment, a train environment, a road or sidewalk environment, a park environment, etc. Interface system 705 may be configured to exchange control information and related data with audio devices in the audio environment in some aspects. The control information and related data may belong to one or more software applications that device 700 is executing in some examples.

[0078] Interface system 705 may be configured to receive or provide a content stream in some aspects. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some cases, the audio data may include spatial data such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0079] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some aspects, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 705 may include one or more interfaces between the control system 710 and a memory system, such as the memory system 715 shown as an option in FIG. 7. However, the control system 710 may include a memory system in some cases. The interface system 705 may, in some aspects, be configured to receive input from one or more microphones in the environment.

[0080] The control system 710 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, and / or discrete hardware components.

[0081] In some embodiments, the control system 710 may be present in multiple devices. For example, in some embodiments, a portion of the control system 710 may be present in a device within one of the environments depicted herein, and another portion of the control system 710 may be present in a device external to the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer). In other examples, a portion of the control system 710 may be present in a device within one of the environments depicted herein, and another portion of the control system 710 may be present in one or more other devices of the environment. For example, a portion of the control system 710 may be present in a device implementing a cloud-based service, such as a server, and another portion of the control system 710 may be present in another device implementing a cloud-based service, such as another server or a memory device. The interface system 705 may also, in some examples, be present in multiple devices.

[0082] In some embodiments, the control system 710 may be configured to at least partially execute the methods disclosed herein. According to some examples, the control system 710 may be configured to implement a method of attenuating a spread signal, a method of distributing a spread signal to other output channels, and the like.

[0083] Some or all of the methods described in this specification may be performed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described in this specification. Such memory devices may include, but are not limited to, random access memory (RAM) devices, read only memory (ROM) devices, and the like. One or more non-transitory media may be present, for example, within the optional memory system 715 and / or control system 710 illustrated in FIG. 7. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented within one or more non-transitory media storing software. The software may, for example, determine or obtain diffusion signal attenuation parameters or generate an output multi-channel signal based on the diffusion signal attenuation parameters. The software may be executable by one or more components of a control system such as, for example, control system 710 of FIG. 7.

[0084] In some examples, the apparatus 700 can include a microphone system 720 as an option shown in FIG. 7. The optional microphone system 720 can include one or more microphones. In some aspects, one or more of the microphones may be part of or associated with another device such as a speaker of a speaker system, a smart audio device, or the like. In some examples, the apparatus 700 may not include the microphone system 720. However, in some such aspects, the apparatus 700 may nevertheless be configured to receive microphone data of one or more microphones in an audio environment via the interface system 710. In some such aspects, a cloud-based implementation of the apparatus 700 may be configured to receive microphone data or a noise metric at least partially corresponding to the microphone data from one or more microphones in an audio environment via the interface system 710.

[0085] According to some aspects, device 700 may include a loudspeaker system 725 as an option shown in FIG. 7. The optional loudspeaker system 725 may include one or more loudspeakers, also referred to herein as "speakers" or more generally "audio playback transducers". In some examples (e.g., cloud-based aspects), device 700 may not include a loudspeaker system 725. In some aspects, device 700 may include headphones. The headphones may be connected or coupled to device 700 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).

[0086] Some aspects of the present disclosure include a system or device configured (e.g., programmed) to execute one or more examples of the disclosed method and a tangible computer-readable medium (e.g., a disk) storing code for performing one or more examples of the disclosed method or its steps. For example, some of the disclosed systems are programmable general-purpose processors, digital signal processors, or microprocessors, or include them, programmed and / or otherwise configured by software or firmware to perform any of various operations on data, including embodiments of the disclosed method or its steps. Such a general-purpose processor is a computer system including, or capable of including, a processing subsystem programmed (and / or otherwise configured) to execute one or more examples of the disclosed method (or its steps) in response to an input device, memory, and data asserted thereto.

[0087] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed or otherwise configured) to perform the necessary processing on an audio signal(s) including execution of one or more examples of the disclosed method. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and memory) programmed and / or otherwise configured by software or firmware to perform any of the various operations including one or more examples of the disclosed method. Alternatively, elements of some embodiments of the system of the present invention may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed method, and the system may also include other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed method may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0088] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code for performing (e.g., executable code for) one or more examples of the disclosed method or steps thereof.

[0089] Although specific embodiments of the present disclosure and applications thereof have been described herein, it will be apparent to those skilled in the art that many variations to the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. While specific forms of the present disclosure have been shown and described, it should be understood that the present disclosure is not limited to the particular embodiments or particular methods described and illustrated.

Claims

1. A method for processing audio, comprising: Receiving a stereo audio signal; Separating the stereo audio signal into a steer signal and a diffusion signal, wherein the steer signal corresponds to the directional content in the stereo audio signal, and the diffusion signal corresponds to the background content in the stereo audio signal; Determining one or more diffusion signal correction parameters based on a current listening context, wherein the one or more diffusion signal correction parameters indicate a ratio of the diffusion signal to be redistributed to one or more output channels in an output multi-channel signal, or a degree of attenuation to be applied to the diffusion signal; Generating the output multi-channel signal based on the steer signal, the diffusion signal, and the one or more diffusion signal correction parameters; Providing the output multi-channel signal to a virtualizer for rendering as a binaural audio signal for playback on a wearable device. A method comprising the above steps.

2. The method according to claim 1, wherein the one or more output channels include at least one of a left channel, a right channel, or a center channel.

3. Generating the output multi-channel signal comprises: Obtaining a spreading matrix; Generating a modified spreading matrix using the spreading matrix and the one or more diffusion signal correction parameters; Generating a diffusion multi-channel signal using the modified spreading matrix; Generating the output multi-channel signal based on the diffusion multi-channel signal and the steer signal. The method according to claim 1, comprising the above steps.

4. The method according to claim 3, wherein the one or more diffusion signal correction parameters are for redistributing the ratio of the diffusion signal to the one or more output channels, and generating the modified spreading matrix comprises determining a matrix dot product of a norm related to the spreading matrix, a matrix representing the diffusion signal correction parameters, a matrix related to the one or more diffusion signal correction parameters, and the spreading matrix.

5. The one or more diffusion signal correction parameters include one diffusion signal redistribution correction parameter indicating redistribution of the diffusion signal in the multi-channel output, and the norm normalizes the energy of the diffusion signal, the method according to claim 4.

6. The one or more diffusion signal correction parameters cause the degree of attenuation applied to the diffusion signal, and generating the output multi-channel signal includes performing energy normalization configured such that the energy of the output multi-channel signal is the same as the energy of the stereo audio signal, the method according to claim 3.

7. The energy normalization is performed by either an upmixer that generates the output multi-channel signal or the virtualizer, the method according to claim 6.

8. The current listening context includes one of a movie content viewing mode, a music listening mode, or a game play mode, the method according to any one of claims 1 to 7.

9. The current listening context is the movie content viewing mode, and the one or more diffusion signal correction parameters are in the range of about 0.8 to 1, the method according to claim 8.

10. The current listening context is the music listening mode, and the one or more diffusion signal correction parameters are in the range of about 0 to 0.2, the method according to claim 8.

11. The current listening context is the game play mode, and the one or more diffusion signal correction parameters have a value smaller than the value associated with the movie content viewing mode, the method according to claim 8.

12. The one or more diffusion signal correction parameters are received from a user of the wearable device, the method according to any one of claims 1 to 11.

13. The one or more diffusion signal correction parameters are received via a user interface, the method according to claim 12.

14. Generating the output multi-channel signal is performed on a companion user device associated with the wearable device, and the virtualizer includes one or more components executed on the wearable device, the method according to any one of claims 1 to 13.

15. The method according to claim 14, further comprising transmitting data from the companion user device to the wearable device via a Bluetooth communication protocol.

16. The method according to any one of claims 1 to 15, wherein the wearable device includes one of earphones or headphones.

17. The method according to any one of claims 1 to 16, wherein the wearable device has one or more sensors that collect sensor data that can be used to generate head tracking information related to a wearer of the wearable device.

18. The method according to claim 17, wherein the virtualizer is configured to render the binaural audio signal based on the output multi-channel signal and the head tracking information.

19. One or more processors; A non-transitory computer-readable medium storing instructions that cause the one or more processors to perform the operations according to claims 1 to 18 when executed by the one or more processors; A system comprising

20. A non-transitory computer-readable medium storing instructions that cause one or more processors to perform the operations according to claims 1 to 18 when executed by the one or more processors.