Rotation of sound fields
By re-centering sound field orientations in the quaternion domain with smoothing filters and stability detection, the method addresses disorientation and gimbal lock issues in audio rendering, ensuring accurate alignment with user movements.
Patent Information
- Application Number
- PCT/US2025/028102
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-05-07
- Publication Date
- 2025-11-13
AI Technical Summary
Existing audio rendering techniques struggle to smoothly and accurately re-center sound field orientations in response to user head movements, particularly when using earbuds with tilt sensors, leading to disorientation and gimbal lock issues.
The method involves re-centering sound field orientations in the quaternion domain, using smoothing filters and stability detection to decouple yaw, pitch, and roll angles, and employing context-dependent reference frame updates based on user movement and device stability.
This approach ensures smooth and accurate audio rendering that aligns with user head movements, preventing gimbal lock and maintaining orientation coherence.
Smart Images

Figure US2025028102_13112025_PF_FP_ABST
Abstract
Description
ROTATION OF SOUND FIELDSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from PCT Application No. PCT / CN2024 / 092023 filed on 9 May 2024, U.S. Provisional Application No. 63 / 658,532, filed on 11 June 2024, and European Patent Application 24183180.9, filed June 19, 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] This disclosure pertains to systems, methods, and media for rotation of sound fields.BACKGROUND
[0003] Listeners of audio content may be interested in listening to immersive audio content in which the center direction of a virtual audio scene may change according to the user’s movement. For example, a listener, wearing headphone or earbuds, may move their head, which may cause changes in the position of audio objects in the virtual scene. However, it can be difficult to render such audio content smoothly and accurately.NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.SUMMARY
[0008] Methods, systems, and media for determining sound field directions are provided. In some embodiments, a method may involve receiving orientation data and acceleration data based on sensor data obtained from sensors disposed in a headset or earbuds worn by a user, wherein the orientation data is in a quaternion format. The method may further involve determining a center direction for a user orientation based at least in part on the orientation data. The method may further involve selecting a center direction for a sound field based at least in part on the center direction for the user orientation. The method may further involve re-centering a reference frame used to render audio content to be played back using the headset or earbuds in accordance with the selected center direction for the sound field, wherein the re-centering is performed in the quaternion domain.
[0009] In some examples, the sensor data represents user orientation in Euler angle format, and further comprising transforming the user orientation in Euler angle to the orientation data in the quaternion format.
[0010] In some examples, determining the center direction for the user orientation comprise smoothing the orientation data. In some examples, smoothing the orientation data comprises applying an exponential filter in the quaternion domain.
[0011] In some examples, the method may further involve determining whether the center direction for the user orientation is stable. In some examples, determining whether the center direction for the user orientation is stable may involve: determining a rotation angle between a first quaternion representing a current user orientation and a second quaternion representing a previous user orientation; and comparing the rotation angle to a threshold. In some examples, re-centering the reference frame is responsive to a determination that the center direction for the user orientation is stable. In some examples, the center direction for the sound field is based on the determination of whether the center direction for the user orientation is stable. In some examples, the center direction for the sound field is selected as the center direction for the user orientation responsive to a determination that the center direction for the user orientation is stable. In some examples, the center direction for the sound field is selected as a time-averaged user orientation direction responsive to a determination that the user orientation is not stable and the user is not walking. In some examples, the center direction for the sound field is selected as a user walking direction responsive to a determination that the user orientation is not stable and the user is walking.
[0012] In some examples, re-centering the reference frame comprises updating a center direction of the reference frame toward the selected center direction for the sound field. In some examples, re-centering the reference frame comprises determining a difference between a current center direction of the reference frame and the selected center direction for the sound field, and wherein a manner in which the center direction of the reference frame is updated is dependent on the difference. In some examples, the method may further involve, responsive to determining the difference exceeds a threshold, updating the center direction at a fixed angular rate. In some examples, the method may further involve, responsive to determining the difference is below a threshold, updating the center direction using exponential filtering.
[0013] In some examples, rendering the audio content is performed without considering roll angle, and selecting the center direction for the sound field may involve: transforming a current quaternion and a previous quaternion to Euler angles; and determining a yaw angle by accumulating yaw angles over a set of previous time samples.
[0014] In some examples, the reference frame corresponds to an orientation of a host device causing playback of the audio content, and wherein the center direction for the sound field is determined based on an orientation of the user determined based on the orientation data with respect to the orientation of the host device.
[0015] In some examples, the reference frame corresponds to the orientation of the user, and wherein the yaw angle of the center direction for the sound field is determined based on a difference in yaw angle between the orientation of the user and an orientation of host device causing the playback of the audio content.
[0016] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.
[0017] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0018] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 illustrates an example system for re-centering and / or rotating sound field orientations in accordance with some embodiments.
[0020] Figure 2 illustrates example graphs that include sound field orientations in yaw, pitch, and roll directions in accordance with some embodiments.
[0021] Figure 3A illustrates an example system for determining yaw in conjunction with a two- dimensional Tenderer device in accordance with some embodiments.
[0022] Figure 3B is a flowchart of an example process for selecting a yaw angle in accordance with the example system shown in Figure 3A.
[0023] Figure 4 illustrates example graphs of yaw, pitch, and roll directions utilizing the system shown in Figure 3 A in accordance with some embodiments.
[0024] Figures 5A and 5B illustrate example systems for performing sound field orientation re- centering on a user device and / or a host device in accordance with some embodiments.
[0025] Figure 6 is a flowchart of an example process for performing sound field orientation re- centering in accordance with some embodiments.
[0026] Figure 7 shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0027] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF EMBODIMENTS
[0028] Listeners of audio content may be interested in listening to immersive audio content. For example, in immersive audio content, the center direction of a virtual audio scene may change according to the user’s movement. For example, a listener, wearing headphone or earbuds, may move their head, which may cause changes in the position of audio objects in the virtual scene. However, it can be difficult to render such audio content smoothly and accurately. For example, conventional techniques may perform sound field orientation re-centering without updating a reference frame, which may result in limited re-centering in the pitch and roll directions. This may cause audio content rendering that does not match the user’s rotational head movements (e.g., in the yaw direction) after re-centering of the sound field orientation, particularly when using earbuds with a tilt sensor. As another example, conventional techniques may couple yaw, pitch, and roll output angles. This may cause, e.g., changes in an audio object’s location with respect to pitch and roll angles, even if the user is only moving their head in the yaw direction. As yet another example, conventional techniques may be susceptible to gimbal lock, in which flipping may occur (e.g., in the yaw direction) when the listener looks straight up or straight down (e.g., with a pitch angle at + / - 90 degrees). This may be disorienting for a user.
[0029] Disclosed herein are techniques for re-centering sound field orientations and / or reference frames. As used herein, a “reference frame” refers to the orientation of axes that center a sound field orientation. A reference frame may be with respect to a user’ s head such that the sound field orientation is aligned with up in the sound field orientation corresponding to an axis pointing out of the top of the user’s head. As another example, a reference frame may be with respect to world coordinates such that the sound field orientation is aligned with up in world coordinates. Using the techniques disclosed herein, re-centering and reference frame updating may be performed in the quaternion domain, which may allow re-centering to be performed smoothly and may inhibit gimbal lock. Moreover, re-centering may be performed in a manner that decouples yaw, pitch, and roll output angles, thereby causing the sound field orientation to be modified in accordance with the listener’s head movement. In some embodiments, re-centering may be performed in a manner that is context dependent such that the reference frame may be considered the user’s head, or the reference frame may be aligned with a user device the listener is using to play back audio content. By selecting the reference frame based on the listening context (e.g., whether the listener is watching the user device, whether the listener is in motion, etc.), re-centering of the sound field may occur in a manner aligned with the listener’s current activity.
[0030] FIG. 1 illustrates an example system for re-centering a sound field orientation in accordance with some embodiments. As illustrated, orientation information may be provided to a smoothing block 102. The orientation data may be obtained from one or more accelerometers and / or one or more gyroscopes associated with (e.g., disposed in) a pair of headphones or earbuds worn by a user. The input orientation may be in any suitable rotation format, e.g., quaternion, Euler angles, etc. However, note that, if the input orientation is not in a quaternion format, the orientation may be transformed to the quaternion format.
[0031] Smoothing block 102 may be configured to smooth the input orientation with respect to time, which will in turn allow a determined sound field center direction to be smooth with respect to time, which will in turn inhibit discontinuous jumps in the sound field orientation. Smoothing block 102 may be configured to perform an exponential filter to smooth input orientation samples.
[0032] Given an input orientation represented in quaternion format as q(k), where k is the current sample, the input quaternion format can be updated based on the previous input orientation by:
[0033] Note that the above equation accounts for negative and positive values having the same meaning in quaternion format. Additionally, note that q(k) ■ q(k-l ) represents the dot product of two quaternions, and sign() represents the signum function represented by:
[0034] The klhsample of the input orientation may be represented by: q( / c) = aq(k - 1) + (1 - a)qupdate(k)
[0035] In the equation given above, a is a smoothing factor, where T is a smoothing time constant. The smoothing time constant r may have a value within a range of about 5 seconds to 20 seconds. For example, in some implementations, r may be 10 seconds. In this way, the input orientation may be smoothed using the smoothing time constant over time, which may lessen discontinuous jumps in input orientation with respect to successive time samples.
[0036] The smoothed input orientation may be provided to stable detector block 104. Stable detector block 104 may be configured to utilize the smoothed input orientation to determine whether the listener is currently in a stable orientation, or, whether the listener is currently in movement. In some embodiments, the listener may be considered to be in a stable orientation if the difference (e.g., the angular difference) between two successive user orientations (sometimes referred to herein as “poses”) is less than a stability threshold. Given a current user orientation represented in quaternion format as q2 and a previous user orientation represented in quaternion format as qi, the angle between the rotation angle between the two quaternions may be determined, which is generally represented herein as On . The listener may be determined to be stable if 6new is less than the stability threshold, and may determined to be not stable (i.e., in movement) if 6„ew is greater than or equal to the stability threshold. Example values of the stability threshold are 2 degrees, 5 degrees, 7 degrees, 10 degrees, etc.0 0
[0037] By way of example, a quaternion q may be represented as q = [cos-, sin - ■ u]. The rotation between the current quaternion q2 and the previous quaternion qi may be represented as q„ew, and may be determined by:In which u is a unit vector [x, y, z].
[0038] The first component of qnewmay be represented by: new(O) = <7i(0) • Q2(0) + ^(1) • q2(l) + ^(2) • <?2(2) + ^(3) • q2(3)
[0039] The rotation angle between the two quaternions, 0new, may be determined by:
[0040] The determination of stability may be provided as input to selector block 106. Separately, as indicated in FIG. 1, acceleration data (which may be obtained from one or more accelerometers disposed in the headphones and / or earbuds) may be provided to walking detector block 110, which may be configured to determine whether or not the listener is currently walking or running. In some implementations, walking detector block 110 may additionally determine a direction the listener is walking or running. In some embodiments, walking detector block 110 may determine whether the listener is currently walking or running, and / or the direction of movement, based on a direction of acceleration orthogonal to a movement direction in which at least a portion of the user acceleration data is anti-periodic over a period of time. In other words, in some embodiments, walking detector block 110 may utilize periodicities and anti-periodicities in the acceleration data to determine whether or not the listener is walking or running (e.g., traveling in a linear straight forward manner), and, optionally, the direction of movement. The determination of whether the listener is walking or running, and optionally, the movement direction, may be provided to selector block 106.
[0041] Selector block 106 may be configured to select the target center direction for the sound field orientation. In other words, selector block 106 may be configured to select the center direction of the reference frame associated with the sound field orientation to which the reference frame and therefore the sound field orientation will be rotated.
[0042] In some embodiments, selector block 106 may select the target center direction based on whether the listener is in motion, and, in some embodiments, based on the type of motion. For example, in a case in which the listener is not in motion (e.g., is stable), the center direction may be the user-facing or listener-facing direction, which may be determined based at least in part on the orientation data. As another example, in an instance in which the listening is in motion but is not walking or running (e.g., the listening is moving their head to look around, or is otherwise moving in a non-linear or non-forward direction manner), the center direction may be selected asa short-time averaged or time-smoothed direction over recent orientation samples. As yet another example, in an instance in which walking detector block 110 indicates that the listener is walking or running, the center direction may be the direction of movement.
[0043] Note that selector block 106 may represent the target center direction in quaternion format. In instances in which the target center direction corresponds to a walking or running direction, the walking or running direction may be represented as an Euler angle, which may then be transformed to quaternion format.
[0044] In some embodiments, a determination may be made of whether the listener is in a “captured” state or a “non-captured” state. As used herein, a captured state refers to the difference between the current center direction and the target center direction being less than a threshold difference. A non-captured state refers to a difference between the current center direction and the target center direction being greater than or equal to the threshold difference. In some embodiments, the threshold difference may be, e.g., 0.5 degrees, 2 degrees, 5 degrees, 10 degrees, 15 degrees, 20 degrees, or the like.
[0045] The target center direction may be provided to re-centering block 108. Re-centering block 108 may perform any suitable operations that cause the sound field orientation to be rotated such that the center direction (e.g., the forward-facing direction) associated with the sound field orientation is rotated to be aligned with the target center direction. Accordingly, the reference fame associated with rendered audio content may be re-centered to align with the target center direction.
[0046] Note that rotation may be performed based on whether the listener is in a captured state or a non-captured state.
[0047] In a captured state, rotation may be performed using exponential filtering. In the noncaptured state, rotation may be performed based on a fixed angular rate, which may be represented herein as r. By way of example, the movement between the current center direction (represented herein as q^d) and the target center direction qtarget may be represented by the quaternion qmme, which may be represented by:
[0048] The rotation movement may be represented by a rotation vector, where 0 represents the angle of rotation of the sound field orientation:
[0049] In the equation given above, [x, y, z] represents the unit vector. The angle 0 may be split into n steps such that 0 = n * 0inc, where 0mcrepresents the angle of rotation over one incremental rotation step. After one incremental rotation step, the quaternion representation of the sound field orientation may be determined by:Qinc = [x, y, z] * 9 inc
[0050] The center direction of the sound field orientation in quaternion format may be represented by: fwd 0fwd * inc
[0051] Given a re-centering angular rate of r, the time to perform re-centering may be represented by t = -. Note that rotation is performed in the quaternion domain. The rotation vector [x, y, z] is used to perform segmentation in the quaternion domain, where segmentation is performed using 9inc.
[0052] FIG. 2 illustrates an example graph of yaw, pitch, and roll angles responsive to various user head movements in accordance with the re-centering techniques described above in connection with FIG. 1. As illustrated, during time window 202, stability of the user may be detected, e.g., based on yaw, pitch, and roll angles remaining within a predetermined range and / or changing less than a predetermined threshold. Responsive to detecting user stability, in time window 204, re-centering of the sound field orientation may occur. Note that yaw and pitch angles may change as part of the re-centering process. Additionally, note that the state during time window 204 is non-captured, until the center of the sound field orientation is within a predetermined angle of the target direction. The state changes from captured during time window 202 to non-captured during time window 204 due to a new target center direction that exceeds a threshold difference to the current center direction. During time window 206, the state changes to captured. Note that the re-centering process still occurs via exponential filtering during time window 206. During time window 208, the reference frame is re-centered, and yaw, pitch, and roll angles all change because the orientations depicted in FIG. 2 are in world coordinates and the reference frame has been updated to a new position during time window 206 During time window 210, the yaw, pitch, and roll angles do not change. During time window 212, the state changes tonon-captured, and re-centering of the sound field orientation occurs again. Note that re-centering occurs in time window 212 and is triggered due to stability of the user detected in time window 210, and identification of a new target center direction.
[0053] When utilizing a two-dimensional tenderer that renders spatial audio content with regard to two dimensions, typically, yaw angle and pitch angle are used to perform the rendering while roll angle is ignored. In such cases, if the listener rotates or moves their head around some combination of axes (rather than just one axis) and / / or the pitch angle is close to + / - 90 degrees (e.g., the listener is looking straight up or straight down), unwanted flipping may occur in the yaw direction. This unwanted flipping occurs due to gimbal lock, which occurs in the transformation from quaternions to Euler angles.
[0054] Disclosed herein are techniques to avoid unwanted flipping due to gimbal lock. In particular, rather than directly determining yaw angle from the quaternion domain after re- centering the sound field orientation, the yaw angle is determined based on an accumulation of yaw angles between successive time samples. This technique may introduce a bias compared to directly determining the yaw angle, particularly when the listener rotates their head around a single axis. To compensate for this bias, the yaw angle may be adjusted during re-centering toward the target yaw angle, which may be performed in a manner which is imperceptible to the listener.
[0055] FIG. 3A illustrates example components that may be used to determine yaw angle by accumulating successive yaw angles in accordance with some embodiments. As illustrated, transformation and calculation block 302 is configured to take, as input, a previous quaternion (e.g., associated with a previous time sample k-1) and a current quaternion (e.g., associated with a current time sample k). Transformation and calculation block 302 may transform the current quaternion and the previous quaternion to Euler angles, and may calculate the current yaw angle using intrinsic accumulation, e.g., by accumulative successive yaw angles. Transformation and calculation block 302 may be configured to output the transformed current yaw, the transformed previous yaw, and the accumulated yaw.
[0056] By way of example, the rotation between the klhsample and the k-lthsample may be determined by:
[0057] The yaw angle between the kth sample and the k-l,hsample may be determined in quaternion domain using the equation given above. The final yaw angle for the k,hsample may be determined by accumulating yaw over previous samples. For example, the final yaw angle at the k111sample may be determined by:
[0058] Note that, in the quaternion domain, the 0thsample refers to the reference frame, whereas in Euler angles, the 0thsample refers to the initial sample at which the start of yaw accumulation begins.
[0059] The transformed current yaw, the transformed previous yaw, and the accumulated yaw may be provided to selector block 306. Selector block 306 may additionally receive input from stable detector block 304. Note that stable detector block 304 may operate similarly to stable detector block 104 as shown in and described above in connection with FIG. 1, and may be configured to output an indication of whether or not the listener is stable. Selector block 306 may then output a yaw angle based on the accumulated yaw angle and whether or not the listener is stable. Selector block 306 may additionally perform exponential filtering to compensate for bias due to the accumulation of yaw angles.
[0060] FIG. 3B is a flowchart of an example process 350 for selecting a yaw angle in accordance with some embodiments of the disclosed subject matter. For example, process 350 may be performed by transformation and calculation block 302 and / or selector block 306 as shown in and described above in connection with FIG. 3A. In some implementations, blocks of process 350 may be executed by one or more processors and / or control systems of one or more devices (e.g., a user device such as a mobile phone, processors of headphones or earbuds, etc.). An example of such as a control system is control system 710 of FIG. 7. In some embodiments, blocks of process 350 may be performed in an order other than what is shown in FIG. 3B. In some embodiments, two or more blocks of process 350 may be executed substantially in parallel. In some embodiments, one or more blocks of process 350 may be omitted.
[0061] Process 350 can begin at 352 by receiving the previous quaternion and the current quaternion indicating user orientation. For example, the previous quaternion may be associated with a previous time sample k-1, and the current quaternion may be associated with the current time sample k.
[0062] At 354, process 300 can determine the previous yaw angle (e.g., transformed from the previous quaternion), the current accumulated yaw angle (e.g., accumulated based on previous successive samples), and the transformed current yaw angle (e.g., transformed from the current quaternion). The transformed previous and current yaw angles may be determined by transforming the corresponding quaternion to Euler angles to determine the previous and current yaw angles.
[0063] At 356, process 300 can determine whether the user is stable. Stability may be determined by comparing the previous yaw angle to the current yaw angle. For example, if the difference is less than a predetermined stability threshold (e.g., 2 degrees, 5 degrees, 10 degrees, etc.) the user may be determined to be stable.
[0064] If, at 356, process 350 determines that the user is not stable (“no” at 356), process 350 can select the current accumulated yaw angle at 358. For example, the current accumulated yaw angle may be the yaw angle accumulated over the previous k samples. Process 350 may then end.
[0065] Conversely, if, at 356, process 350 determines that the user is stable (“yes” at 356), process 350 can proceed to 360 and can determine if the product of the previous yaw angle and the accumulated yaw is less than 0. If, at 360, process 350 determines that the product of the previous yaw angle and the accumulated yaw is less than 0 (“yes” at 360), process 350 can proceed to 362 and can select the previous yaw angle. Note that the product of the previous yaw and the accumulated yaw indicates whether the previous yaw and the accumulated yaw have the same sign, which indicates whether the previous yaw and the accumulated yaw are in the same direction. For example, if the product is less than 0 (e.g., the product is negative), the previous yaw and the accumulated yaw are in different directions (e.g., one is left and the other is right). In such cases, the previous yaw angle is selected so that the re-centering direction is not switched.
[0066] Conversely, if at 360, process 350 determines that the product of the previous yaw angle and the accumulated yaw is greater than or equal to 0 (“no” at 360), process 350 can proceed to 364 and can determine whether to smooth the accumulated yaw angles. In some embodiments, process 350 can determine that the accumulated yaw angles are to be smoothed responsive to determining that an absolute value of the accumulated yaw is greater than the absolute value of the previous yaw. In other words, process 350 can automatically determine how to compensate bias introduced by the accumulation. If, at 364, process 350 determines that the accumulated yaw angles are to be smoothed (“yes” at 364), process 350 can proceed to 366 and can provide the smoothed accumulated yaw angles. For example, the accumulated yaw angles may be smoothed using an exponential filter.
[0067] Conversely, if, at 364, process 350 determines that the accumulated yaw angles are not to be smoothed (“no” at 364), process 350 can proceed to 368 and can provide the raw accumulated yaw angle, e.g., without performing any smoothing.
[0068] FIG. 4 illustrates an example graph of yaw, pitch, and roll angles using a conventional technique compared to the technique shown in and described above in connection with FIG. 3A that utilizes yaw angle accumulation. As illustrated, during time window 402, flipping in the yaw direction occurs using the conventional technique due to gimbal lock. In particular, as illustrated by the pitch and roll angles, the listener is rotating their head in a manner that does not correspond to rotation around a single x ,y or z axis, which causes gimbal lock using the conventional techniques. However, as illustrated during time window 402, using the techniques described above which calculate yaw angle using intrinsic accumulation over successive samples, the yaw angle does not flip. Bias compensation is applied during time window 404 to the yaw angle, which causes the yaw angle to be brought towards the target yaw angle using exponential filtering.
[0069] In some embodiments, re-centering and / or modifying a sound field orientation may be performed based on user head orientation, user device position / orientation, and / or a combination of user head orientation and user device position / orientation. As used herein, the user device may correspond to the user device that renders and plays back audio content and / or video content, such as a mobile phone, a tablet computer, a vehicle computer, a smart television, etc. Note that the user device is sometimes referred to herein as “a host device.” In some embodiments, the user’s headphones and / or earbuds may be paired with the user device, e.g., via a BLUETOOTH connection or other short-range wireless communication channel.
[0070] In some embodiments, whether user head orientation and / or user device orientation are used to perform re-centering and / or modification of the sound field orientation may depend on the listening context. For example, in an instance in which the listener is watching video content on a user device (e.g., a mobile phone, a tablet computer, etc.), the listener may want the user device position to correspond to the reference frame such that the sound field orientation is modified with the user device positioned at the center of the reference frame. In such instances, re-centering may be performed based on the listener’s head orientation, e.g., using sensors of the headphones and / or earbuds. Continuing with this example, position and / or orientation information obtained from sensors of the user device may be used to determine whether or not the user device is stable. If the user device is stable (e.g., in instances in which the user device is not in motion, such as when the user device is sitting on a table), re-centering and / or modification of the sound field orientationmay be performed based only on the listener’s head orientation. Conversely, if it is determined that the user device is in motion, a “combined pose” may be determined which corresponds to the user’s orientation with respect to the reference frame defined by the user device. In some embodiments, the combined pose may be determined by multiplying, in the quaternion domain, the listener’s head orientation represented in the quaternion domain by the user device orientation represented in the quaternion domain. Re-centering and / or sound field orientation may then be performed based on the combined pose.
[0071] FIG. 5A illustrates an example system that utilizes the user device as the center of the reference frame in accordance with some embodiments. The system shown in FIG. 5A may be used in instances in which the reference frame is to be specified by the position and / or orientation of the user device, for example, in instances in which the listener is watching video content presented on the user device. As illustrated, user orientation information (specified as user quaternion information) and user acceleration information may be provided to re-centering block 502. The user orientation information and the user acceleration information may be obtained using sensors disposed in headphones and / or earbuds worn by the user. An example implementation of a re-centering block is shown in and described above in connection with FIG. 1. Re-centering block 502 may be configured to determine an updated sound field orientation based on the user’s orientation and / or acceleration (which may be used to determine whether or not the listener is static, or is walking / running).
[0072] The output of re-centering block 502 may be provided to combine pose block 504. Combine pose block 504 may be configured to receive, as input, both the output of re-centering block 502 and device orientation information, which is represented in quaternion format. The device orientation information may be obtained from one or more sensors (e.g., one or more accelerometers and / or gyroscopes) of the user device. Combine pose block 504 may be configured to determine whether or not the user device is static or is in motion. If the user device is static, combine pose block 502 may simply pass the user orientation to re-centering block 506. Conversely, if the user device is not static (e.g., the user is moving their mobile device around), combine pose block 502 may determine the user’s orientation with respect to the user device’s reference frame. The combined pose may be determined by multiplying, in the quaternion domain, the user quaternion by the device quaternion.
[0073] Re-centering block 506 may receive, as input, the combined pose and acceleration information associated with the user device, and may perform re-centering based on the combinedpose and the acceleration information. Note that an example implementation of re-centering block 506 is shown in and described above in connection with FIG. 1 , however, in instances in which re- centering is performed with respect to the user device being the reference frame (as in the example shown in FIG. 5A), the target center direction may be specified with respect to the user device, and the combined pose may be the input orientation (in the quaternion domain) with which re- centering is performed. Note that, with respect to FIG. 1 , acceleration data may be obtained from the user device.
[0074] In some embodiments (e.g., in some listening contexts), it may be desirable for the user’s head to correspond to the reference frame regardless of device orientation or position. For example, such a listening context may be an instance in which the user device is in an unwanted position or has erratic movement, like in a pocket. In such cases, the sound field orientation may be modified with respect to a reference frame centered at the user’s head. In some embodiments, in instances in which the user’s head orientation is to correspond to the reference frame, re- centering may be performed based on the user’s head orientation and / or acceleration, where re- centering may involve determining updated pitch and roll angles for the sound field orientation. However, the yaw angle may be determined based on both the orientation and / or acceleration of the user device and the yaw angle of user’s head orientation. In some embodiments, determination of yaw angle may be based on the accumulated yaw angle technique shown in and described above in connection with FIGS. 3 A and 3B, which may prevent flipping in the yaw direction by avoiding gimbal lock.
[0075] FIG. 5B illustrates an example system for utilizing a reference frame that corresponds to a user head orientation in accordance with some embodiments. As illustrated, a re-centering block 552 may receive user head orientation information (represented as the user quaternion) and user acceleration information. The user head orientation and user acceleration information may be obtained from one or more sensors disposed in or on headphones or earbuds worn by the user. Re- centering block 552 may be configured to determine updated pitch and roll angles to re-center the sound field. An example implementation of re-centering block 552 is shown in and described above in connection with FIG. 1. Note that re-centering block 552 may determine yaw angles, as well, which may be provided to device orientation block 554, and may output the updated pitch and roll angles directly.
[0076] Device orientation block 554 may be configured to receive device orientation information (represented as the device quaternion) and device acceleration information. Thedevice orientation information and the device acceleration information may be obtained from one or more sensors disposed in the user device, e.g., one or more gyroscopes and / or one or more accelerometers. If the device is determined to be static (e.g., based on the device orientation information and / or the device accelerometer information), the yaw angle may be determined to be the previous yaw angle. However, if the device is determined to not be static (e.g., the user device is in motion), device orientation block 554 may be configured to determine the combined yaw angle. The combined yaw angle may be determined by transforming the user quaternion and the device quaternion to Euler angles, and then determining a difference between the user yaw angle and the device yaw angle. Re-centering may then be performed based on the pitch and roll angles determined based on the user head orientation and based on the combined yaw that accounts for a difference in yaw angle between the user’s head and the device’s orientation.
[0077] FIG. 6 is a flowchart of an example process 600 for performing sound field orientation modification and re-centering in accordance with some embodiments. In some embodiments, blocks of process 600 may be performed using one or more processors and / or one or more control systems associated with user headphones or earbuds and / or with a user device used to play audio content. An example of such a control system is control system 710 of FIG. 7. In some embodiments, blocks of process 600 may be performed in an order other than what is shown in FIG. 6. In some embodiments, two or more blocks of process 600 may be executed substantially in parallel. In some embodiments, one or more blocks of process 600 may be omitted.
[0078] Process 600 can begin at 602 by receiving orientation data and acceleration data based on sensor data obtained from sensors disposed in a headset or earbuds worn by a user. The orientation data may be in a quaternion format. In instances in which the obtained orientation data is not in a quaternion format (e.g., in instances in which the orientation data is represented in Euler angles), the orientation data may be transformed to quaternion format. The one or more sensors may include one or more accelerometers and / or one or more gyroscopes disposed in or on the headphones or earbuds.
[0079] At 604, process 600 can determine a center direction for a user orientation based at least in part on the orientation data. For example, as shown in and described above in connection with FIG. 1, in some embodiments, the center direction for the user orientation may be determined by smoothing user orientation data with respect to time. Such smoothing may be implemented using an exponential filter.
[0080] At 606, process 600 can select a center direction for a sound field based at least in part on the center direction for the user orientation. For example, as shown in and described above in connection with FIG. 1, in some embodiments, the center direction for the sound field may be determined based at least in part on whether the user is determined to be stable in orientation and / or position, or whether the user is currently in motion. As another example, the center direction for the sound field may be determined based at least in part on whether the user is moving in a linear forward manner (e.g., walking / running), or whether the user is moving in a non-walking or nonrunning manner (e.g., whether the user is looking around). As a more particular example, in an instance in which the user is not moving, the center direction may correspond to the forward direction aligned with user head orientation (e.g., the direction the listener is looking). In an instance in which the listener is walking, the center direction may correspond to the direction the listener is walking. In an instance in which the user is moving but not in a linear forward manner, the center direction for the sound field orientation may correspond to a time-averaged direction of the user’ s head movement. Note that the selected center direction may correspond to a target center direction of the sound field. Additionally, note that the center direction may be quaternion. In some embodiments, the yaw, pitch, and roll angles may be determined with respect to the user device corresponding to the reference frame (e.g., as shown in and described above in connection with FIG. 5A). Alternatively, in some embodiments, the pitch and roll angles may be determined based on the user’s head orientation whereas the yaw angle may he determined based on the difference between the user’s head orientation and the device orientation allowing the user’s head to correspond to the reference frame.
[0081] At 608, process 600 can re-center a reference frame used to render audio content to be played back using the headset or earbuds in accordance with the selected center direction for the sound field. Re-centering may be performed in the quaternion domain. For example, in some embodiments, process 600 may modify the orientation of the sound field to align with the target center direction determined at 606. In some embodiments, the orientation of the sound field may be modified with a fixed angular rate, e.g., in instances in which the difference between the current orientation and the target orientation exceeds a predetermined threshold (referred to herein as a “non-captured state”). Alternatively, in some embodiments, the orientation may be modified using exponential filtering, e.g., in instances in which the difference between the current orientation and the target orientation is less than the predetermined threshold (referred to herein as the “captured state”).
[0082] Figure 7 is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 7 are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 700 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 700 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0083] According to some alternative implementations the apparatus 700 may be, or may include, a server. In some such examples, the apparatus 700 may be, or may include, an encoder. Accordingly, in some instances the apparatus 700 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 700 may be a device that is configured for use in “the cloud,” e.g., a server.
[0084] In this example, the apparatus 700 includes an interface system 705 and a control system 710. The interface system 705 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 705 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 700 is executing.
[0085] The interface system 705 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0086] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, suchas one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 705 may include one or more interfaces between the control system 710 and a memory system, such as the optional memory system 715 shown in Figure 7. However, the control system 710 may include a memory system in some instances. The interface system 705 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0087] The control system 710 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0088] In some implementations, the control system 710 may reside in more than one device. For example, in some implementations a portion of the control system 710 may reside in a device within one of the environments depicted herein and another portion of the control system 710 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 710 may reside in a device within one environment and another portion of the control system 710 may reside in one or more other devices of the environment. For example, a portion of the control system 710 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 710 may reside in another device that is implementing the cloudbased service, such as another server, a memory device, etc. The interface system 705 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0089] In some implementations, the control system 710 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 710 may be configured for re-centering reference frames, rending audio content, playing back audio content, or the like.
[0090] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 715 shown in Figure 7 and / or in the control system 710. Accordingly, various innovative aspects ofthe subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, be executable by one or more components of a control system such as the control system 710 of Figure 7.
[0091] In some examples, the apparatus 700 may include the optional microphone system 920 shown in Figure 7. The optional microphone system 720 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 700 may not include a microphone system 720. However, in some such implementations the apparatus 700 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 710. In some such implementations, a cloud-based implementation of the apparatus 700 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 710.
[0092] According to some implementations, the apparatus 700 may include the optional loudspeaker system 725 shown in Figure 7. The optional loudspeaker system 725 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 700 may not include a loudspeaker system 725. In some implementations, the apparatus 700 may include headphones. Headphones may be connected or coupled to the apparatus 700 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0093] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0094] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0095] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.
[0096] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.Various aspects of the present invention may be appreciated from the following enumerated example embodiments (EEEs):1. A method for determining sound field directions, the method comprising: receiving orientation data and acceleration data based on sensor data obtained from sensors disposed in a headset or earbuds worn by a user, wherein the orientation data is in a quaternion format; determining a center direction for a user orientation based at least in part on the orientation data;selecting a center direction for a sound field based at least in part on the center direction for the user orientation; and re-centering a reference frame used to render audio content to be played back using the headset or earbuds in accordance with the selected center direction for the sound field, wherein the re-centering is performed in the quaternion domain.2 . The method of EEE 1 , wherein the sensor data represents user orientation in Euler angle format, and further comprising transforming the user orientation in Euler angle to the orientation data in the quaternion format.3. The method of EEE 1 or 2, wherein determining the center direction for the user orientation comprise smoothing the orientation data.4. The method of any one of EEE 1-3, further comprising determining whether the center direction for the user orientation is stable.5. The method of EEE 4, wherein determining whether the center direction for the user orientation is stable comprises: determining a rotation angle between a first quaternion representing a current user orientation and a second quaternion representing a previous user orientation; and comparing the rotation angle to a threshold.6. The method of EEE 4 or 5, wherein re-centering the reference frame is responsive to a determination that the center direction for the user orientation is stable.7. The method of any one of EEE 4-6, wherein the center direction for the sound field is based on the determination of whether the center direction for the user orientation is stable.8. The method of EEE 7, wherein the center direction for the sound field is selected as the center direction for the user orientation responsive to a determination that the center direction for the user orientation is stable.9. The method of EEE 7, wherein the center direction for the sound field is selected as a time- averaged user orientation direction responsive to a determination that the user orientation is not stable and the user is not walking.10. The method of EEE 7, wherein the center direction for the sound field is selected as a user walking direction responsive to a determination that the user orientation is not stable and the user is walking.11. The method of any one of EEE 1-10, wherein re-centering the reference frame comprises updating a center direction of the reference frame toward the selected center direction for the sound field.12. The method of EEE 11, wherein re-centering the reference frame comprises determining a difference between a current center direction of the reference frame and the selected center direction for the sound field, and wherein a manner in which the center direction of the reference frame is updated is dependent on the difference.13. The method of any one of EEE 1-12, wherein rendering the audio content is to be performed without considering roll angle, and wherein selecting the center direction for the sound field comprises: transforming a current quaternion and a previous quaternion to Euler angles; and determining a yaw angle by accumulating yaw angles over a set of previous time samples.14. The method of any one of EEE 1-13, wherein the reference frame corresponds to an orientation of a host device causing playback of the audio content, and wherein the center direction for the sound field is determined based on an orientation of the user determined based on the orientation data with respect to the orientation of the host device.15. The method of any one of EEE 1-13, wherein the reference frame corresponds to the orientation of the user, and wherein the yaw angle of the center direction for the sound field is determined based on a difference in yaw angle between the orientation of the user and an orientation of host device causing the playback of the audio content.
Claims
CLAIMS1. A method for determining sound field directions, the method comprising: receiving orientation data and acceleration data based on sensor data obtained from sensors disposed in a headset or earbuds worn by a user, wherein the orientation data is in a quaternion format; determining a center direction for a user orientation based at least in part on the orientation data; selecting a center direction for a sound field based at least in part on the center direction for the user orientation; and re-centering a reference frame used to render audio content to be played back using the headset or earbuds in accordance with the selected center direction for the sound field, wherein the re-centering is performed in the quaternion domain, wherein the reference frame corresponds to an orientation of a host device causing playback of the audio content, and wherein the center direction for the sound field is determined based on an orientation of the user determined based on the orientation data with respect to the orientation of the host device.
2. The method of claim 1 , wherein the sensor data represents user orientation in Euler angle format, and further comprising transforming the user orientation in Euler angle to the orientation data in the quaternion format.
3. The method of claim 1 or 2, wherein determining the center direction for the user orientation comprise smoothing the orientation data.
4. The method of any one of claims 1-3, further comprising determining whether the center direction for the user orientation is stable.
5. The method of claim 4, wherein determining whether the center direction for the user orientation is stable comprises: determining a rotation angle between a first quaternion representing a current user orientation and a second quaternion representing a previous user orientation; and comparing the rotation angle to a threshold.
6. The method of claim 4 or 5, wherein re-centering the reference frame is responsive to a determination that the center direction for the user orientation is stable.
7. The method of any one of claims 4-6, wherein the center direction for the sound field is based on the determination of whether the center direction for the user orientation is stable.
8. The method of claim 7, wherein the center direction for the sound field is selected as the center direction for the user orientation responsive to a determination that the center direction for the user orientation is stable.
9. The method of claim 7, wherein the center direction for the sound field is selected as a time-averaged user orientation direction responsive to a determination that the user orientation is not stable and the user is not walking.
10. The method of claim 7, wherein the center direction for the sound field is selected as a user walking direction responsive to a determination that the user orientation is not stable and the user is walking.
11. The method of any one of claims 1-10, wherein re-centering the reference frame comprises updating a center direction of the reference frame toward the selected center direction for the sound field.
12. The method of claim 11, wherein re-centering the reference frame comprises determining a difference between a current center direction of the reference frame and the selected center direction for the sound field, and wherein a manner in which the center direction of the reference frame is updated is dependent on the difference.
13. The method of any one of claims 1-12, further comprising rendering the audio content without considering roll angle, and wherein selecting the center direction for the sound field comprises: transforming a current quaternion and a previous quaternion to Euler angles; and14. determining a yaw angle by accumulating yaw angles over a set of previous time samples..
15. The method of any one of claims 1-13, wherein the reference frame corresponds to the orientation of the user, and wherein the yaw angle of the center direction for the sound field is determined based on a difference in yaw angle between the orientation of the user and an orientation of host device causing the playback of the audio content.
Citation Information
Patent Citations
Inertially stable virtual auditory space for spatial audio applications
US20210400420A1
Adaptive Audio Centering for Head Tracking in Spatial Audio Applications
US20220103965A1
Sensor data prediction
US20240147180A1
Inertial sensor fusion orientation correction
US9068843B1
Sound field rotation
WO2023146909A1