Spatialized audio using dynamic head tracking
Headphones with dynamic head tracking adjust virtual sound sources to maintain spatialized audio during frequent head turns, ensuring a consistent and comfortable experience by aligning audio frames with the user's head orientation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2026-04-02
AI Technical Summary
Current headphones fail to provide a consistently comfortable spatialized audio experience for users engaged in activities involving frequent head turns, such as walking or running, as they often resort to a 'fixed mode' that destroys the auditory illusion of spatialized audio.
Headphones equipped with sensors and controllers that dynamically track the user's head movements, adjusting the virtual sound stage to maintain spatialized audio by rotating virtual sound sources relative to the user's head orientation, ensuring a consistent and comfortable experience during activities that require frequent head turns.
The solution maintains the illusion of spatialized audio by dynamically tracking the user's head, providing a consistent and comfortable spatialized audio experience even during activities that involve frequent head movements.
Smart Images

Figure 2026510309000001_ABST
Abstract
Description
[Technical Field]
[0001] (Cross-reference of related applications) This application claims priority to U.S. Nonprovisional Patent Application No. 18 / 182,290, filed on March 10, 2023, entitled “Spatialized Audio With Dynamic Head Tracking,” which is incorporated herein by reference in its entirety. [Background technology]
[0002] This disclosure relates, in general terms, to a system and method for providing spatialized audio using dynamic head tracking. [Overview of the project] [Means for solving the problem]
[0003] All the examples and features mentioned below can be combined in any technically feasible way.
[0004] In one embodiment, a pair of headphones comprises a sensor that outputs a sensor signal representing the orientation of the user's head, and a controller that receives the sensor signal, the controller programmed to output a spatialized audio signal to a pair of electroacoustic transducers for conversion to a spatialized acoustic signal, the spatialized acoustic signal being perceived by the user as originating from a virtual sound stage containing at least one virtual sound source, each virtual sound source in the virtual sound stage being positioned at a different location from the electroacoustic transducers, and being perceived relative to an audio frame of the virtual sound stage, the audio frame being aligned with the reference axis of the user's head. The system includes a controller positioned in a first location, the controller being further programmed to determine from the sensor signal whether the characteristics of the user's head satisfy at least one predetermined condition, the at least one predetermined condition including whether the orientation of the user's head is outside a predetermined angular boundary, and if the controller determines that the characteristics of the user's head do not satisfy at least one predetermined condition, the controller is programmed to maintain the audio frame in the first position, and if the controller determines that the orientation of the user's head is outside a predetermined angular boundary, the controller is programmed to rotate the position of the audio frame around a rotation axis to reduce the angular offset with respect to the reference axis of the user's head.
[0005] For example, rotating the position of an audio frame involves rotating the audio frame so that it aligns with the reference axis of the user's head when the user's head turns.
[0006] In one example, while the user's head is experiencing increasing angular acceleration, the angular velocity of the audio frame's rotation is at least partially based on the user's head's angular velocity.
[0007] In one example, while the user's head is experiencing decreasing angular acceleration, the angular velocity of the audio frame's rotation is selected so that the audio frame aligns with a predicted position on the user's head's reference axis when the user's head turns.
[0008] In one example, the predicted position of the user's head's reference axis when the user's head turns is updated with each sample that has decreasing angular acceleration.
[0009] In one example, at least one predetermined condition further includes whether the angle jerk of the user's head exceeds a predetermined threshold, and at least one predetermined condition is satisfied if the orientation of the user's head is outside a predetermined angle boundary or if the angle jerk of the user's head exceeds a predetermined threshold.
[0010] In one example, if the controller determines that the user's head orientation is outside a predetermined angular boundary, it is programmed to rotate the angular boundary around a rotation axis along with the audio frame, and the position of the audio frame is rotated around the rotation axis for a predetermined period after the user's head orientation is back inside the angular boundary, in order to reduce the angular offset with respect to the user's head reference axis.
[0011] In one example, if it is determined that the user's head orientation is outside a predetermined angular boundary, a second angular boundary, narrower than the first, is rotated around the axis of rotation along with the audio frame, and the position of the audio frame is rotated around the axis of rotation to reduce the angular offset with the user's head reference axis for a predetermined period of time until the user's head orientation is within the second angular boundary.
[0012] In one example, at least one virtual sound source includes a first virtual sound source and a second virtual sound source, the first virtual sound source is located at a first position, the second virtual sound source is located at a second position, and the first and second positions are relative to an audio frame.
[0013] In one example, maintaining an audio frame in a first position involves rotating the audio frame at a rate adjusted to eliminate sensor drift.
[0014] For example, a sensor that outputs a sensor signal may include multiple sensors that output multiple signals.
[0015] In another example, a method for providing spatialized audio includes outputting a spatialized audio signal to a pair of electroacoustic transducers for conversion to a spatialized acoustic signal, based on a sensor signal representing the orientation of the user's head, wherein the spatialized acoustic signal is perceived by the user as originating from a virtual sound stage including at least one virtual sound source, each virtual sound source in the virtual sound stage being positioned at a different position from the position of the electroacoustic transducers, and perceived with respect to an audio frame of the virtual sound stage, the audio frame being positioned at a first position aligned with the reference axis of the user's head, and determining from the sensor signal whether the characteristics of the user's head satisfy at least one predetermined condition, which includes whether the orientation of the user's head is outside a predetermined angular boundary, and if it is determined that the orientation of the user's head is outside the predetermined angular boundary, rotating the position of the audio frame around a rotation axis to reduce the angular offset with respect to the reference axis of the user's head.
[0016] For example, rotating the position of an audio frame involves rotating the audio frame so that it aligns with the reference axis of the user's head when the user's head turns.
[0017] In one example, while the user's head is experiencing increasing angular acceleration, the angular velocity of the audio frame's rotation is at least partially based on the user's head's angular velocity.
[0018] In one example, while the user's head is experiencing decreasing angular acceleration, the angular velocity of the audio frame's rotation is selected so that the audio frame aligns with a predicted position on the user's head's reference axis when the user's head turns.
[0019] In one example, the predicted position of the user's head's reference axis when the user's head turns is updated with each sample that has decreasing angular acceleration.
[0020] In one example, at least one predetermined condition further includes whether the angle jerk of the user's head exceeds a predetermined threshold, and at least one predetermined condition is satisfied if the orientation of the user's head is outside a predetermined angle boundary or if the angle jerk of the user's head exceeds a predetermined threshold.
[0021] In one example, if it is determined that the user's head orientation is outside a predetermined angular boundary, the angular boundary rotates around the axis of rotation along with the audio frame, and the position of the audio frame is rotated around the axis of rotation to reduce the angular offset with the reference axis of the user's head for a predetermined period after the user's head orientation is inside the angular boundary.
[0022] In one example, if it is determined that the user's head orientation is outside a predetermined angular boundary, a second angular boundary, narrower than the first, is rotated around the axis of rotation along with the audio frame, and the position of the audio frame is rotated around the axis of rotation to reduce the angular offset with the user's head reference axis for a predetermined period of time until the user's head orientation is within the second angular boundary.
[0023] In one example, at least one virtual sound source includes a first virtual sound source and a second virtual sound source, the first virtual sound source is located at a first position, the second virtual sound source is located at a second position, and the first and second positions are relative to an audio frame.
[0024] In one example, maintaining an audio frame in a first position involves rotating the audio frame at a rate adjusted to eliminate sensor drift.
[0025] For example, a sensor that outputs a sensor signal may include multiple sensors that output multiple signals.
[0026] Details of one or more implementations are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from this specification and the drawings, as well as from the claims. [Brief explanation of the drawing]
[0027] In drawings, the same reference numeral generally refers to the same part across different drawings. Furthermore, drawings are not necessarily to scale; rather, they generally focus on illustrating principles in various aspects.
[0028] [Figure 1] A block diagram of a pair of headphones configured to provide spatialized audio is shown as an example. [Figure 2] An example shows a top view of the user's head and the virtual soundstage perceived by the user at two different points in time. [Figure 3] An example shows a top view of a user's head wearing a pair of headphones, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4A] An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4B] An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4C] An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4D]An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4E] An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 4F] An example shows a top view of the user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering the dynamic phase of the spatialized audio mode. [Figure 5A] An example shows a top view of the head of a user wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering and exiting the dynamic phases of the spatialized audio mode. [Figure 5B] An example shows a top view of the head of a user wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering and exiting the dynamic phases of the spatialized audio mode. [Figure 5C] An example shows a top view of the head of a user wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular boundary for entering and exiting the dynamic phases of the spatialized audio mode. [Figure 6] An example shows a top view of a user's head wearing a pair of headphones, the virtual soundstage perceived by the user, and the angular jerk and threshold for entering the dynamic phase of the spatialized audio mode. [Figure 7A]An example shows a top view of a user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular jerk and threshold for entering the dynamic phase of the spatialized audio mode. [Figure 7B] An example shows a top view of a user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular jerk and threshold for entering the dynamic phase of the spatialized audio mode. [Figure 7C] An example shows a top view of a user's head wearing a pair of headphones at various points in head rotation, the virtual soundstage perceived by the user, and the angular jerk and threshold for entering the dynamic phase of the spatialized audio mode. [Figure 8A] An example shows a top view of a user's head wearing a pair of headphones, the virtual soundstage perceived by the user, the angular acceleration of the user's head, and the angular velocity of the virtual soundstage at various points in a head turn. [Figure 8B] An example shows a top view of a user's head wearing a pair of headphones, the virtual soundstage perceived by the user, the angular acceleration of the user's head, and the angular velocity of the virtual soundstage at various points in a head turn. [Figure 8C] An example shows a top view of a user's head wearing a pair of headphones, the virtual soundstage perceived by the user, the angular acceleration of the user's head, and the angular velocity of the virtual soundstage at various points in a head turn. [Figure 9] An example shows a timing diagram of various signals and values associated with dynamic head tracking resulting from user head rotation and a virtual soundstage. [Figure 10A] A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Figure 10B]A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Figure 10C] A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Figure 10D] A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Figure 10E] A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Figure 10F] A flowchart of Method 1000 is shown, which provides spatialized audio using a virtual soundstage that dynamically tracks the user's head movements. [Modes for carrying out the invention]
[0029] Current headphones that offer spatialized audio fail to deliver a consistently comfortable experience to users engaged in activities involving frequent head turns, such as walking, running, and cycling. To improve this, such headphones often offer a "fixed mode" in which the audio is fixed to the user's head, meaning the audio is always rendered in front of the user (i.e., without head tracking). While this effectively solves the problems encountered with spatialized audio when engaging in these activities, it also fails to deliver truly spatialized audio. Rendering the audio in front of the user's head effectively destroys the auditory illusion of spatialized audio, leading to a "collapse" of the externalized audio inside the user's head, meaning the user perceives the audio as originating from speakers.
[0030] Therefore, there is a need for headphones that can provide spatialized audio using head tracking to maintain a consistent and comfortable experience for users engaged in activities that require frequent head turns.
[0031] Figure 1 shows an exemplary pair of headphones 100 configured to generate a spatialized audio signal that creates a virtual soundstage that dynamically tracks the rotation of the user's head when certain predetermined conditions are met. In the illustrated example, the headphones 100 include earcups 102, 104. Each earcup 102, 104 includes an electroacoustic transducer 106, 108 (also called a speaker) for converting an incoming signal into an acoustic signal. The earcup 102 further houses a controller 110 and a sensor 112 configured to generate a sensor signal representing the orientation of the user's head. Based on the sensor signal, the controller 110 generates a spatialized audio signal from an incoming audio signal, such as music or spoken content, to the electroacoustic transducers 106, 108, which convert the spatialized audio signal into a spatialized acoustic signal that is perceived by the user as occurring at one or more locations in space different from the location of the electroacoustic transducer. (In this example, the controller 110 is connected to the electroacoustic transducers 106, 108 via wires extending through the headband 114.) The received audio signal may include any suitable audio source, including multichannel and / or object audio, as well as general audio content (not limited to music), and audio for general contexts (audio for video, communications, games, etc.).
[0032] For the sake of brevity and to highlight more relevant aspects of the headphones 100, certain features of the block diagram in Figure 1, such as the Bluetooth® system-on-a-chip and battery, have been omitted. Furthermore, while Figure 1 shows a pair of over-ear headphones, it should be understood that the headphones 100 could be any appropriate form factor, including in-ear headphones, on-ear headphones, open-ear headphones, earphones, etc.
[0033] The controller 110 includes a processor 116 and a memory 118 for storing program code to be executed by the processor 116 to perform various functions for providing spatialized audio as described herein, including, optionally, steps of method 1000 as described below. It should be understood that the processor 116 and memory 118 of the controller 110 do not need to be located in the same housing as part of a dedicated integrated circuit, but can be located in separate housings. Furthermore, the controller 110 may include multiple physically separate memories for storing program code necessary for its functions, and may include multiple processors 116 for executing the program code. The various components of the controller 110 also do not need to be located in the same ear cup (or ancillary components in other headphone form factors), but can be distributed between ear cups. For example, each of the ear cups 102 and 104 may include a processor and memory that work together to perform various functions for providing spatialized audio as described herein, and the processors and memories in both ear cups 102 and 104 form a controller.
[0034] As described above, sensor 112 generates a sensor signal representing the orientation of the user's head. In one example, sensor 112 is an inertial measurement unit used for head tracking; however, it should be understood that sensor 112 can be implemented as any sensor suitable for measuring the orientation of the user's head. Furthermore, sensor 112 may include multiple sensors that work together to generate the sensor signal. In fact, the inertial measurement unit itself typically includes multiple sensors, such as an accelerometer, gyroscope, and / or magnetometer, that work together to generate the sensor signal. The sensor signal representing the orientation of the user's head may be a data signal that directly represents the orientation, for example, as changes in pitch, roll, and yaw, or it may include other data from which the orientation can be derived, such as specific forces and angular velocities of the user's head. In addition, the sensor signal itself may consist of multiple sensor signals, such as when multiple separate sensors are used to measure the orientation of the user's head. In one example, separate inertial measurement units can be positioned on the ear cups 102, 104 (or other accessory parts of the headphone form factor), or at any other suitable location, to form a sensor 112 together, and signals from the separate inertial measurement units form a sensor signal.
[0035] The controller 110 may be configured to generate audio signals in one or more modes. These modes include, for example, an active noise reduction mode or a hear-through mode. In addition, the controller 110 may generate spatialized audio in a mode in which the virtualized soundstage is fixed in space (e.g., in front of the user) and does not change the perceived position of the virtualized soundstage in response to the user's head movements, or changes it only in response to the user facing in a direction greater than a predetermined angular rotation away from the virtual soundstage for a predetermined period of time. For the purposes of this disclosure, this is referred to as the “room-fixed” mode, referring to the fact that the virtual soundstage is perceived as being fixed in place in the room. Further details regarding the room-locked mode are described in U.S. Patent Application No. 16 / 592,454, “SYSTEMS AND METHODS FOR SOUND SOURCE VIRTUALIZATION,” filed on 3 October 2019, and U.S. Patent Application No. 63 / 415,783, “SCENE RECENTERING,” filed on 13 October 2022, both of which are incorporated herein by reference. As described above, the room-locked mode is best suited for relatively stationary users, such as those sitting at a desk, and is not suitable for active users, such as those walking or running. To address this, the controller 110 is programmed to operate in a “head-locked” mode by user selection or by some trigger condition, which maintains the virtual sound stage at a fixed point in space until a predetermined condition is met (similar to the room-locked mode, referred to in this disclosure as the “static phase” of the head-locked mode), and once the predetermined condition is met, the virtual sound stage can be dynamically rotated to follow the rotation of the user’s head (referred to in this disclosure as the “dynamic phase” of the head-locked mode). The details of the head-locked mode will be described in more detail with reference to Figures 2 to 10.
[0036] Referring to Figure 2, the user's head 202 is shown moving from a first orientation at time t0 to a second orientation, indicated by 202', at time t1. Headphones 100 are omitted from Figure 2 so that various features and angles can be seen more clearly, but it should be understood that headphones 100 are worn on the user's head 202 at both times t0 and t1. The user perceives a virtual sound stage 204 based on the spatialized acoustic signal generated by headphones 100. The virtual sound stage 204 includes one or more virtualized speakers (also called "virtualized sound sources"), shown here as virtualized speakers 206, 208. Although two virtual speakers are shown, it will be understood that in various alternative examples, any number of virtualized speakers can be created from the spatialized acoustic signal (according to the spatialized audio signal generated by controller 110). Furthermore, while the virtualized speakers 206 and 208 are shown symmetrically positioned in front of the user's head 202 (i.e., around the longitudinal axis extending from the Z-axis to point P in Figure 2), it should be understood that in various examples, the virtualized speakers may be asymmetrically distributed with respect to the front of the user's head. The generation of virtualized speakers is generally understood, so a more detailed explanation is omitted here.
[0037] For the purposes of this disclosure and for simplicity, the positions of virtual sound stages and virtual speakers will often be discussed as if the virtual sound stage or virtual speaker were physically positioned at a given location. Even if not explicitly stated, it should be understood that the position of a virtual sound stage or virtual speaker is merely a perceived position. In other words, to the extent that a virtual sound stage or virtual speaker is described as having a physical position in space, it refers only to the perceived position of the virtual sound stage or virtual speaker, and not to its actual position in space.
[0038] The positions of virtualized speakers 206 and 208 may be the positions in which they were initialized. Speaker initialization may occur, for example, when the user first selects a spatialized audio mode (such as room-fixed mode or head-fixed mode), and they may initially be placed in front of the user (for example, as shown in Figure 2), but virtualized speakers 206 and 208 may be initialized in other locations. In fact, the initial positions (and number) of virtualized speakers may be selected by the user, for example, via a dedicated mobile application or web interface accessible through a computer or mobile device.
[0039] As described above, after the virtualized speaker is initialized, the controller 110 maintains the virtual sound stage 204 in the same position in the static phase until one or more predetermined conditions are met. When at least one of the predetermined conditions is met, the spatialized audio signal may be adjusted by the controller 110 in the dynamic phase so that the virtual sound stage 204 rotates to track the movement of the user's head. The movement of the virtual sound stage 204 is shown in Figure 2 by the rotation of the virtual sound stage 204, which is shown as virtual sound stage 204'. More specifically, when the user's head 202 rotates to a second orientation at time t1, the virtual sound stage 204 tracks the position of the user's head 202, but typically with some delay. Thus, at time t1, the virtual sound stage 204' tracks the user's head 202' in angle. This slight delay is desirable as it maintains the illusion of the virtual sound stage. If the virtual soundstage perfectly tracks the user's head without any delay, the perception of the virtual soundstage "collapses" within the user's head, as there is no longer any distinction between the movement of the user's head and the perceived movement of the virtual speaker.
[0040] The rotation of the virtual sound stage 204 is achieved by the rotation of each virtual speaker 206, 208 (since the virtual sound stage 204 is entirely composed of a collection of virtual speakers). In other words, the virtual sound stage 204 rotates around the Z axis at an angle α from time t0 to time t1. c When rotating angularly, each virtual speaker 206, 208 similarly rotates around the Z axis by an angle α from its initial position at time t0 to its position at time t1. c The virtual soundstage 204 rotates. However, for the purposes of this disclosure, the rotation of the virtual soundstage 204 is described with respect to a single reference point, shown as point A in Figure 2 and referred to as the “audio frame”. Thus, the virtual soundstage that dynamically tracks the user’s head is described with respect to the rotation of audio frame A that follows the rotation of the user’s head 200. In the example shown, audio frame A is initialized to be aligned with the longitudinal axis ZP of the user’s head 202. However, it should be understood that since audio frame A is the reference point for describing the rotation of the virtual soundstage 204, any suitable reference axis, i.e., a reference axis on which the angular offset between the user’s head 202 and the virtual soundstage 204 can be measured, may be used.
[0041] The virtual soundstage 204 (and consequently, each virtual speaker 206, 208) rotates angularly around the Z-axis. Typically, the Z-axis corresponds to the axis around which the user's head rotates; otherwise, the virtual soundstage 204 would not be perceived as remaining at a fixed distance from the user throughout the rotation. In practice, this axis can be approximated to any point where the change in distance throughout the rotation is not noticed by the user.
[0042] The virtual sound stage 204, which tracks the user's head movements, achieves this by rotating the virtual sound stage 204 to reduce the angular offset between the virtual sound stage 204 and the orientation of the user's head. This is achieved by reducing the turn angle α of the user's head 202. turn(i.e., the angle between the orientation of the user's head 202 at time t0 and the orientation of the user's head 202' at time t1) and the rotation angle α of the virtual sound stage c (i.e., the angle between the orientation of the virtual sound stage 204 at time t0 and the orientation of the virtual sound stage 204' at time t1) is the offset angle α off as shown in FIG. 2. Therefore, the moving direction of the virtual sound stage 204 is selected to reduce the large angular offset between the turn angle α turn and the rotation angle α c .
[0043] When at least one of the predetermined conditions is satisfied, the virtual sound stage starts to move toward the current orientation of the user's head 202 at an angular velocity of μ dps degrees per second (the selection of the value of μ dps can be a dynamic process and will be described in more detail below). The rotation angle α c can be given by integrating the value of μ dps per sample from time t0 to t1.
[0044]
Equation
[0045] Similarly, the turn angle α turn of the user's head 202 can be obtained by integrating the measured angular velocity Ω z (in degrees per second) of the user's head measured by the controller 110 at each sample from time t0 to t1 according to the input from the sensor 112.
[0046]
Equation
[0047] Therefore, the angle offset α off at time t is the angular velocity Ω zand the angular velocity μ of sound stage 204 dps This can be obtained by integrating the difference between and and summing the result with the offset that existed at time t0.
[0048]
number
[0049] Equation (3) is the offset at time t0 and the rotation angle α of the user's head 202 from time t0 to t1. turn and the rotation angle α of the virtual sound stage 204 c It can be rewritten as the sum of the difference between and .
[0050]
number
[0051] As described above, the virtual sound stage 204 remains fixed in space unless certain predetermined conditions are met. Such predetermined conditions may be, for example, an angular boundary (such as a wedge or cone) placed around the user's head to determine whether the user's head is turning beyond a predetermined maximum value, and whether the angular jerk of the user's head exceeds a threshold to determine whether the user's head is turning rapidly, or an early indication of rotation exceeding the angular boundary. Other predetermined conditions are also conceivable and within the scope of this disclosure.
[0052] Referring to Figure 3, the first of such predetermined conditions, namely angle α, max The angular boundary is shown as α. max This is represented as a two-dimensional cone (also called a wedge) corresponding to whether the user's head is moving within the maximum allowable yaw (i.e., rotating to its maximum possible range) before the virtual sound stage 204 is adjusted to realign audio frame A with the longitudinal axis ZP of the user's head. Angular boundary α maxFor example, it can cover a maximum rotation of 50°, but its width is a design choice, and in other examples, it may have other appropriate values. In this example, the 2D cone has its vertex at the axis of rotation Z, and as a result, the longitudinal axis ZP is perfectly aligned with the angular boundary α. max It is either inside or completely outside.
[0053] Figure 3 shows the angle boundary α when the user's head is facing the angle boundary. max The longitudinal axis ZP is shown as the reference axis used to determine whether or not the angle α is exceeded, but any suitable reference axis or reference point can be used to compare the orientation of the user's head to the angular boundary. For example, a reference point P representing the front of the user's head can be used instead of the longitudinal axis ZP. However, the reference point does not have to lie on the longitudinal axis ZP, and the same reference axis used to determine the offset or alignment of the virtual sound stage 204 can be used instead of the angular boundary α. max It does not need to be used in the same way as it is used to compare orientations relative to . However, if different reference axes or reference points on the longitudinal axis are used, the angular boundary α should be taken into account to account for the difference in the initial direction or position of the reference axis / point used. max It may be necessary to adjust its position or orientation.
[0054] However, using the same reference axis for offsetting the virtual sound stage 204 and detecting when the user's head orientation exceeds the angular boundary means that the angular offset α off This allows the user's head 202 to be used as a proxy for the reference axis. In other words, it determines the alignment between the virtual sound stage 204 and the user's head 202, and the user's head 202 is at the angular boundary α. max If the same reference axis is used to determine whether or not it exceeds the angle boundary α, then the user's head 202 is at the angular boundary α max While remaining inside, angle offset α off The angular boundary α max This is equal to the distance from the center, allowing for a certain degree of computational economy. (This assumes that the angular boundaries are symmetrically arranged around the reference axis, which is a typical case.)
[0055] It should be further understood that other suitable shapes may be used for angular boundaries. For example, a three-dimensional cone may be used instead of a two-dimensional cone to determine whether the pitch of the user's head exceeds the boundary of the vertical dimension (i.e., the pitch of the user's head).
[0056] The user's head reaches the maximum angle boundary α max While remaining inside, the controller 110 operates in the static phase of head-fixed mode, which means the user perceives the virtual sound stage 204 as fixed in space. This is equivalent to the operation in room-fixed mode described above. Generally, during this period, the angular velocity μ of the virtual sound stage 204 dps This is either 0 degrees / second or held at a very low value to compensate for the drift of sensor 112 (thus the user does not perceive the operation of the virtual sound stage 204).
[0057] The user's head is at the angular boundary α max If it is determined to be outside of it, the angular velocity μ of the virtual sound stage 204 dps is the angle offset α off In the direction of reducing, that is, the angular velocity Ω of the user's head 202 z This increases in the direction of α. This is shown in detail in Figures 4A to 4F, which together show the rotation of the user's head and, consequently, the response of the virtual sound stage 204. In Figure 4A, the orientation of the user's head, as indicated by the longitudinal axis ZP, is pointed upward on the page, and the angular boundary α max Located within, therefore, the controller 110 remains in a static phase, and the virtual sound stage 204 remains in its fixed position. In Figure 4B, the user's head begins to turn left within the page, but the longitudinal axis ZP is at the angular boundary α maxBecause it remains internal, the controller 110 stays in a static phase, and the user continues to perceive the virtual soundstage 204 as fixed in the same position. (It should be understood that adjustments to the spatialized audio signal are required based on changes in the orientation of the user's head 202 so that the virtual soundstage 204 is perceived as existing in the same position in space.)
[0058] In Figure 4C, the longitudinal axis ZP is the angular boundary α max This occurs, and in response, the controller 110 enters the dynamic phase of the head-fixing mode. However, Figure 4C shows that the orientation of the user's head is at the angular boundary α. max Measured to be greater than , thus representing the first sample where the virtual soundstage 204 has not yet begun to track the movement of the user's head 202. As shown in Figure 4D, the user's head 202 continues to turn to the left, and the virtual soundstage 204 has begun to track the movement of the user's head and is also shifting to the left, but there is a slight delay between the angle of audio frame A and the longitudinal axis ZP, i.e., the offset angle α off There is an angular boundary α. max This is the same rotation angle α as the virtual sound stage 204 is rotated. c It is similarly rotated around the Z-axis. In Figure 4E, the angular boundary α continues to rotate. max The longitudinal axis ZP is again at the angular boundary α max As shown inside, it catches up with and overtakes the longitudinal axis ZP. In Figure 4F, the sound stage (and angular boundary α) max ) has reached the end position of the user's head turn, and therefore the virtual sound stage 204 is again aligned with the longitudinal axis ZP and offset angle α offThe offset is rendered to 0° or within a predetermined tolerance. (Generally, for the purposes of this disclosure, the offset only needs to be within a predetermined range, which is a design choice that determines how precisely the virtual soundstage 204 is aligned with the user's head 202. Typically, the predetermined range is chosen to be a value that is imperceptible to the user in order to maintain the perception that the audio frame has been adjusted to its previous position relative to the user's head.)
[0059] The controller continues the dynamic phase until a second predetermined condition is met. In one example of such a second predetermined condition, the virtual sound stage 204 is positioned so that the user's head is at the angular boundary α max After returning to the starting position, the movement of the user's head 202 is tracked for a predetermined duration. For example, as shown in Figure 4E, the longitudinal axis ZP is at the angular boundary α max It has just returned to (angle boundary α max As the virtual sound stage 204 rotates toward it, a predetermined period begins before tracking of the user's head stops and the virtual sound stage 204 locks into place (i.e., enters a static phase). In one example, the predetermined period may be 0.5 seconds, but other appropriate time lengths may be used. The predetermined period may be chosen to allow tracking of the user's head 202 to continue until the virtual sound stage 204 is realigned with the longitudinal axis ZP, as shown in Figure 4F.
[0060] In an alternative example, a second predetermined condition is established to determine when the virtual sound stage 204 is aligned with the longitudinal axis ZP, the angular boundary α max It could be a separate angular boundary that is narrower than α. In other words, in this example, the angular boundary α maxThe first angle boundary is used to determine when the virtual sound stage 204 begins tracking the user's head 202 (i.e., enters the dynamic phase), while the second angle boundary is used to determine when the virtual sound stage 204 stops tracking the user's head 202 and becomes fixed in space again (i.e., enters the static phase). An example of this is shown in Figures 5A and 5B, which show the user's head turned to the left (as described in relation to Figures 4A-4F), the virtual sound stage 204, and the angle boundary α. max and a narrower α min However, the user's head is at the angular boundary α max This indicates that tracking to the left begins in response to a turn beyond α. In Figure 5A, the virtual sound stage 204 tracks the turn of the user's head 202, but is not yet aligned with the longitudinal axis ZP. The longitudinal axis ZP is aligned with the angular boundary α max and angular boundary α min It is located on both sides of the
[0061] In Figure 5B, the user's head 202 is at the angular boundary α max It is inside, but the angular boundary α min It is not internal, and therefore the virtual soundstage continues to track the movement of the user's head 202. In Figure 5C, the longitudinal axis ZP is the angular boundary α min Located within, the virtual sound stage 204 ceases tracking the user's head 202, and the controller 110 returns to its static phase. Angular boundary α min The width is a design choice depending on how precisely the virtual sound stage 204 is aligned with the longitudinal axis ZP. Generally, the angular boundary α min The narrower the gap, the more closely the virtual sound stage 204 is aligned with the front of the user's head point P.
[0062] Referring now to Figure 6, a second example of predetermined conditions for enabling tracking of the virtual sound stage 204 is shown. In this example, instead of determining whether the user's head has crossed an angular boundary, the controller 110 determines whether the user's head has begun to turn rapidly in one direction or another by determining whether the angular jerk, i.e., the rate of change of the user's head acceleration, has exceeded a threshold. In Figure 6, the angular jerk of the user's head 202 is represented by a labeled curved arrow.
number
[0063] Figures 7A to 7C show the adjustment of the virtual soundstage after detecting angular jerk of the user's head 202 that exceeds a predetermined threshold. In Figure 7A, the user's head 202 is represented by a labeled curved arrow.
number
[0064] Generally, monitoring the angular jerk of the user's head 202 is useful for early detection of head turns, but it will (by design) not detect slower movements, even movements that would cause the user to rotate sharply to the left or right. Therefore, the angular jerk of the user's head 202 is used to determine the angular boundary α max It is thought to be used in conjunction with the above, and either an angular jerk exceeding a threshold or the user's head orientation exceeding an angular boundary is sufficient to enter the dynamic phase and adjust the position of the virtual sound stage 204. However, it is conceivable that either the angular boundary condition or the angular jerk threshold condition may be used as the sole predetermined condition for initiating user head tracking. Furthermore, it should be understood that any appropriate predetermined condition may be used to detect or predict user head rotation beyond a predetermined range, instead of or in addition to the two methods described above.
[0065] Angular velocity μ of virtual soundstage 204 dps This is the angular velocity Ω of the user's head. zThis can be based on the following. Generally, when soundstage tracking is triggered, the goal is to quickly cancel large head rotations (typically so that the user perceives the soundstage primarily in front of the user's head 202) and eliminate the most unpleasant artifact of recentering the virtual soundstage, while also allowing for some delay so that the illusion of the virtual soundstage is preserved. The applicant also understands that delaying the user's head after the user's head has stopped moving is usually an unpleasant experience for the soundstage. In other words, the audio frame A of the virtual soundstage 204 should ideally be aligned with the longitudinal axis ZP (or other reference as an axis) when the user's head stops. Thus, to track the movement of the user's head, the controller 110 adjusts the alignment of the virtual soundstage 204 with the longitudinal axis ZP based on the movement of the user's head 202, but to coincide with the end of the head turn, thereby adjusting the angular velocity μ of the virtual soundstage 204. dps It can be adjusted dynamically.
[0066] To achieve these goals, the controller 110 controls the angular velocity μ of the virtual sound stage 204 according to two distinct stages: (1) when the user's head is accelerating and (2) when the user's head is decelerating. dps This can be dynamically adjusted. Figures 8A to 8C show μ at different stages of the head turn. dps This shows the process of selecting. In Figure 8A, the user's head has moved to the left, and a labeled curved arrow appears.
number
number
[0067] In Figure 8B, the user's head is still moving to the left, but has begun to decelerate, and is indicated by a labeled curved arrow pointing to the right.
number
[0068] This predicts the time when the user's head turn will be completed, and the virtual sound stage 204 determines the remaining angular offset α between its current position and the predicted orientation of the user's head 202 at the end of the turn. off μ traverses dps,2This can be achieved by setting the following: For example, if h0 represents the time from the current sample to the end of the head turn, the virtual soundstage is set to the existing angular offset α. off And the additional angular offset α that occurs from the current time t to the time t+h0 when the head turn is completed. off And must be compensated for (i.e., traversed).
[0069] The angular velocity of the user's head at a future time h can be approximated by a linear approximation as follows:
number
[0070]
number
[0071] Therefore, the linear approximation at future time h can be rewritten as follows:
[0072]
number
[0073] Therefore, the additional angle that occurs from the current time to the time t+h0 when the head turn is completed can be approximated as follows:
[0074]
number
[0075] The angular velocity μ of the virtual sound stage 204 at future time h.dps,2 (t) is a linear approximation Ω z It can be assumed that t+h0) exists, and therefore it can be written as a linear approximation as follows.
[0076]
number
[0077] Therefore, the angle compensated (traversed) by the virtual sound stage 204 from the current time to the time t+h0 when the head turn ends can be approximated as follows:
[0078]
number
[0079] μ dps,2 The angular velocity of (t) is given by equation 9, which is the angular offset α that existed between equation 11 and the current time t. off It can be set to cancel and. In other words, μ dps,2 (t) is selected as follows:
[0080]
number
[0081] μ dps,2 By solving for (t), the angular velocity that causes the virtual sound stage 204 to arrive in front of the user's head 202 when the head turn is complete is obtained as follows:
[0082]
number
[0083] This angular velocity is the angular velocity Ω of the user's head. z To adjust for the changes, it can be recalculated for each input sample.
[0084] Regardless of the stage of the user's head turn, the dynamic angular velocity is μ dps,1 and μ dps,2 and the baseline angular velocity μ baseline which may include the calculated values of. The baseline angular velocity μ baseline may be added to account for a very slow head turn or the user's head returning to the angular boundary α max due to a residual offset. The baseline angular velocity μ baseline ensures that the virtual sound stage 204 does not deviate far from the longitudinal axis Z-P (or other reference axis), or that the residual angular offset is eliminated. In one example, the baseline angular velocity μ baseline may be 20 degrees per second, although other suitable values are contemplated herein.
[0085] Referring to FIG. 9, an exemplary timing diagram of various signals and values associated with the rotation of the head turn is shown to demonstrate the detection of the condition causing entry into the dynamic phase and the resulting adjustment of the angular velocity μ dps . Starting from the upper plot in FIG. 9, a signal representing the rotation angle α z of the user's head 202 about the Z-axis is shown. The upper plot in FIG. 9 also shows the angular boundary α max represented as a shaded horizontal boundary. As shown in FIG. 9, at point 1a, the rotation angle α z exits the angular boundary α max , and as a result, the controller 110 enters the dynamic phase and increases the angular velocity μ dps of the virtual sound stage 204 from a zero value to a non-zero value. (Although a zero value of μ dps is shown in FIG. 9, it should be understood that in practice, a non-zero value of the angular velocity μ dps may be maintained during the static phase to eliminate sensor drift. The initial dynamic phase is shown here as the first of two vertically shaded regions, and the second shaded region represents the second dynamic phase.) Immediately after 1a, the angular velocity μ dps is the angular velocity Ω zIt increases based on this. At point 2a, the angular acceleration of the user's head 202 begins to decrease, signaling the end of the user's head turn. Based on the predicted end of the head turn, the angular velocity μ dps The rotation angle α of the user's head is rapidly increased, and as a result, the virtual sound stage 204 realigns with the reference axis (e.g., longitudinal axis) of the user's head when the user's head stops. At point 3a, the rotation angle α of the user's head 202 around the Z axis z This triggers a predetermined period of time, and the angular boundary α max Returning to the interior, at its end, the dynamic phase ends, represented as point 4a, and the angular velocity μ of the virtual sound stage 204 is reached. dps It reaches zero again.
[0086] At point 1b, the second dynamic phase begins, which is the result of the user's head 202 angular jerk exceeding threshold t, both of which are shown in the central plot of Figure 9. When the user's head 202 angular jerk exceeds threshold t, the dynamic phase begins as an early detection of the user's head turn. Thus, the dynamic phase begins when the user's head 202 exits and reaches the angular boundary α at point 3b. max This continues until it returns to (as shown in the upper plot), at which point the predetermined period begins again and ends at point 4b at the end of the second dynamic phase. The angular velocity μ of the virtual sound stage 204 in the lower plot during the second dynamic phase. dps Looking at it, at point 1b, the angular velocity μ dps It becomes non-zero again, and the angular velocity Ω of the user's head 202 z It has a value based on . At point 2b, the user's head 202 begins to decelerate, and as a result the angular velocity μ dps The angular velocity μ increases rapidly, allowing the virtual sound stage 204 to realign with the reference axis when the user's head turns are complete. dps It will be distributed.
[0087] Referring here to Figures 10A to 10D, flowcharts of Method 1000 for providing spatialized audio with a virtual soundstage that dynamically tracks the movement of the user's head are shown. Method 1000 can be implemented by a controller (e.g., controller 110) contained in a pair of headphones (e.g., headphones 100). The controller may include one or more processors and one or more non-temporary storage media that store programs for execution by the one or more processors. For example, the controller may include two microcontrollers, each containing a processor and memory, located in the earcups of the headphones and working in coordination to perform the steps of Method 1000. The headphones may further include at least one pair of electroacoustic transducers that receive audio signals from the controller and convert them into acoustic signals.
[0088] In step 1002, a sensor input signal or a selection of mode operation is received. In one example, the headphones (and in particular the controller) may operate in two or more operating modes, which include different spatialized audio modes. These modes include, for example, a room-fixed mode, in which the virtual soundstage is perceived as being fixed in a specific position in space, and does not move in response to the user's head movement, except in certain examples when the user's head is turned away from the virtual soundstage for at least a predetermined period of time. In the head-fixed mode, as will be described in more detail below (and in relation to Figures 1 to 9), the virtual soundstage is fixed in a specific position in space until certain predetermined conditions are met, at which point the virtual soundstage rotates to track the user's head movement.
[0089] The sensor input signal may be from a sensor such as an inertial measurement unit that can provide input indicating user activity, such as walking or running, for which the head tracking mode is more suitable (other suitable types of sensors such as accelerometers and gyroscopes are intended). The sensor input may be received from the same sensor that detects the orientation of the user's head, or from different sensors. In yet another example, the sensor signal may be mediated by a secondary device such as a mobile phone or a wearable such as a smartwatch that includes a sensor. Alternatively, the input may be received from the user, for example, using a dedicated application, or via a web interface, or via a button on headphones or other input, and selected directly between modes.
[0090] Step 1004 is a decision block indicating the selection of an operating mode or whether the sensor signal meets the requirements for head-locking mode. Upon receiving the input of an operating mode selection from the user, this is typically sufficient without requiring any further action or decision to meet the requirements. However, the sensor signal requires specific analysis to determine whether it indicates an activity worthy of switching to head-locking mode. Such analysis may, for example, determine whether the user has walked a predetermined number of steps within a predetermined period, or whether the user has completed a predetermined number of head turns within a predetermined period (such as those determined by measuring the reference axis of the user's head relative to an angular boundary). Other appropriate means of determining whether the user is engaged in an activity that will be supported (i.e., made more comfortable) through the implementation of head-locking mode are also contemplated. In fact, many tests already exist for identifying when the user is engaged in an activity or a particular type of activity, and any such appropriate test may be used. In an alternative example, step 1004 may be performed by a secondary device (e.g., a mobile phone or wearable) that, in response to the analysis of sensor signals by the secondary device, instructs the controller to enter head-locking mode or to notify that a specific activity is occurring.
[0091] If the requirements for head-locked mode are not met, in step 1006, the controller operates in room-locked mode, in which the virtual soundstage remains fixed, except in the narrow situation where the user is facing different directions for an extended period of time, as described above. As stated above, further details regarding room-locked mode are described in U.S. Patent Applications 16 / 592,454 and 63 / 415,783, and these disclosures are incorporated herein by reference. Furthermore, it should be understood that while room-locked mode is enumerated as the sole alternative to head-locked mode, head-locked mode can be one of any number of potential modes, which may or may not be spatialized audio modes. If the requirements for head-locked mode are met, the method proceeds to step 1008, shown in Figure 10B.
[0092] In step 1008, the spatial audio signal is output by the controller to a pair of electroacoustic transducers for conversion to a spatialized acoustic signal, based on a sensor signal representing the orientation of the user's head. The spatialized acoustic signal is perceived as originating from a virtual sound stage containing at least one virtual sound source, each of which is perceived by the user as being located at a different position from the electroacoustic transducers. The virtual sound sources are also referenced to an audio frame of the virtual sound stage, which is located at a first position and aligned with the reference axis of the user's head. In other words, the audio frame is used as a singularity to describe the position and rotation of the virtual sound stage. In one example, the longitudinal axis of the user's head may be used as the reference axis, but the reference axis may be any suitable axis for determining the angular offset between the user's head and the virtual sound stage as the user's head and the virtual sound stage rotate in the manner described below.
[0093] The current position depends on how and when the head-lock mode was selected in step 1004, since the spatialized audio signal may be initialized in step 1008 or earlier, such as in relation to the room-lock mode. In the former case, the spatialized audio signal is initialized in step 1008 and is therefore determined automatically according to user input or according to the direction the user is facing in step 1008. In the latter case, the first position may depend on the position of the audio frame determined in relation to the room-lock mode, which may be the location where the room-lock mode was initialized or adjusted.
[0094] Step 1010 is a decision block that determines whether the characteristics of the user's head satisfy at least one predetermined condition. Such predetermined conditions may be, for example, whether the orientation of the user's head is outside a predetermined angular boundary, or whether the angular jerk of the user's head exceeds a predetermined threshold.
[0095] A brief reference to Figure 10C shows an example of step 1010, including steps 1018 and 1020. Step 1018 is a determination block representing determining whether the user's head orientation exceeds a predetermined angular boundary. The angular boundary may be used to determine whether the user's head has rotated beyond a predetermined range. The angular boundary may, in one example, be a two-dimensional cone, but a three-dimensional cone and other appropriate shapes are intended. A reference axis or reference point, which may be the same as or different from the reference axis for determining the alignment of the audio frame, may be compared to the angular boundary to determine whether the user's head orientation has rotated beyond a predetermined range. An embodiment of comparing the reference axis (here, the longitudinal axis of the user's head) with the angular boundary is shown in Figures 4 and 5. If it is determined that the user's head orientation does not exceed the predetermined angular boundary, the method proceeds to step 1020. If it is determined that the user's head orientation exceeds the angular boundary, the method proceeds to step 1014.
[0096] Step 1020 is a decision block representing the determination of whether the user's head angle jerk exceeds a predetermined threshold. In this example, the user's head angle jerk is compared to the threshold as an early indication of a turn that is likely to exceed the angle boundary. An example of comparing the angle jerk to the threshold is shown and explained in relation to Figures 6 and 7. If it is determined that the user's head angle jerk does not exceed the predetermined threshold, the method proceeds to step 1012 (i.e., continues the static phase). If it is determined that the user's head angle jerk exceeds the predetermined threshold, the method proceeds to step 1014 (i.e., begins the dynamic phase).
[0097] Therefore, angular jerk functions as another predetermined condition that can trigger the dynamic phase. While angular boundaries and angular jerk thresholds are described, it should be understood that other examples of appropriate predetermined conditions, i.e., conditions that indicate or predict at least a predetermined degree of head turn, are contemplated. It is also contemplated that only one such predetermined condition, e.g., only the angular boundary or only the angular jerk threshold, may be used.
[0098] If it is determined that at least one predetermined condition is not met, in step 1012, the audio frame is maintained in its current position. In other words, in the example above, it is determined that the user's head is within the angular boundary and the angular jerk is below the threshold. Maintaining the audio frame in its current position is implemented by rendering at least one virtual sound source so that it is perceived as fixed in space regardless of the user's head movement (in practice, this requires adjusting the spatialized acoustic signal based on the detected change in the orientation of the user's head so that the virtual sound source is perceived as fixed in space). Therefore, if it is determined that the user's head is relatively still (e.g., the user is sitting at a desk), the user will perceive the virtual soundstage as fixed in space.
[0099] Returning to Figure 10B, in step 1014, if at least one predetermined condition is met, the position of the audio frame is rotated around the axis of rotation to reduce the angular offset with respect to the reference axis of the user's head. In other words, the virtual soundstage can be rotated around the axis of rotation of the user's head (or an axis positioned approximately around the axis of rotation of the user's head) to reduce the offset introduced by the rotation of the user's head. As illustrated in relation to Figure 10D, this rotation may continue until the audio frame is aligned with the reference axis. Aligning the audio frame with the reference axis involves rotating the virtual soundstage by the same angle of rotation as the user's head has been rotated. Thus, if the user's head has been rotated 15° from its initial position, the virtual soundstage will also be rotated 15° to align with the user's head. Furthermore, since the virtual soundstage consists of at least one virtual sound source, the rotation of the audio frame is achieved by the respective rotation of each virtual sound source. Thus, a 15° rotation of the virtual soundstage is achieved by a 15° rotation of each virtual sound source.
[0100] Temporarily referring to Figure 10D, steps 1022-1026 show method steps that illustrate in more detail how the rotation in step 1014 is achieved, and more specifically, how the rotation speed of the virtual sound stage (and therefore the rotation speed of each virtual sound source) in step 1014 is selected. Step 1022 is a decision block that indicates whether the user's head has increasing angular acceleration. If it is determined that the user's head has increasing angular acceleration, the method proceeds to step 1024, where the angular velocity of the audio frame rotation is based at least in part on the velocity of the user's head. In one example, the angular velocity may be set to a scale value of the angular velocity of the user's head, and the scale value is based on the acceleration of the user's head (as described in relation to equation (5)). The value of the angular velocity of the virtual sound stage is selected to rotate the virtual sound stage at a pace that follows the user's head but allows for some delay so as not to break the illusion of the virtual sound stage. However, the angular velocity of the virtual soundstage may be further added to the baseline velocity to account for a distracting, delayed virtual soundstage when the user's head is slowly turning.
[0101] However, if it is determined in step 1022 that the user's head does not have angular acceleration, then in step 1026, the angular velocity of the audio frame's rotation is selected such that the audio frame aligns with the predicted position on the reference axis when the user's head finishes turning. This can be achieved by first predicting the time when the user's head will stop turning (by assuming that the current deceleration continues) and the angular rotation the user's head will move from its current position to that time. The angular velocity of the virtual sound stage (and therefore each virtual sound source) can then be selected so that the virtual sound stage aligns with the user's head at the predicted time, which means that this will compensate for both the existing offset between the virtual sound stage and the user's head and the additional angle the user's head moves between the current time and the predicted time (as described in relation to equation (5) above).
[0102] Returning to Figure 10B, step 1016 is a decision block that determines whether the characteristics of the user's head satisfy at least one second predetermined condition. The at least one second predetermined condition is a condition that signals the end of the dynamic phase of the head-fixed mode. Figures 10E and 10F provide examples of the second predetermined condition. Specifically, Figure 10E and step 1028 are decision blocks that determine whether a predetermined period of time has elapsed since the orientation of the user's head (indicated by the reference axis or reference point of the user's head) entered within the angular boundary again. The controller is programmed to rotate the angular boundary around the axis of rotation along with the audio frame (i.e., the angular boundary rotates by the same amount as the user's head from its initial position). In practice, this means that the angular boundary overtakes the reference axis or reference point of the user's head before the audio frame is realigned to the (alignment) reference axis. The reference axis or reference point returning to the angular boundary may initiate a predetermined period, after which the controller exits the dynamic phase of the head-locked mode and returns to step 1010. Before the predetermined period initiated by the reference axis or reference point returning to the angular boundary (typically achieved by the angular boundary overtaking the reference axis) expires, the method returns to step 1014 and continues rotating the virtual sound stage (and angular boundary).
[0103] Figure 10F and step 1030 are a decision block that determines whether the orientation of the user's head (indicated by the reference axis or reference point of the user's head) is within a second angular boundary. The controller is programmed to rotate the angular boundary and the second angular boundary around the axis of rotation along with the audio frame (so both the angular boundary and the second angular boundary rotate by the same amount as the user's head from its initial position). The second angular boundary may be narrower than the angular boundary that caused the dynamic phase to begin in step 1010 (as described, for example, in relation to Figure 5A). In this example, rather than waiting for an additional predetermined period, when the second angular boundary is entered, the method returns to step 1010 to signal sufficient alignment with the reference axis or reference point of the audio frame, and thus the end of the dynamic phase. Before the reference axis enters the second angular boundary, the method returns to step 1014 and continues rotating the virtual stage (and the second angular boundary).
[0104] Looking at Figure 10B, it can be understood that method 1000 returns to step 1010 to determine again whether or not to enter the dynamic phase. If the current sample does not meet the predetermined conditions, the audio frame is maintained at its current position. Its current position after exiting the dynamic phase is the position to which the audio frame arrived for the dynamic phase. If the current one meets at least one predetermined condition, the virtual sound stage is rotated, and each sample continues to rotate until at least one second predetermined condition is reached, at which point method returns to 1010.
[0105] Although not shown in Method 1000, there are additional conditions for completely exiting the head tracking phase. This can be achieved, for example, by returning each sample or periodically returning to step 1004 to determine whether the requirements for the head-fixed mode are still met. If it is determined that the requirements are no longer met, the method may enter room-fixed mode or some other mode.
[0106] The functions or parts thereof described herein, and various modifications thereof (hereinafter, “Functions”) may be implemented, at least in part, through computer program products, for example, computer programs tangibly embodied in information carriers such as one or more non-temporary machine-readable media or storage devices for execution by or control thereof by one or more data processing devices, for example, programmable processors, computers, multiple computers, and / or programmable logical components.
[0107] Computer programs can be written in any form of programming language, including compiled or interpreted languages, and can be deployed as standalone programs or in any form, including modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs can be deployed to run on one computer or on multiple computers at one location, or they can be distributed across multiple locations and interconnected by a network.
[0108] The operations associated with implementing all or part of the functionality may be carried out by one or more programmable processors that execute one or more computer programs to perform the functions of the calibration process. All or part of the functionality may be implemented as special-purpose logic circuits, such as FPGAs and / or ASICs (Application-Specific Integrated Circuits).
[0109] Examples of processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Generally, a processor receives instructions and data from read-only memory, random-access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.
[0110] While several embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily recall various other means and / or structures for performing and / or achieving functions, and / or results, and one or more advantages described herein, and each of such variations and / or modifications will be considered to fall within the scope of the embodiments of the present invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are illustrative, and that actual parameters, dimensions, materials, and / or configurations will depend on the specific application or the application in which the teachings of the present invention are used. Those skilled in the art will be able to recognize or confirm many equivalents to the specific embodiments of the present invention described herein using only ordinary experiments. Therefore, it should be understood that the embodiments described herein are presented merely as examples, and embodiments of the present invention can be practiced in ways different from those specifically described and claimed, within the scope of the appended claims and their equivalents. Embodiments of the present invention in this disclosure relate to each of the individual features, systems, articles, materials, and / or methods described herein. Furthermore, any combination of two or more such features, systems, articles, materials, and / or methods is included within the scope of the invention of this disclosure, provided that such features, systems, articles, materials, and / or methods are not mutually inconsistent.
Claims
1. A pair of headphones, A sensor that outputs a sensor signal representing the orientation of the user's head, A controller that receives the sensor signal, and is programmed to output a spatialized audio signal to a pair of electroacoustic converters for conversion to a spatialized sound signal based on the sensor signal, The spatialized acoustic signal is perceived by the user as originating from a virtual sound stage containing at least one virtual sound source, each virtual sound source in the virtual sound stage is positioned at a different location from the position of the electroacoustic transducer, and is perceived relative to the audio frame of the virtual sound stage, the audio frame being positioned at a first position aligned with the reference axis of the user's head, and the controller Equipped with, The controller is further programmed to determine from the sensor signal whether the characteristics of the user's head satisfy at least one predetermined condition, the at least one predetermined condition including whether the orientation of the user's head is outside a predetermined angular boundary. If it is determined that the characteristics of the user's head do not satisfy the at least one predetermined condition, the controller is programmed to maintain the audio frame in the first position. If the controller determines that the orientation of the user's head is outside the predetermined angular boundary, it is programmed to rotate the position of the audio frame around the rotation axis to reduce the angular offset of the user's head from the reference axis. A pair of headphones.
2. The pair of headphones according to claim 1, wherein rotating the position of the audio frame includes rotating the position of the audio frame so that it aligns with the reference axis of the user's head when the user's head turns out.
3. The pair of headphones according to claim 2, wherein, while the user's head has increasing angular acceleration, the angular velocity of rotation of the audio frame is at least partially based on the angular velocity of the user's head.
4. A pair of headphones according to claim 3, wherein, while the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that the audio frame aligns with the predicted position of the reference axis of the user's head when the turn of the user's head is completed.
5. The pair of headphones according to claim 4, wherein the predicted position of the reference axis of the user's head when the turn of the user's head is completed is updated with each sample in which the user's head has decreasing angular acceleration.
6. The pair of headphones according to claim 1, wherein the at least one predetermined condition further includes whether the angle jerk of the user's head exceeds a predetermined threshold, and the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angle boundary or the angle jerk of the user's head exceeds the predetermined threshold.
7. A pair of headphones according to claim 1, wherein, upon determining that the orientation of the user's head is outside the predetermined angular boundary, the controller is programmed to rotate the angular boundary together with the audio frame around the rotation axis, the position of the audio frame being rotated around the rotation axis for a predetermined period after the orientation of the user's head is again inside the angular boundary, in order to reduce the angular offset of the user's head with respect to the reference axis.
8. The pair of headphones according to claim 1, wherein, when it is determined that the orientation of the user's head is outside the predetermined angular boundary, a second angular boundary narrower than the angular boundary is rotated together with the audio frame around the rotation axis, and the position of the audio frame is rotated around the rotation axis for a predetermined period of time until the orientation of the user's head is within the second angular boundary, in order to reduce the angular offset of the user's head with respect to the reference axis.
9. The pair of headphones according to claim 1, wherein the at least one virtual sound source includes a first virtual sound source and a second virtual sound source, the first virtual sound source being positioned at a first position, the second virtual sound source being positioned at a second position, and the first and second positions being relative to the audio frame.
10. Maintaining the audio frame in a first position includes rotating the audio frame at a rate adjusted to eliminate sensor drift, the pair of headphones according to claim 1.
11. The pair of headphones according to claim 1, wherein the sensor that outputs the sensor signal includes a plurality of sensors that output a plurality of signals.
12. A method for providing spatialized audio, Based on a sensor signal representing the orientation of the user's head, a spatialized audio signal is output to a pair of electroacoustic transducers for conversion into a spatialized audio signal, wherein the spatialized audio signal is perceived by the user as originating from a virtual sound stage including at least one virtual sound source, each virtual sound source in the virtual sound stage is positioned at a different position from the position of the electroacoustic transducers, is perceived with respect to an audio frame of the virtual sound stage, and the audio frame is positioned at a first position aligned with the reference axis of the user's head. The method involves determining from the sensor signal whether the characteristics of the user's head satisfy at least one predetermined condition, wherein the at least one predetermined condition includes whether the orientation of the user's head is outside a predetermined angular boundary. If it is determined that the orientation of the user's head is outside the predetermined angular boundary, the position of the audio frame is rotated around the rotation axis in order to reduce the angular offset of the user's head with respect to the reference axis, Methods that include...
13. The method according to claim 12, wherein rotating the position of the audio frame includes rotating the position of the audio frame so that it aligns with the reference axis of the user's head when the user's head turns out.
14. The method according to claim 13, wherein, while the user's head has an increasing angular acceleration, the angular velocity of the rotation of the audio frame is at least partially based on the angular velocity of the user's head.
15. The method according to claim 14, wherein, while the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that the audio frame aligns with the predicted position of the reference axis of the user's head when the turn of the user's head is completed.
16. The method according to claim 15, wherein the predicted position of the reference axis of the user's head when the turn of the user's head is completed is updated for each sample in which the user's head has decreasing angular acceleration.
17. The method according to claim 12, wherein the at least one predetermined condition further includes whether the angle jerk of the user's head exceeds a predetermined threshold, and the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angle boundary or the angle jerk of the user's head exceeds the predetermined threshold.
18. The method according to claim 12, wherein, when it is determined that the orientation of the user's head is outside a predetermined angular boundary, the angular boundary rotates with the audio frame around the rotation axis, and the position of the audio frame is rotated around the rotation axis for a predetermined period of time after the orientation of the user's head is inside the angular boundary to reduce the angular offset of the user's head with respect to the reference axis.
19. The method according to claim 12, wherein, if it is determined that the orientation of the user's head is outside the predetermined angular boundary, a second angular boundary narrower than the angular boundary is rotated together with the audio frame around the rotation axis, and the position of the audio frame is rotated around the rotation axis for a predetermined period of time until the orientation of the user's head is within the second angular boundary, in order to reduce the angular offset of the user's head with respect to the reference axis.
20. The method according to claim 12, wherein the at least one virtual sound source includes a first virtual sound source and a second virtual sound source, the first virtual sound source is positioned at a first position, the second virtual sound source is positioned at a second position, and the first position and the second position are relative to the audio frame.
21. The method according to claim 12, wherein maintaining the audio frame in a first position includes rotating the audio frame at a rate adjusted to eliminate sensor drift.
22. The method according to claim 12, wherein the sensor that outputs the sensor signal includes a plurality of sensors that output a plurality of signals.