Spatial audio implemented with dynamic head tracking

Through head tracking technology in the headphone, the position of the virtual sound field is dynamically adjusted to follow the user's head rotation, solving the problem of poor spatial audio experience in frequent head rotation activities, and achieving the stability and pleasure of maintaining auditory illusions during activities.

CN120569983APending Publication Date: 2025-08-29BOSE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480008675.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-10
Filing Date
2024-02-05
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Existing headphones fail to provide a consistent and pleasant spatial audio experience when users perform frequent head rotation activities, resulting in the auditory illusion of audio perception that is derived from the speakers and destroys the spatial audio.

Method used

Through head tracking technology, the position of the virtual sound field is dynamically adjusted to follow the user's head rotation, the sensor detects the head orientation and adjusts the rotation speed and direction of the audio frame under predetermined conditions, keeping the virtual sound source aligned with the reference axis of the user's head.

Benefits of technology

The auditory illusion of maintaining spatial audio during frequent head rotation activities is achieved, providing a consistent and pleasant audio experience, avoiding the collapse of audio perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120569983A_ABST
    Figure CN120569983A_ABST
Patent Text Reader

Abstract

An apparatus and method for providing spatialized audio implemented with dynamic head tracking, the apparatus and method comprising: in a static phase, providing a spatialized acoustic signal to a user, the spatialized acoustic signal being perceived as originating from a virtual sound field at a first location; and upon determining that one or more predetermined conditions are satisfied, rotating the virtual sound field to track movement of the user's head, the one or more predetermined conditions may include whether the user's head has exceeded an angular limit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. non-provisional patent application serial number 18 / 182,290, filed on March 10, 2023, and entitled “Spatialized Audio With Dynamic Head Tracking,” which is incorporated herein by reference in its entirety. Background Art

[0003] The present disclosure generally relates to systems and methods for providing spatialized audio implemented with dynamic head tracking. Summary of the Invention

[0004] All examples and features mentioned below can be combined in any technically possible way.

[0005] According to one aspect, a pair of headphones includes: a sensor that outputs a sensor signal representing the orientation of a user's head; and a controller that receives the sensor signal and is programmed to output a spatialized audio signal to a pair of electroacoustic transducers for conversion into a spatialized acoustic signal based on the sensor signal, wherein the spatialized acoustic signal is perceived by the user as originating from a virtual sound field including at least one virtual source, each virtual source of the virtual sound field being perceived as being located at a respective position different from a position of the electroacoustic transducer and referenced to an audio frame of the virtual sound field, the audio frame being arranged to be aligned with a reference axis of the user's head. at a first position; wherein the controller is further programmed to determine, based on the sensor signal, whether a characteristic of the user's head satisfies at least one predetermined condition, the at least one predetermined condition including whether the orientation of the user's head is outside a predetermined angular limit, wherein, upon determining that the characteristic of the user's head does not satisfy the at least one predetermined condition, the controller is programmed to maintain the audio frame at the first position, wherein upon determining that the orientation of the user's head is outside the predetermined angular limit, the controller is programmed to rotate the position of the audio frame about a rotation axis to reduce an angular offset from the reference axis of the user's head.

[0006] In an example, rotating the position of the audio frame includes rotating the position of the audio frame to align with the reference axis of the user's head when the rotation of the user's head ends.

[0007] In an example, when the user's head has increased angular acceleration, the angular velocity of the rotation of the audio frame is based at least in part on the angular velocity of the user's head.

[0008] In an example, when the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that when the rotation of the user's head ends, the audio frame will be aligned with the predicted position of the reference axis of the user's head.

[0009] In an example, the predicted position of the reference axis of the user's head is updated at each sampling time when the user's head has a decreasing angular acceleration when the rotation of the user's head ends.

[0010] In an example, the at least one predetermined condition also includes whether the angular jerk of the user's head exceeds a predetermined threshold, wherein the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angular limit or the angular jerk of the user's head exceeds the predetermined threshold.

[0011] In an example, upon determining that the orientation of the user's head is outside the predetermined angular limit, the controller is programmed to rotate the angular limit along with the audio frame about the rotation axis, wherein after the orientation of the user's head is within the angular limit, the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period.

[0012] In an example, when it is determined that the orientation of the user's head is outside the predetermined angular limit, a second angular limit narrower than the angular limit is rotated about the rotation axis together with the audio frame, wherein the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period until the orientation of the user's head is within the second angular limit.

[0013] In an example, the at least one virtual source includes a first virtual source and a second virtual source, the first virtual source being disposed at a first position and the second virtual source being disposed at a second position, wherein the first position and the second position are referenced to the audio frame.

[0014] In an example, maintaining the audio frame at the first position includes rotating the audio frame at a rate modulated to cancel drift of the sensor.

[0015] In an example, the sensor that outputs the sensor signal includes a plurality of sensors that output a plurality of signals.

[0016] According to another example, a method for providing spatialized audio includes: outputting a spatialized audio signal to a pair of electroacoustic transducers based on a sensor signal representing the orientation of a user's head to convert into a spatialized acoustic signal, wherein the spatialized acoustic signal is perceived by the user as originating from a virtual sound field including at least one virtual source, each virtual source of the virtual sound field being perceived as being located at a corresponding position different from the position of the electroacoustic transducer and referring to an audio frame of the virtual sound field, the audio frame being set at a first position aligned with a reference axis of the user's head; determining whether a characteristic of the user's head satisfies at least one predetermined condition based on the sensor signal, the at least one predetermined condition including whether the orientation of the user's head is outside a predetermined angular limit, and when it is determined that the orientation of the user's head is outside the predetermined angular limit, rotating the position of the audio frame around a rotation axis to reduce the angular offset from the reference axis of the user's head.

[0017] In an example, rotating the position of the audio frame includes rotating the position of the audio frame to align with the reference axis of the user's head when the rotation of the user's head ends.

[0018] In an example, when the user's head has increased angular acceleration, the angular velocity of the rotation of the audio frame is based at least in part on the angular velocity of the user's head.

[0019] In an example, when the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that when the rotation of the user's head ends, the audio frame will be aligned with the predicted position of the reference axis of the user's head.

[0020] In an example, the predicted position of the reference axis of the user's head is updated at each sampling time when the user's head has a decreasing angular acceleration when the rotation of the user's head ends.

[0021] In an example, the at least one predetermined condition also includes whether the angular jerkiness of the user's head exceeds a predetermined threshold, wherein the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angular limit or the angular jerkiness of the user's head exceeds the predetermined threshold.

[0022] In an example, upon determining that the orientation of the user's head is outside the predetermined angular limit, the angular limit is rotated along with the audio frame about the rotation axis, wherein after the orientation of the user's head is within the angular limit, the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period.

[0023] In an example, when it is determined that the orientation of the user's head is outside the predetermined angular limit, a second angular limit narrower than the angular limit is rotated about the rotation axis together with the audio frame, wherein the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period until the orientation of the user's head is within the second angular limit.

[0024] In an example, wherein the at least one virtual source includes a first virtual source and a second virtual source, the first virtual source is set at a first position, and the second virtual source is set at a second position, wherein the first position and the second position are referenced to the audio frame.

[0025] In an example, wherein maintaining the audio frame at the first position includes rotating the audio frame at a rate modulated to cancel drift of the sensor.

[0026] In an example, the sensor that outputs the sensor signal includes a plurality of sensors that output a plurality of signals.

[0027] The details of one or more implementations are discussed in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In the drawings, like reference characters generally refer to the same parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the various aspects.

[0029] Figure 1 Depicted is a block diagram of a pair of headphones configured to provide spatialized audio, according to an example.

[0030] Figure 2 Depicted is a top view of a user's head and the virtual sound field perceived by the user at two different points in time, according to an example.

[0031] Figure 3 Depicted is a top view of a user's head wearing a pair of headphones, a virtual sound field perceived by the user, and angular bounds for a dynamic stage for entering a spatialized audio mode, according to an example.

[0032] Figures 4A to 4F Depicted is a top view of a user's head wearing a pair of headphones, the virtual sound field perceived by the user, and angular bounds for dynamic stages entering spatialized audio modes at various points of head rotation, according to an example.

[0033] Figures 5A to 5CDepicted is a top view of a user's head wearing a pair of headphones, the virtual sound field perceived by the user, and the angular bounds of the dynamic phase for entering and exiting a spatialized audio mode at various points of head rotation, according to an example.

[0034] Figure 6 Depicted is a top view of a user's head wearing a pair of headphones, a virtual sound field perceived by the user, and angular jerkiness and thresholds for entering a dynamic phase of a spatialized audio mode, according to an example.

[0035] 7A to 7C Depicted is a top view of a user's head wearing a pair of headphones, the virtual sound field perceived by the user, and angular jerkiness and thresholds for entering a dynamic phase of a spatialized audio mode at various points of head rotation, according to an example.

[0036] Figures 8A to 8C Depicted is a top view of a user's head wearing a pair of headphones, the virtual sound field perceived by the user, the angular acceleration of the user's head, and the angular velocity of the virtual sound field at various points of head rotation, according to an example.

[0037] Figure 9 Depicted are timing diagrams of various signals and values ​​associated with rotation of a user's head and resulting dynamic head tracking of a virtual sound field, according to an example.

[0038] 10A to 10F A flow chart of a method 1000 of providing spatialized audio with a virtual sound field that dynamically tracks a user's head movements is depicted. DETAILED DESCRIPTION

[0039] Current headsets that provide spatialized audio fail to provide a consistent and enjoyable experience for users engaging in activities that involve frequent head movements, such as walking, running, cycling, and the like. To address this issue, such headsets typically offer a "fixed mode" in which the audio is "fixed to the user's head," meaning the audio is always rendered in front of the user (i.e., without head tracking). While this effectively addresses the issues encountered with spatialized audio when engaging in these activities, it also fails to provide true spatialized audio. Rendering audio in front of the user's head effectively destroys the auditory illusion of spatialized audio, resulting in a "collapse" of the externalized audio inside the user's head, meaning the user perceives the audio as originating from the speakers.

[0040] Therefore, there is a need for a headset that can provide spatialized audio through head tracking in a manner that maintains a consistent and enjoyable experience for users engaging in activities that require frequent head movement.

[0041] Figure 1FIGURE 1 shows an example pair of headphones 100 configured to generate spatialized audio signals that produce a virtual sound field that dynamically tracks the rotation of a user's head when certain predetermined conditions are met. In the example shown, the headphones 100 include ear cups 102, 104. Each ear cup 102, 104 includes an electroacoustic transducer 106, 108 (also known as a speaker) for converting received signals into acoustic signals. The ear cup 102 also houses a controller 110 and a sensor 112 configured to generate a sensor signal representing the orientation of the user's head. Based on the sensor signal, the controller 110 generates a spatialized audio signal to the electroacoustic transducers 106, 108 based on the received audio signal (such as music or spoken content). The electroacoustic transducers convert the spatialized audio signal into a spatialized acoustic signal that is perceived by the user as originating from one or more locations in space that are different from the locations of the electroacoustic transducers. (In this example, the controller 110 is connected to the electroacoustic transducers 106, 108 via wires extending through the headband 114.) The received audio signals may include any suitable audio source, including multi-channel and / or object audio, as well as general audio content (not limited to music), and audio for general contexts (audio for video, communication, gaming, etc.).

[0042] For simplicity and to emphasize the more relevant aspects of the headset 100, Figure 1 Certain features of the block diagram, such as, for example, Bluetooth system-on-chip, battery, etc. In addition, although Figure 1 A pair of circumaural headphones are depicted, but it will be appreciated that the headphones 100 may be any suitable form factor, including in-ear headphones, supra-aural headphones, open-back headphones, earbuds, and the like.

[0043] The controller 110 includes a processor 116 and a memory 118 storing program code executed by the processor 116 to perform various functions for providing spatialized audio as described herein, including, where appropriate, the steps of method 1000 described below. It should be understood that the processor 116 and memory 118 of the controller 110 need not be located within the same housing (such as part of an application-specific integrated circuit), but may instead be located in separate housings. Furthermore, the controller 110 may include multiple physically distinct memories to store the program code required for its operation and may include multiple processors 116 for executing the program code. Furthermore, the various components of the controller 110 need not be located within the same earcup (or other matching parts within a headphone form factor), but may be distributed between the earcups. For example, each of the earcups 102 and 104 may include a processor and memory that work in conjunction to perform the various functions for providing spatialized audio as described herein, with the processors and memories within both earcups 102 and 104 forming the controller.

[0044] As described above, sensor 112 generates a sensor signal representing the orientation of the user's head. In the example, sensor 112 is an inertial measurement unit for head tracking; however, it should be understood that sensor 112 can be implemented as any sensor suitable for measuring the orientation of the user's head. In addition, sensor 112 may include multiple sensors that work together to generate the sensor signal. In practice, the inertial measurement unit itself typically includes multiple sensors (e.g., accelerometers, gyroscopes, and / or magnetometers) that work together to generate the sensor signal. The sensor signal representing the orientation of the user's head can be a data signal that directly represents the orientation, for example, as a change in pitch, roll, and yaw, or can include other data from which the orientation can be derived, such as specific forces and angular rates of the user's head. In addition, the sensor signal itself can be composed of multiple sensor signals, such as where multiple separate sensors are used to measure the orientation of the user's head. In an example, separate inertial measurement units may be provided in ear cups 102, 104 (or other matching parts in a headphone form factor) or any other suitable location and together form sensor 112, and signals from the separate inertial measurement units form the sensor signal.

[0045] The controller 110 may be configured to generate audio signals in one or more modes. These modes include, for example, an active noise reduction mode or a hear-through mode. In addition, the controller 110 may generate spatialized audio in a mode in which the virtualized sound field is fixed in space (e.g., in front of the user) and the perceived position of the virtualized sound field does not change in response to movement of the user's head, or only changes in response to the user spending a predetermined period of time facing a direction that is rotated more than a predetermined angle away from the virtual sound field. For the purposes of this disclosure, this will be referred to as a "room-fixed" mode, referring to the fact that the virtual sound field is perceived as fixed in place in the room. Additional details regarding room-fixed mode are described in the following patent applications: U.S. patent application serial number 16 / 592,454, entitled “SYSTEMS AND METHODS FOR SOUND SOURCE VIRTUALIZATION,” filed on October 3, 2019, and published as U.S. Patent Application Publication No. 2020 / 0037097; and U.S. patent application serial number 63 / 415,783, entitled “SCENE RECENTERING,” filed on October 13, 2022, the entire disclosures of which are incorporated herein by reference. As described above, room-fixed mode is best suited for relatively stationary users (such as sitting at a desk) and not for active users (e.g., walking or running). To address this issue, the controller 110 is programmed to operate in a "head-fixed" mode, either by user selection or by some trigger condition, which keeps the virtual sound field at a fixed point in space (similar to the room-fixed mode, referred to in this disclosure as the "static phase" of the head-fixed mode), until a predetermined condition is met, at which point the virtual sound field can dynamically rotate with the rotation of the user's head (referred to in this disclosure as the "dynamic phase" of the head-fixed mode). Figure 2 to Figure 1 0 describes the details of the head fixation mode in more detail.

[0046] Go to Figure 2 , showing the user's head 202 moving from a first orientation at time t0 to a second orientation at time t1, indicated as 202'. Figure 21 so that various features and angles can be more clearly seen, it should be understood that the headset 100 is worn on the user's head 202 at both times t0 and t1. The user perceives a virtual sound field 204 based on the spatialized acoustic signal generated by the headset 100. The virtual sound field 204 includes one or more virtualized speakers (also referred to as "virtualized sources"), which are depicted here as virtualized speakers 206, 208. Although two virtual speakers are shown, it should be understood that in various alternative examples, any number of virtualized speakers may be created from the spatialized acoustic signal (according to the spatialized audio signal generated by the controller 110). Furthermore, although the virtualized speakers 206, 208 are shown as being symmetrically disposed in front of the user's head 202 (i.e., about the longitudinal axis, in FIG. 1 ), the virtualized speakers 206, 208 are symmetrically disposed in front of the user's head 202 (i.e., about the longitudinal axis, in FIG. 1 ). Figure 2 , extending from the Z axis to point P), but it should be understood that in various examples, the virtualized speakers may be asymmetrically distributed relative to the front of the user's head. The production of virtualized speakers is generally understood, so a more detailed description will be omitted here.

[0047] For the purposes of this disclosure, and for simplicity, the positions of virtual sound fields and virtual speakers will generally be discussed as if the virtual sound field or virtual speakers were physically located at a given location. Even if not explicitly stated, it should be understood that the positions of the virtual sound field or virtual speakers are merely perceived positions. In other words, to the extent that a virtual sound field or virtual speaker is described as having a physical position in space, that physical position refers only to the perceived position of the virtual sound field or virtual speaker, not to its actual position in space.

[0048] The positions of the virtualized speakers 206, 208 may be positions where the virtualized speakers 206, 208 are initialized. Initialization of the speakers may occur, for example, when a user first selects a spatialized audio mode (such as a room-fixed mode or a head-fixed mode) and may be initially placed in front of the user (e.g., Figure 2 ), but it is contemplated that the virtualized speakers 206, 208 may be initialized elsewhere. Indeed, it is contemplated that the initial positions (and number) of the virtualized speakers may be selected by the user, such as through a dedicated mobile application or a web interface accessible via a computer or mobile device.

[0049] As described above, after the virtualized speakers are initialized, the controller 110 maintains the virtual sound field 204 in the same position in the static phase until one or more predetermined conditions are met. Once at least one of the predetermined conditions is met, the controller 110 may adjust the spatialized audio signal in the dynamic phase so that the virtual sound field 204 rotates to track the movement of the user's head. The movement of the virtual sound field 204 is Figure 2, is represented by the rotation of the virtual sound field 204, which is depicted as virtual sound field 204'. More specifically, when the user's head 202 rotates to the second orientation at t1, the virtual sound field 204 tracks the position of the user's head 202, but typically with some lag. Thus, at time t1, the virtual sound field 204' is trailing the user's head 202' at a certain angle. Some lag here is desirable because it maintains the illusion of the virtual sound field. If the virtual sound field perfectly tracked the user's head without any lag, the perception of the virtual sound field would "collapse" inside the user's head because there would no longer be a distinction between the movement of the user's head and the perceived movement of the virtual speakers.

[0050] The rotation of the virtual sound field 204 is achieved by rotating each virtual speaker 206, 208 (because the virtual sound field 204 is completely composed of a collection of virtual speakers). In other words, when the virtual sound field 204 is angularly rotated about the Z axis by an angle α from time t0 to time t1 c Each virtual speaker 206, 208 is also rotated around the Z axis by an angle α from its initial position at time t0. c to its position at time t1. However, for the purposes of this disclosure, the rotation of the virtual sound field 204 is about a single reference point (at Figure 2 2 and is referred to as an "audio frame"). Thus, the virtual sound field that dynamically tracks the user's head is described with respect to the rotation of audio frame A as it follows the rotation of the user's head 200. In the example shown, audio frame A is initialized to be aligned with the longitudinal axis ZP of the user's head 202. However, since audio frame A is a reference point for describing the rotation of the virtual sound field 204, it should be understood that any suitable reference axis may be used, i.e., the angular offset between the user's head 202 and the virtual sound field 204 may be measured from the reference axis.

[0051] The virtual sound field 204 (and therefore each virtual speaker 206, 208) rotates angularly around the Z axis. Typically, the Z axis corresponds to the axis about which the user's head rotates, otherwise the virtual sound field 204 would not be perceived as remaining at a fixed distance from the user throughout the rotation. In practice, this axis can be approximated by any point where the distance changes imperceptibly to the user during the rotation.

[0052] The virtual sound field 204 tracks the user's head movements by rotating the virtual sound field 204 to reduce the angular offset between the virtual sound field 204 and the orientation of the user's head. Figure 2 is shown as the offset angle α off , the offset angle is the rotation angle α of the user's head 202 turn(ie, the angle between the orientation of the user's head 202 at time t0 and the orientation of the user's head 202' at time t1) and the rotation angle α of the virtual sound field c (ie, the angle between the orientation of the virtual sound field 204 at time t0 and the orientation of the virtual sound field 204' at time t1). Therefore, the direction of movement of the virtual sound field 204 is selected to reduce the rotation angle α turn With rotation angle α c Large angle offset between.

[0053] Once at least one of the predetermined conditions is met, the virtual sound field starts to be transmitted at μ per second. dps degrees toward the current orientation of the user's head 202 (μ dps The selection of the value of can be a dynamic process and will be described in more detail below). The value of can be selected by performing a t-test on each sampling time μ from time t0 to t1. dps The value of is integrated to give the rotation angle α c .

[0054]

[0055] Similarly, the user's head angular velocity Ω measured by the controller 110 at each sampling from time t0 to t1 based on the input from the sensor 112 can be calculated. z (in degrees / second) to find the rotation angle α of the user's head 202 turn .

[0056]

[0057] Therefore, the angular velocity of the user's head 202 from time t0 to t1 and the angular velocity μ of the sound field 204 can be calculated by dps The angular offset α at time t is found by integrating the difference between off .

[0058]

[0059] Equation (3) can be rewritten as the offset of time t0 and the rotation angle α of the user's head 202 from time t0 to t1: turn The rotation angle α with respect to the virtual sound field 204 c The sum of the differences between .

[0060] α off (t1) = a off (t0)+α turn (t0,t1)-α c (t0,t1) (4)

[0061] As described above, when certain predetermined conditions are not met, the virtual sound field 204 remains fixed in space. Such predetermined conditions can be, for example, angular limits (such as a wedge or cone) set around the user's head to determine whether the user's head has rotated beyond a predetermined maximum, and whether the angular jerk of the user's head exceeds a threshold to determine whether the user's head is rotating rapidly, which is an early indication that the rotation will exceed the angular limits. Other predetermined conditions are also contemplated and within the scope of the present disclosure.

[0062] Go to Figure 3 , showing the first of such predetermined conditions (shown as angle α max Angle limit α max Depicted as a two-dimensional cone (also called a wedge) that corresponds to whether the user's head has traveled the maximum allowed yaw (i.e., rotated to the greatest possible extent) before the virtual sound field 204 is adjusted to realign the audio frame A with the longitudinal axis ZP of the user's head. max A maximum value of rotation of, for example, 50° may be covered, although its width is a design choice and, in other examples, may have other suitable values. In this example, the apex of the two-dimensional cone is at the rotation axis Z, so that the longitudinal axis ZP is completely within the angular limit α max within or completely outside the angular limits.

[0063] Although Figure 3 The longitudinal axis ZP is depicted as being used to determine whether the orientation of the user's head has exceeded an angular limit α. max The reference axis of the user's head may be any suitable reference axis or reference point for comparing the orientation of the user's head to the angular limits. In one example, a reference point P representing the front of the user's head may be used instead of the longitudinal axis ZP. However, the reference point need not be located on the longitudinal axis ZP, nor need it be used in the same manner as used to compare the orientation to the angular limits α. max The same reference axis is used for comparison to determine the offset or alignment of the virtual sound field 204. However, if a different reference axis or reference point for the longitudinal axis is used, the angle limit α may need to be adjusted. max position or orientation to resolve differences in the initial direction or position of the reference axes / points used.

[0064] However, using the same reference axis for the offset of the virtual sound field 204 and detecting when the orientation of the user's head exceeds an angular limit allows the use of an angular offset α off As a proxy for the reference axis of the user's head 202. In other words, if the same reference axis is used to determine the alignment of the virtual sound field 204 with the user's head 202 and to determine whether the user's head 202 exceeds the angular limit α max, then when the user's head 202 is kept within the angle limit α max When the angle is offset α off Equal to the angle limit α max The distance from the center of the axis is thus reduced, allowing for some computational economy. (This assumes that the angular limits are set symmetrically around the reference axis, although this is often the case.)

[0065] It should be further understood that other suitable shapes of angle limits may be used. For example, a three-dimensional cone may be used instead of a two-dimensional cone to determine whether the pitch of the user's head 202 exceeds the limit in the vertical dimension (ie, the pitch of the user's head).

[0066] When the user's head is kept at the maximum angle limit α max When the user is in the room, the controller 110 operates in the static phase of the head-fixed mode, which means that the user perceives the virtual sound field 204 as being fixed in space. This is equivalent to operating in the room-fixed mode described above. Generally speaking, during this period, the angular velocity μ of the virtual sound field 204 is dps is 0 degrees / second or kept at a very low value to compensate for drift in the sensor 112 (so the user will not perceive any movement in the virtual sound field 204).

[0067] After determining the user's head is within the angle limit α max When the angular velocity μ of the virtual sound field 204 is dps When reducing the angle offset α off , i.e., in the direction of the angular velocity of the user's head 202. Figures 4A to 4F As shown in more detail, Figures 4A to 4F Together they depict the rotation of the user's head, and subsequently the response of the virtual sound field 204. Figure 4A , the orientation of the user's head (as represented by the longitudinal axis ZP) is pointing upward on the page and at the angle limit α max , and thus the controller 110 remains in the static phase and the virtual sound field 204 remains in its fixed position. Figure 4B In the example, the user's head has begun to turn left on the page, but the longitudinal axis ZP remains within the angle limit α max , so the controller 110 remains in the static phase, and the user continues to perceive the virtual sound field 204 as fixed in the same location. (It will be appreciated that the spatialized audio signal needs to be adjusted based on changes in the orientation of the user's head 202 so that the virtual sound field 204 is perceived as existing in the same location in space.)

[0068] exist Figure 4C , the longitudinal axis ZP has exited the angular limit α max, and in response, the controller 110 enters the dynamic phase of the head fixed mode. However, Figure 4C represents the first sample where the orientation of the user's head is measured to exceed the angle limit α max , and therefore the virtual sound field 204 has not yet started tracking the movement of the user's head 202. Figure 4D As shown, the user's head 202 continues to turn to the left, and the virtual sound field 204 has begun to track the movement of the user's head (similarly shifting to the left), although there is some lag between the angle of the audio frame A and the longitudinal axis ZP, namely the offset angle α off . Angle limit α max Similarly, the virtual sound field 204 rotates around the Z axis by the same rotation angle α c .exist Figure 4E In the figure, the angle limit α max Continuing its rotation, it has caught up with and surpassed the longitudinal axis ZP, so that the longitudinal axis ZP is again at the angular limit α max .exist Figure 4F In the sound field (and angle limit αmax ) has reached the end position of the user's head rotation, and therefore the virtual sound field 204 is aligned with the longitudinal axis ZP again, presenting an offset angle α off 0° or within some predetermined tolerance. (Typically, for the purposes of this disclosure, to be considered “aligned,” the offset need only be within a predetermined degree, which is a design choice that dictates how closely the virtual sound field 204 is aligned with the user’s head 202. Typically, the predetermined degree is chosen to be imperceptible to the user, so as to preserve the perception that the audio frame has been adjusted to its previous position relative to the user’s head.)

[0069] The controller will continue in the dynamic phase until a second predetermined condition is met. In one example of such a second predetermined condition, when the user's head returns to the angle limit α max Thereafter, the virtual sound field 204 tracks the movement of the user's head 202 for a predetermined length of time. Figure 4E As shown, the longitudinal axis ZP has just returned to the angular limit α max (Because the angle limit α max The predetermined time period is initiated before tracking of the user's head 202 stops and the virtual sound field 204 is fixed in position (i.e., enters a static phase). In one example, the predetermined time period may be 0.5 seconds, although other suitable time lengths may be used. The predetermined time period may be selected so as to allow continued tracking of the user's head 202 until the virtual sound field 204 is aligned with the longitudinal axis ZP again, as in Figure 4F shown.

[0070] In an alternative example, the second predetermined condition may be greater than the angle limit αmax A narrow, single angular limit is established to determine when the virtual sound field 204 is aligned with the longitudinal axis ZP. In other words, in this example, the angular limit α max The narrower angular bounds are used to determine when the virtual sound field 204 begins tracking the user's head 202 (i.e., enters the dynamic phase), but the narrower angular bounds are used to determine when the virtual sound field 204 stops tracking the user's head 202 and becomes fixed in space again (i.e., enters the static phase). Figure 5A and Figure 5B An example of this is shown in Figure 5A and Figure 5B Shows that the user's head has turned to the left (as combined with Figures 4A to 4F described), the virtual sound field 204 and the angle limit α max and a narrower α min In response to the user's head turning beyond the angle limit α max and start tracking to the left. Figure 5A In the example, the virtual sound field 204 tracks the rotation of the user's head 202, but is not yet aligned with the longitudinal axis ZP. The longitudinal axis ZP is within the angle limit α. max and angle limit α min Beyond both.

[0071] exist Figure 5B In the example, the user's head 202 is within the angle limit α max Within but not within the angle limit α min , so the virtual sound field continues to track the movement of the user's head 202. Figure 5C In the longitudinal axis ZP, the angle limit α min , and thus the virtual sound field 204 stops tracking the user's head 202, and the controller 110 enters the static phase again. min The width of is a design choice depending on how closely the virtual sound field 204 is aligned with the longitudinal axis ZP. min The narrower it is, the more closely the virtual sound field 204 is aligned with point P in front of the user's head.

[0072] Now go to Figure 6 , shows a second example of a predetermined condition for enabling tracking of the virtual sound field 204. In this example, the controller 110 determines whether the user's head has begun to turn rapidly in one direction or another by determining whether the angular jerk (i.e., the rate of change of acceleration of the user's head) has exceeded a threshold, rather than determining whether the user's head has exceeded an angular limit. Figure 6 In FIG, the angular jerk of the user's head 202 is indicated by a curved arrow (marked as ), where the length of the curved arrow represents the value of angular jerk relative to a threshold value, represented by the dashed line labeled T. If the angular jerk exceeds the threshold value, then the user's head is turning rapidly in a certain direction, indicating that the user's head will soon exceed the angular limit. Therefore, angular jerk represents an early indication of head rotation that requires adjustment of the position of the virtual sound field 204. Angular jerk can be received directly from the sensor 112, but more typically can be calculated by comparing the change in orientation from one sample to the next; however, any suitable method for calculating angular jerk can be used.

[0073] 7A to 7C Depicted is the adjustment of the virtual sound field after detecting an angular jerk of the user's head 202 that exceeds a predetermined threshold. Figure 7A In FIG. 1 , the user's head 202 has begun to turn to the left, with the angular jerk being indicated by the curved arrow (marked as ) indicates but has not yet exceeded the threshold; therefore, the controller 110 remains in the static phase and the virtual sound field is perceived as fixed in place. Figure 7B A first sample is depicted, where the measured jerkiness exceeds a threshold value and thus the position of the virtual sound field 204 has not yet started to adjust. Figure 4C As the angular jerk exceeds the threshold value T, the position of the virtual sound field 204 begins to adjust to track the movement of the user's head 202. The virtual sound field 204 can continue to track the user's head 202 until a second predetermined condition is met. Examples of the second predetermined condition include tracking until the user's head 202 returns to the angular limit or returns to the angular limit for a predetermined period of time, as described in conjunction with Figures 4 and 5.

[0074] Generally speaking, monitoring the angular jerkiness of the user's head 202 is useful for early detection of head turns, but it will not (by design) detect slower movements, or even movements that cause the user to rotate significantly to the left or right. Therefore, the angular jerkiness of the user's head 202 is assumed to be within the angular bounds α. max The angular jerk condition and the angular jerk threshold condition are used in conjunction with each other, where the angular jerk condition exceeds a threshold or the orientation of the user's head exceeds an angular limit, which is sufficient to enter the dynamic phase and adjust the position of the virtual sound field 204; however, it is contemplated that either the angular limit condition or the angular jerk threshold condition can be used as the only predetermined condition for initiating tracking of the user's head. It should also be understood that any suitable predetermined condition for detecting or predicting a rotation of the user's head exceeding a predetermined degree can be used in addition to or as an alternative to the above two methods.

[0075] Angular velocity μ of the virtual sound field 204 dps The angular velocity Ω of the user's head can be z. Generally speaking, when sound field tracking is triggered, the goal is to quickly eliminate large head rotations (such that the user typically perceives the sound field to be primarily in front of the user's head 202) and eliminate the most unpleasant artifacts of re-centering the virtual sound field, while also allowing for some lag so that the illusion of a virtual sound field is preserved. Applicants also recognize that once the user's head stops moving, the sound field lags behind the user's head, which is generally an unpleasant experience. In other words, ideally, the audio frame A of the virtual sound field 204 should be aligned with the longitudinal axis ZP (or other reference axis) when the user's head stops. Therefore, in order to track the movement of the user's head, the controller 110 can dynamically adjust the angular velocity μ of the virtual sound field 204 based on the movement of the user's head 202 but timing the alignment of the virtual sound field 204 with the longitudinal axis ZP to coincide with the end of the head rotation. dps .

[0076] To achieve these goals, the controller 110 may dynamically adjust the angular velocity μ of the virtual sound field 204 according to two separate stages: dps : (1) when the user's head is accelerating, and (2) when the user's head is decelerating. Figures 8A to 8C Depicts the selection of μ at different stages of head rotation dps process. Figure 8A In , the user's head moves to the left and accelerates in that direction, as shown by the curved arrow (marked ) is indicated. At this stage, the angular velocity μ of the virtual sound field 204 is dps The value (expressed as μ dps,1 , for the angular velocity during the acceleration phase) based on the angular velocity Ω of the user's head z , so that the virtual sound field 204 tracks the user's head 202 while allowing a certain amount of hysteresis that allows the user to experience some spatial cues of the head turning relative to the virtual sound field 204 (and prevents the perception of the virtual sound field 204 from "collapsing"). In one example, this can be achieved according to the following equation:

[0077] μ dps,1 =f turn,acceleration *|Ω z | (5)

[0078] where f turn,acceleration Represents the scaling factor applied to the angular velocity of the user's head and is a design choice.

[0079] exist Figure 8B , the user's head is still moving to the left, but has begun to slow down, as indicated by the curved arrow pointing to the right (marked (It should be understood that deceleration is acceleration in a different direction. As described herein, deceleration is relative to the direction of the initial acceleration.) Once the user's head begins to decelerate (indicating that the user's head turn is nearing the end), the final orientation of the user's head at the end of the turn is predicted, which is indicated by dashed line 202. p is represented and used to select μ dps,2 (angular velocity in the deceleration phase) such that the virtual sound field 204 arrives in front of the user's head 202 when the head rotation is completed (i.e., the virtual sound field 204 arrives in front of the user's head 202 and the user completes the head rotation approximately simultaneously).

[0080] This can be done by predicting when the user's head turn will be complete and setting μ dps,2 is achieved so that the virtual sound field 204 passes through the remaining angular offset α between its current position and the predicted orientation of the user's head 202 at the end of the rotation off For example, if h0 represents the time from the current sample to the end of the head turn, the virtual sound field must compensate (i.e., traverse) the existing angular offset α off , and an additional angular offset α off The time accumulated from the current time t to the time t+h0 when the head turns are completed.

[0081] The angular velocity of the user's head at the future time h can be approximated by a linear approximation:

[0082]

[0083] (This linear approximation has been truncated to the second term. However, it will be appreciated that this Taylor series, as well as any other series described in this disclosure, can be extended to any number of terms.) Assuming that at the end of the user's head rotation, the angular velocity is zero (assuming the user's head has stopped), the linear approximation for t+h0 can be written as follows:

[0084]

[0085] Therefore, the linear approximation at future time h can be rewritten as:

[0086]

[0087] Therefore, the additional angle from the current time to the time t+h0 at which the head turns end can be approximated as:

[0088]

[0089] It can be assumed that the angular velocity μ of the virtual sound field 204 at the future time h is dps,2 (t) has a linear distribution Ω z (t+h0), and can therefore be written as a linear approximation:

[0090]

[0091] And therefore, from the current time to the time t+h0 at which the head turn ends, the angle compensated (traversed) by the virtual sound field 204 can be approximated as follows:

[0092]

[0093] μ can be set dps,2 The angular velocity of (t) is such that Equation 9 eliminates Equation 11 and the angular offset α existing at the current time t off In other words, μ dps,2 (t) is chosen such that:

[0094]

[0095] Solving for μ dps,2 (t) yields the angular velocity that causes the virtual sound field 204 to arrive in front of the user's head 202 when the head turn is complete:

[0096]

[0097] This angular velocity can be recalculated for each incoming sample to account for the angular velocity Ω of the user's head z Adjust for changes.

[0098] Regardless of the phase of the user's head rotation, the dynamic angular velocity may include a baseline angular velocity μ baseline , the baseline angular velocity is related to μ dps,1 and μ dps,2 The baseline angular velocity μ can be added. baseline To solve the problem of very slow head rotation or the user's head returning to the angle limit α with residual offset max The baseline angular velocity μ baseline Ensure that the virtual sound field 204 does not deviate too far from the longitudinal axis ZP (or other reference axis) or eliminate residual angular offset. In the example, the baseline angular velocity μ baseline It may be 20 degrees / second, but other suitable values ​​are also contemplated herein.

[0099] Go to Figure 9 , shows an example timing diagram of various signals and values ​​associated with the rotation of the head turn to demonstrate the detection of the conditions caused by entering the dynamic phase and the resulting response to the angular velocity μ dps Adjustment. Figure 9 Starting from the top graph of FIG, it shows the rotation angle α of the user's head 202 around the Z axis. z signal. Figure 9The top graph also depicts the angle limit α max , the angular limit is represented as a shaded horizontal limit. Figure 9 As shown, at point 1a, the rotation angle α z Exit angle limit α max , causing the controller 110 to enter the dynamic phase and change the angular velocity μ of the virtual sound field 204 dps Increments from zero to non-zero. (Although Figure 9 μ is represented in dps It will be appreciated that, in practice, the angular velocity μ may be maintained during the static phase. dps The initial dynamic phase is depicted here as the first of two vertical shaded areas, with the second shaded area representing the second dynamic phase.) Following 1a, the angular velocity μ dps Based on the angular velocity Ω of the user's head 202 z At point 2a, the angular acceleration of the user's head 202 begins to decrease, thereby signaling the end of the head rotation to the user. Based on the predicted end of the head rotation, the angular velocity μ dps The rotation angle α of the user's head 202 around the Z axis is increased suddenly so that when the user's head stops, the virtual sound field 204 will be realigned with the reference axis of the user's head (e.g., the longitudinal axis). z Returned to angle limit α max Internally, a predetermined time period is triggered, at the end of which (indicated as point 4a), the dynamic phase ends and the angular velocity μ of the virtual sound field 204 is dps Reaching zero again.

[0100] At point 1b, the second dynamic phase begins, this time due to the angular jerk of the user's head 202 exceeding the threshold t, both in Figure 9 As shown in the middle graph of FIG. , the dynamic phase begins as soon as the angular jerk of the user's head 202 exceeds the threshold value t as an early detection of the user's head rotation. Thus, the dynamic phase continues until the user's head 202 exits and returns to the angular limit α at point 3b. max (as shown in the top graph), at which point the predetermined time period begins again and ends at the end of the second dynamic phase at point 4b. During the second dynamic phase, see the angular velocity μ of the virtual sound field 204 in the bottom graph. dps , at point 1b, the angular velocity μ dps becomes non-zero again, with a value based on the angular velocity Ω of the user's head 202 z At point 2b, the user's head 202 begins to decelerate, resulting in an angular velocity μ dpsThe rapid increase in angular velocity μ is distributed in a way that allows the virtual sound field 204 to realign with the reference axis at the end of the user's head rotation. dps .

[0101] Now go to 10A to 10D , a flowchart of a method 1000 for providing spatialized audio with a virtual sound field that dynamically tracks the movement of a user's head is shown. Method 1000 can be implemented by a controller (e.g., controller 110) included in a pair of headphones (e.g., headphones 100). The controller may include one or more processors and one or more non-transitory storage media storing programs for execution by the one or more processors. For example, the controller may include two microcontrollers (including processors and memories), the two microcontrollers being respectively disposed in ear cups of the headphones, the two microcontrollers working in conjunction to perform the steps of method 1000. The headphones may also include at least one pair of electroacoustic transducers that receive audio signals from the controller and convert the audio signals into acoustic signals.

[0102] At step 1002, a sensor input signal or selection of a mode of operation is received. In examples, the headset (and in particular the controller) is operable in more than one operating mode, including different spatialized audio modes. These modes include, for example, a room-fixed mode in which the virtual sound field is perceived as fixed to a specific location in space that does not move in response to movement of the user's head unless, in some examples, the user's head has turned away from the virtual sound field for at least a predetermined period of time. In head-fixed mode, as will be described in more detail below (and as will be described in conjunction with Figures 1 to 9 As described above, the virtual sound field is fixed to a specific position in space until certain predetermined conditions are met, at which point the virtual sound field rotates to track the movement of the user's head.

[0103] The sensor input signal may be, for example, input from a sensor such as an inertial measurement unit, which may provide input indicative of a user activity (such as walking or running) for which the head tracking mode is more suitable (other suitable types of sensors are contemplated, such as accelerometers, gyroscopes, etc.). The sensor input may be received from the same sensor that detects the orientation of the user's head, or from a different sensor. In yet another example, the sensor signal may be mediated by an auxiliary device that includes a sensor, such as a mobile phone or a wearable device such as a smartwatch. Alternatively, input may be received from the user (e.g., using a dedicated application or through a web interface, or through a button or other input on the headset) to directly select between modes.

[0104] Step 1004 is a decision block indicating whether the selection of an operating mode or the sensor signal satisfies the requirements for head-fixed mode. Upon receiving input from the user selecting an operating mode, this is typically sufficient, without requiring any further action or decision. However, the sensor signal requires some analysis to determine whether it indicates activity worthy of switching to head-fixed mode. Such analysis could, for example, determine whether the user has taken a predetermined number of steps within a predetermined time period, or whether the user has completed a predetermined number of head turns within a predetermined time period (e.g., as determined by measuring a reference axis of the user's head relative to angular limits). Other suitable measures for determining whether the user is engaging in an activity that would be facilitated (i.e., made more comfortable) by implementing head-fixed mode are contemplated; indeed, numerous tests exist for identifying when a user is engaging in an activity or a particular type of activity, and any such suitable test could be used. In an alternative example, step 1004 could be performed by an auxiliary device (e.g., a mobile phone or wearable device). After analyzing the sensor signal, the auxiliary device could direct the controller to enter head-fixed mode or otherwise notify the controller that a particular activity is occurring.

[0105] If the requirements for head-fixed mode are not met, then at step 1006 the controller operates in room-fixed mode, in which the virtual sound field remains fixed, as described above, except in confined situations where the user faces different directions for extended periods of time. As mentioned above, additional details regarding room-fixed mode are described in U.S. patent application serial numbers 16 / 592,454 and 63 / 415,783, the disclosures of which are incorporated herein by reference. Furthermore, while room-fixed mode is listed as the only alternative to head-fixed mode, it should be understood that head-fixed mode can be one of any number of potential modes, which may or may not be spatialized audio modes. If the requirements for head-fixed mode are met, then the method proceeds to step 1008, as Figure 10B shown.

[0106] At step 1008, the controller outputs a spatial audio signal to a pair of electroacoustic transducers based on the sensor signal representing the orientation of the user's head to convert into a spatialized acoustic signal. The spatialized acoustic signal is perceived as originating from a virtual sound field including at least one virtual source, each of which is perceived by the user as being located at a different position than the position of the electroacoustic transducer. The virtual source also references an audio frame of the virtual sound field, which is set at a first position and aligned with a reference axis of the user's head. In other words, the audio frame serves as a singularity to describe the position and rotation of the virtual sound field. In one example, the longitudinal axis of the user's head may be used as the reference axis, however, the reference axis may be any suitable axis for determining the angular offset between the user's head and the virtual sound field when the user's head and the virtual sound field are rotated in the manner described below.

[0107] The location of the current position depends on how and when the head-fixed mode was selected in step 1004, as the spatialized audio signal may be initialized at step 1008, or the spatialized audio signal may be initialized earlier (such as in conjunction with the room-fixed mode). In the former case, the spatialized audio signal is initialized at step 1008 and is therefore determined based on user input or automatically based on the direction the user is facing in step 1008. In the latter case, the first position may depend on the position of the audio frame determined in conjunction with the room-fixed mode, which may be the position at which the room-fixed mode was initialized or the position at which the room-fixed mode was adjusted.

[0108] Step 1010 is a decision block that determines whether the characteristics of the user's head meet at least one predetermined condition. Such predetermined conditions can be, for example, whether the orientation of the user's head is outside a predetermined angular limit or whether the angular jerkiness of the user's head exceeds a predetermined threshold.

[0109] Briefly go to Figure 10C , shows an example of step 1010, including steps 1018 and 1020. Step 1018 is a decision box representing determining whether the orientation of the user's head exceeds a predetermined angular limit. The angular limit can be used to determine whether the user's head has rotated beyond a predetermined degree. In one example, the angular limit can be a two-dimensional cone, but three-dimensional cones and other suitable shapes are also conceivable. A reference axis (which can be the same or different from the reference axis for determining the alignment of the audio frame) or a point can be compared with the angular limit to determine whether the orientation of the user's head has rotated beyond a predetermined degree. Figures 4 and 5 depict examples of comparing a reference axis (here, the longitudinal axis of the user's head) with the angular limit. When it is determined that the orientation of the user's head does not exceed the predetermined angular limit, the method proceeds to step 1020. When it is determined that the orientation of the user's head exceeds the angular limit, the method proceeds to step 1014.

[0110] Step 1020 is a decision block that indicates whether the angular jerkiness of the user's head exceeds a predetermined threshold. In this example, the angular jerkiness of the user's head is compared to the threshold as early evidence of a rotation that may exceed the angular limit. Figure 6 An example of comparing the angular jerk to a threshold is depicted and described in FIG7 . If it is determined that the angular jerk of the user's head does not exceed the predetermined threshold, the method proceeds to step 1012 (i.e., continues in the static phase). If it is determined that the angular jerk of the user's head exceeds the predetermined threshold, the method proceeds to step 1014 (i.e., begins the dynamic phase).

[0111] Thus, angular jerk serves as another predetermined condition that can trigger the dynamic phase. While angular limits and angular jerk thresholds have been described, it should be understood that other examples of suitable predetermined conditions can be envisioned, i.e., other examples that indicate or predict at least a predetermined degree of head turn. It is also envisioned that only one such predetermined condition can be used, e.g., only angular limits or only angular jerk thresholds.

[0112] When it is determined that at least one predetermined condition is not met, then at step 1012, the audio frame is maintained at its current position. In other words, in the above example, when it is determined that the user's head is within the angular limits and the angular jerkiness is below the threshold, maintaining the audio frame at the current position is achieved by rendering at least one virtual source such that the at least one virtual source is perceived as fixed in space regardless of the movement of the user's head (in practice, this requires adjusting the spatialized acoustic signal in such a way that the virtual source is perceived as fixed in space based on the detected change in the orientation of the user's head). Therefore, when it is determined that the user's head is relatively stationary (for example, the user is sitting at a table), the user will perceive the virtual sound field as being fixed in space.

[0113] return Figure 10B , at least one predetermined condition is satisfied, then at step 1014, the position of the audio frame is rotated around the rotation axis to reduce the angular offset from the reference axis of the user's head. In other words, the virtual sound field can be rotated around the rotation axis of the user's head (or an axis approximately located around the rotation axis of the user's head) to reduce the offset introduced by the rotation of the user's head. Figure 10D As described, this rotation can continue until the audio frame is aligned with the reference axis. Aligning the audio frame with the reference axis includes rotating the virtual sound field by the same rotation angle as the rotation of the user's head. Therefore, if the user's head has been rotated 15° from its initial positioning, the virtual sound field is also rotated 15° to align it with the user's head. Furthermore, since the virtual sound field is composed of at least one virtual source, the rotation of the audio frame is achieved by a corresponding rotation of each virtual source. Therefore, a rotation of the virtual sound field by 15° is achieved by a rotation of each virtual source by 15°.

[0114] Temporarily go to Figure 10D , method steps are shown in steps 1022-1026, which show in more detail how the rotation of step 1014 is accomplished, and more specifically, how the rotation rate of the virtual sound field at step 1014 (and therefore the rotation rate of each virtual source) is selected. Step 1022 is a decision box indicating whether the user's head has increased angular acceleration. When it is determined that the user's head has increased angular acceleration, the method proceeds to step 1024, in which the angular velocity of the rotation of the audio frame is based at least in part on the velocity of the user's head. In an example, the angular velocity may be set to a scaled value of the angular velocity of the user's head, which is based on the acceleration of the user's head (as described in conjunction with equation (5)). The value of the angular velocity of the virtual sound field is selected to rotate the virtual sound field at a pace that follows the user's head but allows a certain amount of lag so that the illusion of the virtual sound field does not collapse. However, the angular velocity of the virtual sound field may be further added to the baseline velocity to address the problem of a lagging virtual sound field that is distracting when the user's head turns slowly.

[0115] However, when it is determined at step 1022 that the user's head has no angular acceleration, then at step 1026 the angular velocity of the rotation of the audio frame is selected such that when the rotation of the user's head ends, the audio frame will be aligned with the predicted position of the reference axis. This can be achieved by first predicting the time at which the user's head will stop rotating (assuming the current deceleration continues) and the angular rotation that the user's head will have traveled from its current position to that time. The angular velocity of the virtual sound field (and therefore each virtual source) can then be selected such that at the predicted time, the virtual sound field will be aligned with the user's head, meaning that it will compensate for the existing offset between the virtual sound field and the user's head, as well as the additional angle that the user's head will have traveled between the current time and the predicted time (as described in conjunction with equation (5) above).

[0116] return Figure 10B Step 1016 is a decision block for determining whether the characteristics of the user's head satisfy at least one second predetermined condition. The at least one second predetermined condition is a condition for signaling the end of the dynamic phase of the head fixed mode. Figure 10E and Figure 10F An example of the second predetermined condition is provided. Specifically, Figure 10EStep 1028 is a decision block that determines whether a predetermined time period has elapsed after the orientation of the user's head (as indicated by a reference axis or point on the user's head) is again within the angular limits. The controller is programmed to rotate the angular limits along with the audio frame about the rotation axis (that is, the angular limits are rotated by the same amount as the user's head from the initial positioning of the user's head). In practice, this means that the angular limits will exceed the reference axis or point on the user's head before the audio frame is again aligned with (aligned with) the reference axis. Returning to the reference axis or point of the angular limits may initiate a predetermined time period, after which the controller exits the dynamic phase of the head-fixed mode and returns to step 1010. Prior to the expiration of the predetermined time period initiated by the reference axis or point returning to within the angular limits (typically achieved by the angular limits exceeding the reference axis), the method returns to step 1014 to continue rotating the virtual sound field (and the angular limits).

[0117] Figure 10F Step 1030 is a decision block that determines whether the orientation of the user's head (as indicated by a reference axis or point on the user's head) is within a second angular bound. The controller is programmed to rotate the angular bound and the second angular bound along with the audio frame about the rotation axis (so that both the angular bound and the second angular bound are rotated from the initial position of the user's head by the same amount as the user's head). The second angular bound may be narrower than the angular bound that caused the dynamic phase to be entered at step 1010 (e.g., as combined with Figure 5A In this example, rather than waiting for an additional predetermined period of time, entering the second angular limit signals sufficient alignment of the audio frame with the reference axis or point, and thus signals the end of the dynamic phase, by returning to step 1010. Before the reference axis enters the second angular limit, the method returns to step 1014 to continue rotating the virtual field (and the second angular limit).

[0118] Check Figure 10B , it can be seen that method 1000 returns to step 1010 to again determine whether to enter the dynamic phase. When the current sample does not meet the predetermined condition, the audio frame remains at its current position. After exiting the dynamic phase, the current position of the audio frame is the position the audio frame has reached due to the dynamic phase. When the current sample does meet at least one predetermined condition, the virtual sound field rotates and continues to rotate with each sample until at least one second predetermined condition is met, at which point the method returns to 1010.

[0119] Additional conditions for completely exiting the head tracking phase are not shown in method 1000. This can be achieved, for example, by returning to step 1004 at each sampling or periodically to determine whether the requirements for head-fixed mode continue to be met. If it is determined that these requirements are no longer met, the method can enter room-fixed mode or some other mode.

[0120] The functionality described herein, or portions thereof, and various modifications thereof (hereinafter referred to as "functionality") may be implemented at least in part via a computer program product (e.g., a computer program tangibly embodied in an information carrier, such as one or more non-transitory machine-readable media or storage devices) for execution by one or more data processing apparatuses (e.g., a programmable processor, a computer, multiple computers, and / or programmable logic components) or to control the operation of the one or more data processing apparatuses.

[0121] A computer program can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed on one computer or executed on multiple computers at one site or distributed across multiple sites and interconnected by a network.

[0122] The actions associated with implementing all or part of the functionality may be performed by one or more programmable processors executing one or more computer programs to perform the functionality of the calibration process. All or part of the functionality may be implemented as dedicated logic circuitry, such as an FPGA and / or an ASIC (application-specific integrated circuit).

[0123] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The components of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.

[0124] Although several invention embodiments have been described and shown herein, a person of ordinary skill in the art will readily envision a variety of other components and / or structures for performing the functions described herein and / or obtaining one or more of the results and / or advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the invention embodiments described herein. More generally, a person of ordinary skill in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications in which the teachings of the present invention are used. A person of ordinary skill in the art will recognize or be able to determine many equivalents to the specific invention embodiments described herein using only routine experiments. Therefore, it should be understood that the above embodiments are presented by way of example only, and within the scope of the appended claims and their equivalents, the invention embodiments may be practiced in a manner different from that specifically described and claimed. The invention embodiments disclosed herein relate to each individual feature, system, article, material, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials and / or methods is within the inventive scope of the present disclosure if such features, systems, articles, materials and / or methods are not mutually inconsistent.

Claims

1. A pair of headphones, comprising: a sensor that outputs a sensor signal representing an orientation of the user's head; and a controller that receives the sensor signal, the controller being programmed to output a spatialized audio signal to a pair of electroacoustic transducers for conversion into a spatialized acoustic signal based on the sensor signal, wherein the spatialized acoustic signal is perceived by the user as originating from a virtual sound field comprising at least one virtual source, each virtual source of the virtual sound field being perceived as being located at a respective position different from a position of the electroacoustic transducer and referenced to an audio frame of the virtual sound field, the audio frame being disposed at a first position aligned with a reference axis of the user's head; wherein the controller is further programmed to determine, based on the sensor signal, whether a characteristic of the user's head satisfies at least one predetermined condition, the at least one predetermined condition comprising whether the orientation of the user's head is outside of predetermined angular limits, wherein, upon determining that the characteristic of the user's head does not satisfy the at least one predetermined condition, the controller is programmed to maintain the audio frame at the first position, Wherein, upon determining that the orientation of the user's head is outside the predetermined angular limits, the controller is programmed to rotate the position of the audio frame about a rotation axis to reduce the angular offset from the reference axis of the user's head.

2. A pair of headphones according to claim 1, wherein Rotating the position of the audio frame includes rotating the position of the audio frame to align with the reference axis of the user's head when the rotation of the user's head ends.

3. A pair of headphones according to claim 2, wherein: When the user's head has increased angular acceleration, the angular velocity of the rotation of the audio frame is based at least in part on the angular velocity of the user's head.

4. A pair of headphones according to claim 3, wherein: When the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that when the turning of the user's head ends, the audio frame will be aligned with the predicted position of the reference axis of the user's head.

5. A pair of headphones according to claim 4, wherein the predicted position of the reference axis of the user's head is updated at each sampling time when the user's head has a decreasing angular acceleration when the rotation of the user's head ends.

6. A pair of headphones according to claim 1, wherein the at least one predetermined condition further includes whether the angular jerkiness of the user's head exceeds a predetermined threshold, wherein the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angular limit or the angular jerkiness of the user's head exceeds the predetermined threshold.

7. A pair of headphones according to claim 1, wherein Upon determining that the orientation of the user's head is outside the predetermined angular limits, the controller is programmed to rotate the angular limits together with the audio frame about the rotation axis, wherein after the orientation of the user's head is again within the angular limits, the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period.

8. A pair of headphones according to claim 1, wherein when it is determined that the orientation of the user's head is outside the predetermined angular limit, a second angular limit narrower than the angular limit is rotated about the rotation axis together with the audio frame, wherein the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined period of time until the orientation of the user's head is within the second angular limit.

9. A pair of headphones according to claim 1, wherein the at least one virtual source includes a first virtual source and a second virtual source, the first virtual source is set at a first position, and the second virtual source is set at a second position, wherein the first position and the second position are referenced to the audio frame.

10. The pair of headphones of claim 1, wherein maintaining the audio frame at a first position comprises rotating the audio frame at a rate modulated to cancel drift of the sensor.

11. A pair of headphones according to claim 1, wherein the sensor that outputs the sensor signal comprises a plurality of sensors that output a plurality of signals.

12. A method for providing spatialized audio, the method comprising: outputting a spatialized audio signal to a pair of electroacoustic transducers for conversion into a spatialized acoustic signal based on a sensor signal representing an orientation of a user's head, wherein the spatialized acoustic signal is perceived by the user as originating from a virtual sound field including at least one virtual source, each virtual source of the virtual sound field being perceived as being located at a respective position different from a position of the electroacoustic transducer and referenced to an audio frame of the virtual sound field, the audio frame being disposed at a first position aligned with a reference axis of the user's head; determining, based on the sensor signal, whether a characteristic of the user's head satisfies at least one predetermined condition, the at least one predetermined condition including whether the orientation of the user's head is outside predetermined angular limits, and Upon determining that the orientation of the user's head is outside the predetermined angular limits, the position of the audio frame is rotated about a rotation axis to reduce the angular offset from the reference axis of the user's head.

13. The method according to claim 12, wherein: Rotating the position of the audio frame includes rotating the position of the audio frame to align with the reference axis of the user's head when the rotation of the user's head ends.

14. The method according to claim 13, wherein When the user's head has increased angular acceleration, the angular velocity of the rotation of the audio frame is based at least in part on the angular velocity of the user's head.

15. The method according to claim 14, wherein When the user's head has decreasing angular acceleration, the angular velocity of the rotation of the audio frame is selected so that when the turning of the user's head ends, the audio frame will be aligned with the predicted position of the reference axis of the user's head. 16 . The method of claim 15 , wherein the predicted position of the reference axis of the user's head is updated at each sampling time when the user's head has a decreasing angular acceleration when the rotation of the user's head ends.

17. A method according to claim 12, wherein the at least one predetermined condition also includes whether the angular jerkiness of the user's head exceeds a predetermined threshold, wherein the at least one predetermined condition is satisfied if the orientation of the user's head is outside the predetermined angle limit or the angular jerkiness of the user's head exceeds the predetermined threshold.

18. The method according to claim 12, wherein: Upon determining that the orientation of the user's head is outside the predetermined angular limits, the angular limits are rotated together with the audio frame about the rotation axis, wherein after the orientation of the user's head is within the angular limits, the position of the audio frame is rotated about the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period.

19. The method according to claim 12, wherein: Upon determining that the orientation of the user's head is outside the predetermined angular limit, a second angular limit narrower than the angular limit is rotated around the rotation axis together with the audio frame, wherein the position of the audio frame is rotated around the rotation axis to reduce the angular offset from the reference axis of the user's head for a predetermined time period until the orientation of the user's head is within the second angular limit.

20. The method of claim 12, wherein the at least one virtual source comprises a first virtual source and a second virtual source, the first virtual source being set at a first position and the second virtual source being set at a second position, wherein the first position and the second position refer to the audio frame.

21. The method of claim 12, wherein maintaining the audio frame at a first position comprises rotating the audio frame at a rate modulated to cancel drift of the sensor.

22. The method of claim 12, wherein the sensor that outputs the sensor signal comprises a plurality of sensors that output a plurality of signals.

Citation Information

Patent Citations

  • Systems and methods for sound source virtualization

    US20200037097A1