Re-centering the scene
By adjusting virtual sound sources relative to an anchor position based on head orientation and gaze direction, the systems and methods ensure a stable and immersive spatialized audio experience, addressing the issue of shifting sound stages in response to user movements.
Patent Information
- Application Number
- JP2025521177
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-10-28
AI Technical Summary
Existing spatialized audio systems fail to effectively adjust the perceived location of virtual sound sources in response to significant changes in a user's head orientation or gaze direction, leading to a shifting sound stage that can disrupt the immersive audio experience.
The systems and methods adjust the location of virtual sound sources relative to an anchor position that adapts to the user's head orientation and gaze direction, using inertial measurement units to detect changes and implement slow or rapid adjustments based on the duration of the orientation change.
Maintains a stable virtual sound stage perception by re-centering the audio experience in front of the user, ensuring a consistent and immersive audio experience even with significant changes in head orientation or gaze direction.
Smart Images

Figure 2025535777000001_ABST
Abstract
Description
[Technical Field]
[0001] The term "spatialized audio" can refer to a variety of audio or acoustic experiences, and in some spatialized audio, it can refer to a simulated experience of one or more virtual out-louds (e.g., loudspeakers) delivered to a user or listener, typically via headphones, earbuds, or other suitable wearable audio devices. Such virtualized speaker experiences are intended to be heard or perceived by the user as occurring at a location in the user's environment, rather than forming part of the headphones themselves. For example, while traditional stereo listening on headphones may sound as if the audio is coming from inside the user's head, a stereo-spatialized audio experience may sound as if there are left and right (virtual) loudspeakers in front of the user. There are numerous techniques for achieving such an experience, at least one of which is disclosed in U.S. Patent Application No. 16 / 592,454, filed October 3, 2019, entitled "SYSTEMS AND METHODS FOR SOUND SOURCE VIRTUALIZATION," published as U.S. Patent Application Publication No. 2020 / 0037097. There is a need for various adjustments to the perceived location of virtual sound sources within spatialized audio systems, methods, and processes. Summary of the Invention [Means for solving the problem]
[0002] The systems and methods disclosed herein are directed to audio rendering systems, methods, and applications. In particular, the disclosed systems and methods are directed to audio systems and methods that generate audio that is perceived by a user or listener as coming from (or being generated at) a virtual location within the user's realm, even though no actual sound source is present at the virtual location. Various systems and methods herein may generate spatialized sound from multiple virtual locations, such as a virtualized multi-channel surround sound system.
[0003] The systems and methods herein establish a location, possibly centered in front of the user / listener, as an anchor location around which various virtual sound source locations are established. For example, the location where the center channel of a multi-channel audio system is perceived to be originating may be the anchor location, while a virtual left speaker and a virtual right speaker may be rendered by the audio system and method to be perceived as originating from locations to the left and right of the anchor location. Similarly, virtual rear speakers, virtual height channels, virtual object audio sources (e.g., whose perceived locations move around the listener, such as by tracking a virtual object), and / or other virtual sources and / or reference locations used by various systems and methods may be appropriate. In various embodiments, the anchor location may be any suitable location and may or may not be associated with a particular perceived virtual sound source location. In some embodiments, the anchor location may be defined relative to a user of the system or method.
[0004] A system, method, and computer-readable medium are disclosed that detect a user's head orientation, determine an anchor position from the detected head orientation, detect changes in head orientation, and adapt the anchor position to the detected changes in head orientation.
[0005] In various embodiments, the anchor position is slowly adapted when the head orientation change is within the angle limits.
[0006] In some embodiments, the anchor position is rapidly adapted when the head orientation change exceeds the angle limit.
[0007] In certain embodiments, if the change in head orientation exceeds the angle limit, a hold time is imposed, and if the head orientation exceeds the angle limit for longer than the hold time, the anchor position may be rapidly adapted.
[0008] Further aspects, embodiments, and advantages of these exemplary aspects and embodiments are discussed in detail below. The embodiments disclosed herein may be combined with other embodiments in a manner consistent with at least one of the principles disclosed herein, and references to "an example," "some examples," "an alternate example," "various examples," "one example," etc. are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described may be included in at least one embodiment. Appearances of such terms herein do not necessarily all refer to the same embodiment. [Brief explanation of the drawings]
[0009] Various aspects of at least one embodiment are discussed below with reference to the accompanying drawings, which are not intended to be drawn to scale. These drawings are included to provide illustration and a further understanding of various aspects and embodiments, and are incorporated into and constitute a part of this specification, but are not intended to define the limitations of the invention. In the drawings, identical or nearly identical components shown in various figures may be represented by like reference characters or numerals. For clarity, not all components may be labeled in every figure. In the drawings,
[0010] [Figure 1] FIG. 1 is a schematic diagram of an exemplary listener scenario. [Figure 2] 1 is a schematic diagram of an exemplary spatial audio scenario; [Figure 3] FIG. 1 is a schematic diagram of an exemplary spatial audio centering scenario. DETAILED DESCRIPTION OF THE INVENTION
[0011] Aspects of the present disclosure are directed to systems and methods.
[0012] As used herein, the term "headphones" is intended to mean any sound-generating device configured to provide acoustic energy to each of a user's left and right ears, providing some separation or control over what reaches each ear without being heard by the opposite ear. Such devices are often fitted around, over, in, or adjacent to a user's ears to radiate acoustic energy into the user's ear canal. Headphones may also be referred to as earphones, earpieces, earbuds, or earcups and may be wired or wireless. Headphones may be integrated into another wearable device, such as a headset, helmet, hat, hood, smart glasses, or clothing. As used herein, the term "headphones" is also intended to include other form factors capable of providing binaural acoustic energy, such as headrest speakers in an automobile or other vehicle. Further examples include neck-worn devices, eyewear, or other structures that may be hooked around the ears or otherwise configured to be positioned adjacent to a user's ears. Thus, various examples may include open-ear configurations as well as over-ear or around-ear configurations. Headphones may include an acoustic driver that converts audio signals into acoustic energy. The acoustic driver may be housed in earcups or earbuds, or may be open-ear, or associated with other structures as described, such as a headrest. The headphones may be a single, standalone unit or may be one of a pair of headphones, such as one headphone for each ear.
[0013] FIG. 1 schematically illustrates a user 100 receiving sound from a sound source 102. As described above, head-related transfer functions (HRTFs) characterizing how the user 100 receives sound from various directions may be calculated or stored in memory and are represented by arrows as left HRTF 104L and right HRTF 104R (collectively or generally, HRTFs 104). HRTFs 104 are defined at least in part based on the user's orientation relative to incoming acoustic waves emanating from the sound source, as indicated by angle θ. That is, angle θ represents the relationship of the direction in which the user 100 is facing relative to the direction from which the sound is arriving (represented by the dashed line). The directionality of the sound produced by sound source 102 may be defined by a radiation pattern that varies with angle α, which represents the relationship between the principal (or axial) direction in which the sound source 102 is producing sound and the direction in which the user 100 is located.
[0014] For example, spatialized audio, which processes audio signals so that sound is perceived as coming from a virtual sound source (e.g., sound source 102) even when nothing is physically generating the sound from that location, can be simulated in a number of ways. In some embodiments, one or more HRTFs 104 may be applied at an angle θ. In various embodiments, the directionality of reflections from real or virtual reflective surfaces (e.g., walls or other objects in a physical or virtual space) may be considered. In such embodiments, virtual reflected sounds arrive from different angles and with different arrival times, each of which may be simulated by additional signal components representing such reflections. In certain embodiments, the directionality of the virtual sound source (e.g., sound source 102), which is the radiation pattern of the sound source, may also be considered. As mentioned above, at least one example of a system for spatializing audio to one or more virtual sound sources may be found in U.S. patent application Ser. No. 16 / 592,454, filed October 3, 2019, entitled "SYSTEMS AND METHODS FOR SOUND SOURCE VIRTUALIZATION."
[0015] Regardless of the various methods of processing audio signals to simulate virtual sound sources, the location of the virtual sound sources should remain relatively fixed as the head of user 100 moves around. Therefore, the systems and methods herein may use various sensors and methods to detect the orientation of the head of user 100, and may take into account changes in directivity and reflection angles, as well as HRTFs for the radiation pattern, as appropriate.
[0016] However, in various embodiments, it may be desirable to adjust the perceived position of the virtual sound source. For example, various systems and methods according to the systems and methods herein may position a virtual sound source directly in front of the user 100 to play a center channel (or a phantom center channel not present in the audio source but, for example, “center” content derived from a left-right stereo pair), as shown in FIG. 2 , position another virtual sound source to the left of the front to play left channel content, and position yet another virtual sound source to the right of the front to play front channel content. Such a configuration may work well while maintaining the position of the virtual sound source, which means adjusting for small movements of the user's head to account for changes in the virtual audio direction of, for example, directional, reflection, and radiation patterns. However, if the user 100 decides to reposition to look in a different overall direction, it may be desirable to have the various positions of the virtual sound sources adjust to the user's new “normal” or front-facing position. For example, if a user is walking and makes a right turn, the front, left, and right channel content will be perceived to remain in place, now all to one side of the user rather than in front of them. Various systems and methods herein make adjustments to the selected position of the virtual sound source to account for significant, substantially permanent changes in the direction in which the user 100 is facing, e.g., gaze direction, as opposed to the user 100 momentarily looking left or right.
[0017] Thus, the systems and methods herein adjust the location of a virtual sound source in response to a user's or listener's head movements. Spatialized audio systems and methods virtualize incoming signals so that a user may perceive one or more sounds as coming from a fixed location, and such signals must be adjusted as the user moves their head to maintain the perception. Such adjustments due to changes in angle of arrival (generally described above with respect to FIG. 1 ) are not the adjustments discussed here. Instead, in addition to their perceptual adjustments to maintain the virtual sound source location, systems and methods according to the systems and methods herein also adjust the virtual sound source location in response to longer-term changes in the user's head orientation, e.g., gaze direction. According to various embodiments, one or more virtual sound source positions may be positioned or selected (by the system or by user preference) relative to an anchor position, and as the user's head orientation or gaze direction changes, the anchor position may adapt to move according to the user's gaze direction, so that if the user permanently changes the direction they are facing, the sound stage of the virtual sound sources around the user will adapt to the new orientation; for example, in various embodiments, the virtual center channel will adjust to remain in front of the user and the other virtual channels will adjust accordingly.
[0018] 2 shows user 100 listening to spatialized audio having a virtual left speaker 200L, a virtual right speaker 200R, and a virtual center speaker 200C (collectively, virtual speakers 200). The head orientation of user 100 may be detected using any number of systems, sensors, and methods, and in some embodiments may be determined or assisted by an inertial measurement unit (IMU). Thus, a gaze direction 210 may be determined, and the position of virtual center speaker 200C may be located directly in front of user 100.
[0019] As the user's 100 head moves, through normal small changes in orientation, the virtual signals are adjusted, as described above, to maintain the perceived position of the virtual speakers 200. However, according to various embodiments, if the change in the user's 100 head orientation persists for a short period of time, the systems and methods herein may move the positions of the virtual speakers 200 so that they are again re-centered in front of the user 100.
[0020] In various embodiments, systems and methods may select an anchor position, and the position of virtual speaker 200 may be relative to the anchor position. In some embodiments, the anchor position may coincide with the position of virtual center speaker 200C, but in other embodiments, this need not be the case. Various systems and methods may spatialize additional virtual speaker channels (e.g., rear left, rear right, height channels, etc.) and / or may spatialize additional or other virtual sound sources, such as moving virtual sound sources (e.g., the sound of a motorcycle passing on the left or the sound of an airplane flying overhead, etc.). According to various embodiments, the position of each of these sound sources may be characterized relative to a single anchor position that adjusts to changing orientations of user 100 in systems and methods according to the systems and methods herein.
[0021] FIG. 3 shows an anchor position 300, which in this example is established directly in front of the user 100 and which may, in some cases, but not necessarily, be aligned with the position of the virtual center speaker 200C (see FIG. 2).
[0022] According to various embodiments, as the head of user 100 moves (e.g., the orientation of the head of user 100 changes), anchor position 300 is slowly adjusted to move so that it is again centered in front of user 100. In particular embodiments, slow adjustment may mean that it may take about 3 seconds for anchor position 300 to move, although other time frames and / or time constants are contemplated herein.
[0023] For example, suppose user 100 is looking at a computer display and imagines the virtual sound stage to be centered on the computer display (e.g., the virtual center channel is in front of the right, with the left virtual channel on the left and the right virtual channel on the right, respectively). Then, user 100 turns his head (e.g., about 10 degrees) to look at an adjacent display. The systems and methods herein adjust anchor position 300 accordingly, and after about three seconds, the virtual sound stage is again in front of user 100 and centered on the adjacent display. In this way, if user 100 simply glances around at the adjacent display and then looks behind, anchor position 300 will begin to adjust for a short period of time, but then readjust to remain centered on the original computer display. User 100 may not even perceive the moving virtual sound stage in this case.
[0024] Now consider that there is an additional display, such as an adjacent wall, that requires user 100 to rotate his / her head 90 degrees. If user 100 only looks at the additional display for a short period of time, it may be desirable not to adjust anchor position 300 at all. However, if user 100 turns to look at the additional display for an extended period of time, it may be desirable to adjust anchor position 300 more quickly (e.g., to more quickly recenter the virtual sound stage in front of user 100). In fact, user 100 may rotate his / her entire body, such as by rotating in an office chair, to look at the additional display for an extended period of time.
[0025] Thus, the angle β on either side of the user's 100 line of sight defines an angular limit boundary 220 that defines the range of head orientations over which the systems and methods herein can slowly adapt the anchor position 300.
[0026] According to various embodiments, the angle β defining the boundary 220 may be less than or equal to 40 degrees. In particular embodiments, the angle limit may be defined by an angle β of approximately 30 degrees, 20 degrees, 10 degrees, or 5 degrees. Other angle limits may be applicable to various systems and methods.
[0027] When the user's 100 head orientation goes outside the angle limits, various embodiments herein do not slowly adapt the anchor positions 300. They may adapt more quickly, or they may stop adapting for a hold time, such as to "make sure" that a large change is more permanent, and then adapt more quickly (e.g., more quickly than slow adaptation).
[0028] The faster the anchor position 300 adapts, the more quickly the anchor position is adjusted to re-center the virtual sound stage in front of the user 100. In particular embodiments, rapid adjustment may mean that it takes only about one second for the anchor position 300 to move, although other time frames and / or time constants are contemplated herein.
[0029] In some embodiments, when the user's 100 head orientation goes outside of the angle limits, a hold time may be imposed on adapting the anchor position 300. During the hold time, the anchor position 300 may not adapt at all, but instead may remain fixed in place. In particular embodiments, the hold time may be approximately 3 seconds.
[0030] In various embodiments, the angular limits may define a wedge- or pie-shaped range of positions in a two-dimensional plane, for example, defined relative to gravity or up / down. In other words, certain systems and methods herein may only be concerned with left / right head rotation and not with looking up or down. In other embodiments, the systems and methods may accommodate all changes in head orientation, and thus the angular limits may define a cone.
[0031] The method and apparatus embodiments discussed herein are not limited in their application to the details of construction and the arrangement of components set forth in the above description or illustrated in the accompanying drawings. The methods and apparatus of the present invention may be implemented in other embodiments and may be practiced or carried out in various ways. Examples of specific implementations are provided herein for illustrative purposes only and are not intended to be limiting. In particular, functions, components, elements, and features discussed in connection with any one or more embodiments are not intended to be excluded from a similar role in any other embodiments.
[0032] Additionally, the phraseology and terminology used herein are for descriptive purposes and should not be considered limiting. Any reference herein to system and method examples, components, elements, acts, or functions in the singular may also encompass embodiments that include the plural, and any reference herein to any example, component, element, act, or function in the plural may also encompass examples that include only the singular. Thus, singular or plural references are not intended to limit the disclosed systems or methods, their components, acts, or elements. The use herein of "including," "comprising," "having," "containing," "involving," and variations thereof, is meant to encompass the items listed below and equivalents thereof, as well as additional items. References to "or" may be interpreted as inclusive, such that any term described with "or" may refer to either one, more than one, or all of the listed terms. Unless the context reasonably suggests otherwise, references to front, back, left, right, up, down, top, bottom, and length and width are for convenience of description and are not intended to limit the present systems and methods, or components thereof, to any one positional or spatial orientation.
[0033] Having described several aspects of at least one embodiment, it will be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be part of this disclosure and are intended to be within the scope of the invention. Accordingly, the foregoing description and drawings are by way of example only, with the scope of the invention to be determined from proper construction of the appended claims and their equivalents.
Claims
1. 1. A method for adapting an anchor position to a relative position of one or more virtual loudspeakers, comprising: Detecting a user's head orientation; determining the anchor position from the detected head orientation; detecting a change in head orientation; and slowly adapting the anchor position to the detected head orientation change when the head orientation change is within an angle limit.
2. The method of claim 1 , wherein the angle limit is less than or equal to one of 40 degrees, 30 degrees, 20 degrees, 10 degrees, and 5 degrees from a previous head orientation.
3. 2. The method of claim 1, wherein slowly adapting the anchor position to changes in the detected head orientation comprises adjusting the anchor position to the detected head orientation on a time scale of about 3 seconds.
4. The method of claim 1 , further comprising freezing the anchor position when the head orientation change exceeds the angle limit.
5. 5. The method of claim 4, further comprising rapidly adapting the anchor position to the head orientation when the amount of change in the head orientation exceeds the angle limit for a duration exceeding a certain amount of time.
6. The method of claim 5 , wherein the amount of time is 3 seconds or greater.
7. The method of claim 5 , wherein rapidly adapting the anchor position to the head orientation comprises adjusting the anchor position to the detected head orientation on a timescale of about 1 second.
8. 1. An apparatus comprising: an acoustic transducer; an inertial measurement unit (IMU); and A processor coupled to the acoustic transducer, the processor configured to perform the method of claim 1.
9. A non-transitory computer-readable medium encoded with instructions that, when executed by a processor, cause the processor to perform the method of claim 1.
Citation Information
Patent Citations
Head tracking for mobile applications
JP2012518313A
Head Tracking Using Adaptive Criteria
JP2019523510A
Sound processing device and sound processing method
JP2021005822A
Moving body position estimation device and moving body position estimation method
JP2021156600A
Signal processing device and method, and program
WO2019009085A1