Sound Field Rotation

The method of sound field rotation adjusts audio presentation based on user activity and head orientation to maintain spatial accuracy, addressing the issue of jarring experiences when moving, and ensuring smooth audio object positioning.

JP2026035579APending Publication Date: 2026-03-04DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025179365
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-09
Filing Date
2025-10-24
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing audio rendering technologies struggle to accurately present spatial audio content when a user is moving, leading to a jarring experience due to incorrect rendering of audio objects relative to the user's head orientation.

Method used

A method for determining sound field rotation based on user activity status and head orientation, adjusting the sound field to maintain audio objects in their intended spatial positions relative to the user, using sensors and processors to update the sound field rotation in response to changes in head orientation and activity.

Benefits of technology

Ensures smooth and accurate spatial presentation of audio objects, maintaining their intended positions relative to the user's head even as they move, thereby enhancing the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035579000001_ABST
    Figure 2026035579000001_ABST
Patent Text Reader

Abstract

A method, system, and medium for determining the rotation of a sound field are provided. [Solution] A method for determining the rotation of a sound field includes step 302 of determining a user's activity status, step 304 of determining the user's head orientation using at least one of one or more sensors, and step 306 of determining a target direction based on the activity status and the user's head orientation, and determining a rotation of a sound field used to present an audio object via headphones based on the target direction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 303,201, filed January 26, 2022, and U.S. Provisional Patent Application No. 63 / 479,078, filed January 9, 2023, each of which is incorporated herein by reference in its entirety.

[0002] Technical Field The present disclosure relates to systems, methods, and media for sound field rotation. [Background technology]

[0003] background Audio content intended to be presented in various spatial contexts can be difficult to render, for example, when a user is moving while wearing headphones. Rendering such audio content incorrectly can be undesirable, as it can create a jarring experience for the listener.

[0004] Notation and Nomenclature Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any sound-producing transducer (or set of transducers). A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter). These transducers may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds may undergo different processing in different circuit branches coupled to different transducers.

[0005] Throughout this disclosure, including the claims, the phrase "operating on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) is used broadly to indicate performing an operation directly on the signal or data, or performing an operation on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation on it).

[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the remaining XM inputs are received from external sources) may also be referred to as a decoder system.

[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), pipelined processors for audio or other sound data, and Includes digital signal processors that are programmed and / or otherwise configured to perform processing, programmable general purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the Invention [Means for solving the problem]

[0008] overview In some embodiments, a method for determining a sound field rotation includes: (a) determining a user's activity status; (b) determining a user's head orientation using at least one of the one or more sensors; (c) determining a target direction based on the activity status and the user's head orientation; and (d) determining a rotation of a sound field used to present an audio object via headphones based on the target direction.

[0009] In some examples, the method may further include (e) repeating (a) through (d) such that the rotation of the sound field is updated over time based on the user's activity status and changes in the user's head orientation.

[0010] In some examples, the activity situation includes at least one of walking, running, non-walking and non-running movement, or minimal movement. In some examples, the activity situation includes walking or running, and the target direction is determined based on a direction in which the user is walking or running. In some examples, the activity situation includes non-walking and non-running movement, and the target direction is determined based on a direction in which the user was facing within a predetermined previous time window. In some examples, the predetermined previous time window is within a range of approximately 0.2 seconds to 3 seconds. In some examples, the activity situation includes minimal movement, and the target direction is determined based on a direction in which the user was facing within a predetermined previous time window. In some examples, the predetermined previous time window is longer than the predetermined previous time window used to determine the target direction associated with activity situations of non-walking and non-running movement. In some examples, the predetermined previous time window used to determine the target direction associated with activity situations of minimal movement is within a range of approximately 3 seconds to 10 seconds. In some examples, the direction the user was facing is determined using a motion tolerance threshold, the motion tolerance threshold being in a range of approximately 2 degrees to 20 degrees. In some examples, the rotation of the sound field includes an incremental rotation toward the target direction. In some examples, the incremental rotation is based at least in part on angular velocity measurements obtained from a user device. In some examples, the user device is substantially stationary relative to the headphones worn by the user. In some examples, the user device provides audio content to headphones.

[0011] In some examples, the activity status of the user is determined based at least on sensor data obtained from one or more sensors located in or on headphones worn by the user.

[0012] In some examples, the orientation of the user's head is determined using at least one sensor located in or on headphones worn by the user.

[0013] In some examples, the headphones include earbuds.

[0014] In some examples, the method further includes, after (d), causing an audio object to be rendered based on the determined rotation of the sound field. In some examples, the method further includes causing the rendered audio object to be presented through the headphones.

[0015] In some embodiments, an apparatus configured to implement any of the above methods is provided.

[0016] In some examples, one or more non-transitory media are provided that store software, the software configured to perform a method according to any of the methods described above. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a schematic diagram of a system for sound field rotation, according to some embodiments.

[0018] [Figure 2] FIG. 2 is a schematic diagram of a system for determining sound field orientation based on activity, according to some embodiments.

[0019] [Figure 3] FIG. 3 is a flowchart of an exemplary process for determining how to rotate a sound field based on the orientation of a user's head, according to some embodiments.

[0020] [Figure 4] FIG. 4 is a flowchart of an exemplary process for determining a sound field orientation based on a user's head orientation, according to some embodiments.

[0021] [Figure 5] FIG. 5 is an exemplary graph illustrating sound field orientation for various activity situations, according to some embodiments.

[0022] [Figure 6] FIG. 6 is a flowchart of an exemplary process for determining sound field orientation in static activity situations, according to some embodiments.

[0023] [Figure 7] FIG. 7 shows a block diagram illustrating example components of an apparatus capable of implementing various aspects of the disclosure.

[0024] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0025] Detailed Description of the Embodiments Audio content may be scene-based. For example, an audio content creator may create audio content that includes various audio objects intended to be rendered and played back to create the perception of being present in a particular spatial location for a user. As an example, the audio content may include primary vocals intended to be perceived as being "in front" of the listener. Additionally or alternatively, the audio content may include various audio objects (e.g., secondary instruments, sound effects, etc.) intended to be perceived as being present to the side of the listener, above the listener, etc. Achieving a spatial presentation of audio objects that is faithful to the audio creator's intent can be challenging, especially when the audio content is rendered to and / or presented through headphones. For example, a listener may move their head in various ways (e.g., to look at things) and / or move their body in ways that change head orientation when wearing headphones. Furthermore, Additionally, rendering audio content to strictly follow the user's head orientation when the user is moving can lead to unintended perceptual results. For example, when a user is walking or running, the user may occasionally turn their head to look around, e.g., to check for cars when crossing a road. Continuing with this example, rotating the sound field to correspond to changes in the user's head orientation may be jarring to the listener; in such situations, it may be desirable to continue rendering audio objects intended to be "in front" of the listener in a manner that is perceived to be in front of the listener's body (e.g., in the direction the listener is walking or running), even when the listener occasionally turns their head away from the direction of movement. Conversely, in instances where the user is not moving in a straight line and / or substantially forward (e.g., when performing household chores and / or engaging in another activity involving frequent twisting and / or turning in various directions), it may be desirable to rotate the sound field to correspond to the user's head orientation.

[0026] Techniques for determining sound field rotation based on a listener's activity status are disclosed herein. In some embodiments, a listener's current activity status can be determined. As used herein, "activity status" refers to a characterization of the current activity in which the user is engaged, particularly a characterization of the user's movement during the current activity. Examples of activity status include walking, running, or other movement involving substantially linear or forward movement, minimal movement, and non-walking or non-running movement. Note that, as used herein, the term "listening status" is used interchangeably with the term "activity status." The activity status can be determined based on one or more sensors located in or on headphones worn by the user. The headphones can be over-the-head headphones, earbuds, or the like. A target direction of the listener can be determined. As used herein, "target direction" refers to an azimuth angle relative to a vertical axis representing the user's head that corresponds to a target azimuth orientation of the sound field. The target direction can depend on the current activity status in which the listener is engaged. For example, for a given user's head orientation, the target direction may be different if the user is walking or running compared to when the user is in a minimal movement activity situation. Rotation information may then be determined such that the audio object is presented based on the user's current orientation and the target direction (which is therefore dependent on the current activity situation). The rotation information may be determined to avoid abrupt discontinuities in the rendering of the audio object by smoothly rotating the sound field toward the target direction.

[0027] 1 is a block diagram of an exemplary system 100 for sound field rotation according to some embodiments. As shown, system 100 includes a user orientation system 102, a sound field orientation system 104, and a sound field rotation system 106. It should be noted that various components of system 100 may be implemented by one or more processors located in or on headphones worn by a user, one or more processors or controllers of a user device paired with the headphones (e.g., a user device presenting content played by the headphones), etc. An example of such a processor or controller is shown in and described below in connection with FIG. 7.

[0028] In some embodiments, the user orientation system 102 may be configured to determine the listener's head orientation. For example, the listener's head orientation may be considered the direction in which the listener's nose is pointing. The listener's head orientation is generally referred to herein as α nose The orientation of the listener's head may be determined using one or more inertial sensors, for example, located in or on the listener's headphones. The inertial sensors may include one or more accelerometers, one or more gyroscopes, one or more magnetometers, etc. In some examples, the orientation of the listener's head may be determined relative to an external reference frame.

[0029] As shown, system 100 may include a sound field directing system 104. Sound field directing system 104 may be configured to determine a forward-facing direction of the sound field, generally referred to herein as α fwdIn some implementations, the forward-facing direction of the sound field may be determined based on the current listener's situation or the current activity in which the listener is engaged. Examples of activities include walking, running, cycling, riding in a vehicle such as a car or bus, remaining substantially stationary (e.g., while watching television, reading a book, etc.), or engaging in non-walking or non-running movements (e.g., loading the dishwasher, unpacking groceries, various household chores, non-walking or non-running athletic activities such as weightlifting, etc.). A more detailed technique for determining the forward-facing direction of the sound field based on the current listener's activity is shown in FIG. 2 and FIGS. 4-6 and described below in connection with those figures.

[0030] As shown, system 100 may include a sound field rotation system 106. Sound field rotation system 106 may be configured to rotate the sound field using the listener's head orientation determined by user orientation system 102 and the forward-facing direction of the sound field determined by sound field orientation system 104. For example, the sound field may be rotated so that the listener perceives various audio objects, when rendered based on the rotated sound field, as being in front of the listener's head, even as the listener moves around. A more detailed technique for rotating the sound field is shown in and described below in connection with FIG. 3.

[0031] In some implementations, the sound field orientation facing forward (here generally α fwd The current listener context (denoted as ) may be determined based on a determination of a current listener context. The current listener context may indicate, for example, a characterization of a current listener activity. In particular, the characterization of a current listener activity may indicate whether the listener is currently moving and / or the type of movement the user is engaged in. The listener status or activity status may include categorizing the listener's activity into one of a set of possible listener statuses. For example, the set may include 1) walking or running, 2) substantially stationary, and 3) non-walking and non-running. In some embodiments, a set of context-dependent sound field orientation directions may be determined, where each context-dependent sound field orientation direction corresponds to a possible listener status. In other words, in some embodiments, multiple possible sound field orientation directions may be determined. A sound field orientation direction may then be selected from the set of possible sound field orientation directions based on the current listener status. After the sound field orientation direction is selected, the final sound field orientation (e.g., used to rotate the sound field as described above in connection with FIG. 1) may be determined by smoothing the selected sound field orientation based on recent previous sound field orientations. In some embodiments, the smoothing may be based on the current listener situation.

[0032] 2 is a block diagram of an exemplary system 200 for determining a forward-facing sound field orientation according to some embodiments. It should be noted that system 200 is an example implementation of sound field orientation system 104 shown in and described above in connection with FIG. 1. As shown, system 200 includes a listener context determination block 202, a context-dependent orientation determination block 204, an orientation selection block 206, and an orientation smoothing block 208. It should be noted that various components of system 200 may be implemented by one or more processors located in or on headphones worn by a user, one or more processors or controllers of a user device paired with the headphones (e.g., a user device presenting content played by the headphones), etc. Examples of such processors or controllers are shown in FIG. 7. 7 and described below in conjunction with FIG.

[0033] In some embodiments, the listener state determination block 202 may be configured to determine a current listener state. As noted above, examples of listener states include walking, running, cycling, riding in a vehicle such as a car or bus, being substantially stationary (e.g., while watching television, reading a book, etc.), and / or engaging in a non-walking or non-running movement (e.g., loading or unloading the dishwasher, unpacking groceries, doing yard work, etc.). In some embodiments, the listener state determination block 202 may determine the current listener state based on sensor data from one or more inertial sensors, such as one or more accelerometers, one or more gyroscopes, one or more magnetometers, etc. The inertial sensors may be located in or on headphones worn by the listener. In some implementations, the listener state may be determined by subjecting the sensor data to one or more machine learning models trained to output or classify the input sensor data into one of a set of possible listener states.

[0034] The context-dependent orientation determination block 204 may be configured to determine a set of multiple context-dependent sound field orientations. For example, the context-dependent orientation determination block 204 may determine three possible sound field orientations: 1) a walking or running sound field orientation, 2) a static sound field orientation, and 3) a non-walking and non-running sound field orientation. In some embodiments, as illustrated in FIG. 2, the context-dependent sound field orientation may be determined based on the current listener's context. More detailed techniques for determining multiple context-dependent sound field orientations are shown in and described below in connection with FIGS. 4-6.

[0035] The orientation selection block 206 may select one of the context-dependent sound field orientations generated by the context-dependent orientation determination block 204 based on the current listener context determined by the listener context block 202. For example, when the current listener context is walking or running, the orientation selection block 206 may be configured to select a context-dependent sound field orientation corresponding to the walking or running sound field orientation.

[0036] The orientation smoothing block 208 may be configured to smooth the selected context-dependent sound field orientation based on the previous sound field orientation. For example, the orientation smoothing block 208 may use a first smoothing technique to smooth the selected context-dependent sound field orientation in response to determining that the current listener's situation is the same as the previous listener's situation, and may use a second smoothing technique to smooth the selected context-dependent sound field orientation in response to determining that the current listener's situation is different from the previous listener's situation. As another example, the orientation smoothing block 208 may smooth the sound field orientation based on the difference between the current target direction and a smoothed representation of the sound field orientation at the previous time sample. Smoothing the sound field orientation based on the previous sound field orientation may make the sound field rotation perceptually smooth, i.e., the sound field may not appear to jump from one orientation to another when the sound field is rotated. Note that in some embodiments, the sound field orientation may be smoothed based on the current listener's situation. For example, the sound field orientation may be rotated at a first slew rate if the current listener situation is walking or running, at a second slew rate if the current listener situation is one in which the listener is substantially stationary, and at a third slew rate if the current listener situation is one in which the listener is not walking or running. By utilizing different smoothing techniques for different listener situations, the sound field may be rotated in a manner desirable for the listener. More detailed techniques for smoothing the sound field orientation are shown in and described below in connection with Figures 4 and 5.

[0037] In some implementations, the sound field orientation (e.g., indicated by the front-facing orientation of the sound field) The orientation of the sound field may be determined based on the current orientation of the user's head and the current listener situation. The determined sound field orientation may then be used to identify rotation information (e.g., a rotation angle, etc.) to be utilized to rotate the sound field according to the sound field orientation. In some embodiments, the rotation information may then be used to render the audio object. For example, rendering the audio object may include modifying audio data associated with the audio object so that the audio object, when presented, is spatially perceived at a spatial location relative to the user's frame of reference that corresponds to its intended spatial location (e.g., as specified by a content creator).

[0038] FIG. 3 is a flowchart of an exemplary process 300 for rotating a sound field according to some embodiments. In some implementations, the blocks of process 300 may be performed by a processor or controller. Such a processor or controller may be part of the headphones and / or part of a mobile device paired with the headphones, such as a mobile phone, tablet computer, laptop computer, etc. An example of such a processor or controller is shown in FIG. 7 and described below in connection with FIG. 7. In some embodiments, the blocks of process 300 may be performed in an order other than that shown in FIG. 3. In some implementations, two or more blocks of process 300 may be performed substantially in parallel. In some embodiments, one or more blocks of process 300 may be omitted.

[0039] The process 300 can begin at 302 by determining the user's head orientation. The user's head orientation is generally referred to herein as M FU Here, M FU is an n-dimensional matrix (e.g., a 3x3 matrix) that indicates orientation data for n axes. For example, M FUmay be a three-dimensional matrix indicating orientation data relative to the X, Y, and Z axes. The orientation of the user's head may be determined using one or more sensors located in or on headphones worn by the user. The sensors may include one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof. Note that the orientation of the user's head, M FU F can be a function of time samples, generally denoted herein as k, as well as other parameters described below. s Assuming a sample rate of k, the time instant associated with sample k can be determined by:

number

[0040] In some implementations, the orientation of the user's head may be used to determine the direction in which the user's nose is pointing, which is generally referred to herein as α nose For example, in some embodiments, a vector N represents the direction in which the user's nose is pointing in a fixed external frame. F is the user's head orientation M FU For example, in some embodiments, N F can be determined by:

number

[0041] Continuing with this example, in some embodiments, the azimuth angle (herein α nose ) is expressed as N F The vector may be determined based on the first and second elements. nose can be determined by:

number

[0042] When the user's head is tilted, the azimuth angle (α nose ) is the angular rotation information (generally referred to herein as ω ) of the user's head about the Z axis (e.g., the axis that exits the top of the user's head and points vertically) ZU (k)). The angular rotation information may be obtained from any subset or all of the sensors used to determine the orientation of the user's head. The user's head may be determined to be tilted in response to determining that the user's head is tilted by more than a predetermined threshold (e.g., more than 20 degrees from vertical, more than 30 degrees from vertical, etc.). In some embodiments, in cases where the user's head is determined to be tilted from vertical (e.g., the user's head is not substantially upright), the azimuth angle of the user's nose may be determined by:

number

[0043] At 304, the process 300 can determine the sound field orientation. The sound field orientation can be a forward-facing direction of the sound field. The sound field orientation is generally referred to herein as α fwd The sound field orientation may be determined based on the current listener's situation. More detailed techniques for determining the sound field orientation based on the current listener's situation are shown in and described below in connection with FIGS. 4-6. Note that in some embodiments, the sound field orientation may be determined based at least in part on the user's head orientation determined in block 302.

[0044] At 306, process 300 can determine rotation information for rotating the sound field based on the user's head orientation. In some embodiments, the rotation information is an n-dimensional matrix (generally referred to herein as M) that indicates the rotation of a virtual reference frame relative to a fixed external frame to achieve the sound field orientation determined in block 304. VF In some embodiments, the rotation information may describe a rotation about the Z axis, which corresponds to the axis projecting from the user's head. For example, in some embodiments, the rotation information (e.g., MVF matrix) can be determined by:

number

[0045] At 308, the process 300 can cause the audio object to be rendered according to the rotation information. For example, in some embodiments, the process 300 generates an n-dimensional matrix (generically referred to herein as M) that indicates the rotation of the virtual scene containing the audio object relative to the user's head. VU The virtual head position of the user can be determined. The rotation of the scene (including one or more audio objects) may be determined based on the orientation of the user's head (e.g., determined above in block 302) and the rotation information determined in block 306 indicating the rotation of the virtual scene relative to a fixed external frame. For example, the rotation M of the virtual scene relative to the user's head VU can be determined by:

number

[0046] M VF After determining the matrix, the position of the audio object in the fixed external frame is determined based on the position of the audio object in the virtual frame and M, which indicates the rotation of the virtual scene relative to the fixed external frame. VF For example, the (x v ,y v ,z v ), the position of the audio object relative to the fixed frame (herein (x f ,y f ,z f ) can be determined by:

number

[0047] In this specification, the direction in which the user's nose is pointing (for example, α nose ) and the desired forward-facing direction of the sound field (e.g., α fwd ) is α UV which corresponds to the angle by which the virtual sound field into which the audio object is to be rendered should be rotated relative to the user's frame of reference. Thus, the audio object's position can be expressed as Z V axis and with respect to the virtual reference frame V, UV The image can be rendered by rotating it by an angle of .

[0048] Process 300 can then loop back to block 302. In some implementations, process 300 can continuously loop through blocks 302-308, thereby continuously updating the rotation of the sound field based on the listener's situation. For example, process 300 can rotate the orientation of the sound field when the user is walking or running in a first manner (e.g., a state in which the target direction corresponds to the direction in which the user is walking or running), and then, in response to determining that the user has changed activity (e.g., to a listener situation with minimal movement or to non-running and non-walking movement), process 300 can rotate the sound field based on the updated activity situation or listener situation (e.g., to correspond to the direction in which the user was most recently looking). In this manner, process 300 can adaptively respond to the listener's situation by adaptively rotating the sound field based on both the listener's situation and the user's current orientation.

[0049] In some implementations, the rotation of the sound field may be determined based on the current listener's situation. For example, the current listener's situation can be used to determine the target direction. As an example, if the current listener's situation is that the user is walking or running (e.g., moving forward in a substantially straight line), the target direction may be the current direction in which the listener is walking, running, or moving. As another example, if the current listener's situation is that the listener is engaged in minimal movement (e.g., sitting or standing still), the target direction may be the direction in which the listener has been moving for a recent time window (e.g., about 3 seconds to 10 seconds). The target direction may correspond to the direction the listener was facing during a recent time window (e.g., within a 10-second time window). As yet another example, if the current listener situation is one in which the listener is moving but not walking or running (e.g., engaged in a non-walking or non-running activity such as housework), the target direction may correspond to the direction the listener was facing during a recent time window (e.g., within a time window of approximately 0.2 seconds to 3 seconds). Note that the recent time window used to determine the target direction for non-walking and non-running motion activities may be relatively shorter than the time window used to determine the target direction for a static or minimally moving listener situation. Note that in some implementations, multiple target directions may be determined, each corresponding to a different possible listener situation. These directions may be referred to herein as azimuth directions and indicate the azimuth direction in which the listener is interested and / or focusing. In some embodiments, for example, Based on a determination of a current listener's situation or a current listener's activity, a target direction may be determined from a set of candidate target directions. A sound field rotation direction may then be determined based on the selected target direction. For example, the sound field rotation may be determined in a manner that smooths the rotation of the sound field toward the selected target direction. The smoothing may be performed by considering whether the current listener's situation is different from the previous listener's situation in order to rotate the sound field more smoothly and improve discontinuous rotation of the sound field.

[0050] 4 is a flowchart of an exemplary process 400 for determining sound field rotation according to some embodiments. For example, using the notation generally used herein, the process 400 determines the azimuthal sound field orientation, or generally referred to herein as α fwd 4 may be utilized to determine the front-facing direction of a sound field, referred to as the front-facing direction of the sound field. As a more specific example, FIG. 4 illustrates an example technique that may be used, for example, in block 304 of FIG. 3, to determine the direction of a sound field. The blocks of process 400 may be performed, for example, by one or more processors or controllers located in or on headphones, or by a processor or controller of a user device (e.g., a mobile phone, tablet computer, laptop computer, desktop computer, smart television, video game system, etc.) paired to headphones worn by a user. In some embodiments, the blocks of process 400 may be performed in an order other than that shown in FIG. 4. In some implementations, two or more blocks of process 400 may be performed substantially in parallel. In some implementations, one or more blocks of process 400 may be omitted.

[0051] Process 400 can begin at 402 by determining a set of context-dependent orientation directions. Each context-dependent orientation direction may indicate a possible target direction of the user. For example, each context-dependent orientation direction may indicate a direction the user is primarily looking, facing, moving toward, etc. Each context-dependent orientation direction may correspond to a particular listener situation or activity situation. For example, the set of context-dependent orientation directions may include a first direction corresponding to a first activity, a second direction corresponding to a second activity, a third direction corresponding to a third activity, etc. The set of context-dependent orientation directions may include any suitable number of directions, such as 1, 2, 3, 5, 10, 20, etc.

[0052] In one example, the set of context-dependent orientation directions may include a first direction corresponding to a listener situation or activity of walking or running (or any other substantially linear forward movement activity), a second direction corresponding to a listener situation or activity of minimal movement, and a third direction corresponding to a listener situation or activity of non-walking or non-running movement (e.g., performing housework). Note that minimal movement activities generally relate to minimal movement relative to the listener's frame of reference. For example, a listener riding a bus or other vehicle may be determined to have a minimal movement listener situation while riding the bus or other vehicle, even if the vehicle itself is moving, in response to determining that the listener is substantially stationary while riding the bus or other vehicle. It is conceivable.

[0053] The set of context-dependent orientation directions is a running or walking target direction (generally referred to herein as α walk In cases where the target direction includes a target direction (denoted as ), the running or walking target direction may be determined as the current direction in which the listener is walking or running. The current direction in which the listener is walking or running may be determined using various techniques, such as, for example, using one or more sensors located in or on headphones worn by the listener. The one or more sensors may include one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof.

[0054] The set of context-dependent orientation directions is the minimum motion target direction (generally referred to herein as α staticIn cases where the minimum motion target direction includes a time window (represented as , the minimum motion target direction) corresponding to the minimum movement of the listener, the minimum motion target direction may be determined as the direction in which the listener was primarily facing within a recent time window corresponding to the minimum movement of the listener. Example time windows include 3 to 10 seconds, 2 to 12 seconds, 5 to 15 seconds, 5 to 20 seconds, etc. In some embodiments, the minimum motion target direction may be determined using information indicative of the movement of a device (e.g., a mobile phone, a tablet computer, a laptop computer, etc.) to which the headphones are paired. For example, the movement of the paired device may indicate the movement of a vehicle in which the listener is currently riding. By accounting for the movement of the paired device, the sound field may be rotated in a manner that takes into account, for example, the turning of a vehicle in which the listener is riding. A more detailed technique for determining the minimum motion target direction is shown in FIG. 5 and described below in connection with FIG. 5.

[0055] In cases where the set of context-dependent orientation directions includes non-walking and non-running motion (e.g., motion that is not substantially linear and / or forward, but includes leaning, turning, etc.), the non-walking and non-running target direction may be determined as the direction the user is currently facing or was facing in a recent time window. Example time windows include 0.2 seconds to 3 seconds, 0.1 seconds to 4 seconds, etc. Note that the time window used to determine the target direction for non-walking and non-running motion may be shorter than the time window used to determine the target direction for a listener situation with minimal motion. In some embodiments, the time window used to determine the target direction for non-walking and non-running motion may be an order of magnitude shorter than the time window used to determine the target direction for a listener situation with minimal motion.

[0056] At 404, process 400 can determine a current listener situation. The current listener situation can be determined based on sensor data obtained from one or more sensors disposed in or on headphones worn by the listener. The one or more sensors can include one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof. The sensor data can include the listener's current movement, current direction of movement, current orientation, etc. In some implementations, the current listener situation can be determined by applying the sensor data to a trained machine learning model configured to output a classification indicating a plausible current listener situation from a set of possible current listener situations. Note that the set of possible current listener situations can correspond to the set of situation-dependent orientation directions. For example, if the set of situation-dependent orientation directions includes a walking or running target direction, a minimal-motion target direction, and a non-walking or non-running motion target direction, the set of possible current listener situations can include walking or running, minimal movement, and non-walking or non-running movement.

[0057] In some embodiments, the current listener situation may be determined based at least in part on data from a user device paired with headphones worn by the user. For example, data from the user device may include whether the user is currently or recently interacted with the user device, current movement information provided by a motion sensor or GPS sensor in the user device (which may indicate that the user device is currently in a moving vehicle, that the user device is currently moving in a direction and / or at a speed that suggests the listener is running or walking, etc.). In some embodiments, the current listener situation may be determined based at least in part on data obtained from a microphone, camera, or other sensor of a paired user device. User devices may include mobile devices (e.g., mobile phones, tablet computers, laptop computers, vehicle entertainment systems, etc.) and non-mobile devices (e.g., desktop computers, televisions, video game systems, etc.).

[0058] At 406, process 400 may select a context-dependent orientation direction from the set of context-dependent orientation directions. For example, process 400 may select a context-dependent orientation direction that corresponds to the current listener's context determined at block 404. As an example, if the current listener's context corresponds to a walking or running activity, process 400 may select a context-dependent orientation direction that corresponds to a walking or running activity (e.g., may correspond to the direction in which the listener is currently walking or running, as determined at block 402). The selected context-dependent orientation direction is generally referred to herein as α target The selected context-dependent orientation direction can be considered as the selected target direction.

[0059] At 408, the process 400 selects a selected azimuthal direction (e.g., α target (k)) and the previous sound field direction (e.g., α fwd (k-1)) exceeds a predetermined threshold. For example, in some embodiments, process 400 may determine whether the difference between the selected context-dependent azimuth angle (e.g., α target (k)) and the previous sound field direction (e.g., α fwd In one example, the process 400 may determine the difference γ(k) by:

number

[0060] In the formula given above, the ModC() function may function to remove the effects of periodic angle wrap-around by transforming the given angle to be in the range [-180, 180). In one example, ModC(x) may be determined by:

number

[0061] In some embodiments, process 400 may determine that the current listener's situation is different from the previous listener's situation if γ(k) exceeds a predetermined threshold, which may be ±3 degrees, ±5 degrees, ±10 degrees, ±15 degrees, etc.

[0062] If at 408 process 400 determines that the difference between the selected azimuthal direction and the previous sound field orientation exceeds a predetermined threshold (“yes” at 408), process 400 may proceed to block 410 and determine the sound field orientation based on a maximum rate of change of the sound field orientation. For example, in some embodiments, process 400 may determine the sound field rotation angular velocity based on a maximum allowable angular velocity of the target direction. In some embodiments, the sound field rotation angular velocity (generally referred to herein as δ) may be determined based on a maximum allowable angular velocity of the target direction. fwd The sound field rotation angular velocity can be determined as a fixed angular velocity based on the maximum allowable angular velocity. Note that the sound field rotation angular velocity is determined by the direction of the sound field per sample period (e.g., α fwd ) can be used to indicate the change in the direction of the sound field (e.g., α fwd ) may be changed smoothly (e.g., based on changes in the user's head orientation and / or changes in the listener's situation) and without jumps in the sound field orientation that a listener may perceive as discontinuous and / or jumpy. In one example, the sound field rotation angular velocity may be referred to herein generally as ω cap , which may be a constant that depends on the maximum allowable angular velocity, expressed as: Exemplary maximum allowable angular velocities are 10 degrees / second, 30 degrees / second, 50 degrees / second, etc. In one example, the sound field rotation rate may be determined by:

number

[0063] Next, the sound field direction α fwd (k) is the previous sound field direction (e.g., α fwd (k-1)) and the sound field rotation angular velocity (e.g., δ fwd (k)). For example, in some implementations, the orientation of the sound field at the current time may be determined by changing the orientation of the previous sound field based on the determined sound field rotation angular velocity. As an example, in some implementations, the orientation of the sound field at the current time may be determined by:

number

[0064] Conversely, if at 408 the process 400 determines that the selected azimuthal direction and the previous sound field orientation do not exceed the predetermined threshold (“no” at 408), the process 400 proceeds to block 412 and determines the previous sound field orientation (e.g., α fwd (k-1)) in a selected azimuth direction (e.g., α target (k)), the direction of the sound field at the current time (e.g., α fwd (k)) can be determined. For example, the sound field rotation angular velocity δ fwd (k) is the difference between the selected azimuthal direction and the previous sound field orientation (e.g., γ(k)) and the smoothing time constant (generally τ cap In some embodiments, τ cap can be in the range of about 0.1 to 5 seconds. cap Example values ​​for include 0.1 seconds, 0.5 seconds, 2 seconds, 5 seconds, etc. The smoothing time constant may ensure that the updated sound field orientation is changed from the previous sound field orientation in a relatively smooth manner that approximately tracks the changing situational target direction. In one example, the sound field rotation angular velocity may be determined by:

number

[0065] Similar to what was discussed above in relation to block 410, the sound field orientation α fwd (k) is then calculated based on the previous sound field orientation (e.g., α fwd (k-1)) and the sound field rotation angular velocity (e.g., δ fwd (k)). For example, in some implementations, the sound field orientation at the current time may be determined by changing the previous sound field orientation based on the determined sound field rotation angular velocity. In particular, in some implementations, the sound field orientation may be determined based on the rate of change of the sound field orientation per sample field (e.g., δ fwd (k)) by updating the previous value of the sound field orientation based on (k). As an example, in some implementations, the sound field orientation at the current time may be determined by:

number

[0066] FIG. 5 is an example of a chat illustrating the change in sound field orientation based on the listener's situation, according to some embodiments. Note that curve 502 (az in FIG. 5) fast ) represents the user's orientation as a function of time, and curve 504 (denoted as az in FIG. 5) target ) represents the azimuth direction selected based on the current listener situation (e.g., the selected target direction), and curve 506 (represented as az in FIG. 5 ) represents the azimuth direction selected based on the current listener situation (e.g., the selected target direction). fwd ) represents the determined sound field orientation. During time period 508, the listener is in a static or minimally moving listening situation. Thus, the selected azimuth direction (represented by curve 504) remains stationary during time window 508, even though the user's head orientation (illustrated by curve 502) moves within approximately ±20 degrees. Because there is no change in the selected azimuth direction during time window 508, the sound field orientation (represented by curve 506) accurately tracks the selected azimuth direction (note that curves 504 and 506 overlap during time window 508).

[0067] Referring to time window 510, the listener's situation changes to a non-walking or non-running motion activity. Note that there is a relatively large and sudden change in the user's head orientation, as represented by curve 502 within time window 508. Because the target direction during a non-walking or non-running motion activity approximately tracks the user's head orientation, the selected azimuth direction (represented by curve 504) approximately tracks the user's head orientation (note that curves 502 and 504 substantially overlap during time window 510). However, during time window 510, because the difference between the selected azimuth direction (represented by curve 504) and the previous sound field orientation (represented by curve 506) exceeds a threshold, the sound field orientation is incrementally adjusted toward the selected azimuth direction. In particular, note that curve 506 slopes toward curve 504 during time window 510. The time window during which the sound field is incrementally oriented towards a selected azimuthal direction may be referred to herein as the "approach state" or "non-capture state."

[0068] Referring to time window 512, the sound field orientation (represented by curve 506) now coincides with the selected azimuth direction (represented by curve 504). Therefore, the sound field orientation is then adjusted approximately to coincide with the most recent direction the user was facing within the previous time window. Note that during time window 512, the sound field orientation is adjusted smoothly. The time window is sometimes referred to herein as the "main state" or "capture state."

[0069] In some embodiments, a context-dependent orientation direction for a static or minimally moving listening situation may be determined such that the orientation direction corresponds to the direction the user was facing within a recent time window. In some embodiments, a smoothed representation of the user's head orientation (e.g., smoothed over time) may be determined, and the context-dependent orientation direction may approximately track the smoothed representation of the user's head orientation, thereby allowing the orientation direction to smoothly track the user's orientation. In situations where the user's head orientation changes rapidly by more than a predetermined threshold, the smoothed representation of the user's head orientation may jump in a discontinuous manner. However, in such cases, the orientation direction may be determined based on previous orientation directions to provide smoothing of the sound field orientation regardless of the discontinuous jumps in the user's head orientation. In some embodiments, the smoothed representation of the user's head orientation may be determined based on the angular velocity of a paired user device (e.g., a paired mobile phone, a paired tablet computer, a paired laptop computer, etc.). For example, a paired user device may be paired with headphones worn by a user to present, e.g., video or audio content. By utilizing angular velocity measurements from the paired device, the orientation of the user's head and resulting azimuth direction may be altered to account for changes in the user's orientation relative to an external (e.g., fixed) frame of reference. For example, utilizing angular velocity measurements from the paired device may allow the sound field to be rotated when the user is riding in a vehicle that has turned or otherwise changed direction.

[0070] FIG. 6 is a flowchart of an example process 600 for determining a context-dependent azimuth direction for a static listening situation, according to some embodiments. In some embodiments, the blocks of process 600 may be performed by a processor or controller of headphones worn by a listener or a processor or controller of a user device paired with headphones worn by a listener. An example of such a processor or controller is shown in FIG. 7 and described below in connection with FIG. 7. In some implementations, the blocks of process 600 may be performed in an order other than the order shown in FIG. 6. In some embodiments, two or more blocks of process 600 may be performed substantially in parallel. In some embodiments, one or more blocks of process 600 may be omitted.

[0071] Process 600 may begin at block 602 by determining the current user's head orientation, which is generally referred to herein as α nose As noted above, the current user's head orientation may be determined from one or more sensors located in or on headphones worn by the user. The one or more sensors may include one or more accelerometers, one or more gyroscopes, one or more magnetometers, or any combination thereof.

[0072] In some embodiments, at 604, the process 600 can determine orientation information of a paired user device (e.g., paired with headphones worn by a listener). Examples of paired user devices include mobile phones, tablet computers, laptop computers, etc. The orientation information can be represented as ω dev,x , ω dev,y , and ω dev,zThe angular velocity information may include the angular velocity of the user device, represented by angular velocity components about the x-, y-, and z-axes of a fixed external frame, represented as . The angular velocity information may indicate whether the user device is moving relative to the external (e.g., fixed) frame. For example, such angular velocity information may indicate movement as the vehicle in which the listener is riding moves and / or turns. The angular velocity information may indicate whether the user device is currently being used and / or handled by the listener. The azimuthal contribution of the laser device can be determined, where the azimuthal direction is ω dev (k). In one example, the azimuthal contribution can be determined by:

number

[0073] In the above formula, ω dev_max_tilt represents the maximum allowable angular velocity of the companion device's tilt. F (k)3 is the vector T F (k), this vector represents the direction in which the top of the user's head is pointing relative to a fixed external reference frame.

[0074] If the companion device is not handled during normal vehicle movement, the vehicle's turning angular velocity is ω dev,Z (k). In some embodiments, ω dev,X (k) and ω dev,Y Excessive tilt of the device, indicated by a large non-zero value of (k), may indicate that the device is being handled by a user. In some embodiments, the threshold ω for detecting tilt of the companion device dev_max_tilt The value of ω can be in the range of about 10 degrees / second to 200 degrees / second. dev_max_tilt is 60 degrees / second.

[0075] At 606, the process 600 generates a smoothed representation of the user's head orientation (generally referred to herein as α followFor example, in some embodiments, the process 600 can first determine the current user head orientation (e.g., α nose (k)) and a smoothed representation of the user's head orientation at a previous time point (generally referred to herein as α follow For example, the difference between the k and k k k − 1 k ... follow The difference, denoted (k), can be determined by:

number

[0076] As discussed above in connection with FIG. 4, the ModC function may be bounded by difference angles within the range [-180 degrees, 180 degrees).

[0077] A smoothed representation of the user's head orientation may then be determined based on the difference between the user's current head orientation and a smoothed representation of the user's head orientation at a previous time point. For example, if the difference is less than a predetermined threshold (generally δ follow_max If the difference is less than the predetermined threshold (e.g., δ), the smoothed representation of the user's head orientation may be modified to approximately track the difference using a smoothing time constant. Conversely, if the difference is greater than a predetermined threshold, the smoothed representation of the user's head orientation may be set to jump discontinuously to the current user's head orientation. follow_max ) are 2 degrees, 5 degrees, 10 degrees, 15 degrees, 20 degrees, 25 degrees, etc. In one example, a smoothed representation of the user's head orientation may be determined by:

number

[0078] In the above formula, τ follow is a value in the range of approximately 5 to 50 seconds, and may have a value of, for example, 10 seconds, 20 seconds, 30 seconds, etc.

[0079] In the case where orientation information for the paired user device is obtained in block 604, process 600 may utilize the orientation information to determine a smoothed representation of the user's head orientation. For example, in some embodiments, process 600 may include an orientation contribution to the angular velocity of the paired user device in the smoothed representation of the user's head orientation. As a more specific example, the smoothed representation of the user's head orientation may be incrementally adjusted toward the user's current orientation and the direction of movement of the paired user device. In one example, the smoothed representation of the user's head orientation may be determined by:

number

[0080] At 608, process 600 can determine an azimuthal sound field orientation for a static (e.g., minimally moving) listening situation based on the smoothed representation of the user's head orientation, which is generally referred to herein as α static , and may be included in the set of context-dependent azimuthal sound field orientations as shown in and described in connection with FIG. 4. For example, the difference (generally referred to herein as δ ) between the user's current head orientation and a smoothed representation of the user's head orientation at a previous time instant may be follow (k)) is determined by a predetermined threshold (e.g., a movement tolerance threshold, generally referred to herein as δ follow_max If the azimuthal sound field orientation is less than ∑ i = 1 ⁢ ...

number

[0081] In cases where orientation information from paired user devices is obtained (e.g., at block 604), the azimuthal sound field orientation may be determined based at least in part on the azimuthal contribution of the angular velocity of the paired user devices. In some embodiments, the angular velocity of the paired user devices is determined when the difference between the user's head orientation at the current time and a smoothed representation of the user's head orientation at a previous time is less than a predetermined threshold (e.g., δ follow_max ) may be utilized only if the angular velocity of the paired user device exceeds a predetermined threshold (e.g., δ). In other words, the angular velocity of the paired user device may be used to update the azimuthal sound field orientation when there is a relatively large change in the user's head orientation. In one example, the difference between the user's head orientation at the current time and a smoothed representation of the user's head orientation at a previous time is less than a predetermined threshold (e.g., δ follow_max ), the azimuthal sound field orientation can be determined by:

number

[0082] As mentioned above, the azimuthal sound field orientation α for a static listening situation static can then be used as a possible situation-dependent azimuthal sound field orientation that can be selected depending on the current listener situation, as shown in FIG. 4 and described above in connection with FIG.

[0083] 7 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types and number of elements shown in FIG. 7 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 700 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 700 may be or include a television, one or more components of an audio system, a mobile device (such as a mobile phone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0084] According to some alternative implementations, apparatus 700 may be or include a server. In some such examples, apparatus 700 may be or include an encoder. Thus, in some examples, apparatus 700 may be a device configured for use in an audio environment, such as a home audio environment, while in other examples, apparatus 700 may be a device configured for use in the "cloud," e.g., a server.

[0085] In this example, device 700 includes an interface system 705 and a control system 710. The interface system 705, in some implementations, may be configured to communicate with one or more other devices in an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a road or sidewalk environment, a park environment, etc. The interface system 705, in some implementations, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, may relate to one or more software applications that device 700 is executing.

[0086] In some implementations, the interface system 705 may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, a voice signal. In some examples, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0087] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. 7. In some implementations, interface system 705 may include one or more devices for implementing a user interface, such as a microphone system. In some examples, interface system 705 may include one or more interfaces between control system 710 and a memory system, such as optional memory system 715 shown in FIG. 7. However, control system 710 may include a memory system in some examples. Interface system 705, in some implementations, may be configured to receive input from one or more microphones in the environment.

[0088] The control system 710 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0089] In some implementations, control system 710 may reside in more than one device. For example, in some implementations, a portion of control system 710 may reside in a device in one of the environments illustrated herein, while another portion of control system 710 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone, or a tablet computer). In other examples, a portion of control system 710 may reside in a device within one environment, while another portion of control system 710 may reside in one or more other devices in the environment. For example, a portion of control system 710 may reside in a device implementing a cloud-based service, such as a server, while another portion of control system 710 may reside in another device implementing the cloud-based service, such as another server, memory device, etc. Also, interface system 705 may reside in more than one device in some examples.

[0090] In some implementations, the control system 710 may be configured to perform, at least in part, the methods disclosed herein. According to some examples, the control system 710 may be configured to implement a method for determining a user's orientation, a method for determining a user's listening situation, etc.

[0091] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices, such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 715 shown in FIG. 7 and / or control system 710. Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include instructions, for example, for determining a direction of movement, determining a direction of movement based on a direction orthogonal to the direction of movement, etc. The software may be executable by one or more components of a control system, such as control system 710 of FIG. 7.

[0092] In some examples, the device 700 may include an optional microphone system 720 shown in FIG. 7. The optional microphone system 720 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, or the like. In some examples, the device 700 may not include the microphone system 720. However, some such devices may include a microphone system 720. In implementations, device 700 may nevertheless be configured to receive microphone data for one or more microphones in the audio environment via interface system 710. In some such implementations, a cloud-based implementation of device 700 may be configured to receive microphone data, or noise metrics that at least partially correspond to the microphone data, from one or more microphones in the audio environment via interface system 710.

[0093] According to some implementations, device 700 may include an optional loudspeaker system 725 shown in FIG. 7. Optional loudspeaker system 725 may include one or more loudspeakers, which may also be referred to herein as "speakers" or, more generally, "audio reproduction transducers." In some examples (e.g., cloud-based implementations), device 700 may not include loudspeaker system 725. In some implementations, device 700 may include headphones. Headphones may be connected or coupled to device 700 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).

[0094] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media (e.g., disks) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can include a programmable general-purpose processor, digital signal processor, or microprocessor that can be programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor can be or include a computer system that includes input devices, memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0095] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal, including performing one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, the system also including other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0096] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code for performing (e.g., executable code) one or more examples of the disclosed methods or steps thereof.

[0097] Specific embodiments of the present disclosure and applications of the present disclosure have been described herein, and It will be apparent to those skilled in the art that many variations of the above embodiments and applications are possible without departing from the scope of the disclosure as described and claimed herein. While certain forms of the present disclosure have been illustrated and described, it is to be understood that the present disclosure is not limited to the specific embodiments described and shown or to the specific methods described.

Claims

[Claim 1] 1. A method for determining sound field rotation, comprising: (a) determining a current activity state of a user, the current activity state indicating whether the user is currently moving and the type of movement the user is engaged in; (b) determining an orientation of the user's head using at least one sensor of the one or more sensors; (c) determining a target direction based on the activity status and the user's head orientation; (d) determining a rotation of a sound field used to present an audio object through headphones based on the target direction; The method includes: