A spatial soundscape adaptive control method with head tracking

By performing time alignment and filtering on head posture data, a head state sequence is constructed and multi-scale decomposition is performed to generate a perceptual factor vector, construct a spatially perceptual stable potential field, and adjust the sound source parameters. This solves the problem of unstable sound image localization in head tracking spatial audio control methods under complex motion scenarios, and improves sound image stability, response balance, and adaptive capability.

CN122093735APending Publication Date: 2026-05-26SHENZHEN XINGMAN SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XINGMAN SMART TECH CO LTD
Filing Date
2026-04-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing head-tracking spatial audio control methods struggle to distinguish between conscious rotation and unconscious jitter in complex motion scenarios, leading to unstable sound image localization. Furthermore, a single control parameter cannot simultaneously address both response speed and positioning accuracy.

Method used

By acquiring head posture data, performing time alignment and filtering, a head state sequence is constructed, and multi-scale decomposition is performed to generate spatial perturbation factors, motion trend factors, and gaze lock factors. A spatially perceptual stable potential field is constructed, and sound source parameters are adjusted to generate a set of target soundscape parameters.

Benefits of technology

It improves the stability and response balance of the audio-visual system, enhances the system's adaptability, reduces computational complexity, and is suitable for real-time implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093735A_ABST
    Figure CN122093735A_ABST
Patent Text Reader

Abstract

This invention discloses a spatial soundscape adaptive control method with head tracking, relating to the fields of spatial audio processing and human-computer interaction. The method includes: acquiring head posture data and performing time alignment and filtering to obtain a head state sequence; based on the head state sequence, performing multi-scale decomposition within a preset time window to obtain a multi-scale behavioral feature set; constructing a perceptual factor vector representing the head motion state based on the multi-scale behavioral feature set; constructing a spatial perceptual stable potential field based on the perceptual factor vector and the current head posture; and constraining the sound source parameters in the spatial soundscape within the sound image change tolerance range defined by the spatial perceptual stable potential field to generate a target soundscape parameter set. By distinguishing between conscious rotation and unconscious jitter, dynamic adaptive control is achieved, balancing sound image stability and response sensitivity, enhancing the immersive experience, and is computationally efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spatial audio processing and human-computer interaction technology, specifically to a spatial soundscape adaptive control method with head tracking. Background Technology

[0002] With the development of virtual reality, augmented reality, and immersive multimedia technologies, spatial audio technology has gradually become an important means to enhance the user's immersive experience. By controlling the direction, distance, and acoustic characteristics of sound sources, a real or virtual auditory environment can be reconstructed in three-dimensional space, giving users a more spatial and directional audio experience. In this process, head tracking technology has been widely introduced to acquire changes in the user's head posture, thereby enabling the sound field to dynamically adjust according to the user's viewpoint.

[0003] Currently, mainstream head-tracking spatial audio technologies typically employ inertial measurement units or visual sensors to acquire the user's head posture information and then render the audio signal in both ears based on this information. In practical applications, the system needs to adjust the sound image position in real time according to head rotation to achieve a sound field following effect.

[0004] However, existing head-tracking spatial audio control methods still have certain limitations when facing complex motion scenarios. For example, when users are walking, running, or in a vibrating environment, head posture data often contains a mixture of conscious rotation and involuntary shaking. Existing methods struggle to effectively distinguish between these two motion components, potentially leading to unnecessary swaying or response delays in sound image localization. Furthermore, users have different needs for sound image stability in different usage scenarios, and a single control parameter cannot simultaneously address both response speed and localization accuracy.

[0005] To improve the adaptability of spatial audio systems, it is necessary to conduct more detailed analysis of head movement characteristics and to dynamically adjust the control strategy of soundscape parameters based on real-time movement status. Summary of the Invention

[0006] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a spatial audio-visual adaptive control method with head tracking to solve the above-mentioned technical problems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a spatial soundscape adaptive control method with head tracking, comprising: S1: Acquire head pose data output by the head tracking device, and perform time alignment and filtering on the head pose data to construct a head state sequence; S2: Based on the head state sequence, perform multi-scale decomposition processing within a preset time window to obtain a multi-scale behavioral feature set that reflects the head movement characteristics. S3: Based on a multi-scale behavioral feature set, construct a perceptual factor vector representing the head movement state. The perceptual factor vector includes: spatial perturbation factor, motion trend factor, and gaze lock factor. S4: Based on the perception factor vector and the current head posture, construct a spatial perception stable potential field to characterize the tolerance distribution of acoustic image changes in different spatial directions. S5: Based on the spatially perceived stable potential field and the head state sequence, within the tolerance range of sound image changes defined by the spatially perceived stable potential field, the sound source parameters in the spatial soundscape are constrained and adjusted to generate the target soundscape parameter set.

[0008] The present invention is further configured such that the process of acquiring head pose data and constructing a head state sequence in S1 includes: Acquire raw data from the head tracking device, including head attitude angles, angular velocities, and corresponding timestamps; By resampling the original data according to a unified time base, sequence data with consistent time intervals is formed. The resampled sequence data is filtered to suppress noise and maintain the continuity of attitude changes; A head state sequence is constructed based on the filtered sequence data, the head state sequence including the smoothed filtered head attitude angle and angular velocity.

[0009] The present invention is further configured such that S2 includes: The head state sequence is decomposed in the frequency domain within a preset time window to distinguish between high-frequency and low-frequency components, and low-frequency motion components and high-frequency motion components are extracted respectively. By performing energy statistical analysis on high-frequency motion components, high-frequency energy characteristics that characterize the degree of head shaking can be obtained. By performing trend fitting analysis on low-frequency motion components, trend slope characteristics that characterize the head orientation movement trend are obtained.

[0010] The present invention is further configured such that S2 further includes: By performing a stability assessment method on the head state sequence, a stability index characterizing the degree of head stability is obtained. The high-frequency energy characteristics, trend slope characteristics, and stability indicators are fused to generate a multi-scale behavioral feature set.

[0011] The present invention is further configured such that S3 includes: By normalizing the high-frequency energy in the multi-scale behavioral feature set and using threshold mapping and interval partitioning methods to hierarchically represent the normalization results, a spatial perturbation factor for describing the degree of spatial perturbation is obtained. By performing directional consistency analysis and amplitude constraint processing on the trend slope in the multi-scale behavioral feature set, and using the trend maintenance assessment method to characterize its stability, a motion trend factor for describing head movement trends is obtained. By performing temporal continuity analysis on stability indices in a multi-scale behavioral feature set and smoothing them using fluctuation characteristics within a sliding window, a gaze lock factor is obtained to describe the degree of gaze concentration.

[0012] The present invention is further configured such that S3 further includes: By performing a uniform scale transformation on spatial disturbance factors, motion trend factors, and gaze lock factors, and combining them in a preset order, a perception factor vector is generated.

[0013] The present invention is further configured such that S4 includes: Construct a spatial angle domain model with the head orientation as a reference based on the current head posture; By discretizing the spatial angle domain, multiple spatial direction units are formed; The initial distribution values ​​of each spatial orientation unit are assigned based on the head posture to form the initial distribution of the spatial sensing potential field. The initial distribution values ​​are generated by a preset distribution function based on the head orientation.

[0014] The present invention is further configured to perform multi-factor modulation processing on the initial distribution of the spatial sensing potential field based on the sensing factor vector to generate an intermediate distribution, wherein the multi-factor modulation processing specifically includes: Based on the spatial perturbation factor, the extent of expansion of the initial distribution of the spatially sensed potential field is adjusted by the distribution range adjustment method; Based on the motion trend factor, the center direction of the initial distribution of the spatially sensed potential field is offset by the direction offset adjustment method. Based on the gaze lock factor, the initial distribution of the spatial perception potential field is concentrated near the current head orientation through a local contraction adjustment method.

[0015] The present invention is further configured to generate a stable spatial sensing potential field for characterizing the sound-image change tolerance distribution by performing a smoothing constraint process on the intermediate distribution of the spatial sensing potential field after multi-factor modulation. The smoothing constraint processing includes: Continuity constraints are applied to the distribution changes between adjacent spatial orientation units to suppress abrupt changes in the intermediate distribution; By smoothing the intermediate distribution, the impact of instantaneous changes on the stability of the potential field can be reduced. Boundary constraints are applied to the smoothed intermediate distribution to ensure that it remains within a preset effective range.

[0016] The present invention is further configured such that S5 includes: Based on the spatially perceived stable potential field, the range of acoustic image change tolerance corresponding to each spatial direction is determined by performing interval mapping processing on the potential field distribution values ​​corresponding to the spatial direction units. The sound source parameters of each sound source in the spatial soundscape are obtained based on the preset sound source management module, and the coordinate transformation of each sound source parameter is performed in combination with the head state sequence to determine the spatial position relationship of each sound source relative to the current head posture. The sound source parameters include: sound source direction, smoothing coefficient, maximum angular velocity, damping coefficient, sound source weight, and HRTF selection index. Based on the tolerance range of sound image changes corresponding to each spatial direction, the parameters of each sound source are subject to constraint adjustments, wherein the constraint adjustments include: The direction of the sound source is adjusted by offset constraint method based on the allowable offset range defined by the current head posture, angular velocity and spatial perception stable potential field; The smoothing coefficient is adaptively adjusted based on the spatial perturbation factor and gaze lock factor in the perception factor vector through a coefficient mapping adjustment method. The maximum angular velocity is limited by the distribution characteristics of the spatially sensed stable potential field and is restricted by the amplitude constraint adjustment method. The damping coefficient is adjusted using a damping adjustment method based on the sound source direction deviation and spatial disturbance factor; The sound source weight is adjusted based on the distance of the sound source direction relative to the current head orientation, using a weighting method. The HRTF selection index is determined by a lookup mapping method based on the adjusted sound source direction. The sound source parameters, after being constrained and adjusted, are integrated to generate a set of target soundscape parameters.

[0017] This invention provides a spatial soundscape adaptive control method with head tracking. The method comprises: S1: acquiring head posture data output from a head tracking device and performing time alignment and filtering on the head posture data to construct a head state sequence; S2: based on the head state sequence, performing multi-scale decomposition processing within a preset time window to obtain a multi-scale behavioral feature set reflecting head movement characteristics; S3: based on the multi-scale behavioral feature set, constructing a perception factor vector representing the head movement state, the perception factor vector including: spatial perturbation factor, motion trend factor, and gaze lock factor; S4: based on the perception factor vector and the current head posture, constructing a spatial perception stable potential field to represent the sound-image change tolerance distribution in different spatial directions; S5: based on the spatial perception stable potential field and the head state sequence, within the sound-image change tolerance range defined by the spatial perception stable potential field, constraining the sound source parameters in the spatial soundscape to generate a target soundscape parameter set. The beneficial effects include: Improved audio-visual stability and response balance: By decomposing head movements at multiple scales, it effectively distinguishes between conscious rotation and unconscious shaking, maintaining a rapid response while suppressing audio-visual sway caused by disturbances, thus achieving a balance between stability and sensitivity.

[0018] Enhanced Adaptability and Immersive Experience: Based on the dynamic modulation of the spatial potential field by the perceptual factor, the system can adaptively adjust the sound and image control strategy according to the user's motion state (degree of disturbance, motion trend, degree of gaze) to maintain a natural and comfortable auditory experience in different scenarios.

[0019] Reduced computational complexity, suitable for real-time implementation: Discretized potential field modeling and table lookup operations are used to avoid complex iterative calculations. The overall algorithm is lightweight and efficient, and easy to run in real time on embedded devices.

[0020] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 The flowchart illustrates a spatial soundscape adaptive control method with head tracking, as an exemplary embodiment of the present invention. Detailed Implementation

[0022] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0023] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0024] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0025] Example: A spatial soundscape adaptive control method with head tracking, such as Figure 1 As shown, it includes: S1: Acquire head pose data output by the head tracking device, and perform time alignment and filtering on the head pose data to construct a head state sequence; S2: Based on the head state sequence, perform multi-scale decomposition processing within a preset time window to obtain a multi-scale behavioral feature set that reflects the head movement characteristics. S3: Based on a multi-scale behavioral feature set, construct a perceptual factor vector representing the head movement state. The perceptual factor vector includes: spatial perturbation factor, motion trend factor, and gaze lock factor. S4: Based on the perception factor vector and the current head posture, construct a spatial perception stable potential field to characterize the tolerance distribution of acoustic image changes in different spatial directions. S5: Based on the spatially perceived stable potential field and the head state sequence, within the tolerance range of sound image changes defined by the spatially perceived stable potential field, the sound source parameters in the spatial soundscape are constrained and adjusted to generate the target soundscape parameter set.

[0026] The present invention is further configured such that the process of acquiring head pose data and constructing a head state sequence in S1 includes: Acquire raw data from the head tracking device, including head attitude angles, angular velocities, and corresponding timestamps; By resampling the original data according to a unified time base, sequence data with consistent time intervals is formed. The resampled sequence data is filtered to suppress noise and maintain the continuity of attitude changes; A head state sequence is constructed based on the filtered sequence data. This head state sequence includes smoothed and filtered head attitude angles and angular velocities. Specifically, raw data is acquired from the inertial measurement unit (IMU) of the head tracking device. This raw data includes three-axis angular velocities measured by a gyroscope, three-axis accelerations measured by an accelerometer, and a timestamp corresponding to each data point. Before data acquisition, the gyroscope is zero-biased using a static acquisition method: gyroscope data is continuously acquired for 10 seconds while the device is stationary, and the arithmetic mean of the outputs of each axis is calculated as the zero-bias value and stored in non-volatile memory. During actual data retrieval, the raw angular velocity measurement value is subtracted from this zero-bias value to obtain the calibrated angular velocity data. Then, the raw data undergoes time alignment processing, with a target sampling frequency set to 100 Hz, meaning the time interval between adjacent data points is 10 milliseconds. Starting with the timestamp of the first valid data packet, a series of target time points are generated at 10-millisecond intervals. For each target time point, the angular velocity and acceleration values ​​at that moment are calculated using linear interpolation. The two nearest original sampling points before and after the target time point are found, and their values ​​are weighted and averaged according to the time distance ratio between the two original points. This ensures that the closer the target time point is to the original point, the closer the interpolation result is to the value of that point. For target time points located at the boundary of the data sequence and for which it is not possible to obtain adjacent sampling points before and after simultaneously, the nearest neighbor data is used for assignment processing. An angular velocity sequence and an acceleration sequence with consistent time intervals are obtained through resampling. Next, the resampled data is filtered using a moving average filtering method with a filter window length of 5 sampling points. For the angular velocity value at each target time point, the angular velocity values ​​of 2 points before and after that time point, plus the angular velocity value itself (5 points in total), are taken, and their arithmetic mean is calculated as the filtered angular velocity value. For the head attitude angles, the pitch and roll angles are calculated based on the acceleration data, and the yaw angle is determined through attitude fusion processing combined with the angular velocity data, thus obtaining complete head attitude angle information. Then, the head attitude angle sequence is also subjected to a 5-point moving average filtering process to obtain the filtered head attitude angles. The filtering process is applied to both the angular velocity sequence and the head attitude angle sequence. Finally, the filtered head attitude angles and angular velocities are organized into a head state sequence in chronological order. The state vector at each time point contains the yaw angle, pitch angle, roll angle, and angular velocity value at that time. This sequence is stored in array form for subsequent processing modules to call.

[0027] The present invention is further configured such that S2 includes: The head state sequence is decomposed in the frequency domain within a preset time window to distinguish between high-frequency and low-frequency components, and low-frequency motion components and high-frequency motion components are extracted respectively. By performing energy statistical analysis on high-frequency motion components, high-frequency energy characteristics that characterize the degree of head shaking can be obtained. By performing trend fitting analysis on low-frequency motion components, trend slope characteristics that characterize the head orientation movement trend are obtained.

[0028] By performing a stability assessment method on the head state sequence, a stability index characterizing the degree of head stability is obtained. High-frequency energy characteristics, trend slope characteristics, and stability indices are fused to generate a multi-scale behavioral feature set. Specifically, the behavioral analysis window length is first set to 100 milliseconds. Since the sampling frequency of the head state sequence is 100 Hz, meaning one data point is contained every 10 milliseconds, each behavioral analysis window contains head state data from 10 consecutive time points. The head state at each time point includes yaw angle, pitch angle, roll angle, and angular velocity values. This behavioral analysis window slides forward with a period of 100 milliseconds, acquiring the latest window data for subsequent processing. Taking the yaw angle as an example, the pitch angle and roll angle are processed in the same way, and the same processing procedure is performed on the yaw angle, pitch angle, and roll angle respectively. Next, the yaw angle data within the window undergoes frequency band separation processing using a dual-filter differential method: First, the original yaw angle sequence is subjected to strong smoothing, specifically using a 5-point moving average. This means that for each moment within the window, the arithmetic mean of the yaw angle value is taken from the two points before and after that moment, plus the value itself, for a total of five points, as the low-frequency motion component. For moments at the beginning and end of the window where five points cannot be obtained, the average value is calculated using available data points. Then, the corresponding low-frequency motion component value is subtracted from the original yaw angle value to obtain the high-frequency motion component. This frequency domain decomposition processing is equivalently implemented using the aforementioned time-domain filtering and differential method. Then, high-frequency energy characteristics are calculated based on the high-frequency motion component. The instantaneous energy value is obtained by squaring the high-frequency motion component value at each moment within the window. The sum of all instantaneous energy values ​​is then divided by the number of data points within the window, i.e., 10, to obtain the average high-frequency energy value. This value characterizes the degree of head shaking. For example, in typical application scenarios, the average high-frequency energy value is usually less than 5 (degrees²) when the user is standing still, but may rise to over 20 (degrees²) when the user is walking on a bumpy surface. Simultaneously, based on the low-frequency motion components, the trend slope characteristics are calculated. The least squares method is used to linearly fit the low-frequency motion component sequence within the window. The coefficient of the first-order term obtained from the fitting is the trend slope, measured in degrees per second. This value characterizes the speed and direction of head orientation movement; a positive value indicates turning to the right, and a negative value indicates turning to the left. The absolute value represents the turning speed. For example, in a typical application scenario, when a user turns their head to the right at a constant speed of 30 degrees per second, the trend slope is approximately 30 degrees per second. Furthermore, a stability index is calculated based on the original yaw angle data. The variance of the yaw angle values ​​at each moment within the window is calculated. The variance calculation steps are as follows: first, calculate the arithmetic mean of the 10 yaw angle values ​​within the window; then, calculate the square of the difference between each yaw angle value and the mean; sum all the squared values ​​and divide by 10 to obtain the variance value. The smaller the variance value, the more stable the head. For example, in a typical application scenario, when a user is focused on a single point, the variance value is usually less than 1 (degree²). The larger the variance value, the more violent the head fluctuations. Finally, the high-frequency energy characteristics, trend slope characteristics, and stability indicators are organized in order into a multi-scale behavioral characteristic set.

[0029] The present invention is further configured such that S3 includes: By normalizing the high-frequency energy in the multi-scale behavioral feature set and using threshold mapping and interval partitioning methods to hierarchically represent the normalization results, a spatial perturbation factor for describing the degree of spatial perturbation is obtained. By performing directional consistency analysis and amplitude constraint processing on the trend slope in the multi-scale behavioral feature set, and using the trend maintenance assessment method to characterize its stability, a motion trend factor for describing head movement trends is obtained. By performing temporal continuity analysis on stability indices in a multi-scale behavioral feature set and smoothing them by combining fluctuation characteristics within a sliding window, a gaze lock factor is obtained to describe the degree of gaze concentration. A perception factor vector is generated by uniformly scaling the spatial disturbance factor, motion trend factor, and gaze lock factor, and then combining them in a preset order. Specifically, three feature values ​​are first obtained from the multi-scale behavioral feature set output in step S2: high-frequency energy, trend slope, and stability index. These are obtained by fusing the feature values ​​corresponding to yaw angle, pitch angle, and roll angle using a weighted average method, with the weights for each direction set by a preset weight or an equal weight. The high-frequency energy is expressed in degrees squared, representing the degree of head shaking; the trend slope is expressed in degrees per second, representing the speed and direction of head orientation movement; and the stability index is expressed in degrees squared, representing the degree of head stability. Next, the spatial perturbation factor is calculated. Before normalizing the high-frequency energy, threshold mapping and interval division are performed on the high-frequency energy: the high-frequency energy is divided into multiple intervals according to a preset threshold, and different intervals are assigned corresponding hierarchical weights to obtain the hierarchical high-frequency energy characterization value; on this basis, normalization is then performed, specifically by dividing the high-frequency energy by the sum of the high-frequency energy and a preset constant, denoted as K1, with a value of 10 degrees squared. K1 is a calibration parameter preset based on the statistical distribution range of head movement characteristics. The value of this constant is determined according to the distribution range of high-frequency energy in typical application scenarios. When the high-frequency energy is equal to 10 degrees squared, the spatial perturbation factor is 0.5; when the high-frequency energy approaches 0, the spatial perturbation factor approaches 0; when the high-frequency energy is much greater than 10 degrees squared, the spatial perturbation factor approaches 1. Then, a motion trend factor is calculated. This factor quantifies the user's intention to move their head in a directional manner, with a value ranging from -1 to 1. Before performing nonlinear mapping, a directional consistency analysis and amplitude constraint are applied to the trend slope: the consistency of the trend slope sign over multiple consecutive time windows is statistically analyzed to determine directional stability, and the trend slope amplitude is constrained to suppress abrupt changes. Simultaneously, the continuity of trend changes is evaluated using a trend persistence assessment method to obtain a stable trend representation value. Based on this, nonlinear mapping is then performed, specifically implemented as follows: The trend slope is divided by a preset scaling factor, and then a hyperbolic tangent transformation is applied to the quotient. This scaling factor is denoted as K2 and has a value of 50 degrees per second. K2 is a calibration parameter preset according to the statistical range of head rotation speed. It is set according to the normal head rotation speed range so that most normal rotation speeds can be mapped to the sensitive range of the hyperbolic tangent function. When the trend slope is 50 degrees per second, the motion trend factor is approximately 0.76; when the trend slope is 100 degrees per second, the motion trend factor is approximately 0.96; when the trend slope is negative, the motion trend factor is correspondingly negative.Next, the gaze lock factor is calculated. This factor quantifies head stability and ranges from 0 to 1. Before exponential decay processing, a time continuity analysis is performed on the stability index, and smoothing is done by combining the fluctuation characteristics within a sliding window: a time-continuous stability characterization value is obtained by performing a moving average or weighted smoothing on the stability index within multiple consecutive time windows. On this basis, exponential decay processing is then performed. Specifically, the exponential function value of the ratio of the negative stability index to a preset decay constant is calculated. This decay constant is denoted as K3 and has a value of 1 degree squared. K3 is a calibration parameter preset according to the statistical range of head stability. It is set based on the typical variance range of the head in a stable state. When the stability index is 0 degrees squared, the gaze lock factor is 1; when the stability index is 1 degree squared, the gaze lock factor is approximately 0.37; and when the stability index is 2 degrees squared, the gaze lock factor is approximately 0.14. Finally, the calculated spatial perturbation factor, motion trend factor, and gaze lock factor are combined into a perception factor vector in the order of spatial perturbation factor first, motion trend factor in the middle, and gaze lock factor last. This vector fully describes the three key dimensions of the current user's head movement: the degree of environmental disturbance, the orientation movement trend, and the gaze stability. The specific calculation method of the above-mentioned spatial perturbation factor, motion trend factor, and gaze lock factor is a preferred implementation. In other implementations, different normalization functions, nonlinear mapping functions, or attenuation functions can be used to achieve the corresponding functions according to the actual application requirements.

[0030] The present invention is further configured such that S4 includes: Construct a spatial angle domain model with the head orientation as a reference based on the current head posture; By discretizing the spatial angle domain, multiple spatial direction units are formed; The initial distribution values ​​of each spatial orientation unit are assigned based on the head posture to form the initial distribution of the spatial sensing potential field. The initial distribution values ​​are generated by a preset distribution function based on the head orientation. Based on the sensing factor vector, the initial distribution of the spatial sensing potential field is subjected to multi-factor modulation processing to generate an intermediate distribution. Specifically, the multi-factor modulation processing includes: Based on the spatial perturbation factor, the extent of expansion of the initial distribution of the spatially sensed potential field is adjusted by the distribution range adjustment method; Based on the motion trend factor, the center direction of the initial distribution of the spatially sensed potential field is offset by the direction offset adjustment method. Based on the gaze lock factor, the initial distribution of the spatial perception potential field is concentrated near the current head orientation through a local contraction adjustment method. By smoothing and constraining the intermediate distribution of the spatial sensing potential field after multi-factor modulation, a stable spatial sensing potential field for characterizing the tolerance distribution of acoustic-image changes is generated. The smoothing constraint processing includes: Continuity constraints are applied to the distribution changes between adjacent spatial orientation units to suppress abrupt changes in the intermediate distribution; By smoothing the intermediate distribution, the impact of instantaneous changes on the stability of the potential field can be reduced. The smoothed intermediate distribution is then subjected to boundary constraints to ensure it remains within a preset effective range. Specifically, firstly, angular spatial discretization is performed, limiting the spatial angle domain with the head as the reference to a horizontal angle domain. The horizontal angle range from -180 degrees to +180 degrees is divided into 36 continuous and non-overlapping spatial direction units at 10-degree intervals. Each unit corresponds to a discretized angle value; for example, -175 degrees represents the center angle of the range from -180 degrees to -170 degrees, and +175 degrees represents the center angle of the range from +170 degrees to +180 degrees, and so on until the entire angle range is covered. Next, an initial distribution is constructed. Taking the yaw angle in the current head posture as the center direction, a Gaussian distribution function is used to initially assign values ​​to each spatial direction unit. The standard deviation parameter of the Gaussian distribution is denoted as σ0, with a value of 30 degrees. This value indicates that the tolerance is lowest when the acoustic image is directly in front of the head, and the tolerance is higher the further away from the center direction. The initial distribution is calculated as follows: for the angle value of each spatial direction unit, the absolute value of the difference between it and the current head yaw angle is calculated. The square of this difference is negative, and it is divided by twice the square of the standard deviation parameter σ0. The result is then taken as the initial potential value of the unit. The above calculation method for the initial distribution is in Gaussian form. In specific implementation, an exponential decay form can also be selected. Both forms are feasible and do not affect the overall feasibility of the scheme.Then, multi-factor modulation processing is performed. Based on the perception factor vector output in step S3, the initial distribution is adjusted, specifically including three parallel modulation operations: 1. Adjusting the expansion range of the distribution based on the spatial disturbance factor. The preset expansion coefficient is denoted as a, with a value of 2.0. The standard deviation σ1 of the expanded distribution is equal to σ0 multiplied by 1 plus a and the product of the spatial disturbance factor. The larger the spatial disturbance factor, the more diffuse the distribution, indicating that the sound image is allowed to change over a larger range when the environmental disturbance is severe; 2. Adjusting the center direction offset of the distribution based on the motion trend factor. The preset offset coefficient is denoted as b, with a value of 15 degrees. The offset center angle is equal to the current yaw angle plus b and the product of the motion trend factor. When the motion trend factor is positive, the center offsets to the right; when it is negative, it offsets to the left. The larger the absolute value, the larger the offset, indicating that when the user has a clear intention to turn their head. The center of the time field shifts in advance in the direction of rotation; 3. Adjust the degree of contraction of the distribution range based on the gaze lock factor. The preset contraction coefficient is denoted as c and has a value of 0.5. The standard deviation of the final distribution after modulation, σ2, is equal to σ1 multiplied by 1 minus c and the product of the gaze lock factor. The closer the gaze lock factor is to 1, the tighter the distribution contraction, indicating that the sound image is more firmly locked near the current direction when the user focuses on the gaze; The above three modulation operations are performed in sequence according to a unified distribution update process: with the initial distribution as input, the distribution width is first adjusted according to the spatial disturbance factor to obtain the first intermediate distribution; then, with the first intermediate distribution as input, the center direction is shifted according to the motion trend factor to obtain the second intermediate distribution; finally, with the second intermediate distribution as input, the distribution width is contracted according to the gaze lock factor to obtain the modulated intermediate distribution. Next, a smoothing constraint is applied, imposing a continuity constraint on the modulated intermediate distribution between adjacent spatial orientation cells. Specifically, a three-point moving average method is used: for each spatial orientation cell, the potential values ​​of itself and its two adjacent cells (to the left and right) are taken, and the arithmetic mean is calculated as the smoothed potential value of that cell. For cells located near the negative 180 degrees and positive 180 degrees of the boundary, since the left and right adjacent cells are incomplete, only the average value of the available adjacent cells is taken. For example, the leftmost cell only takes the average value of itself and one cell to its right (to the right). This process can suppress the adjacent cells. The distribution undergoes abrupt changes in potential values. Simultaneously, the smoothed distribution is recursively smoothed over time. The smoothed distribution at the current moment is weighted and averaged with the stable potential field output at the previous moment using a 7:3 weighting. That is, the final potential value at the current moment equals 0.7 multiplied by the current smoothed value plus 0.3 multiplied by the potential value at the previous moment, thus reducing the impact of instantaneous changes on the stability of the potential field. Finally, the smoothed distribution is subject to boundary constraints, limiting the potential values ​​of all spatial direction units to the range of 0 to 1. If the potential value of a unit is less than 0, it is set to 0; if it is greater than 1, it is set to 1, ensuring that the distribution remains within a preset effective range.After the above four steps of discretization, initial distribution construction, multi-factor modulation and smoothing constraint processing, the spatial sensing stable potential field at the current time is obtained. This potential field stores the potential values ​​corresponding to 36 spatial direction units in the form of an array, which fully characterizes the tolerance distribution of acoustic image changes in each spatial direction, and serves as the input data for step S5 for subsequent use.

[0031] The present invention is further configured such that S5 includes: Based on the spatially perceived stable potential field, the range of acoustic image change tolerance corresponding to each spatial direction is determined by performing interval mapping processing on the potential field distribution values ​​corresponding to the spatial direction units. The sound source parameters of each sound source in the spatial soundscape are obtained based on the preset sound source management module, and the coordinate transformation of each sound source parameter is performed in combination with the head state sequence to determine the spatial position relationship of each sound source relative to the current head posture. The sound source parameters include: sound source direction, smoothing coefficient, maximum angular velocity, damping coefficient, sound source weight, and HRTF selection index. Based on the tolerance range of sound image changes corresponding to each spatial direction, the parameters of each sound source are subject to constraint adjustments, wherein the constraint adjustments include: The direction of the sound source is adjusted by offset constraint method based on the allowable offset range defined by the current head posture, angular velocity and spatial perception stable potential field; The smoothing coefficient is adaptively adjusted based on the spatial perturbation factor and gaze lock factor in the perception factor vector through a coefficient mapping adjustment method. The maximum angular velocity is limited by the distribution characteristics of the spatially sensed stable potential field and is restricted by the amplitude constraint adjustment method. The damping coefficient is adjusted using a damping adjustment method based on the sound source direction deviation and spatial disturbance factor; The sound source weight is adjusted based on the distance of the sound source direction relative to the current head orientation, using a weighting method. The HRTF selection index is determined by a lookup mapping method based on the adjusted sound source direction. The constrained sound source parameters are integrated to generate a target soundscape parameter set. Specifically, firstly, based on the spatially perceived stable potential field output in step S4, the tolerance range of sound image changes corresponding to each spatial direction is determined. The spatially perceived stable potential field stores the potential values ​​corresponding to 36 spatial direction units in array form. Each potential value is between 0 and 1. The smaller the potential value, the higher the tolerance of sound image changes in that direction; the larger the potential value, the lower the tolerance, i.e., the sound image is more tightly constrained. During interval mapping processing, the potential value interval is first divided into multiple level intervals, including low potential value intervals, medium potential value intervals, and high potential value intervals. Different intervals correspond to different levels of sound image change tolerance, and each interval is mapped to the corresponding allowable offset range based on preset mapping rules. Then, the initial parameters of each sound source in the spatial soundscape are obtained. Each sound source has an initial direction angle, an initial smoothing coefficient, an initial maximum angular velocity, an initial damping coefficient, an initial sound source weight, and an initial HRTF selection index. Combined with the current head yaw angle and angular velocity in the head state sequence output in step S1, the initial direction angle of each sound source is converted into a relative direction angle relative to the current head orientation. The conversion method is to subtract the current head yaw angle from the initial direction angle of the sound source to obtain the azimuth angle of the sound source relative to the head orientation. Next, based on the tolerance range of sound image changes corresponding to each spatial direction, the parameters of each sound source are constrained and adjusted: 1. Sound source direction adjustment: The adjustment of the sound source direction is based on the current head yaw angle, the current angular velocity, and the allowable offset range defined by the spatial perception stable potential field. The angular velocity and sound source direction are both represented based on the same head reference coordinate system to ensure consistency in compensation calculations. First, the relative direction angle of the sound source is predicted and compensated based on the current angular velocity. The preset prediction time coefficient is 10 milliseconds. The compensation amount is equal to the current angular velocity multiplied by the prediction time coefficient. The compensation amount is added to the relative direction angle of the sound source to obtain the predicted sound source direction. Then, based on the predicted sound source direction, the potential value of the corresponding spatial direction unit in the spatial perception stable potential field is found. To avoid the allowable offset range being too small, the potential value is adjusted accordingly. If the abnormal increase is significant, a lower limit ε is set, with ε set to 0.05. When the potential value is less than this lower limit, calculation is performed based on ε. The allowable offset range is calculated based on this potential value. The allowable offset range is equal to the preset allowable offset coefficient divided by the potential value. The preset allowable offset coefficient is 5 degrees. The smaller the potential value, the larger the allowable offset range, and vice versa. Finally, the predicted sound source direction is restricted to the range between the current head yaw angle and the allowable offset range. If the predicted sound source direction exceeds this range, the range boundary value is taken; if it is within the range, it remains unchanged, resulting in the adjusted sound source direction. 2. Smoothing coefficient adjustment: The smoothing coefficient is adjusted adaptively based on the spatial disturbance factor and gaze lock factor in the perception factor vector. The preset baseline smoothing coefficient is 0.5, and the smoothing coefficient adjustment coefficient is 0.3. The adjusted smoothing coefficient equals the base smoothing coefficient plus the smoothing coefficient adjustment coefficient multiplied by the spatial disturbance factor, minus the product of the gaze lock factor and 0.2. The calculation result is limited to the range of 0.1 to 0.9. The larger the spatial disturbance factor, the larger the smoothing coefficient, indicating smoother changes in sound image parameters when environmental disturbances are severe. The larger the gaze lock factor, the smaller the smoothing coefficient, indicating more sensitive response of sound image parameters when focused on the gaze. 3. Maximum angular velocity adjustment: The maximum angular velocity adjustment is limited in amplitude according to the distribution characteristics of the spatial perception stable potential field. By statistically analyzing the potential values ​​of all spatial direction units in the spatial perception stable potential field, the arithmetic mean is calculated to obtain the average potential value. This average potential value is used to characterize the overall spatial constraint strength. The preset maximum angular velocity base value is 150 degrees per second, and the maximum angular velocity adjustment coefficient is... The value is set to 200 degrees per second. The adjusted maximum angular velocity is equal to the maximum angular velocity reference value minus the maximum angular velocity adjustment coefficient multiplied by the average potential value. The calculation result is limited to the range of 30 degrees per second to 150 degrees per second. The larger the average potential value, the tighter the overall potential field, and the smaller the maximum angular velocity, preventing the sound image from moving rapidly under tight constraints. 4. Damping coefficient adjustment: The damping coefficient is adjusted according to the sound source direction deviation and the spatial disturbance factor. The absolute value of the deviation between the adjusted sound source direction and the predicted sound source direction is calculated to obtain the direction deviation. The preset damping coefficient reference value is 0.2, and the damping coefficient adjustment coefficient is 0.3. The adjusted damping coefficient is equal to the damping coefficient reference value plus the damping coefficient adjustment coefficient multiplied by the direction deviation, multiplied by 1, plus the spatial disturbance factor. The calculation result is limited to 0.1 to 0.Within the range of 8, the larger the directional deviation, the larger the damping coefficient; the larger the spatial disturbance factor, the larger the damping coefficient. This indicates that damping is increased to suppress sound image oscillation when the directional deviation is large or the environmental disturbance is severe. 5. Sound source weight adjustment: The sound source weight is adjusted based on the distance of the sound source direction relative to the current head orientation. The absolute value of the difference between the adjusted sound source direction and the current head yaw angle is calculated to obtain the distance angle. The preset weight attenuation coefficient is 45 degrees. The sound source weight is equal to the negative square of the distance angle with the natural constant e as the base, divided by the exponential function value of twice the square of the weight attenuation coefficient. This calculation method makes the weight of the sound source closer to the current head orientation closer to 1, and the weight of the sound source farther away gradually attenuates to close to 0. 6. HRTF selection index determination: The HRTF selection index is determined based on the adjustment. The sound source direction is determined using a lookup mapping method. First, the adjusted sound source direction is mapped to a range of -180 degrees to +180 degrees. If the angle exceeds this range, it is adjusted by adding or subtracting 360 degrees. Then, the HRTF data corresponding to the closest discrete angle in the preset HRTF database is searched. The HRTF database stores head-related transfer function pairs for 72 discrete directions from -180 degrees to +175 degrees at 5-degree intervals. If the adjusted sound source direction matches a discrete angle, the corresponding HRTF index is directly selected. If they do not match, a one-dimensional linear interpolation method is used to take the HRTF data of two adjacent discrete angles in that direction and perform a weighted average. The weights are determined according to the distance ratio between directions, resulting in the interpolated HRTF data, and the corresponding index position is recorded. Finally, the sound source direction, smoothing coefficient, maximum angular velocity, damping coefficient, sound source weight, and HRTF selection index, after the above-mentioned constrained adjustments, are integrated into a target soundscape parameter set in a preset order. This set is stored in data packet form and used as the output of step S5 for the audio rendering module to drive subsequent spatial audio synthesis processing.

[0032] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A spatial soundscape adaptive control method with head tracking, characterized in that, include: S1: Acquire head pose data output by the head tracking device, and perform time alignment and filtering on the head pose data to construct a head state sequence; S2: Based on the head state sequence, perform multi-scale decomposition processing within a preset time window to obtain a multi-scale behavioral feature set that reflects the head movement characteristics. S3: Based on a multi-scale behavioral feature set, construct a perceptual factor vector representing the head movement state. The perceptual factor vector includes: spatial perturbation factor, motion trend factor, and gaze lock factor. S4: Based on the perception factor vector and the current head posture, construct a spatial perception stable potential field to characterize the tolerance distribution of acoustic image changes in different spatial directions. S5: Based on the spatially perceived stable potential field and the head state sequence, within the tolerance range of sound image changes defined by the spatially perceived stable potential field, the sound source parameters in the spatial soundscape are constrained and adjusted to generate the target soundscape parameter set.

2. The spatial audio-visual adaptive control method with head tracking according to claim 1, characterized in that, The process of acquiring head pose data and constructing a head state sequence in S1 includes: Acquire raw data from the head tracking device, including head attitude angles, angular velocities, and corresponding timestamps; By resampling the original data according to a unified time base, sequence data with consistent time intervals is formed. The resampled sequence data is filtered to suppress noise and maintain the continuity of attitude changes; A head state sequence is constructed based on the filtered sequence data, the head state sequence including the smoothed filtered head attitude angle and angular velocity.

3. The spatial audio-visual adaptive control method with head tracking according to claim 1, characterized in that, S2 includes: The head state sequence is decomposed in the frequency domain within a preset time window to distinguish between high-frequency and low-frequency components, and low-frequency motion components and high-frequency motion components are extracted respectively. By performing energy statistical analysis on high-frequency motion components, high-frequency energy characteristics that characterize the degree of head shaking can be obtained. By performing trend fitting analysis on low-frequency motion components, trend slope characteristics that characterize the head orientation movement trend are obtained.

4. The spatial audio-visual adaptive control method with head tracking according to claim 3, characterized in that, S2 further includes: By performing a stability assessment method on the head state sequence, a stability index characterizing the degree of head stability is obtained. The high-frequency energy characteristics, trend slope characteristics, and stability indicators are fused to generate a multi-scale behavioral feature set.

5. The spatial audio-visual adaptive control method with head tracking according to claim 1, characterized in that, S3 includes: By normalizing the high-frequency energy in the multi-scale behavioral feature set and using threshold mapping and interval partitioning methods to hierarchically represent the normalization results, a spatial perturbation factor for describing the degree of spatial perturbation is obtained. By performing directional consistency analysis and amplitude constraint processing on the trend slope in the multi-scale behavioral feature set, and using the trend maintenance assessment method to characterize its stability, a motion trend factor for describing head movement trends is obtained. By performing temporal continuity analysis on stability indices in a multi-scale behavioral feature set and smoothing them using fluctuation characteristics within a sliding window, a gaze lock factor is obtained to describe the degree of gaze concentration.

6. The spatial audio-visual adaptive control method with head tracking according to claim 5, characterized in that, S3 further includes: By performing a uniform scale transformation on spatial disturbance factors, motion trend factors, and gaze lock factors, and combining them in a preset order, a perception factor vector is generated.

7. The spatial audio-visual adaptive control method with head tracking according to claim 1, characterized in that, S4 includes: Construct a spatial angle domain model with the head orientation as a reference based on the current head posture; By discretizing the spatial angle domain, multiple spatial direction units are formed; The initial distribution values ​​of each spatial orientation unit are assigned based on the head posture to form the initial distribution of the spatial sensing potential field. The initial distribution values ​​are generated by a preset distribution function based on the head orientation.

8. A spatial audio-visual adaptive control method with head tracking according to claim 7, characterized in that, Based on the sensing factor vector, the initial distribution of the spatial sensing potential field is subjected to multi-factor modulation processing to generate an intermediate distribution. Specifically, the multi-factor modulation processing includes: Based on the spatial perturbation factor, the extent of expansion of the initial distribution of the spatially sensed potential field is adjusted by the distribution range adjustment method; Based on the motion trend factor, the center direction of the initial distribution of the spatially sensed potential field is offset by the direction offset adjustment method. Based on the gaze lock factor, the initial distribution of the spatial perception potential field is concentrated near the current head orientation through a local contraction adjustment method.

9. A spatial audio-visual adaptive control method with head tracking according to claim 8, characterized in that, By smoothing and constraining the intermediate distribution of the spatial sensing potential field after multi-factor modulation, a stable spatial sensing potential field for characterizing the tolerance distribution of acoustic-image changes is generated. The smoothing constraint processing includes: Continuity constraints are applied to the distribution changes between adjacent spatial orientation units to suppress abrupt changes in the intermediate distribution; By smoothing the intermediate distribution, the impact of instantaneous changes on the stability of the potential field can be reduced. Boundary constraints are applied to the smoothed intermediate distribution to ensure that it remains within a preset effective range.

10. A spatial audio-visual adaptive control method with head tracking according to claim 1, characterized in that, S5 includes: Based on the spatially perceived stable potential field, the range of acoustic image change tolerance corresponding to each spatial direction is determined by performing interval mapping processing on the potential field distribution values ​​corresponding to the spatial direction units. Based on the preset sound source management module, the sound source parameters of each sound source in the spatial soundscape are obtained, and the coordinate transformation processing of each sound source parameter is performed in combination with the head state sequence to determine the spatial position relationship of each sound source relative to the current head posture. The sound source parameters include: sound source direction, smoothing coefficient, maximum angular velocity, damping coefficient, sound source weight and HRTF selection index, where HRTF is the synchronous correlation transfer function. Based on the tolerance range of sound image changes corresponding to each spatial direction, the parameters of each sound source are subject to constraint adjustments, wherein the constraint adjustments include: The direction of the sound source is adjusted by offset constraint method based on the allowable offset range defined by the current head posture, angular velocity and spatial perception stable potential field; The smoothing coefficient is adaptively adjusted based on the spatial perturbation factor and gaze lock factor in the perception factor vector through a coefficient mapping adjustment method. The maximum angular velocity is limited by the distribution characteristics of the spatially sensed stable potential field and is restricted by the amplitude constraint adjustment method. The damping coefficient is adjusted using a damping adjustment method based on the sound source direction deviation and spatial disturbance factor; The sound source weight is adjusted based on the distance of the sound source direction relative to the current head orientation, using a weighting method. The HRTF selection index is determined by a lookup mapping method based on the adjusted sound source direction. The sound source parameters, after being constrained and adjusted, are integrated to generate a set of target soundscape parameters.