Self-adaptive spatial audio rendering method based on head tracking
Through an adaptive spatial audio rendering method based on head tracking, combined with deep learning and complex reverberation models, users' individual differences, real-time and device adaptability problems are solved, and high-quality and real-time spatial audio experience is achieved, especially on TWS headsets.
Patent Information
- Application Number
- CN202510299576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-18
AI Technical Summary
The existing spatial audio rendering technology cannot adapt to individual differences between different users, lacks real-time head tracking mechanisms, cannot truly simulate complex acoustic environments, and performs poorly on TWS headsets.
Adaptive spatial audio rendering method based on head tracking is adopted, and personalized binaural transfer function and audio rendering parameters are generated through deep learning technology and real-time head tracking, and combined with complex reverb models and delay correction, it is optimized for TWS headsets.
It realizes highly personalized, real-time adaptive spatial audio rendering, improves spatial positioning accuracy and immersion, ensures high-quality audio output simultaneous response speed, and expands the application range of portable devices.
Smart Images

Figure CN120343485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adaptive spatial audio rendering methods, and particularly to an adaptive spatial audio rendering method based on head tracking. Background Art
[0002] In recent years, with the booming development of virtual reality, augmented reality, and the gaming industry, spatial audio technology has attracted increasing attention in the industry. Traditional stereo technology can no longer meet the needs of users for immersive audio experiences. Therefore, major technology companies and research institutions have invested resources in developing more advanced spatial audio rendering technologies.
[0003] Currently, the closest prior art mainly uses a method based on head-related transfer functions (HRTF) to achieve spatial audio rendering. This method creates a sense of a three-dimensional sound field for users by simulating the propagation characteristics of sound from the source to the listener's ears. However, this technology still faces many challenges in practical applications.
[0004] Firstly, existing HRTF methods usually use general, preset transfer functions and cannot adapt to the individual differences of different users. As is well known, there are slight differences in the head shape and pinna structure of each person, and these differences will significantly affect the propagation characteristics of sound. Using a unified HRTF will inevitably lead to inaccurate spatial positioning and reduce the realism and immersion of the audio experience.
[0005] Secondly, most existing technologies lack a real-time head tracking mechanism. In application scenarios such as virtual reality, the movement of the user's head will frequently change the position relative to the virtual sound source, and an audio rendering system lacking real-time adjustment cannot reflect these changes in a timely manner, thus destroying the coherence and realism of the audio experience.
[0006] In addition, existing technologies often seem inadequate when dealing with complex acoustic environments. For example, when simulating indoor reverberation effects, overly simplified models are often used, and the propagation characteristics of sound in different spaces cannot be truly restored. This results in the rendered audio lacking details and unable to bring users a sense of being on the scene.
[0007] Finally, with the popularization of TWS earphones, how to achieve a high-quality spatial audio experience on such portable devices has also become an urgent problem to be solved. Most existing technologies have not been optimized for the special acoustic characteristics of TWS earphones, resulting in unsatisfactory performance on such devices. Summary of the Invention
[0008] In view of the above problems, the present invention proposes an adaptive spatial audio rendering method based on head tracking. This method aims to solve a series of problems in the prior art, such as insufficient personalization, poor real-time performance, unrealistic environmental simulation, and poor device adaptability. By introducing innovative elements such as deep learning technology, real-time head tracking, and complex reverberation models, the present invention has successfully achieved highly personalized and real-time adaptive spatial audio rendering.
[0009] The present invention proposes an adaptive spatial audio rendering method based on head tracking, comprising the following steps:
[0010] Step S1: Based on a head tracking device, obtain the user's head orientation and calculate the user's current binaural transfer function;
[0011] Step S2: Based on a deep neural network, train the binaural transfer function for each user according to the user's characteristic information, and generate a personalized binaural transfer function and audio rendering parameters;
[0012] Step S3: Input spatial audio data and render it into left and right channel audio signals;
[0013] Step S4: Perform audio delay correction on the rendered left and right channel audio signals and output the rendered sound effect;
[0014] Step S5: Use an ear canal model to input the audio signal into TWS earphones to complete spatial audio rendering.
[0015] Preferably, the specific steps of step S1 are as follows:
[0016] Step S11: Collect the user's head orientation attitude data, where the head orientation attitude data includes the pitch angle θ and the yaw angle The pitch angle θ is the angle formed by the perpendicular to the head at the center position of the head, and the yaw angle is the angle formed by the horizontal to the head at the center position of the head;
[0017] Step S12: Real-time calculate the user's current binaural transfer function where f is the sound frequency;
[0018] Step S13: Use the head orientation attitude data as the input of the binaural transfer function deep neural network to obtain audio rendering parameters where DNN represents the deep neural network model and P is the audio rendering parameter vector.
[0019] Preferably, the specific steps of step S2 are as follows:
[0020] Step S21: Collect user characteristics U = {G, A, H}, where G is the user's gender, A is the user's age, and H is the head attitude data;
[0021] Step S22: Train a deep learning model M = f(U, HRTF), where M is the trained model and f is a deep learning function;
[0022] Step S23: Adjust the personalized user audio rendering parameter P' = M(U), where P' is the personalized audio rendering parameter.
[0023] Preferably, training the deep learning model in step S22 specifically includes:
[0024] Step S221: Construct the input and output of the deep learning model, where the input of the model is the user feature U and the output is the user's current binaural transfer function
[0025] Step S222: Process the collected user features, head orientation attitude data, and binaural transfer function, filter out the data required for training, eliminate the untrustworthy data, and perform normalization processing on different data types. Divide the data into a training set D train , a validation set D val and a test set D test ;
[0026] Step S223: Iteratively train the model with the training set data, compare the calculation results of the validation set with the expected values, calculate the error E = ||HRTF pred - HRTF true ||, and calculate the derivative through the error Then use the gradient descent algorithm to update the model parameters W where η is the learning rate until the training ends;
[0027] Step S224: Perform a performance test on the trained model to obtain the optimized binaural transfer function HRTF opt and the audio rendering parameter P opt .
[0028] Preferably, the adjustment of the personalized user audio rendering parameter in step S23 specifically includes the following steps:
[0029] Step S231: Through the pre-established deep learning model M, perform parameter calculation on the model according to the user feature information U to obtain the unique parameter P for each user u = M(U);
[0030] Step S232: Let the user fine-tune the default audio rendering parameter P default and save the correspondence table T(U, P) of the user features and the audio rendering parameters;
[0031] Step S233: Take the user feature U and the default parameter P default as the input parameters of the model, and generate the user personalized parameter P personalized = M(U, P default );
[0032] Step S234: Import the user feature U and the personalized parameter P personalized into the audio rendering system to automatically adjust the audio rendering parameters.
[0033] Preferably, the step S3 specifically includes the following steps:
[0034] Step S31: Input the spatial audio data S(x, y, z, t), where x, y, z are spatial coordinates and t is time;
[0035] Step S32: Use the signal direction transformation algorithm based on image acoustic simulation to convert the spatial audio into the left-channel audio signal L(t) and the right-channel audio signal R(t). The specific method is:
[0036]
[0037] where HRTF L and HRTF R are the head-related transfer functions of the left ear and the right ear respectively, and * represents the convolution operation;
[0038] Step S33: Use the time-domain to frequency-domain reverberation algorithm to calculate the reverberation response function H(f, t);
[0039] Step S34: Use the frequency-domain synthesis algorithm to calculate the virtual spatial audio V(f, t) = S(f, t) * H(f, t), where S(f, t) is the time-frequency representation of the input audio;
[0040] Step S35: Render different binaural audios L'(t) = L(t) * P L and R'(t) = R(t) * P R , where P L and P R are the personalized rendering parameters of the left ear and the right ear respectively.
[0041] Preferably, the input of the spatial audio data in the step S31 specifically includes the following steps:
[0042] Step S311: Extract the spatial audio information S(x, y, z, t) in the audio scene;
[0043] Step S312: Convert the virtual sound source in the audio scene into a planar array signal
[0044] Step S313: Render virtual spatial audio, specifically in the following way:
[0045]
[0046] where B L and B R represent the left and right sidelobe response functions respectively, θ m represents the direction angle of the microphone array, HRTF L and HRTF R are the head-related transfer functions for the left and right ears respectively, and * represents the convolution operation.
[0047] Preferably, the step S33 for calculating the reverberation response function specifically includes the following steps:
[0048] Step S331: Simulate reverberation using the image method of acoustics;
[0049] Step S332: Calculate the reverberation response function according to the reverberation time prediction algorithm:
[0050] H(f,t) = G(f) * e -t / τ(f) * e j2πfd(t)t * F(f,t)
[0051] where G(f) is the frequency-dependent gain function, τ(f) is the frequency-dependent decay time constant, fd(t) is the time-varying Doppler frequency shift, and F(f,t) is the frequency- and time-dependent filter function, which are defined as follows:
[0052] G(f) = G0 + α G * f
[0053] τ(f) = τ0 + α τ * f
[0054] fd(t) = fd0 + β d * t
[0055] F(f,t) = 1 + (α F + β F * t) * f
[0056] where G0 is the initial gain, α G is the rate of change of the gain with frequency, τ0 is the initial decay time constant, α τ is the rate of change of the decay time constant with frequency, fd0 is the initial Doppler frequency shift, β d is the rate of change of the Doppler frequency shift with time, and α F and β F are the frequency-related coefficient and time-related coefficient of the filter function respectively.
[0057] Preferably, in step S4, the rendered left and right channel audio signals are corrected for audio time delay, which specifically includes the following steps:
[0058] Step S41: Collect the left and right channel audio data respectively, and calculate the average time delay between the left channel audio signal and the right channel audio signal:
[0059]
[0060] where L[i] is the left channel audio signal, R[i] is the right channel audio signal, N L is the number of samples of the left channel audio signal, N R is the number of samples of the right channel audio signal, t L[i] is the timestamp of the left channel audio signal, t R[i] is the timestamp of the right channel audio signal;
[0061] Step S42: Use the average time delay Δt as a parameter for audio delay correction to align the time of the left and right channel signals:
[0062] L′(t) = L(t + (t)
[0063] R′(t) = R(t - Δt)
[0064] where L′(t) and R′(t) are the corrected left and right channel audio signals.
[0065] Preferably, in step S5, the audio signal is input into the TWS earphone by using an ear canal model, which specifically includes the following steps:
[0066] Step S51: Establish the acoustic transfer function model H TWS (f);
[0067] Step S52: Convert the corrected left and right channel audio signals L′(t) and R′(t) to the frequency domain: L′(f) and R′(f);
[0068] Step S53: Multiply the frequency domain signal by the acoustic transfer function of the TWS earphone:
[0069] L TWS (f) = L′(f) * H TWS (f)
[0070] R TWS (f) = R′(f) * H TWS (f)
[0071] Step S54: The processed frequency domain signals L TWS (f) and R TWS(f) Convert back to the time domain to obtain the left and right channel signals L TWS (t) and R TWS (t) for final output to the TWS earphones.
[0072] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0073] The core of the present invention lies in the ingenious combination of a number of advanced technologies, forming an organic whole. The various modules cooperate with each other, generating a remarkable synergistic effect. For example, the introduction of deep learning technology not only solves the problem of personalized HRTF generation, but also provides efficient computational support for real-time head tracking. The combination of these two technologies greatly improves the positioning accuracy and real-time performance of spatial audio, thus greatly enhancing the user's immersion.
[0074] At the same time, the complex reverberation model proposed by the present invention complements the personalized HRTF, jointly constructing a more realistic virtual acoustic environment. The reverberation model takes into account multiple parameters such as frequency-dependent gain, decay time constant, Doppler shift, etc., and can accurately simulate the sound propagation characteristics in various complex environments. Combined with the personalized HRTF, this not only improves the accuracy of spatial positioning, but also greatly enhances the realism and detail performance of the audio.
[0075] In addition, the present invention also ingeniously solves some technical contradictions in audio processing. For example, high-precision audio rendering usually requires complex calculations, which may lead to a reduction in the system response speed. However, by adopting a deep learning model and an optimized signal processing algorithm, the present invention has successfully achieved a millisecond-level response speed while ensuring high-quality audio output. This unity of high quality and high efficiency provides a solid technical foundation for real-time interactive applications.
[0076] Finally, the present invention has also been specifically optimized for the special needs of TWS earphones. By establishing an acoustic transfer function model of the TWS earphones and seamlessly integrating it into the entire rendering process, the present invention has successfully extended the high-quality spatial audio experience to portable devices. This not only solves the problem of device adaptability, but also greatly expands the application scope of the technology.
[0077] In summary, the present invention has successfully solved many challenges faced by existing spatial audio rendering methods through innovative integration of a number of advanced technologies. It not only comprehensively surpasses traditional methods in technical indicators, but more importantly, it provides users with an unprecedented and highly personalized immersive audio experience. This method is expected to bring revolutionary changes in multiple fields such as virtual reality, augmented reality, gaming, and music appreciation, promoting the rapid development of related industries. Brief Description of the Drawings
[0078] Figure 1This is the overall method logic block diagram of the present invention.
[0079] Figure 2 This is the head tracking algorithm logic block diagram of the present invention.
[0080] Figure 3 This is the personalized binaural transfer function generation algorithm logic block diagram of the present invention.
[0081] Figure 4 This is the spatial audio rendering algorithm logic block diagram of the present invention.
[0082] Figure 5 This is the audio delay correction algorithm logic block diagram of the present invention.
[0083] Figure 6 This is the TWS earphone adaptation algorithm logic block diagram of the present invention. Detailed implementation manners
[0084] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments, and details their specific implementation manners, structures, features and effects as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0085] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0086] Embodiment 1
[0087] Refer to Figure 1-6 , the present invention discloses an adaptive spatial audio rendering method based on head tracking. The purpose of this method is to provide a more realistic and immersive audio experience. The following will describe the specific implementation manners of the present invention in detail.
[0088] First of all, the core of the present invention is to adjust the audio rendering parameters in real time through head tracking technology to adapt to the movement of the user's head. Specifically, the method includes the following steps: obtaining the user's head orientation based on a head tracking device and calculating the current binaural transfer function; using a deep neural network to generate a personalized binaural transfer function and audio rendering parameters according to the user's characteristic information; rendering the spatial audio data into left and right channel audio signals; performing delay correction on the rendered audio signals; and finally inputting the corrected audio signals into the TWS earphones to complete the spatial audio rendering.
[0089] Next, we will gradually and deeply analyze the specific implementation manners of each step.
[0090] In the process of obtaining the user's head orientation, the present invention adopts precise head tracking technology. Specifically, we collect the head orientation attitude data of the user, including the pitch angle θ and the yaw angle The pitch angle θ represents the angle between the midpoint of the head and the vertical direction, while the yaw angle represents the angle between the midpoint of the head and the horizontal direction. These two angles jointly describe the three-dimensional spatial position of the user's head.
[0091] Subsequently, the system calculates the user's current binaural transfer function in real time where f represents the sound frequency. This function describes the propagation characteristics of sound from the source to the listener's ears. It should be noted that the HRTF is dynamically adjusted according to the head position, which ensures the real-time and accuracy of audio rendering.
[0092] An innovation of the present invention is the introduction of a deep neural network to process the head orientation data. We input the head orientation attitude data into a pre-trained deep neural network to obtain audio rendering parameters Here, DNN represents the deep neural network model, and P is a vector containing multiple rendering parameters. In this way, we can quickly and accurately adjust the audio rendering parameters according to the head position.
[0093] Next, the present invention introduces personalized elements. We collect the user features U = {G, A, H}, where G represents the user's gender, A represents the user's age, and H contains the head attitude data. These features are used to train a deep learning model M = f(U, HRTF). In this formula, M is the trained model, and f represents the deep learning function. Through this model, we can generate personalized audio rendering parameters P' = M(U) for each user.
[0094] It is worth mentioning that the training process of the deep learning model is carefully designed. First, we construct the input-output structure of the model, with the input being the user features U and the output being the user's current binaural transfer function Then, we process the collected data, including screening, removing untrusted data, normalization processing, etc. The processed data is divided into a training set D train , a validation set D val and a test set D test .
[0095] In the model training process, we adopt an iterative training method. Specifically, we calculate the error E = ||HRTF pred - HRTF true || between the model output and the true value, and then use the gradient descent algorithm to update the model parameters where η is the learning rate. This process is repeated continuously until the model converges or reaches the preset number of training epochs.
[0096] Adjustment of personalized user audio rendering parameters specifically includes the following steps:
[0097] First, through the pre-established deep learning model M, parameter calculation is performed on the model according to the user feature information U to obtain the unique parameter P for each user u = M(U);
[0098] Then, the user fine-tunes the default audio rendering parameter P default and saves the correspondence table T(U, P) of user features and audio rendering parameters;
[0099] Next, the user feature U and the default parameter P default are used as the input parameters of the model, and the user personalized parameter P personalized = M(U, P default ) is generated through the model;
[0100] Finally, the user feature U and the personalized parameter P personalized are imported into the audio rendering system to automatically adjust the audio rendering parameters.
[0101] Preferably, the present invention also provides a mechanism for adjusting personalized parameters. We allow the user to fine-tune the default audio rendering parameter P default and save the correspondence table T(U, P) of user features and audio rendering parameters. This design enables the system to further optimize the audio effect according to the user's personal preferences.
[0102] In the processing of spatial audio data, the present invention adopts a signal direction transformation algorithm based on image acoustic simulation. This algorithm converts the audio signal S(x, y, z, t) in three-dimensional space into the audio signals L(t) and R(t) of the left and right channels. The mathematical expression of the conversion is as follows:
[0103]
[0104] where HRTF L and HRTF R are the head-related transfer functions of the left ear and the right ear respectively, and * represents the convolution operation. This algorithm can accurately simulate the propagation characteristics of sound in three-dimensional space and provide users with a realistic spatial audio experience.
[0105] Then, the reverberation response function H(f, t) is calculated using the time-domain - frequency-domain reverberation algorithm;
[0106] Next, the virtual spatial audio V(f,t) = S(f,t) * H(f,t) is calculated using the frequency-domain synthesis algorithm, where S(f,t) is the time-frequency representation of the input audio;
[0107] Finally, the head models of different users are used to render different binaural audios L'(t) = L(t) * P L and R'(t) = R(t) * P R , where P L and P R are the personalized rendering parameters for the left and right ears respectively.
[0108] Among them, the input of spatial audio data specifically includes the following steps:
[0109] First, extract the spatial audio information S(x,y,z,t) in the audio scene;
[0110] Then, convert the virtual sound sources in the audio scene into plane array signals
[0111] Finally, render the virtual spatial audio, and the specific method is:
[0112]
[0113] Among them, B L and B R represent the left and right sidelobe response functions respectively, θ m represents the direction angle of the microphone array, HRTF L and HRTF R are the head-related transfer functions for the left and right ears respectively, and * represents the convolution operation.
[0114] Another innovation point of the present invention lies in the introduction of the simulation of the reverberation effect. We use the time-domain - frequency-domain reverberation algorithm to calculate the reverberation response function H(f,t). The specific form of this function is:
[0115] H(f,t) = G(f) * e -t / τ(f) * e j2πfd(t)t * F(f,t)
[0116] Among them, G(f) is the frequency-dependent gain function, τ(f) is the frequency-dependent decay time constant, fd(t) is the time-varying Doppler frequency shift, and F(f,t) is the frequency and time-dependent filter function. These parameters all have their specific definitions:
[0117] G(f) = G0 + α G * f
[0118] τ(f) = τ0 + α τ * f
[0119] fd(t) = fd0 + β d *t
[0120] F(f, t) = 1 + (α F + β F *t)*f
[0121] where G0 is the initial gain, α G is the rate of change of the gain with frequency, τ0 is the initial decay time constant, α τ is the rate of change of the decay time constant with frequency, fd0 is the initial Doppler frequency shift, β d is the rate of change of the Doppler frequency shift with time, α F and β F are the frequency - related coefficient and time - related coefficient of the filtering function, respectively. Through this complex reverberation model, we can simulate the sound effects in various real environments, greatly enhancing the realism and immersion of the audio.
[0122] Finally, to solve the possible audio delay problem, the present invention proposes a delay correction method. We first calculate the average delay Δt between the left and right channel audio signals,
[0123]
[0124] where L[i] is the left - channel audio signal, R[i] is the right - channel audio signal, N L is the number of samples of the left - channel audio signal, N R is the number of samples of the right - channel audio signal, t L[i] is the timestamp of the left - channel audio signal, t R[i] is the timestamp of the right - channel audio signal;
[0125] Then we use this delay to align the audio signals in time. This process can be expressed as:
[0126] L′(t) = L(t + Δt)
[0127] R′(t) = R(t - Δt)
[0128] where L′(t) and R′(t) are the corrected left and right channel audio signals. This delay correction mechanism can effectively eliminate the inter - channel delay caused by hardware or signal processing, further improving the localization accuracy of spatial audio.
[0129] In an embodiment of the present invention, we also consider the particularity of TWS earphones. We establish an acoustic transfer function model H of the TWS earphones TWS(f) and apply it to the final audio processing. Specifically, we multiply the corrected left and right channel audio signals by the acoustic transfer function of the TWS earphones:
[0130] L TWS (f) = L′(f) * H TWS (f)
[0131] R TWS (f) = R′(f) * H TWS (f)
[0132] Then, we convert the processed frequency-domain signals L TWS (f) and R TWS (f) back to the time domain to obtain the left and right channel signals L TWS (t) and R TWS (t) that are finally output to the TWS earphones. This targeted processing ensures that high-quality spatial audio effects can also be obtained on TWS earphones.
[0133] Generally speaking, the present invention realizes highly personalized and real-time adaptive spatial audio rendering by combining various advanced technologies such as head tracking technology, deep learning, and signal processing. This method can not only accurately reflect the impact of the user's head movement on audio perception but also optimize the audio effect according to the user's personal characteristics and preferences. At the same time, by introducing a complex reverberation model and a time delay correction mechanism, the realism and localization accuracy of the audio are further improved. In particular, the targeted optimization of TWS earphones enables this high-quality spatial audio experience to be realized on portable devices, greatly expanding its application scope.
[0134] This head-tracking-based adaptive spatial audio rendering method is expected to bring revolutionary changes in multiple fields such as virtual reality, augmented reality, gaming, and music appreciation, providing users with an unprecedented immersive audio experience.
[0135] To further verify the superiority of the present invention, we designed a set of examples and comparative examples and conducted detailed tests and analyses on them. The following will introduce in detail the specific implementation methods of the examples and comparative examples, as well as the corresponding test results and analyses.
[0136] Example 1: In this example, we fully implemented the above-mentioned head-tracking-based adaptive spatial audio rendering method. Specifically, we used a high-precision head-tracking device to collect the user's head orientation data, including the pitch angle θ and the yaw angle At the same time, we collected the user's personal characteristics, including gender, age, and head size data. Based on these data, we trained a deep neural network model for generating personalized binaural transfer functions (HRTFs) and audio rendering parameters.
[0137] In audio processing, we adopted a signal direction transformation algorithm based on image acoustic simulation and introduced a complex reverberation model. In addition, we also implemented an audio delay correction mechanism to eliminate possible inter-channel delays. Finally, we carried out special optimization for the characteristics of TWS earphones.
[0138] Comparative Example 1: For comparison, we implemented a traditional non-adaptive spatial audio rendering method. This method uses fixed and general HRTFs without considering individual differences and head movements of users. Simple stereo panoramic technology is used for audio processing, which does not include a complex reverberation model and a delay correction mechanism. At the same time, this method has not been specially optimized for TWS earphones.
[0139] Test methods and metrics:
[0140] To comprehensively evaluate the performance of the two methods, we designed the following test metrics and methods:
[0141] 1. Spatial localization accuracy: Let the testers locate the sound source position in a virtual environment and calculate the localization error.
[0142] 2. Head tracking response time: Measure the time delay from head movement to the completion of audio adjustment.
[0143] 3. Audio quality score: Subjectively scored by professional audio engineers, with a range of 1 - 10 points.
[0144] 4. Immersion score: Subjectively scored by the testers, with a range of 1 - 10 points.
[0145] 5. Compatibility with TWS earphones: The sound quality performance on TWS earphones is scored by professional audio engineers, with a range of 1 - 10 points.
[0146] We invited 50 volunteers of different ages and genders to participate in the test, and each volunteer conducted multiple tests using both methods. The following are the detailed data of the test results:
[0147] Test indicators Example 1 Comparative example 1 Spatial positioning accuracy 95.2% 78.6% Head tracking response time 15ms N / A Audio quality score 8.7 7.2 Immersion score 9.1 6.8 TWS headphone adaptability 8.9 6.5
[0148] It can be clearly seen from the test results that the method of the present invention (Example 1) is significantly superior to the traditional method (Comparative Example 1) in all aspects.
[0149] First of all, in terms of spatial localization accuracy, the method of the present invention achieved a high accuracy rate of 95.2%, far exceeding 78.6% of the comparative example. This fully proves that the method based on head tracking and personalized HRTF can greatly improve the localization accuracy of spatial audio. Users can more accurately perceive the position of the sound source, which is crucial for application scenarios such as virtual reality.
[0150] Secondly, the method of the present invention excels in head-tracking response time, completing audio adjustment in just 15 milliseconds. This near-real-time response ensures that the audio effect can seamlessly follow the movement of the user's head, greatly enhancing the immersion and realism. In contrast, traditional methods do not support head tracking and thus cannot provide this dynamic adjustment function.
[0151] In terms of audio quality, the method of the present invention received a high rating of 8.7 from professional audio engineers, significantly superior to the 7.2 of the comparative example. This indicates that complex audio processing algorithms, including reverb models and delay correction, can indeed significantly improve the overall audio quality.
[0152] The gap in immersion scores is even more significant. The method of the present invention obtained a high score of 9.1, while the comparative example was only 6.8. This result fully demonstrates the importance of adaptive rendering and precise spatial positioning in enhancing the user experience. When using the method of the present invention, users can obtain a more realistic and natural audio experience, as if they were truly in the sound environment.
[0153] Finally, in terms of the compatibility of TWS earphones, the method of the present invention also performs excellently, with a score of 8.9, far higher than the 6.5 of the comparative example. This proves that the special optimization for TWS earphones does play a role, enabling a high-quality spatial audio experience to be achieved on portable devices.
[0154] In summary, this set of test results fully verifies the superiority of the method of the present invention. It not only comprehensively surpasses traditional methods in technical indicators, but more importantly, it can provide a significantly improved auditory experience for users. By combining advanced technologies such as head tracking, deep learning, and complex audio processing algorithms, the method of the present invention has successfully achieved highly personalized and real-time adaptive spatial audio rendering. This method is expected to bring revolutionary changes in multiple fields such as virtual reality, augmented reality, gaming, and music appreciation, providing users with an unprecedented immersive audio experience.
[0155] In all tested configurations, we found that when the deep neural network uses 5 fully connected layers with 128 neurons in each layer and the ReLU activation function, the effect is the best. At the same time, setting the sampling rate of the head tracking device to 100Hz and the decay time constant τ0 of the reverb model to 0.5 seconds can achieve the best balance between performance and computational complexity. This configuration can be regarded as the best embodiment of the present invention and can be used as a reference configuration in practical applications.
[0156] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An adaptive spatial audio rendering method based on head tracking, characterized in that, It includes the following steps: Step S1: Based on the head tracking device, obtain the user's head orientation and calculate the user's current binaural transfer function; Step S2: Based on the deep neural network, train the binaural transfer function for each user according to the user feature information, and generate personalized binaural transfer functions and audio rendering parameters; Step S3: Input the spatial audio data and render it into left and right channel audio signals; Step S4: Perform audio delay correction on the rendered left and right channel audio signals and output the rendered sound effect; Step S5: Use the ear canal model to input the audio signal into the TWS earphones to complete the spatial audio rendering.
2. The adaptive spatial audio rendering method based on head tracking according to claim 1, wherein The specific steps of Step S1 include the following steps: Step S11: Collect the head orientation attitude data of the user, where the head orientation attitude data includes the pitch angle θ and the yaw angle The pitch angle θ is the angle formed by the vertical line passing through the center of the head and the head, and the yaw angle is the angle formed by the horizontal line passing through the center of the head and the head; Step S12: Calculate the user's current binaural transfer function in real time where f is the sound frequency; Step S13: Use the head orientation attitude data as the input of the binaural transfer function deep neural network to obtain audio rendering parameters where DNN represents the deep neural network model, and P is the audio rendering parameter vector.
3. The adaptive spatial audio rendering method based on head tracking according to claim 1, wherein The specific steps of Step S2 include the following steps: Step S21: Collect user features U = {G, A, H}, where G is the user's gender, A is the user's age, and H is the head pose data; Step S22: Train the deep learning model M = f(U, HRTF), where M is the trained model and f is the deep learning function; Step S23: Adjust the personalized user audio rendering parameters P' = M(U), where P' is the personalized audio rendering parameter.
4. The adaptive spatial audio rendering method based on head tracking according to claim 3, characterized in that In the training of the deep learning model in Step S22, it specifically includes: Step S221: Construct the input and output of the deep learning model, where the input of the model is the user feature U and the output is the current binaural transfer function of the user Step S222: Process the collected user characteristics, head orientation attitude data, and binaural transfer functions, filter out the data required for training, eliminate the untrustworthy data, perform normalization processing on different data types, and divide the data into a training set D train , a validation set D val , and a test set D test ; Step S223: Iteratively train the model with the training set data, compare the calculation results of the validation set with the expected values, calculate the error E = ||HRTF pred - HRTF true ||, and calculate the derivative through the error Then use the gradient descent algorithm to update the model parameter W where η is the learning rate until the training ends; Step S224: Perform performance testing on the trained model to obtain the optimized binaural transfer function HRTF opt and audio rendering parameter P opt .
5. The adaptive spatial audio rendering method based on head tracking according to claim 3, characterized in that In the adjustment of the personalized user audio rendering parameters in Step S23, it specifically includes the following steps: Step S231: Through the pre-established deep learning model M, calculate the parameters of the model according to the user feature information U, and obtain the unique parameters P for each user u = M(U); Step S232: The user makes fine-tuning on the default audio rendering parameter P default and saves the correspondence table T(U, P) between the user characteristics and the audio rendering parameters; Step S233: Use the user feature U and the default parameter P default as the input parameters of the model, and generate the user personalized parameter P through the model personalized = M(U, P default ); Step S234: Import the user feature U and the personalized parameter P personalized into the audio rendering system to automatically adjust the audio rendering parameters.
6. The adaptive spatial audio rendering method based on head tracking according to claim 1, wherein The specific steps of Step S3 include the following steps: Step S31: Input the spatial audio data S(x, y, z, t), where x, y, z are spatial coordinates and t is time; Step S32: Use the signal direction transformation algorithm based on image acoustics simulation to convert the spatial audio into a left channel audio signal L(t) and a right channel audio signal R(t). The specific method is: where HRTF L and HRTF R are the head-related transfer functions of the left ear and the right ear respectively, and * represents the convolution operation; Step S33: Use the time-domain to frequency-domain reverberation algorithm to calculate the reverberation response function H(f, t); Step S34: Use the frequency-domain synthesis algorithm to calculate the virtual spatial audio V(f, t) = S(f, t) * H(f, t), where S(f, t) is the time-frequency representation of the input audio; Step S35: Render different binaural audios using the head models of different users, L'(t) = L(t) * P L and R'(t) = R(t) * P R , where P L and P R are the personalized rendering parameters for the left and right ears respectively.
7. A method for adaptive spatial audio rendering based on head tracking according to claim 6, wherein In the input of the spatial audio data in Step S31, it specifically includes the following steps: Step S311: Extract the spatial audio information S(x, y, z, t) in the audio scene; Step S312: Convert the virtual sound source in the audio scenario into a planar array signal Step S313: Render the virtual spatial audio. The specific method is: Among them, B L and B R respectively represent the left and right sidelobe response functions, θ m represents the direction angle of the microphone array, HRTF L and HRTF R are the head-related transfer functions for the left ear and the right ear respectively, and * represents the convolution operation.
8. The adaptive spatial audio rendering method based on head tracking according to claim 6, wherein, In the calculation of the reverberation response function in Step S33, it specifically includes the following steps: Step S331: Use the image acoustics method to simulate reverberation; Step S332: Calculate the reverberation response function according to the reverberation time prediction algorithm: H(f,t) = G(f) * e- t / τ(f) * e j2πfd(t)t * F(f,t), Among them, G(f) is the frequency-dependent gain function, τ(f) is the frequency-dependent decay time constant, fd(t) is the time-varying Doppler frequency shift, and F(f, t) is the frequency and time-dependent filter function, which are defined as follows: G(f) = G0 + α G *f, τ(f) = τ0 + α τ *f, fd(t) = fd0 + β d *t, F(f,t) = 1 + (α F + β F * t) * f, Among them, G0 is the initial gain, α G is the change rate of the gain with respect to frequency, τ0 is the initial decay time constant, α τ is the change rate of the decay time constant with respect to frequency, fd0 is the initial Doppler frequency shift, β d is the change rate of the Doppler frequency shift with respect to time, α F and β F are the frequency - related coefficient and the time - related coefficient of the filtering function, respectively.
9. The adaptive spatial audio rendering method based on head tracking according to claim 1, wherein In Step S4, the audio delay correction of the rendered left and right channel audio signals specifically includes the following steps: Step S41: Collect the left and right channel audio data respectively and calculate the average delay between the left channel audio signal and the right channel audio signal; Among them, L[i] is the left-channel audio signal, R[i] is the right-channel audio signal, N L is the number of samples of the left-channel audio signal, N R is the number of samples of the right-channel audio signal, t L[i] is the timestamp of the left-channel audio signal, t R[i] is the timestamp of the right-channel audio signal; Step S42: Use the average delay Δt as the parameter for audio delay correction to align the left and right channel signals in time: L′(t) = L(t + Δt), R′(t) = R(t - Δt), where L′(t) and R′(t) are the corrected left and right channel audio signals.
10. The adaptive spatial audio rendering method based on head tracking according to claim 1, wherein Step S5 inputs the audio signal into the TWS earphones by using the ear canal model, which specifically includes the following steps: Step S51: Establish the acoustic transfer function model H TWS (f); Step S52: Convert the corrected left and right channel audio signals L′(t) and R′(t) to the frequency domain: L′(f) and R′(f); Step S53: Multiply the frequency domain signals by the acoustic transfer function of the TWS earphones: L TWS L(f) = L′(f) * H TWS (f), R TWS R(f) = R'(f) * H TWS (f), Step S54: Convert the processed frequency-domain signals L TWS (f) and R TWS (f) back to the time domain to obtain the left and right channel signals L TWS (t) and R TWS (t) that are finally output to the TWS earphones.
Citation Information
Cited By
Sound field expansion method, audio equipment and computer readable storage medium
CN120769218A