TWS earphone head posture estimation method based on multi-sensor fusion

By using multi-sensor fusion and deep learning algorithms, the accuracy and adaptability issues of head pose estimation in TWS earphones have been solved, achieving high-precision and interference-resistant head pose estimation, thus improving user experience and system performance.

CN120991846APending Publication Date: 2025-11-21COSONIC INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510815409.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing head pose estimation methods for TWS earphones suffer from insufficient accuracy, weak anti-interference ability, and poor adaptability, especially in complex environments and fast-moving scenarios.

Method used

By employing a multi-sensor fusion strategy, combining accelerometers, gyroscopes, magnetometers, and microphones, and processing data through Kalman filters and LSTM neural networks, an adaptive model is established to achieve high-precision, robust, and real-time head pose estimation.

Benefits of technology

It improves the accuracy and stability of head pose estimation, enhances the system's anti-interference ability, can quickly adapt to different users and environments, and supports a more natural human-computer interaction and spatial audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120991846A_ABST
    Figure CN120991846A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of TWS earphone head posture estimation methods, in particular to a TWS earphone head posture estimation method based on multi-sensor fusion, which comprises the following steps: (1) collecting sound signals through a microphone; (2) defining a TWS earphone coordinate system, and taking a left earplug as an original point; (3) obtaining a vector between the left earplug and the right earplug measured by the magnetometer; (4) acquiring accelerated speeds and angular speeds of coordinate axes x, y and z of the left earplug through a micro electro mechanical system inertial sensor; (5) calculating a rotation matrix Rt at the moment t; (6) calculating a yaw angle Yawt at the moment t; calculating a pitch angle Pitcht at the moment t according to the angular velocity of the gyroscope and the angular velocity of the coordinate axis of the left earplug; (7) estimating the real posture of the earphone; the weight of each sensor can be dynamically adjusted according to environmental changes, the advantages of each sensor are utilized to the maximum extent, and meanwhile the defects of each sensor are restrained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of TWS earphone head posture estimation methods, specifically to a TWS earphone head posture estimation method based on multi-sensor fusion. Background Technology

[0002] With the rapid development of wireless technology, true wireless stereo (TWS) earbuds have become increasingly popular in recent years, becoming an indispensable electronic device in consumers' daily lives. However, as users' demands for audio experience continue to rise, the limitations of traditional TWS earbuds in areas such as spatial audio and human-computer interaction are becoming increasingly apparent. Among these, accurately estimating the user's head posture is one of the key technologies for achieving high-quality spatial audio and natural human-computer interaction.

[0003] Currently, the most commonly used head pose estimation methods in the industry mainly rely on a single sensor, such as a gyroscope or accelerometer. While this method is simple to implement, it has several problems. First, a single sensor cannot provide comprehensive motion information, resulting in insufficient estimation accuracy, especially in complex head motion scenarios. Second, these methods are highly sensitive to external environmental interference; for example, in environments with magnetic field interference, gyroscope-based methods may produce severe drift errors. Furthermore, traditional methods typically use simple mathematical models for pose calculation, making it difficult to capture complex head motion patterns, leading to inaccurate estimation results during rapid or irregular movements.

[0004] On the other hand, some existing improvement methods attempt to introduce multi-sensor fusion technology, but they are often limited to simple data superposition or weighted averaging, failing to fully utilize the advantages of each sensor. Furthermore, these methods still rely heavily on traditional filters in their algorithm design, lacking the ability to effectively model complex nonlinear systems.

[0005] Furthermore, existing technologies often lack sufficient adaptability and robustness when facing different users and usage environments. This leads to a significant gap between system performance and ideal conditions in practical applications, impacting user experience.

[0006] In view of the above problems, there is an urgent need for a head pose estimation method that can comprehensively utilize multi-sensor data, possess powerful modeling capabilities, and simultaneously balance accuracy and real-time performance. This invention is an innovative solution proposed to address this technical need. Summary of the Invention

[0007] This invention aims to address the problems of insufficient accuracy, weak anti-interference ability, and poor adaptability in existing TWS earphone head posture estimation methods. Through an innovative multi-sensor fusion strategy, advanced deep learning algorithms, and adaptive model fusion technology, this invention provides a high-precision, robust, and highly adaptable TWS earphone head posture estimation method.

[0008] This invention proposes a head pose estimation method for TWS earphones based on multi-sensor fusion, comprising the following steps:

[0009] (1) Set up an accelerometer, gyroscope, and magnetometer on the TWS earphone and collect sound signals through a microphone; (2) Define the TWS earphone coordinate system with the left earbud as the origin; (3) Obtain the vector between the left and right earbuds measured by the magnetometer; (4) Obtain the acceleration and angular velocity of the left earbud's x, y, and z axes through a microelectromechanical system (MEMS) inertial sensor; (5) Calculate the rotation matrix R at time t based on the gyroscope angular velocity and the relative position of the magnetometer and the left earbud. t (6) Calculate the yaw angle at time t based on the relative position of the gyroscope and the left earpiece and the acceleration of the left earpiece coordinate axis. t Calculate the pitch angle at time t based on the gyroscope angular velocity and the angular velocity of the left earpiece coordinate axis. t (7) Use a Kalman filter to filter the estimated values ​​obtained in steps (3)-(6) to estimate the true posture of the earphone; (8) Collect TWS earphone motion data from different users, perform preprocessing, and divide them into training set, validation set and test set; (9) Use an LSTM neural network to perform head motion pattern recognition, including: a) establishing an LSTM neural network model, training it using the training set, and obtaining the output model; b) testing the output model using the test set and obtaining the test results.

[0010] Preferably, the method for defining the TWS earphone coordinate system in step (2) is as follows:

[0011] Assume the wearer of the TWS earphones is facing the XY plane, with the left earbud as the origin of the coordinate system; the line connecting the right and left earbuds points in front of the earphones and is perpendicular to the normal direction of the front of the human body, the Z-axis is vertically upward, and the X-axis is perpendicular to the forward direction of the earphones.

[0012] Preferably, the method for obtaining the vector between the left and right earbuds in step (3) is as follows:

[0013] V = P right -P left

[0014] Where V is the direction vector, representing the relative position of the right earbud with respect to the left earbud, and P... right Let P be the coordinates of the right earpiece. leftThe coordinates are for the left earpiece.

[0015] Preferably, the method further includes a step of transforming the vector between the left and right earbuds in the left earbud coordinate system to the headphone position coordinate system.

[0016] V′=M*V

[0017] Where V′ is the vector in the headphone position coordinate system, and M is the translation matrix from the left earcup coordinate system to the headphone position coordinate system, expressed as:

[0018]

[0019] Among them, [d x ,d y ] represents the distance between the origin of the left earbud coordinate system and the origin of the headphone position coordinate system, and θ represents the angle between the line connecting the left and right earbuds and the line connecting the horizontal axis of the headphone.

[0020] Preferably, the Kalman filtering algorithm in step (7) includes:

[0021] (a) Update the estimated value at time t using the rotation matrix and motion parameters from the previous time t-1:

[0022]

[0023] in, Let F be the prior estimate at time t. t Here is the state transition matrix. Let B be the posterior estimate at time t-1. t To control the input matrix, u t Control vector;

[0024] (b) Calculate the error covariance matrix:

[0025]

[0026] Among them, P t|t-1 Let P be the prior covariance matrix. t-1|t-1 Let Q be the posterior covariance matrix at time t-1. t The process noise covariance matrix is... For F t The transpose of the matrix;

[0027] (c) Update the error covariance matrix Pt to obtain the estimated value of the rotation matrix Rt at time t:

[0028]

[0029] P t|t =(IK t ·H t )·P t|t-1

[0030] Among them, K t For Kalman gain, H t Let R be the observation matrix. t To observe the noise covariance matrix, z t For the observation vector, Let P be the posterior estimate at time t. t|t Let I be the posterior covariance matrix at time t, and let I be the identity matrix.

[0031] Preferably, the method for establishing the LSTM neural network model in step (9) is as follows: Define the input data feature sequence as {x1, x2, ..., x...} n}, where n is the number of input features, x1 is the current frame input, and y t Let h0 be the hidden state at the input of the first frame, and c0 be the hidden state of the previous frame at the time of the first frame. This can be expressed by the formula:

[0032] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0033] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0034]

[0035] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0036] h t =o t ⊙tanh(C t )

[0037] Among them, f t For the Gate of Oblivion, W f Here is the forget gate weight matrix, [h t -1,x t ] is the input for the forget gate, b f For the forget gate bias; i t For the input gate, W i Let b be the input gate weight matrix. i For input gate bias; As a candidate state, WC Let b be the candidate state weight matrix. C For W C Bias matrix; C t For state update gate; o t For output gate, W o Let b be the output gate weight matrix. o For output gate bias; h t σ represents the output corresponding to frame t; σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, and ⊙ represents element-wise multiplication.

[0038] Preferably, the preprocessing in step (8) includes: performing window truncation processing on all data samples to obtain window data in each behavior data, wherein window truncation processing refers to: using a window with a window sliding step size the same as the sensor sampling rate, sliding the window in all frames of the behavior data, forming a new sequence from each frame sampling point during the sliding process, setting each new sequence as a data packet, and each behavior data containing several data packets as training data.

[0039] Preferably, the method further includes the following steps:

[0040] (a) Let Q be the newly collected dataset, where the number of data packets is less than that of dataset P;

[0041] (b) Define the output model consisting of all data packets in dataset Q as M1, where:

[0042] y1=W1*h t +b1

[0043] Where y1 is the output fully connected layer, W1 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b1 is the bias vector.

[0044] (c) Learn the dataset Q to obtain the output model M1;

[0045] (d) Define the model consisting of all window data in dataset P as M2, where:

[0046] y2=W2*h t +b2

[0047] Where y2 is the output fully connected layer, W2 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b2 is the bias vector;

[0048] (e) Learn from dataset P to obtain output model M2;

[0049] (f) Define the common data between dataset Q and dataset P as dataset R, and use the transfer learning method to learn dataset R to obtain the output model M3.

[0050] Preferably, the method for learning datasets Q and P is as follows: using K-fold cross-validation, where K is a positive integer, typically 5 or 10; for dataset Q, using the trained output model M0 for network training; for dataset P, using the trained output model M1 for network training.

[0051] Preferably, it further includes the following step: (a) during the testing phase, given the input data sequence {x1,x2,…,x} t}, where t represents the input time step, and Y is obtained. t = (yaw, pitch, roll), where yaw is the yaw angle, pitch is the pitch angle, and roll is the roll angle; (b) The Mean Absolute Error (MAE) is used as the evaluation index, and the formula for calculating MAE is: Where y pred Represents the predicted value, y true (c) Fuse models M1, M2, and M3 to obtain the final model M. final =α*M1+β*M2+γ*M3, where α, β, and γ are weight coefficients, and α+β+γ=1,0≤α,β,γ≤1.

[0052] The beneficial effects of this invention are mainly reflected in the following aspects:

[0053] The core of this invention lies in its ingenious solution to several key technical challenges. Firstly, regarding sensor fusion, this invention proposes an adaptive fusion algorithm based on Kalman filtering, effectively addressing the issue of inconsistent data characteristics from different sensors. This algorithm can dynamically adjust the weights of each sensor according to environmental changes, maximizing the utilization of each sensor's strengths while suppressing its weaknesses.

[0054] Secondly, in terms of complex head motion modeling, this invention introduces an LSTM neural network, successfully overcoming the limitation of traditional methods in capturing long-term dependencies. The memory mechanism of the LSTM network enables the system to learn and predict complex head motion patterns, greatly improving the accuracy and stability of the estimation.

[0055] Furthermore, this invention innovatively proposes a multi-model fusion technique that cleverly combines the advantages of general-purpose and personalized models. This method not only improves the overall performance of the system but also enables the system to quickly adapt to different users, solving the problem of poor universality in traditional methods.

[0056] It is worth mentioning that the present invention fully considers the real-time requirements in the algorithm design. By optimizing the calculation process and parameter selection, it successfully controls the response time below the human perception threshold while ensuring high accuracy, effectively solving the contradiction between accuracy and real-time performance.

[0057] Furthermore, this invention has achieved a significant breakthrough in anti-interference capabilities. Through cross-validation and adaptive adjustment of multi-sensor data, the system can maintain high-precision estimation results even in the presence of external interference (such as magnetic field interference), greatly enhancing the system's reliability and applicability.

[0058] In summary, this invention achieves a comprehensive improvement in estimation accuracy, response speed, anti-interference capability, and adaptability through the organic combination of multiple innovative technologies. This enhanced overall performance not only directly improves the user experience of TWS earphones but also lays a solid technical foundation for the realization of advanced functions such as spatial audio and natural human-computer interaction. Furthermore, the method of this invention possesses good scalability and versatility, and is expected to find wide application in related fields such as VR / AR and smart wearables, driving technological progress across the entire industry. Attached Figure Description

[0059] Figure 1 This is a block diagram of the overall method logic of the present invention.

[0060] Figure 2 This is a block diagram of the data preprocessing logic of the present invention.

[0061] Figure 3 This is a block diagram of the multi-sensor fusion (Kalman filtering) logic of the present invention.

[0062] Figure 4 This is a block diagram of the LSTM neural network processing logic of the present invention.

[0063] Figure 5 This is a block diagram of the model fusion logic of the present invention. Detailed Implementation

[0064] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects are described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0066] Example 1

[0067] See Figure 1-5 This invention relates to a head pose estimation method for TWS earphones based on multi-sensor fusion. This method fully utilizes the structural characteristics of TWS earphones and achieves high-precision head pose estimation through the collaborative work of multiple sensors.

[0068] The method of the present invention includes the following steps: (1) setting an accelerometer, a gyroscope, and a magnetometer on the TWS earphone, and collecting sound signals through a microphone; (2) defining a TWS earphone coordinate system with the left earbud as the origin; (3) obtaining the vector between the left and right earbuds measured by the magnetometer; (4) obtaining the acceleration and angular velocity of the left earbud's coordinate axes x, y, and z through a microelectromechanical system (MEMS) inertial sensor; (5) calculating the rotation matrix R at time t based on the gyroscope angular velocity and the relative position of the magnetometer and the left earbud. t (6) Calculate the yaw angle at time t based on the relative position of the gyroscope and the left earpiece and the acceleration of the left earpiece coordinate axis. t Calculate the pitch angle at time t based on the gyroscope angular velocity and the angular velocity of the left earpiece coordinate axis. t (7) Use a Kalman filter to filter the estimated values ​​obtained in steps (3)-(6) to estimate the true posture of the earphone; (8) Collect TWS earphone motion data from different users, perform preprocessing, and divide them into training set, validation set and test set; (9) Use an LSTM neural network to perform head motion pattern recognition, including: a) establishing an LSTM neural network model, training it using the training set, and obtaining the output model; b) testing the output model using the test set and obtaining the test results.

[0069] This multi-sensor fusion method has significant advantages. First, it fully utilizes the binaural structure of TWS earbuds, and by analyzing the relative positional relationship between the left and right earbuds, it can more accurately describe head posture. Second, the collaborative work of multiple sensors can complement each other's strengths and weaknesses, improving overall estimation accuracy. For example, in scenarios where a user quickly turns their head, the accelerometer can provide instantaneous linear acceleration information, the gyroscope can provide angular velocity information, and the magnetometer can provide absolute direction information. Through the fusion of these sensors, we can obtain more comprehensive and accurate head posture information. Preferably, we can set the sampling rate of the accelerometer and gyroscope to 200Hz and the sampling rate of the magnetometer to 50Hz, which can reduce system power consumption while ensuring accuracy.

[0070] In one embodiment of the present invention, it is assumed that the wearer of the TWS earphones is positioned in the XY plane, facing the left earbud, with the left earbud as the origin of the coordinate system. The line connecting the right and left earbuds points forward and is perpendicular to the normal direction of the human body's frontal view. The Z-axis is vertically upward, and the X-axis is perpendicular to the forward direction of the earphones. This coordinate system definition fully considers the wearing characteristics of the earphones, making subsequent calculations more consistent with actual usage scenarios. For example, when the user turns their head, this coordinate system can accurately capture the rotational movement of the head. In practical applications, this coordinate system can be determined through an initial calibration process. Specifically, the user can wear the earphones and keep their head still, facing forward, for about 5 seconds. The system can then establish an initial coordinate system using sensor data during this period.

[0071] In one embodiment of the present invention, the method for obtaining the vector between the left and right earbuds is: V = P right -P left

[0072] Where V is the direction vector, representing the relative position of the right earbud with respect to the left earbud, and P... right Let P be the coordinates of the right earpiece. left Let V be the coordinates of the left earbud. Calculating this vector is crucial for subsequent attitude estimation because it directly reflects the head's orientation and tilt. In practical applications, this vector can be obtained using a magnetometer. For example, assuming the coordinates of the left earbud are (0,0,0) and the right earbud are (10cm,0,0), then V equals (10cm,0,0). This vector can be used not only for attitude estimation but also to detect whether the headphones are worn correctly.

[0073] In one embodiment of the present invention, the method for transforming the vector between the left and right earbuds in the left earbud coordinate system to the headphone position coordinate system includes the following steps: We use the formula V′=M*V, where V′ is the vector in the headphone position coordinate system, and M is the translation matrix from the left earbud coordinate system to the headphone position coordinate system, expressed as: This transformation process takes into account the actual wearing situation of the headphones, making the posture estimation more accurate. For example, when a user wears the headphones at an angle, this transformation can help us correctly understand the true posture of the head. In one embodiment of the present invention, matrix M can be expressed as:

[0074]

[0075] Among them, [d x ,d y [ ] represents the distance between the origin of the left earbud coordinate system and the origin of the headphone position coordinate system, and θ is the angle between the line connecting the left and right earbuds and the line connecting the headphone horizontally. This matrix considers both translation and rotation transformations, and can more accurately describe the wearing state of the headphones.

[0076] In one embodiment of the present invention, the Kalman filtering algorithm in step (7) includes the following steps:

[0077] First, we update the estimate at time t using the rotation matrix and motion parameters from the previous time t-1:

[0078]

[0079] in, Let F be the prior estimate at time t. t Here is the state transition matrix. Let B be the posterior estimate at time t-1. t To control the input matrix, u t Control vector.

[0080] Then, we calculate the error covariance matrix:

[0081]

[0082] Among them, P t|t-1 Let P be the prior covariance matrix. t-1|t-1 Let Q be the posterior covariance matrix at time t-1. t The process noise covariance matrix is... For F t The transpose of .

[0083] Finally, we update the error covariance matrix Pt to obtain the estimated value of the rotation matrix Rt at time t:

[0084]

[0085] P t|t =(IK t ·H t )·P t|t-1

[0086] Among them, K t For Kalman gain, H t Let R be the observation matrix. t To observe the noise covariance matrix, z t For the observation vector, Let P be the posterior estimate at time t. t|t Let I be the posterior covariance matrix at time t, and let I be the identity matrix.

[0087] This filtering process can effectively reduce noise in sensor data and improve the accuracy of attitude estimation. For example, in practical use, when a user turns their head quickly, the raw sensor data may fluctuate, while through Kalman filtering, we can obtain a smoother and more accurate head movement trajectory. In one embodiment of the invention, we can use the process noise covariance matrix Q... t Set as a diagonal matrix, the values ​​of the diagonal elements can be determined based on the sensor's accuracy, typically around 10. -6 Up to 10 -4 Between. Observation noise covariance matrix R t It can also be set as a diagonal matrix, the value of which can be determined experimentally, usually around 10. -2 Between 1 and 1.

[0088] In one embodiment of the present invention, the method for establishing the LSTM neural network model in step (9) is as follows: Define the input data feature sequence as {x1,x2,...,x...} n}, where n is the number of input features, x1 is the current frame input, and y t Let h0 be the hidden state at the input of the first frame, and c0 be the hidden state of the previous frame at the time of the first frame. This can be expressed by the formula:

[0089] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0090] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0091]

[0092] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0093] h t =o t ⊙tanh(C t )

[0094] Among them, f t For the Gate of Oblivion, W f Here is the forget gate weight matrix, [h t -1,x t] is the input for the forget gate, b f For the forget gate bias; i t For the input gate, W i Let b be the input gate weight matrix. i For input gate bias; As a candidate state, W C Let b be the candidate state weight matrix. C For W C Bias matrix; C t For state update gate; o t For output gate, W o Let b be the output gate weight matrix. o For output gate bias; h t σ represents the output corresponding to frame t; σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, and ⊙ represents element-wise multiplication.

[0095] This structural design enables LSTM networks to effectively learn and predict head movement patterns. For example, when a user performs a specific head movement (such as nodding), the LSTM model can capture the temporal features of this movement, thereby accurately predicting subsequent head poses. In one embodiment of this invention, we can use 128 LSTM units and set the input sequence length to 50 (equivalent to 0.25 seconds of data, assuming a sampling rate of 200Hz). This allows us to control the model's complexity and computational cost while ensuring prediction accuracy.

[0096] In one embodiment of the present invention, the preprocessing in step (8) includes: performing window truncation on all data samples to obtain window data in each behavior data, wherein window truncation refers to: using a window with a sliding step size equal to the sensor sampling rate, sliding the window across all frames of the behavior data, forming a new sequence from the sampling points of each frame during the sliding process, setting each new sequence as a data packet, and each behavior data containing several data packets as training data. This preprocessing method can effectively capture the temporal features of user head movements. For example, if we set the window size to 1 second and the sampling rate to 200Hz, then each data packet contains 200 sampling points, which can well reflect the head movement trend in a short period of time. In practical applications, we can use a 50% overlap rate, that is, sliding 100 sampling points each time, which can increase the utilization rate of data while ensuring the continuity between adjacent data packets.

[0097] In one embodiment of the present invention, the following steps are also included:

[0098] (a) Let Q be the newly collected dataset, where the number of data packets is less than that of dataset P;

[0099] (b) Define the output model consisting of all data packets in dataset Q as M1, where:

[0100] y1=W1*h t +b1

[0101] Where y1 is the output fully connected layer, W1 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b1 is the bias vector.

[0102] (c) Learn the dataset Q to obtain the output model M1;

[0103] (d) Define the model consisting of all window data in dataset P as M2, where:

[0104] y2=W2*h t +b2

[0105] Where y2 is the output fully connected layer, W2 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b2 is the bias vector;

[0106] (e) Learn from dataset P to obtain output model M2;

[0107] (f) Define the common data between dataset Q and dataset P as dataset R, and use the transfer learning method to learn dataset R to obtain the output model M3.

[0108] The advantage of this approach lies in its ability to fully utilize existing data and models, improving learning efficiency and accuracy in new scenarios. For example, if we have a general-purpose model M1 trained on a large amount of user data, when a model needs to be customized for new users, we can leverage transfer learning to quickly adapt to the head movement features of the new users. In one embodiment of this invention, we can freeze the parameters of the LSTM layers and only fine-tune the last fully connected layer, which can significantly reduce training time while ensuring the model's generalization ability.

[0109] In one embodiment of the present invention, the method for learning datasets Q and P is as follows: using K-fold cross-validation, where K is a positive integer, typically 5 or 10; training the network using the trained output model M0 for dataset Q; and training the network using the trained output model M1 for dataset P. This claim describes a method for model training using K-fold cross-validation. K is typically 5 or 10. This method can effectively prevent model overfitting and improve the model's generalization ability. For example, if we choose 5-fold cross-validation, we will divide the dataset into 5 parts, using 4 parts as the training set and 1 part as the validation set each time, repeating this 5 times, and finally taking the average performance as the model's evaluation result. In one embodiment of the present invention, we can use an early-stopping strategy, that is, when the performance on the validation set does not improve for 5 consecutive epochs, training is stopped, which can avoid overfitting and save training time.

[0110] Finally, in one embodiment of the invention, the following step is further included: (a) during the testing phase, given the input data sequence {x1,x2,…,x} t}, where t represents the input time step, and Y is obtained. t =(yaw,pitch,roll), where yaw is the yaw angle, pitch is the pitch angle, and roll is the roll angle; (b) The mean absolute error (MAE) is used as the evaluation index, and the formula for calculating MAE is: Where y pred Represents the predicted value, y true (c) Fuse models M1, M2, and M3 to obtain the final model M. final =α*M1+β*M2+γ * M3, where α, β, and γ are weighting coefficients, and α+β+γ=1, 0≤α,β,γ≤1.

[0111] In one embodiment of the invention, the MAE threshold can be set to 5 degrees. This threshold is chosen based on the following considerations: the human perception threshold for head posture changes is typically around 2-3 degrees. Considering potential errors in practical applications, setting the threshold to 5 degrees ensures a good user experience while providing the system with some margin for error. If the MAE exceeds this threshold, we need to readjust the model or collect more training data. This model fusion method can combine the advantages of different models to further improve the accuracy of posture estimation. For example, if model M1 performs well in estimating yaw angle and model M2 performs well in estimating pitch angle, then through appropriate weight fusion, we can obtain a comprehensive model that performs well in all aspects.

[0112] In practical applications, we can determine the optimal weight coefficients using a grid search method. For example, we can set the values ​​of α, β, and γ to the range [0, 0.1, 0.2, ..., 1.0], and then iterate through all possible combinations to select the weight set that performs best on the validation set. Typically, we can set the initial weights to α = 0.4, β = 0.3, and γ = 0.3, and then fine-tune them based on actual performance.

[0113] This multi-sensor fusion-based head pose estimation method for TWS earphones offers several advantages. First, it fully leverages the structural characteristics of TWS earphones, achieving high-precision head pose estimation through the collaborative work of multiple sensors. Second, by employing LSTM neural networks and model fusion technology, this method can effectively learn and predict complex head movement patterns, adapting to different users' habits. Third, the application of a Kalman filter effectively reduces noise in the sensor data, improving the stability and reliability of the estimation results.

[0114] In practical applications, this method can significantly improve the user experience of TWS earbuds. For example, in audio playback scenarios, we can adjust the spatial positioning of the audio in real time based on the user's head posture, creating a more realistic 3D sound experience. When the user turns their head, the position of the sound source changes accordingly, just like in real space. Specifically, if the user turns their head 45 degrees to the left, we can move the sound source 45 degrees to the right accordingly, thus maintaining the absolute position of the sound source in the user's perception.

[0115] Furthermore, this method can also be used to develop novel head-pose-based interaction methods. For example, we can control music playback, pause, or toggle by recognizing specific head movements (such as nodding or shaking). In this case, we can set a threshold, for instance, that when a change in pitch angle is detected to exceed 30 degrees and lasts for more than 0.5 seconds, it is considered a nodding action, thereby triggering the corresponding control command.

[0116] The value of this approach is even more pronounced in virtual reality (VR) or augmented reality (AR) applications. Through accurate head pose estimation, we can achieve more natural and fluid perspective transitions, providing a more immersive user experience. For example, in VR games, when a user turns their head, we can adjust the perspective in real time based on the estimated yaw angle, creating a truly immersive experience. If a user quickly turns their head 180 degrees, the system can update the perspective in less than 50 milliseconds; this low-latency response is crucial for reducing VR sickness.

[0117] In summary, the multi-sensor fusion-based head pose estimation method for TWS earphones proposed in this invention provides strong technical support for expanding the functionality and improving the user experience of TWS earphones through innovative algorithm design and advanced machine learning technology. It can not only be applied to traditional scenarios such as audio playback and human-computer interaction, but also may play an important role in emerging fields such as VR and AR, demonstrating broad application prospects and significant technological value.

[0118] To verify the superiority of this invention, we designed a set of embodiments and comparative examples, and conducted a detailed comparative analysis. The following is a detailed description of the specific contents of the embodiments and comparative examples, as well as the test results.

[0119] Example 1: In this example, we adopted the multi-sensor fusion-based head pose estimation method for TWS earphones proposed in this invention. Specifically, we integrated three sensors—an accelerometer, a gyroscope, and a magnetometer—into the TWS earphone, with sampling rates set to 200Hz, 200Hz, and 50Hz, respectively. We used the coordinate system definition method proposed in this invention and employed a Kalman filter to fuse the sensor data. For the machine learning model, we used 128 LSTM units with an input sequence length of 50. We used 5-fold cross-validation for model training and employed model fusion technology with fusion weights of α = 0.4, β = 0.3, and γ = 0.3.

[0120] Comparative Example 1: In the comparative example, we adopted a traditional head pose estimation method based on a single sensor (using only a gyroscope). We also used an LSTM neural network for model training, but did not employ multi-sensor fusion or model fusion techniques.

[0121] To comprehensively evaluate the performance of the two methods, we designed the following test metrics and detection methods:

[0122] 1. Attitude estimation accuracy: We use an optical motion capture system as the ground truth to calculate the mean absolute error (MAE) between the estimated value and the true value.

[0123] 2. Response time: We measure the time delay from the start of head movement to the system outputting the estimated result.

[0124] 3. Anti-interference capability: We tested the system's performance in an environment with external magnetic field interference.

[0125] 4. Battery life: We measured the battery life of TWS earbuds under continuous use.

[0126] The following is a detailed table of the test results:

[0127]

[0128]

[0129] The test results show that the method proposed in this invention is significantly better than traditional methods in several key indicators.

[0130] First, regarding attitude estimation accuracy, the method of this invention achieves lower MAEs for yaw, pitch, and roll angles, at 2.3°, 1.8°, and 2.1°, respectively, while the MAEs of traditional methods are 5.7°, 4.5°, and 5.2°, respectively. This means that the method of this invention can provide more accurate head attitude estimation, which is crucial for achieving precise spatial audio effects and natural human-computer interaction.

[0131] Secondly, regarding response time, the method of this invention can provide an estimation result in just 15ms, while the traditional method requires 28ms. This faster response speed can significantly improve the user experience, especially in latency-sensitive application scenarios such as VR / AR.

[0132] Furthermore, the method of this invention demonstrates excellent anti-interference capability. Under external magnetic field interference, the yaw angle (MAE) of this invention is only 3.5°, while the conventional method reaches as high as 12.8°. This indicates that the multi-sensor fusion strategy of this invention can effectively resist external interference and ensure stable performance in various complex environments.

[0133] Finally, while the method of this invention is slightly inferior to the conventional method in terms of battery life (5.2 hours vs. 6.5 hours), this slight reduction in battery life is acceptable considering the significant performance improvement. Moreover, with further hardware optimization and algorithm efficiency improvements, we are confident that this gap can be narrowed in future versions.

[0134] In summary, the multi-sensor fusion-based head pose estimation method for TWS earphones proposed in this invention significantly outperforms traditional methods in key indicators such as accuracy, response speed, and anti-interference capability. These advantages enable this method to provide users with more accurate, smooth, and stable head pose estimation, thereby supporting a more natural and immersive user experience. This method performs excellently in both everyday audio playback scenarios and more demanding VR / AR applications.

[0135] These test results fully demonstrate the innovation and practical value of this invention. The multi-sensor fusion strategy not only improves the accuracy of pose estimation but also enhances the system's anti-interference capability. The application of LSTM neural networks enables the system to effectively learn and predict complex head movement patterns. Furthermore, the introduction of model fusion technology further improves the overall performance of the system.

[0136] Based on the above analysis, we can consider Embodiment 1 as the preferred embodiment of the present invention. It fully demonstrates the superiority of the present invention in practical applications, providing strong technical support for the functional expansion and user experience enhancement of TWS earphones.

[0137] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for estimating the head pose of TWS earphones based on multi-sensor fusion, characterized in that, Includes the following steps: (1) Set up an accelerometer, gyroscope, and magnetometer on the TWS earphone and collect sound signals through a microphone; (2) Define the TWS earphone coordinate system with the left earbud as the origin; (3) Obtain the vector between the left and right earbuds measured by the magnetometer; (4) Obtain the acceleration and angular velocity of the left earbud's x, y, and z axes through a microelectromechanical system (MEMS) inertial sensor; (5) Calculate the rotation matrix R at time t based on the gyroscope angular velocity and the relative position of the magnetometer and the left earbud. t (6) Calculate the yaw angle at time t based on the relative position of the gyroscope and the left earpiece and the acceleration of the left earpiece coordinate axis. t Calculate the pitch angle at time t based on the gyroscope angular velocity and the angular velocity of the left earpiece coordinate axis. t (7) Use a Kalman filter to filter the estimated values ​​obtained in steps (3)-(6) to estimate the true attitude of the headphones; (8) Collect TWS earphone motion data from different users, preprocess it, and divide it into training set, validation set and test set. (9) Use LSTM neural network to perform head motion pattern recognition, including: a) establish LSTM neural network model, train it using training set, and obtain output model. b) Test the output model using the test set to obtain the test results.

2. The method according to claim 1, characterized in that, The method for defining the TWS earphone coordinate system in step (2) is as follows: Assume the wearer of the TWS earphones is facing the XY plane, with the left earbud as the origin of the coordinate system; the line connecting the right and left earbuds points in front of the earphones and is perpendicular to the normal direction of the front of the human body, the Z-axis is vertically upward, and the X-axis is perpendicular to the forward direction of the earphones.

3. The method according to claim 1, characterized in that, The method for obtaining the vector between the left and right earbuds in step (3) is as follows: V=P right -P left , Where V is the direction vector, representing the relative position of the right earbud with respect to the left earbud, and P... right Let P be the coordinates of the right earpiece. left The coordinates are for the left earpiece.

4. The method according to claim 3, characterized in that, It also includes the step of transforming the vector between the left and right earbuds in the left earbud coordinate system to the headphone position coordinate system: V′=M*V, Among them, V ′ Let M be the vector in the headphone position coordinate system, and M be the translation matrix from the left earpiece coordinate system to the headphone position coordinate system, expressed as: Among them, [d x ,d y ] represents the distance between the origin of the left earbud coordinate system and the origin of the headphone position coordinate system, and θ represents the angle between the line connecting the left and right earbuds and the line connecting the horizontal axis of the headphone.

5. The method according to claim 1, characterized in that, The Kalman filtering algorithm in step (7) includes: (a) Update the estimated value at time t using the rotation matrix and motion parameters from the previous time t-1: in, Let F be the prior estimate at time t. t Here is the state transition matrix. Let B be the posterior estimate at time t-1. t To control the input matrix, u t Control vector; (b) Calculate the error covariance matrix: Among them, P t|t-1 Let P be the prior covariance matrix. t-1|t-1 Let Q be the posterior covariance matrix at time t-1. t The process noise covariance matrix is... For F t The transpose of the matrix; (c) Update the error covariance matrix P t Obtain the rotation matrix R at time t. t The estimated value: P t|t =(I-K t ·H t )·P t|t-1 , Among them, K t For Kalman gain, H t Let R be the observation matrix. t To observe the noise covariance matrix, z t For the observation vector, Let P be the posterior estimate at time t. t|t Let be the posterior covariance matrix at time t, and I be the identity matrix.

6. The method according to claim 1, characterized in that, The method for establishing the LSTM neural network model in step (9) is as follows: Define the input data feature sequence as {x1, x2, ..., x...} n }, where n is the number of input features, x1 is the current frame input, and y t Let h0 be the hidden state at the input of the first frame, and c0 be the hidden state of the previous frame at the time of the first frame. This can be expressed by the formula: f t =σ(W f ·[h t-1 ,x t ]+b f ) i t =σ(W i ·[h t-1 ,x t ]+b i ) the t =σ(W o ·[h t-1 ,x t ]+b o ) h t =o t ⊙tanh(C t ) Among them, f t For the Gate of Oblivion, W f Here is the forget gate weight matrix, [h t -1,x t ] is the input for the forget gate, b f For the forget gate bias; i t For the input gate, W i Let b be the input gate weight matrix. i For input gate bias; As a candidate state, W C Let b be the candidate state weight matrix. C For W C Bias matrix; C t For state update gate; o t For output gate, W o Let b be the output gate weight matrix. o For output gate bias; h t σ represents the output corresponding to frame t; σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, and ⊙ represents element-wise multiplication.

7. The method according to claim 1, characterized in that, The preprocessing in step (8) includes: performing window truncation on all data samples to obtain window data in each behavior data. Window truncation means: using a window with a sliding step size the same as the sensor sampling rate, sliding the window in all frames of the behavior data, forming a new sequence from each frame sampling point during the sliding process, setting each new sequence as a data packet, and each behavior data containing several data packets as training data.

8. The method according to claim 7, characterized in that, It also includes the following steps: (a) Let Q be the newly collected dataset, where the number of data packets is less than that of dataset P; (b) Define the output model consisting of all data packets in dataset Q as M1, where: y1=W1*h t +b1 Where y1 is the output fully connected layer, W1 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b1 is the bias vector. (c) Learn the dataset Q to obtain the output model M1; (d) Define the model consisting of all window data in dataset P as M2, where: y2=W2*h t +b2 Where y2 is the output fully connected layer, W2 is the weight matrix, and h t b1 is the output of the LSTM hidden state, and b2 is the bias vector; (e) Learn from dataset P to obtain output model M2; (f) Define the common data between dataset Q and dataset P as dataset R, and use the transfer learning method to learn dataset R to obtain the output model M3.

9. The method according to claim 8, characterized in that, The method for learning datasets Q and P is as follows: K-fold cross-validation is used, where K is a positive integer, usually 5 or 10. For dataset Q, train the network using the pre-trained output model M0; For dataset P, train the network using the pre-trained output model M1.

10. The method according to claim 9, characterized in that, It also includes the following steps: (a) During the testing phase, given the input data sequence {x1, x2, ..., x...} t }, where t represents the input time step, and Y is obtained. t =(ya w,pitch,roll), where yaw is the yaw angle, pitch is the pitch angle, and roll is the roll angle; (b) The mean absolute error (MAE) is used as the evaluation index, and the formula for calculating MAE is: Where y pred Represents the predicted value, y true (c) Fuse models M1, M2, and M3 to obtain the final model M. final =α*M1+β*M2+γ * M3, where α, β, and γ are weighting coefficients, and α+β+γ=1, 0≤α,β,γ≤1.