Earphone space audio positioning method integrated with MEMS sensor and control system

By integrating MEMS sensors and advanced algorithms, the accuracy and reliability problems in traditional headphone spatial audio positioning technology are solved, and high-precision, low-latency, personalized sound field rendering and cross-modal immersion experience are achieved, improving the consistency and immersion of users' spatial audio experience.

CN120602831APending Publication Date: 2025-09-05SHENZHEN SHENGJIALI ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510657671.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional headphone spatial audio positioning technology has insufficient accuracy and reliability of MEMS sensors, poor generality of sound field rendering algorithms, low multimodal data fusion efficiency, and lack of personalized audio tuning capabilities, resulting in accumulated posture resolution errors, decreased sound source positioning accuracy and significant differences in user experience.

Method used

The temperature-compensated MEMS sensor group, beamforming technology, double-expanded Kalman filtered attitude solution, convolutional neural network sound field prediction, adaptive wavefield synthesis, personalized HRTF generation model, super-resolution DOA estimation, fuzzy logic sensor evaluation, reinforced learning sound field optimization algorithm, heterogeneous computing module and tactile feedback driver are used to realize sensor data synchronization, attitude solution, sound field reconstruction and personalized audio rendering.

Benefits of technology

It achieves improved accuracy of six-degree-of-freedom head motion data, precise capture of environmental sound sources, noise suppression, improved reliability of posture solution, reduced sound and image positioning errors, and enhanced consistency of user experience, providing a high-precision, low-latency spatial audio immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602831A_ABST
    Figure CN120602831A_ABST
Patent Text Reader

Abstract

The invention discloses an earphone space audio positioning method integrated with an MEMS sensor and a control system, and relates to the field of earphone space audio positioning. A temperature compensation MEMS sensor group and a microphone array are used for collecting motion data and sound wave signals, and signals are enhanced through orthogonal wavelet denoising and beam forming; the attitude is solved by adopting double-extended Kalman filtering and fusing data, and fuzzy logic is introduced to evaluate the credibility of the sensor; predicting a sound field by using a CNN model, and reconstructing the sound field in combination with an adaptive wave field synthesis technology; and based on auditory masking effect optimization balance, the influence of altitude on sound velocity is compensated. Through an ultra-precision sensor and an advanced algorithm, the positioning precision and reliability of earphone space audio are improved, and delay and power consumption are reduced; accurate sound source positioning, personalized sound field adaptation and cross-modal immersion experience in a complex scene are realized, and the audio experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of earphone spatial audio positioning, and in particular to an earphone spatial audio positioning method and control system integrating a MEMS sensor. Background Art

[0002] With the rapid development of virtual reality (VR), augmented reality (AR), and smart audio technologies, headphones, as core devices for human-computer interaction, are becoming increasingly demanding in terms of spatial audio localization capabilities, making them crucial for enhancing user immersion. Traditional headphones rely primarily on fixed algorithms or simple inertial sensors for spatial audio localization, making it difficult to accurately capture head movements and complex environmental changes in real time. For example, in dynamic motion scenes, posture calculation errors based on a single MEMS inertial sensor can quickly accumulate, leading to shifts in sound and image localization. In noisy environments, the sound source localization accuracy of traditional microphone arrays decreases significantly, making it impossible to effectively separate the target sound source from ambient noise.

[0003] In existing technologies, spatial audio positioning in headphones faces multiple challenges: First, the accuracy and reliability of MEMS sensors are insufficient. Traditional sensors are easily affected by factors such as temperature drift and mechanical vibration, resulting in the accumulation of attitude solution errors over time. Second, the universality of sound field rendering algorithms is poor. Traditional head-related transfer functions (HRTFs) rely on standardized databases and cannot adapt to individual head shape differences, resulting in large positioning errors in high-frequency bands. Third, the efficiency of multimodal data fusion is low. Time synchronization errors between inertial data and audio data can easily lead to sound field rendering delays, affecting the user experience. In addition, existing systems lack the ability to perceive and adaptively adjust the wearing status of headphones and dynamic changes in the environment, making it difficult to maintain stable positioning performance in complex scenarios.

[0004] When it comes to sound field reconstruction, traditional algorithms like High-Order Ambisonics (HOA) require high computing resources, suffer from poor real-time performance, and suffer from increased sound and image localization errors with distance. Furthermore, existing headphones generally lack awareness of the user's physiological characteristics (such as auricle shape and fit), making personalized audio adjustments impossible. This results in significant differences in the spatial audio experience for different users. Overcoming the limitations of sensor accuracy, improving the efficiency of multimodal data fusion, and achieving personalized sound field adaptation are key challenges that must be addressed in current headphone spatial audio localization technology. Summary of the Invention

[0005] The present invention proposes a method and control system for spatial audio positioning of headphones with integrated MEMS sensors to solve the problems mentioned in the above-mentioned prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method and control system for spatial audio positioning of headphones with integrated MEMS sensors, comprising:

[0007] Multimodal dynamic calibration acquisition steps: A temperature-compensated MEMS sensor group is built into the earphones, and orthogonal wavelet transform is used to remove sensor noise to construct a six-degree-of-freedom head motion data S = [a, ω, P, T] and the ambient sound pressure level L p The microphone array uses beamforming technology to guide the vector a(θ)=e -j2πfd·u / c Suppress noise power ratio;

[0008] Motion attitude fusion solution steps: A dual extended Kalman filter-based attitude solution algorithm is proposed. The main filter fuses inertial data to estimate the attitude Q, and the auxiliary filter monitors the health status of the sensor. When the gyroscope drift rate exceeds the threshold, the fault isolation mechanism is triggered and the redundant sensor data is switched. The attitude update formula is:

[0009] Real-time sound field reconstruction steps: Construct a sound field prediction model based on convolutional neural network, input the head motion trajectory {S1, S2, ..., S n} and environmental characteristics {L p ,RT60}, output virtual sound source position P(x,y,z) and rendering parameters {G,h(t)}; using adaptive wave field synthesis technology, dynamically adjust the speaker array weight W=H according to the head position -1 D;

[0010] Auditory perception optimization steps: Design a dynamic equalization algorithm based on the human ear's auditory masking effect, and adjust the filter coefficient according to the real-time calculation of the critical band masking threshold T(f). The formula is: S(f) is the signal spectrum, = 10 -6 ; Combined with the pressure sensor data to compensate for the effect of altitude on the speed of sound, the formula is corrected

[0011] Furthermore, the multimodal dynamic calibration acquisition steps include developing a sensor spatiotemporal synchronization algorithm, using FPGA nanosecond clock synchronization, and using adaptive sparse representation technology to compress sensor data; the microphone array introduces super-resolution DOA estimation technology and locates based on the spatial smoothing MUSIC algorithm, the formula is R is the covariance matrix.

[0012] Furthermore, the motion posture fusion solution step includes designing a fuzzy logic-based sensor credibility evaluation model, inputting the accelerometer / gyroscope measurement residual e a ,e g , output weight coefficient ω a ,ω g , the formula is Weighted suppression of faulty sensors; using quaternion differential equations to update attitude in real time, combined with magnetic field compensation algorithm to eliminate geomagnetic field interference, attitude matrix update frequency is 200Hz.

[0013] Furthermore, the real-time sound field reconstruction step includes building a personalized HRTF generation model, generating user-exclusive HRTF based on 3D facial scanning data and deep learning, supporting dynamic head model matching, and automatically switching HRTF subsets according to head movement speed; proposing a reinforcement learning-based sound field optimization algorithm, and adjusting the sound field parameters using the user's subjective immersion score R as the reward function.

[0014] Furthermore, in the super-resolution DOA estimation technology, the number of physical microphones is virtualized from 4 to 16 through the virtual array expansion algorithm, and the formula is R 虚拟 =AR 物理 A H , A is the expansion matrix.

[0015] Furthermore, the triangle membership function formula designed in the fuzzy logic evaluation model is μ(e) = max(0,1-|e| / δ), and the sensor self-test is triggered when the residual exceeds the threshold δ.

[0016] Furthermore, in the reinforcement learning sound field optimization algorithm, the state space S includes head posture, ambient noise, and audio content features, and the action space A includes the sound field parameter adjustment amount. The algorithm is trained through the experience playback mechanism.

[0017] Furthermore, the following modules are also included:

[0018] Ultra-precision sensing module: uses wafer-level packaged MEMS sensors, integrating a three-axis accelerometer, a three-axis gyroscope, and an air pressure sensor; the four-microphone array uses MEMS vector microphones to synchronously collect sound pressure and particle velocity, and single-microphone DOA estimation;

[0019] Heterogeneous computing module: Equipped with an AI accelerator and FPGA co-processing unit. The AI ​​accelerator runs the CNN sound field prediction model, FPGA beamforming, and AWFS algorithms.

[0020] Intelligent Rendering Module: Built-in reconstructed sound field driver, supports 128-channel digital beamforming, and uses GaN power amplifiers to drive the speaker array;

[0021] Dynamic adaptation module: includes a real-time operating system and an adaptive control engine, dynamically switches the sound field mode according to the perception data, and supports user-defined parameter mapping strategies.

[0022] Furthermore, the ultra-precision sensing module includes a laser interferometer displacement sensor to monitor the fit between the earphones and the ear canal in real time, and triggers a wearing calibration reminder when the displacement exceeds the threshold; the flexible pressure sensor array detects the force distribution of the auricle and automatically adjusts the sound parameters to compensate for physical sound insulation differences.

[0023] Furthermore, the intelligent rendering module includes a holographic sound rendering unit, which generates three-dimensional sound holograms based on wave field synthesis technology, supports 5.1 / 7.1 channel decoding and Ambisonics format conversion; the tactile feedback driver generates tactile vibration signals according to the low-frequency components of the sound field.

[0024] Compared with the existing technology, the beneficial effects of the present invention are:

[0025] At the perception layer, an ultra-precision MEMS sensor group and a vector microphone array are used, combined with a spatiotemporal synchronization algorithm and super-resolution DOA technology to achieve accurate capture of six-degree-of-freedom head motion data (accuracy <0.3°) and environmental sound sources. The target sound source positioning accuracy is improved to ±2°, the noise suppression ratio reaches 25dB, and voice and environmental noise are effectively separated.

[0026] At the algorithm level, a dual extended Kalman filter architecture combined with a fuzzy logic sensor evaluation model enables dynamic isolation and redundant switching of sensor failures. Even in the event of a single sensor failure, the attitude solution error can be maintained at <1.2°, improving reliability by 90% compared to traditional solutions. A personalized HRTF generation model based on a convolutional neural network (CNN) generates user-specific sound field parameters through 3D facial scan data, reducing high-frequency (>10kHz) positioning errors by 50%, addressing the lack of adaptability of standardized HRTF libraries. The reinforcement learning sound field optimization algorithm autonomously adjusts rendering parameters based on the user's real-time motion and environmental data, improving subjective immersion by 30% and achieving an audio and video positioning error of <2°.

[0027] At the system level, the heterogeneous computing module integrates an AI accelerator and an FPGA, achieving low-latency processing (<8ms) from data acquisition to sound field rendering, keeping power consumption below 50mW to meet the battery life requirements of wearable devices. The introduction of flexible pressure sensors and haptic feedback drivers monitors the wearer's state in real time and dynamically compensates for physical sound insulation differences. Combined with holographic sound rendering technology, this creates a cross-modal immersive experience that synergizes hearing and touch. Tactile vibrations in the low-frequency band enhance positioning perception, further enhancing the realism of spatial audio.

[0028] Actual tests show that the posture solution delay of the solution of the present invention is less than 3ms and the sound field rendering synchronization error is less than 10ms in dynamic motion scenes (such as running and rapid head rotation); in noisy environments (such as subways and shopping malls), the target voice clarity is improved by 40%, and the sound and image positioning error is reduced by 60% compared with traditional solutions. Through personalized adjustment and adaptive environmental perception, the consistency of spatial audio experience of different users is improved by 75%, which significantly reduces the experience deviation caused by individual differences. Overall, the technology of the present invention breaks through the performance bottleneck of traditional headphone spatial audio positioning, and provides a high-precision, low-latency, and highly robust solution for scenarios such as VR / AR, smart communications, and immersive entertainment. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a schematic block diagram of a method for spatial audio localization of headphones with integrated MEMS sensors proposed by the present invention;

[0030] Figure 2 This is a schematic block diagram of a headphone spatial audio positioning control system with integrated MEMS sensors proposed by the present invention;

[0031] Figure 3 This is a schematic diagram of the comparison of attitude solution errors;

[0032] Figure 4 This is a schematic diagram of the comparison of sound source localization accuracy;

[0033] Figure 5 This is a schematic diagram of system power consumption comparison;

[0034] Figure 6 This is a schematic diagram comparing the positioning errors of different algorithms;

[0035] Figure 7 A schematic diagram showing the comparison of users' subjective ratings;

[0036] Figure 8 Schematic diagram of speech clarity comparison under different noise environments;

[0037] Figure 9 Schematic diagram of positioning error comparison at different altitudes. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0040] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.

[0041] Reference Figures 1 to 9 : A method and control system for spatial audio positioning of headphones with integrated MEMS sensors, comprising:

[0042] Multimodal dynamic calibration acquisition steps: A temperature-compensated MEMS sensor group is built into the headset (accelerometer bias stability <50μg, gyroscope angle random walk <0.005° / √h), and orthogonal wavelet transform is used to remove sensor noise. The six-degree-of-freedom motion data S = [a, ω, P, T] and the ambient sound pressure level L are constructed. p The microphone array uses beamforming technology to guide the vector a(θ)=e -j2πfd·u / c The target sound source signal is enhanced and the noise power ratio (NRR) is increased to 25dB.

[0043] Motion attitude fusion solution steps: A dual extended Kalman filter (DEKF)-based attitude solution algorithm is proposed. The main filter fuses inertial data to estimate the attitude Q, and the auxiliary filter monitors the health status of the sensor. When the gyroscope drift rate is detected to be greater than 0.1° / s, the fault isolation mechanism is triggered and the redundant sensor data is switched. The attitude update formula is: Solution delay <3ms, static attitude error <0.3°, dynamic error <1.2°.

[0044] Real-time sound field reconstruction steps: Construct a sound field prediction model based on convolutional neural network (CNN), input the head motion trajectory {S1, S2, ..., S n} and environmental characteristics {L p ,RT60}, outputs the virtual sound source position P(x,y,z) and rendering parameters {G,h(t)}; uses adaptive wave field synthesis (AWFS) technology to dynamically adjust the speaker array weight W=H according to the head position -1D, sound image positioning error <2°, sound field uniformity >90%.

[0045] Auditory perception optimization steps: Design a dynamic equalization algorithm based on the human ear's auditory masking effect, and adjust the filter coefficient according to the real-time calculated critical band masking threshold T(f). The formula is: S(f) is the signal spectrum, = 10 -6 ; Combined with the pressure sensor data to compensate for the effect of altitude on the speed of sound, the formula is corrected Ensure that the positioning error is less than 3% at different altitudes.

[0046] In this invention, the multimodal dynamic calibration acquisition step includes developing a sensor spatiotemporal synchronization algorithm and introducing super-resolution DOA estimation technology. The parallel processing capabilities of FPGA (field programmable gate array) hardware are utilized to achieve nanosecond-level clock synchronization. A high-precision clock management module is built into the FPGA. By real-time monitoring and fine-tuning the clock signals of each sensor, the timestamp error is controlled to less than 10ns, ensuring near-synchronization when different sensors collect data and avoiding data distortion caused by time deviation. Simultaneously, adaptive sparse representation (ASR) technology is used to compress sensor data. By intelligently analyzing data features, key information is identified and redundant data is removed, achieving a compression ratio of up to 8:1. The compressed data significantly reduces the data volume during transmission, reducing transmission delay to less than 1ms, ensuring that data can be quickly and accurately transmitted to the processing unit. A microphone array incorporates super-resolution DOA (angle of arrival) estimation technology. This technology is based on the spatially smoothed MUSIC algorithm, a high-resolution algorithm that utilizes the orthogonality of the signal and noise subspaces to estimate the signal's direction of arrival. By spatially smoothing the signals received by the microphone array, noise and interference are effectively suppressed, improving signal quality. formula In this equation, R is the covariance matrix, reflecting the correlation between the signals received by the microphones. a(θ) is the array manifold vector, which is related to the signal's angle of arrival, θ. By optimizing this formula, the angle of arrival of the sound can be accurately calculated. This technology improves the positioning accuracy of the microphone array to ±2°, breaking through the traditional Rayleigh limit constraint and enabling headphones to more accurately locate the sound source, providing users with a more realistic and directional audio experience.

[0047] In the present invention, the motion posture fusion solution step includes designing a sensor credibility evaluation model based on fuzzy logic and using quaternion differential equations to update the posture in real time. Designing a sensor credibility evaluation model based on fuzzy logic, using the measurement residuals e of the accelerometer and gyroscope a 、e gAs a key input parameter, the measurement residual reflects the deviation between the actual sensor measurement value and the theoretical expected value, and is an important basis for judging whether the sensor is working properly. Based on fuzzy logic rules, the model performs complex logical judgments and operations on the input measurement residual. Through a specific membership function, the measurement residual is mapped to different fuzzy sets to quantify the reliability of the sensor. The final output weight coefficient ω a ,ω g , the formula is k i These coefficients, determined through extensive experimentation and optimization, adjust the sensitivity of the weight coefficient to measurement residuals. When a sensor failure causes large measurement residuals, the corresponding weight coefficient is automatically reduced, dynamically weighting the faulty sensor to suppress the error, thereby ensuring the measurement accuracy and stability of the entire system. Quaternion differential equations are used to update the posture in real time. Quaternions, as a mathematical tool, can concisely and efficiently describe rotation and posture changes in three-dimensional space. By solving these quaternion differential equations, the system can quickly and accurately calculate the device's current posture information based on real-time sensor data. Furthermore, a magnetic field compensation algorithm is incorporated to mitigate geomagnetic interference. The geomagnetic field can interfere with sensor measurements, causing deviations in posture measurements. The magnetic field compensation algorithm analyzes the magnetic field data measured by the sensors, builds an interference model, and applies corrections to the measurement results. The compensated posture matrix updates at a frequency of up to 200Hz. This high update frequency enables rapid response to device posture changes, meeting the stringent low-latency requirements of posture measurement in VR / AR scenarios and providing users with a smooth, realistic, and immersive experience.

[0048] In this invention, the real-time sound field reconstruction step includes constructing a personalized HRTF generation model and proposing a reinforcement learning (RL)-based sound field optimization algorithm. To construct the personalized HRTF (Head-Related Transfer Function) generation model, 3D facial scanning technology is used to obtain the user's facial geometry data, including features such as auricular shape and facial contours that significantly influence sound propagation. This data is input into a deep learning model, which, by learning from a large number of different facial features and corresponding HRTF data, exploits the complex mapping relationship between facial features and sound localization characteristics to generate a user-specific HRTF. In the high-frequency band (>10kHz), this model can effectively reduce localization error by up to 50%. Furthermore, it supports dynamic head model matching. Built-in sensors monitor head movement speed in real time. When the head movement state changes, the system automatically switches from a pre-established HRTF subset to the set of parameters that best suits the current state, ensuring an accurate audio localization experience under various head movement conditions. A reinforcement learning (RL)-based sound field optimization algorithm is proposed, using the user's subjective immersion score R as the reward function. During the algorithm's operation, the intelligent agent continuously attempts to adjust sound field parameters such as volume distribution and virtual sound source location. After each adjustment, the system collects users' subjective feedback on immersion and converts it into a score, R. If the score improves, it indicates the parameter adjustment was correct, and the algorithm strengthens the adjustment strategy; otherwise, it weakens it. Through a process of continuous trial and error and learning, the algorithm automatically optimizes the sound field parameters. Compared to traditional algorithms, the reinforcement learning algorithm converges three times faster, finding the optimal sound field parameter combination more quickly and creating a more immersive audio environment for users.

[0049] In this invention, a virtual array expansion algorithm is used in the super-resolution DOA estimation technology. The headset is originally equipped with 4 physical MEMS microphones. In order to improve the audio positioning accuracy, a virtual array expansion algorithm is used to virtually expand the number of microphones to 16. Based on signal processing and matrix operations, the formula R 虚拟 =AR 物理 A H In, R 物理 Represents the covariance matrix constructed by the signals collected by the four physical microphones, reflecting the correlation between the audio signals received by the physical microphones. A is a carefully designed expansion matrix that processes and expands the physical microphone signals through specific mathematical operation rules. H is the conjugate transposed matrix of A, ensuring signal processing accuracy and stability during calculations. This matrix operation yields the covariance matrix Rvirtual of the virtual 16 microphones, quadrupling the effective aperture and significantly improving the microphone array's range and accuracy of spatial sound signals. This virtual, expanded microphone array more accurately captures the angle of arrival of sound, providing solid technical support for high-precision spatial audio positioning in headphones and effectively enhancing the user's sense of space and orientation in the audio experience.

[0050] In this invention, a fuzzy logic evaluation model is designed with a triangular membership function, calculated as μ(e) = max(0,1-|e| / δ). e is the residual error, reflecting the degree of deviation between the actual audio localization result and the expected ideal result. δ is the threshold value determined through extensive experimentation and precise calibration, and serves as the critical limit for determining whether to trigger a sensor self-test. When the absolute value of the residual error e exceeds the threshold δ, it indicates that the current audio localization may have significantly deviated. In this case, the output of the triangular membership function triggers the sensor self-test. The sensor self-test comprehensively checks all functional indicators of the MEMS sensor, including the measurement accuracy of the accelerometer, the angular stability of the gyroscope, and the sensitivity of the microphone. A series of sophisticated testing and calibration processes ensure the normal and accurate operation of the sensor. The fuzzy logic evaluation model maintains a false alarm rate of less than 0.1%. The sensor self-test is triggered based on actual abnormal conditions, not false positives. This effectively ensures the reliability and stability of the headphone spatial audio localization system, providing users with a continuously accurate and high-quality audio localization experience.

[0051] In the present invention, the state space S of the reinforcement learning sound field optimization algorithm includes head posture, ambient noise, and audio content features. Head posture is captured by the MEMS sensor built into the headset, including data fusion of a three-axis accelerometer and a three-axis gyroscope, which can provide real-time and accurate feedback on posture information such as head rotation and tilt. Ambient noise is collected by a microphone array, and advanced spectrum analysis technology is used to identify noise characteristics such as frequency and intensity, providing environmental background information for subsequent audio processing. Audio content features are extracted with the help of a deep learning model, which analyzes elements such as rhythm, melody, and timbre based on the time and frequency domain characteristics of the audio. The action space A mainly involves the adjustment amount of sound field parameters, including the adjustment amplitude of volume gain, equalizer parameters, virtual sound source position, etc. The algorithm adopts an experience replay mechanism to store the experience data generated by the intelligent agent during the interaction with the environment in a replay buffer. During training, experience data is randomly sampled for learning to break the correlation between data, prevent the neural network from falling into local optimal solutions, and greatly improve training efficiency. After a large number of experimental verifications, the average convergence step of the algorithm is less than 500 times, and it can quickly reach a stable optimization state, realizing the precise positioning of the spatial audio of the headphones and high-quality sound field effects.

[0052] The present invention also includes the following modules:

[0053] Ultra-precision sensing module: MEMS sensor with wafer-level packaging (size < 2mm 3) and an integrated three-axis accelerometer (noise density <100μg / √Hz) reduce the impact of external interference noise when detecting motion acceleration, accurately capturing subtle acceleration changes. The three-axis gyroscope (angle random walk <0.003° / √h) provides ultra-high stability and accuracy when measuring angle changes, providing reliable data for device posture perception. The air pressure sensor (resolution 0.01hPa) accurately senses changes in ambient air pressure, assisting in accurately determining information such as altitude and height.

[0054] Heterogeneous Computing Module: Equipped with an AI accelerator and FPGA co-processing unit, the AI ​​accelerator boasts 12TOPS of computing power, enabling efficient execution of CNN sound field prediction models, rapid analysis of massive amounts of audio data, and prediction of sound field characteristics. The FPGA implements beamforming and AWFS (Adaptive Wideband Frequency Space) algorithms, optimizing signal processing through hardware programming. These two components work together to minimize overall processing latency to less than 8ms, ensuring real-time audio processing while consuming less than 50mW, achieving a balanced balance of performance and energy efficiency.

[0055] Intelligent Rendering Module: Built-in reconfigurable sound field drivers support 128-channel digital beamforming, enabling flexible shaping of complex three-dimensional sound fields. Utilizing a GaN power amplifier with over 90% efficiency, compared to traditional amplifiers, this delivers robust driving power while significantly reducing energy consumption and heat generation. Driven by this module, the speaker array boasts a dynamic sound pressure level range exceeding 120dB, delivering rich sound layers from the faintest to the loudest, with a total harmonic distortion (THD) of less than 0.1%, ensuring high-fidelity sound reproduction.

[0056] Dynamic Adaptation Module: The dynamic adaptation module includes a real-time operating system (RTOS) and an adaptive control engine. The RTOS features fast response and can promptly process data transmitted by the perception module. Based on the perception data, the adaptive control engine can intelligently determine the current usage scenario, such as commuting, theater, and gaming, and dynamically switch the sound field mode. The mode switching delay is less than 50ms, achieving seamless switching. It also supports user-defined parameter mapping strategies, allowing users to personalize parameters such as the sound field mode according to their personal preferences to meet diverse needs.

[0057] In this invention, the ultra-precision sensing module also includes a laser interferometer displacement sensor and a flexible pressure sensor array. The laser interferometer displacement sensor has a measurement accuracy of ±0.1μm. It uses the principle of laser interferometry to measure the relative displacement between the earphone and the ear canal by emitting a laser beam and analyzing the changes in the interference fringes of the reflected light. During use, the fit of the earphone and the ear canal is monitored in real time. If the displacement exceeds 0.5mm, a wearing calibration reminder is quickly triggered. Significant changes in the fit of the earphone and the ear canal can cause sound leakage and deterioration in sound quality. Timely calibration reminders allow users to adjust the earphone position to ensure a consistently good audio experience. The flexible pressure sensor array has a spatial resolution of 1mm and can precisely detect the force distribution on the auricle. Made of a special flexible material, the sensor conforms to the curve of the auricle, deforming with the movement of the auricle while accurately sensing pressure changes. The collected pressure data can be used to analyze differences in physical sound insulation. For example, different wearing force and angle can cause uneven force on different parts of the auricle, thereby affecting the sound insulation effect. The system will automatically adjust sound parameters, such as gain and equalization, to compensate for the sound quality loss caused by differences in physical sound insulation, so that the headphones can output stable, high-quality sound effects in various wearing states, bringing users consistent high-quality listening enjoyment.

[0058] In the present invention, the intelligent rendering module also includes a holographic sound rendering unit and a tactile feedback driver. The holographic sound rendering unit adopts wave field synthesis technology, an audio processing method based on acoustic principles, and generates extremely realistic three-dimensional sound holograms by calculating the physical phenomena such as the propagation, interference and diffraction of sound waves in three-dimensional space. The unit has powerful decoding capabilities and can perfectly support 5.1 and 7.1 channel decoding to meet the needs of different audio systems. At the same time, it can also realize conversion to the Ambisonics format. As a high-order surround sound technology, Ambisonics can provide a more flexible and immersive audio experience. Through format conversion, it can adapt to more audio devices and application scenarios. The tactile feedback driver focuses on enhancing the immersive experience and accurately captures and analyzes the low-frequency components (frequency less than 200Hz) in the sound field. Using high-precision sensors and complex algorithms, low-frequency sound wave signals are converted into tactile vibration signals. During the signal conversion process, the control of the vibration frequency is extremely strict to ensure that the vibration frequency error is less than 1%. Precise frequency control not only allows users to actually feel the vibrations brought by low-frequency sounds and enhance the sense of immersion, but also achieves enhanced cross-modal positioning through the coordination of vibration signals and sound signals, allowing users to more accurately perceive the direction of the sound source and obtain a more realistic and interactive experience.

[0059] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for spatial audio positioning of headphones with integrated MEMS sensors, characterized in that: The following steps are involved: Multimodal dynamic calibration acquisition steps: A temperature-compensated MEMS sensor group is built into the earphones, and orthogonal wavelet transform is used to remove sensor noise to construct a six-degree-of-freedom head motion data S = [a, ω, P, T] and the ambient sound pressure level L p The microphone array uses beamforming technology to guide the vector a(θ)=e -j2πfd·u / c Suppress noise power ratio; Motion attitude fusion solution steps: A dual extended Kalman filter-based attitude solution algorithm is proposed. The main filter fuses inertial data to estimate the attitude Q, and the auxiliary filter monitors the health status of the sensor. When the gyroscope drift rate exceeds the threshold, the fault isolation mechanism is triggered and the redundant sensor data is switched. The posture update formula is Real-time sound field reconstruction steps: Construct a sound field prediction model based on convolutional neural network, input the head motion trajectory {S1, S2, ..., S n } and environmental characteristics {L p ,RT60}, output virtual sound source position P(x,y,z) and rendering parameters {G,h(t)}; using adaptive wave field synthesis technology, dynamically adjust the speaker array weight W=H according to the head position -1 D; Auditory perception optimization steps: Design a dynamic equalization algorithm based on the human ear's auditory masking effect, and adjust the filter coefficient according to the real-time calculation of the critical band masking threshold T(f). The formula is: S(f) is the signal spectrum, = 10 -6 ; Combined with the pressure sensor data to compensate for the effect of altitude on the speed of sound, the formula is corrected 2. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 1, characterized in that: The multimodal dynamic calibration acquisition steps include developing a sensor spatiotemporal synchronization algorithm, using FPGA nanosecond clock synchronization, and using adaptive sparse representation technology to compress sensor data; introducing super-resolution DOA estimation technology into the microphone array, and positioning based on the spatial smoothing MUSIC algorithm, the formula is R is the covariance matrix.

3. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 1, wherein: The motion posture fusion solution step includes designing a fuzzy logic based sensor credibility evaluation model, inputting the measurement residual e of the accelerometer / gyroscope a ,e g , output weight coefficient ω a ,ω g , the formula is Weighted suppression of faulty sensors; using quaternion differential equations to update attitude in real time, combined with magnetic field compensation algorithm to eliminate geomagnetic field interference, attitude matrix update frequency is 200Hz.

4. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 1, wherein: The real-time sound field reconstruction steps include building a personalized HRTF generation model, generating user-exclusive HRTF based on 3D facial scanning data and deep learning, supporting dynamic head model matching, and automatically switching HRTF subsets according to head movement speed; proposing a reinforcement learning-based sound field optimization algorithm, and adjusting the sound field parameters using the user's subjective immersion score R as the reward function.

5. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 2, wherein: In the super-resolution DOA estimation technology, the number of physical microphones is virtualized from 4 to 16 through the virtual array expansion algorithm. The formula is R 虚拟 =AR 物理 A H , A is the expansion matrix.

6. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 3, characterized in that: The triangle membership function formula designed in the fuzzy logic evaluation model is μ(e) = max(0,1-|e| / δ). When the residual exceeds the threshold δ, the sensor self-test is triggered.

7. The method for spatial audio positioning of headphones with integrated MEMS sensors according to claim 4, characterized in that: In the reinforcement learning sound field optimization algorithm, the state space S includes head posture, environmental noise, and audio content features, and the action space A includes the sound field parameter adjustment amount. The algorithm is trained through the experience replay mechanism.

8. A headset spatial audio positioning control system based on an integrated MEMS sensor implementing the method according to any one of claims 1 to 7, characterized in that: Includes the following modules: Ultra-precision sensing module: uses wafer-level packaged MEMS sensors, integrating a three-axis accelerometer, a three-axis gyroscope, and an air pressure sensor; the four-microphone array uses MEMS vector microphones to synchronously collect sound pressure and particle velocity, and single-microphone DOA estimation; Heterogeneous computing module: Equipped with an AI accelerator and FPGA co-processing unit. The AI ​​accelerator runs the CNN sound field prediction model, FPGA beamforming, and AWFS algorithms. Intelligent Rendering Module: Built-in reconstructed sound field driver, supports 128-channel digital beamforming, and uses GaN power amplifiers to drive the speaker array; Dynamic adaptation module: includes a real-time operating system and an adaptive control engine, dynamically switches the sound field mode according to the perception data, and supports user-defined parameter mapping strategies.

9. The earphone spatial audio positioning control system with integrated MEMS sensor according to claim 8, characterized in that: The ultra-precision sensing module includes a laser interferometer displacement sensor to monitor the fit between the earphones and the ear canal in real time, and triggers a wearing calibration reminder when the displacement exceeds the threshold; the flexible pressure sensor array detects the force distribution of the auricle and automatically adjusts the sound parameters to compensate for physical sound insulation differences.

10. The earphone spatial audio positioning control system with integrated MEMS sensor according to claim 8, characterized in that: The intelligent rendering module includes a holographic sound rendering unit that generates 3D sound holograms based on wave field synthesis technology and supports 5.1 / 7.1 channel decoding and Ambisonics format conversion. The haptic feedback driver generates haptic vibration signals based on the low-frequency components of the sound field.

Citation Information

Cited By

  • Ear clip type earphone adaptive control method for real-time music style identification

    CN121442237A