Full spherical surface HRTF acquisition method based on arc-shaped support
By combining a ring-shaped support with a rotating seat, and using pulse positioning and regularized frequency domain deconvolution techniques, the problems of low sampling efficiency and poor signal-to-noise ratio in existing HRTF acquisition schemes are solved. This enables the generation of high-precision, low-cost global surface HRTF datasets and supports personalized spatial audio rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-03
AI Technical Summary
Existing HRTF acquisition solutions suffer from low sampling efficiency, incomplete spherical coverage, poor measurement signal-to-noise ratio, high equipment cost, and inability to generate usable datasets in real time, resulting in inaccurate virtual sound image localization and degraded sound quality.
By employing a combined approach based on a ring-shaped support and a rotating seat, and utilizing pulse positioning and regularized frequency domain deconvolution techniques, along with spherical interpolation and symmetric completion, we can achieve rapid acquisition of global surface HRTF and generation of high signal-to-noise ratio datasets.
It enables low-cost, rapid acquisition of high-precision global surface HRTF datasets, supports personalized spatial audio rendering, is suitable for measurements in ordinary semi-anechoic chambers, and reduces equipment complexity and cost.
Smart Images

Figure CN121792918A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial audio and auditory perception technology, and in particular to a method, measurement system and computer-readable storage medium for rapid acquisition of global face head-related transfer function (HRTF) based on an arc-shaped support. Background Technology
[0002] Head-related transfer function (HRTF) is the cornerstone of spatial audio, and its accuracy directly determines the positioning accuracy and sound field breadth of virtual sound images. However, existing HRTF acquisition solutions generally adopt the methods of "fixed circular track + artificial head rotation" or "spherical robot multi-point sampling": the former can only sparsely sample points in the horizontal plane, with serious loss in the elevation direction; the latter, although it can achieve three-dimensional coverage, is complex, expensive, and has a long measurement cycle, making it difficult to quickly mass-produce for consumer products.
[0003] Due to limitations in measurement paths, most existing public datasets are distributed as "quasi-strips" with 48 horizontal points and an elevation angle of no more than ±30°, resulting in large gaps in the area above and below the head. When these missing data are directly used for dual-channel rendering, the sound image is prone to "jumps" or "collapses," significantly reducing the sound field expansion effect. At the same time, the phase error introduced by the interpolation completion process will exacerbate interaural crosstalk and reduce sound quality.
[0004] In addition, traditional frequency sweep deconvolution is susceptible to environmental noise interference at low frequencies and prone to overshoot at high frequencies, resulting in a low HRTF signal-to-noise ratio, which further affects the clarity and positioning stability of virtual surround sound.
[0005] Therefore, how to complete high signal-to-noise ratio HRTF acquisition of the global plane (elevation ±90°, azimuth 0–360°) in a conventional semi-anechoic chamber at low cost and in a short period of time, and generate a continuous, void-free dataset that can be directly used for sound field expansion and crosstalk cancellation, has become a key problem that current spatial audio technology urgently needs to solve. Summary of the Invention
[0006] 1. The problem to be solved
[0007] The main objective of this invention is to provide a method and system for rapid acquisition of global surface HRTF based on the coupling of a ring support and a rotating seat, in order to solve the technical problems of low sampling efficiency, incomplete spherical coverage, poor measurement signal-to-noise ratio, high equipment cost, and inability to generate SOFA files in real time in the prior art; at the same time, a high SNR HRTF is obtained in one step through "pulse positioning + regularized frequency domain deconvolution", and a hole-free, 1°×1° dense grid is formed by spherical interpolation and symmetrical completion, so as to realize personalized spatial sound rendering that can be used immediately after measurement.
[0008] 2. Technical Solution
[0009] To solve the above problems, the present invention adopts the following technical solution.
[0010] A global surface HRTF measurement method based on a ring-shaped support, characterized by comprising:
[0011] Step S10: In a free field or a space with good sound absorption, the speaker unit is installed on a ring bracket that can slide along an arc guide rail, while the subject or artificial head wears binaural microphones and sits in a horizontally rotatable chair.
[0012] Step S20: By coupling the two-dimensional angle of the sliding of the ring bracket with the rotation of the seat, the spherical 4π spherical degree sound source azimuth coverage is achieved.
[0013] Step S30: Control the speaker to play a combination signal of pulse excitation and exponential sweep frequency excitation in sequence, wherein the pulse is located before the sweep frequency and is used to mark the start time of the sweep frequency.
[0014] Step S40: Use binaural microphones to collect echo signals. First, use matched filtering to locate the arrival time t0 of the pulse, and then use a fixed offset Δ to obtain the start time t of the frequency sweep. start =t0+Δ;
[0015] Step S50, for t start The subsequent frequency sweep response is subjected to regularized frequency domain filtering and one-step deconvolution to directly obtain the high signal-to-noise ratio HRTF of the left and right ears;
[0016] Step S60: Perform spherical interpolation and left-right symmetry completion on the obtained azimuth-elevation two-dimensional HRTF data to generate a continuous, hole-free global surface HRTF dataset;
[0017] Step S70: Encapsulate the dataset and measurement metadata into a SOFA file according to the AES69 standard to complete the output.
[0018] The following sections elaborate on the implementation details and mathematical models.
[0019] The system hardware and coordinate system design are as follows:
[0020] The radius of the circular arc guide rail is approximately 1.2m, driven by a stepper motor, and the positioning accuracy is ≤0.02°.
[0021] The rotating seat rotates around the z-axis (top of head - bottom of feet) with a positioning accuracy of ≤0.02°.
[0022] As long as the distance from the speaker to both ears is kept basically constant and the angle error is ≤1°, there are no strict requirements on the absolute radius of the guide rail or the absolute height of the seat, so it can be adapted to spaces of different sizes.
[0023] Coordinate definition: Establish spherical coordinates with the midpoint of the line connecting the left and right ear canal entrances as the origin O. θ∈[-90°,+90°] is the elevation angle. This is the azimuth angle.
[0024] The transmitted signal design employs a "narrow pulse + exponential frequency sweep" waveform splicing method.
[0025] s(t)=δ(t)+sweep(tT gap )
[0026] δ(t) width is 1 sample; T gap ≥2ms is used to isolate early reflections; sweep(t) is a logarithmic frequency sweep with a duration T = 1–2s, a frequency band of 20Hz–20kHz, and an instantaneous frequency of:
[0027]
[0028] Complex analytic form:
[0029]
[0030] A is set to 75–80 dB SPL, with a peak-to-average power ratio of <6 dB to avoid speaker overload.
[0031] Subsample estimation of pulse arrival time
[0032] (1) Pre-filtering: downsampled to 12kHz, 100Hz high-pass filter, Hamming window segmentation.
[0033] (2) Matched filtering:
[0034] R xy [n] = ∑y[k]δ[kn]
[0035] Peak Index
[0036] (3) Parabolic interpolation:
[0037]
[0038] Pulse arrives (floating-point sample) Error ≤ ±0.3 samples.
[0039] (4) Frequency sweep start point: t start =t0+Δ,Δ=T gap ·F s
[0040] Perform regularized frequency domain deconvolution (RFIR) and truncate y(n) from t. start Starting with a length T·Fs, perform FFT simultaneously to obtain Y(k) and X(k). Then perform Tikhonov filtering:
[0041]
[0042] Pn(k) is the noise power during the silent segment, β∈[1×10]. -4 1×10 -2 Adjustable.
[0043] The time-domain HRIR was obtained as h[n] = IFFT{H(k)}, retaining the first 512 points, with independent processing for the left and right ears. The measured SNR was improved by ≥15dB compared to the traditional time-domain pulse method.
[0044] Spherical interpolation and symmetry completion
[0045] (1) Interpolation kernel: Corrected spherical spline or Slepian band-limited kernel
[0046]
[0047] (2) Dense grid with equal angles of 1°×1°: 181×360=65160 points.
[0048] (3) Symmetrical from left to right: The mirror image fills in the blind spot behind the seat.
[0049] (4) Extrapolation of poles: θ=±90° uses first-order Taylor extrapolation to ensure that there are no voids at the top / bottom of the sphere.
[0050] After calibration and error suppression, the laser ranging accuracy is maintained at |r-1.2m|≤5mm, and the amplitude introduced by the distance error is <0.04dB, which is negligible.
[0051] Infrared optical positioning monitors head pitch / yaw in real time, pausing and re-measuring when the pitch / yaw exceeds ±2°.
[0052] Speaker frequency response pre-calibration: at θ = 0°, S was measured ref (k), and all subsequent H(k) divided by S ref (k) eliminates the fluctuations of the unit itself.
[0053] 3. Beneficial effects
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] (1) This invention has a simple structure and low cost, requiring only a circular arc guide rail and a rotating seat. It eliminates the need for a six-degree-of-freedom robotic arm or a spherical truss, reducing overall cost and size, and allowing it to be deployed in stores, laboratories, and even homes.
[0056] (2) This invention can stably acquire HRTF data from all angles of the global surface, which can be used for a wider range of application scenarios;
[0057] (3) The HRTF measurement device of the present invention has higher accuracy and better universality than the prior art. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of an HRTF measurement method according to one embodiment of the present invention;
[0059] Figure 2 This is a schematic diagram of the structure of an HRTF measuring device according to one embodiment of the present invention. Detailed Implementation
[0060] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0061] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0062] refer to Figure 1 This document illustrates the flow chart of an embodiment of the HRTF measurement method according to this application. This embodiment includes an HRTF measurement device comprising a binaural microphone, a matched filter module, a regularized frequency domain impulse estimation module, and a loudspeaker. It is understood that this embodiment may include multiple binaural microphones, matched filter modules, regularized frequency domain impulse estimation modules, and loudspeakers. This embodiment does not limit the number of binaural microphones, matched filter modules, regularized frequency domain impulse estimation modules, and loudspeakers; those skilled in the art can set the number of these modules according to actual application conditions.
[0063] An HRTF measurement method according to this embodiment includes the following steps:
[0064] Step S100: Obtain different angles of the global surface by changing the angle of the bracket and rotating the seat; play pulse and frequency sweep signals through the speaker; and record the signals using the binaural microphones.
[0065] Subjects sit in a horizontally rotatable chair, wearing calibrated binaural microphones. Their heads are monitored in real-time by an optical positioning system to ensure the interaural axis is substantially aligned with the chair's rotation axis. The system drives a ring-shaped support that interacts with the chair in two dimensions, forming a spatial grid covering the entire sphere. Each point sequentially emits a combination of "Dirac pulse + logarithmic sweep" signals. The pulse is used for subsample time alignment, and the sweep is used for high signal-to-noise ratio deconvolution. The binaural microphones record synchronously, laser ranging maintains a constant sound source radius, and the recording duration includes guard intervals and margins. The entire process completes the global original signal database within tens of minutes.
[0066] Step S200: Estimate and solve the HRTF using matched filtering and regularized frequency domain filtering;
[0067] The recorded data stream is transmitted to the matched filtering module in real time: using the transmitted pulse as a template, the sub-sample level start time is obtained by cross-correlation peak parabolic interpolation; the frequency band after the start time is truncated, and the complex HRTF is obtained by one-step deconvolution after regularization frequency domain filtering, retaining a finite length HRIR, and the left and right ears are processed independently. The measured frequency band SNR is significantly improved compared with the traditional time domain pulse method.
[0068] Step S300: Global HRTF data is encapsulated into a SOFA file according to the AES69 standard.
[0069] By integrating the acquired HRTF dataset, it is packaged into the SOFA format, a universal format for spatial audio, to facilitate a wider range of applications. In a preferred embodiment, the entire system is deployed in a space with good sound absorption conditions. An arc-shaped guide rail and a rotating seat form a two-dimensional linkage mechanism, allowing the speaker to slide along the arc. The subject sits in the rotating seat and wears binaural microphones. The system monitors head posture in real time using optical positioning to ensure that the ear axis is substantially aligned with the rotation axis; laser ranging maintains a constant sound source radius in real time. Before measurement begins, the system automatically returns to zero and generates a global surface grid table.
[0070] During the measurement, the speaker sequentially plays a combination of "pulse and sweep frequency" signals. The pulse segment is used for sub-sample-level time alignment, and the sweep frequency segment is used for high signal-to-noise ratio deconvolution. Upon reaching each grid point, the system first completes angle positioning and distance calibration before triggering playback and recording. If the head offset exceeds a threshold, the system automatically pauses and prompts for repositioning. After all grid points are recorded, the data stream is sent to the matched filtering module. Cross-correlation peak interpolation accurately locates the pulse start point, and then the sweep frequency segment is truncated for regularized frequency domain filtering, resulting in a global surface high signal-to-noise ratio (HRTF) signal-to-noise ratio signal-to-noise ratio signal-to-noise ratio (HRTF) in one step.
[0071] The global surface HRTF dataset is processed through spherical interpolation and symmetric completion to form a continuous, hole-free, dense grid. It is then encapsulated into a SOFA file according to the AES69 standard, supporting minimum phase decomposition and digital signatures. The system also supports multi-speaker cascading, environmental compensation, and real-time quality monitoring: multiple speakers can be arranged on an arc to simultaneously play different sweep frequency sequences, further shortening measurement time; room impulse responses can be collected on-site, generating a compensation matrix and correcting the HRTF in real time; SNR checks and overflow detection can be performed on the impulse segment and sweep frequency segment respectively, automatically re-measuring unqualified points.
[0072] The entire measurement process requires no complex robotic arm, featuring a simple structure, small size, and low cost, making it suitable for laboratory, store, and even home scenarios. After measurement, users can directly import the SOFA file into the rendering engine to achieve a personalized spatial audio experience. The specific implementation methods of each functional module of the HRTF measurement device described in this embodiment can be referred to the relevant descriptions in the foregoing method embodiments, and will not be repeated here. The HRTF measurement device described above is mainly from the perspective of functional modules. To further enrich the implementation, this application also provides an HRTF measurement device, supplementing the technical solution from a hardware perspective. This device includes a memory for storing computer programs; and a processor for executing the computer program to implement the steps of any of the above-described HRTF measurement methods.
[0073] The processor may include one or more processing cores, such as a 4-core or 8-core processor. The processor can be implemented in hardware forms such as DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array), and may also consist of a main processor and a coprocessor: the main processor handles data processing in the wake-up state (i.e., CPU), while the coprocessor handles low-power data processing in the standby state. In some embodiments, the processor may also integrate a GPU (Graphics Processing Unit) for rendering display content; or integrate an AI processor for performing machine learning-related calculations.
[0074] The memory may include one or more computer-readable storage media, and may be non-transitory memory. The memory may also include high-speed random access memory (RAM) and non-volatile memory, such as disk storage devices, flash memory devices, etc. In this embodiment, the memory is at least used to store a computer program, which, after being loaded and executed by the processor, implements the aforementioned portable HRTF measurement method. The memory may also store an operating system (such as Windows, Unix, Linux) and related data (such as test result data).
[0075] The specific implementation of each functional module of the HRTF measurement device described in this embodiment of the invention can be referred to the foregoing method embodiments, and will not be repeated here. As can be seen from the above, the HRTF measurement device provided in this embodiment of the invention can improve the accuracy of test results; and through an adjustable speaker, it provides the optimal auditory experience for different users.
[0076] The embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts can be referred to interchangeably. For the apparatus embodiments, since they correspond to the method embodiments, the descriptions are relatively concise; relevant parts can be found in the method section.
[0077] Those skilled in the art will further recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example in terms of functionality. Whether a function is implemented in hardware or software depends on the application of the technical solution and design constraints. Those skilled in the art can use different methods to implement specific applications, but this does not exceed the scope of the present invention.
[0078] The HRTF measurement method provided in this application has been described in detail above. Specific examples are used to illustrate the principles and implementation methods of this invention. The above embodiments are only for the purpose of helping to understand the method and its core ideas. Those skilled in the art can make various improvements and modifications to this application without departing from the principles of this invention, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0079] The examples described herein are merely illustrative of preferred embodiments and are not intended to limit the scope or concept of the invention. Various modifications and improvements made to the technical solutions of this invention by those skilled in the art without departing from the inventive concept should fall within the protection scope of this invention.
Claims
1. A method for obtaining global surface HRTF based on an arc-shaped support, characterized in that, Simultaneous execution within a semi-anechoic chamber: The speaker was mounted on an arc-shaped bracket that could slide along the arc, and the subject or artificial head wearing binaural microphones sat in a horizontally rotatable chair. By coupling the two degrees of freedom of the bracket sliding and the seat rotation, the azimuth and elevation angles of the speaker relative to the two ears can traverse the global surface sampling points; Pulse signals and exponentially swept frequency signals were played sequentially, and the binaural responses were collected. The response signal is deconvolved using regularized frequency domain impulse estimation to obtain the high signal-to-noise ratio (HRTF) for each sampling point; By using spherical interpolation to complete the unmeasured regions, a continuous and hole-free global surface HRTF dataset is finally generated.
2. The method according to claim 1, characterized in that, in, The sliding of the arc-shaped bracket and the rotation of the seat are controlled by the same trigger signal phase lock, and the angular positioning error is ≤0.5°.
3. The method according to claim 1, characterized in that, in, The pulse signal is a Dirac pulse or a maximum length sequence, and the frequency sweep signal is an exponential frequency sweep, used to improve the signal-to-noise ratio of deconvolution.
4. The method according to claim 1, characterized in that, in, The regularized frequency domain pulse estimation adopts Tikhonov regularization, and the regularization coefficient λ is adaptively set according to the mean of |U(k)|2 to suppress frequency overshoot.
5. The method according to claim 1, characterized in that, in, Spherical interpolation uses bilinear or cubic spline interpolation, and the spatial resolution of the dataset after interpolation is ≤1°.
6. The method according to any one of claims 1 to 5, characterized in that, Further includes: The left and right ear symmetry of the HRTF dataset is checked. If the difference is lower than a preset threshold, the dataset is compressed and stored by mirror copying. The final dataset is packaged in SOFA format for use by virtual reality, spatial audio rendering, or hearing aids.
7. The method according to any one of claims 1 to 5, characterized in that, Further steps include calibrating the mechanical transmission error of the arc-shaped bracket and seat before measurement, and correcting the theoretical position of the speaker in real time.
8. The method according to any one of claims 1 to 5, characterized in that, Further steps include performing band-limited filtering and DC removal on the response signal before deconvolution to suppress ambient noise and system drift.
9. An HRTF measurement system, characterized in that, It includes an arc-shaped support, a rotatable seat, a speaker, binaural microphones, and a processor, the processor being configured to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The system contains a program that, when executed, implements the global surface HRTF acquisition method according to any one of claims 1 to 8.