A personalized head-related transfer function upsampling method and system
Through deep learning methods combined with extrapolation and general feature extraction network, the minimum phase reconstruction technology is used to achieve efficient upsampling of HRTF under extremely sparse measurement conditions, solving the problem of large HRTF measurement error in the prior art, improving the accuracy and stability of personalized HRTF, and improving the auditory positioning effect of spatial audio playback.
Patent Information
- Application Number
- CN202510365758.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The prior art requires HRTF measurements in many directions, with large errors and unstable under extremely sparse conditions, making it difficult for personalized HRTFs to provide accurate auditory positioning perception in spatial audio playback.
Deep learning method is adopted to achieve upsampling of sparse measurement HRTF to target spatial resolution HRTF through extrapolated feature extraction network, general feature extraction network and prediction network, combined with minimum phase reconstruction, head-related transmission function amplitude spectrum and binaural time difference.
Full-space HRTF is obtained with lower error upsampling under extremely sparse conditions, which improves the accuracy and stability of personalized HRTF and improves the auditory positioning effect of spatial audio playback.
Smart Images

Figure CN120111429B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of spatial audio technology, and specifically relates to a personalized head-related transfer function upsampling method and system. Background Art
[0002] Head Related Transfer Function (HRTF) plays an important role in spatial audio technology. HRTF characterizes the process of sound waves passing from the sound source through physiological structures such as the head, auricle and torso to the human eardrum. The differences in physiological parameters of different individuals lead to different HRTFs. Using personalized HRTFs in spatial audio playback can effectively improve the playback effect. However, personalized HRTFs usually need to be measured using special equipment in an anechoic environment, which is very time-consuming, making personalized HRTFs difficult to truly apply. At present, some methods attempt to use HRTF amplitude spectra measured in a small number of directions to upsample or predict personalized HRTFs for the entire space. However, existing upsampling methods require HRTF measurements in more directions, and the errors are large and unstable under extremely sparse conditions, making it difficult for upsampled HRTFs to provide users with accurate and personalized auditory positioning perception, thereby limiting their application in spatial audio playback. Summary of the Invention
[0003] The purpose of this application is to overcome the defects of the prior art that requires HRTF measurement in multiple directions, has large errors and is unstable under extremely sparse conditions.
[0004] To achieve the above objectives, this application proposes a personalized head-related transfer function upsampling method, including:
[0005] The sparsely measured head-related transfer function amplitude spectrum and interaural time difference of the subject are input into the trained upsampling model, and the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution are output;
[0006] The head-related transfer function amplitude spectrum at the minimum phase is reconstructed using the minimum phase, and the upsampled head-related transfer function is obtained by splicing it with the interaural time difference of the target spatial resolution;
[0007] The upsampling model includes an extrapolated feature extraction network, a general feature extraction network and a prediction network; wherein,
[0008] The extrapolation feature extraction network is used to obtain corresponding extrapolation features from the sparsely measured head-related transfer function;
[0009] The universal feature extraction network is used to obtain corresponding universal features from the universal head-related transfer function;
[0010] The prediction network is used to input the fused extrapolated features and universal features, and output the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution.
[0011] As an improvement to the above method, the method for obtaining the head-related transfer function amplitude spectrum and interaural time difference includes:
[0012] The envelope is extracted by low-frequency filtering the original head-related transfer function, and the interaural time difference (ITD) (θ) is obtained by calculating the interaural cross-correlation peak delay:
[0013]
[0014] Where θ represents the sampling azimuth; τ represents the target interaural time difference; p L represents the left ear head-related transfer function envelope after low-pass filtering; p R represents the envelope of the right ear head-related transfer function after low-pass filtering; T represents the sampling time of the head-related transfer function;
[0015] The original head-related transfer function is subjected to fast Fourier transform to obtain the head-related transfer function amplitude spectrum of N sampling frequency points.
[0016] As an improvement to the above method, the extrapolation feature extraction network includes:
[0017] Two extrapolation encoders take the sparsely measured head-related transfer function amplitude spectrum and interaural time difference as input, respectively, and output corresponding extrapolated features; the extrapolation encoders include a convolutional layer, a batch normalization layer, and an activation layer; the extrapolated features characterize the relationship between the low spatial resolution head-related transfer function distribution and the high spatial resolution head-related transfer function distribution; the number of input layers of the extrapolation encoders is the density of the head-related transfer function sampling grid.
[0018] As an improvement to the above method, the general feature extraction network includes:
[0019] Two universal encoders take a universal head-related transfer function amplitude spectrum and interaural time difference as input, respectively, and output corresponding universal features; the universal encoder includes a convolutional layer, a batch normalization layer, and an activation layer; the universal features represent the distribution of high-resolution head-related transfer functions common to different subjects; the number of universal encoder input layers is the density of the head-related transfer function target sampling grid.
[0020] As an improvement to the above method, the prediction network includes a convolutional layer, a batch normalization layer and an activation layer.
[0021] As an improvement to the above method, the method of reconstructing the head-related transfer function amplitude spectrum at the minimum phase by using the minimum phase includes:
[0022]
[0023] Where θ is the azimuth angle; φ is the elevation angle; f is the frequency; ψ min (θ, φ, f) represents the minimum phase obtained; Head-related transfer function magnitude spectrum representing the target spatial resolution;
[0024] The minimum phase is used to restore the head-related transfer function amplitude spectrum of the target spatial resolution to the head-related transfer function amplitude spectrum under the minimum phase.
[0025] The present application also provides a personalized head-related transfer function upsampling system, which is implemented based on the above method, and the system includes:
[0026] An upsampling module is used to input the sparsely measured head-related transfer function amplitude spectrum and interaural time difference of the subject into the trained upsampling model, and output the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution;
[0027] The reconstruction module is used to reconstruct the head-related transfer function amplitude spectrum under the minimum phase using the minimum phase, and to obtain the upsampled head-related transfer function by splicing it with the interaural time difference of the target spatial resolution.
[0028] Compared with the prior art, the advantages of this application are:
[0029] Compared with existing methods, the technical solution of the present application can obtain full-space HRTF with lower error by upsampling under extremely sparse spatial measurement HRTF. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Shown is a flowchart of a personalized head-related transfer function upsampling method. DETAILED DESCRIPTION
[0031] The technical solution of this application is described in detail below with reference to the accompanying drawings.
[0032] Example 1
[0033] like Figure 1 As shown, the present application provides a personalized head-related transfer function upsampling method, including:
[0034] Step S101 : obtaining sparsely measured HRTF amplitude spectrum and interaural time difference (ITD) using head-related impulse responses measured in a small number of directions.
[0035] Specifically, the envelope is extracted after low-frequency filtering of the original HRIR with a cutoff frequency of 3 kHz, and the interaural time difference (ITD) is obtained by calculating the peak delay of the binaural cross-correlation, as shown in the following formula.
[0036]
[0037] Where θ represents the sampling azimuth, τ represents the target ITD, and p L represents the left ear HRIR envelope after low-pass filtering, p R represents the right ear HRIR envelope after low-pass filtering; T represents the sampling time of the head-related transfer function.
[0038] Then, the HRTF amplitude spectrum of N sampling frequency points is obtained by performing fast Fourier transform on the original HRIR.
[0039] Step S102: Obtain corresponding extrapolated features from the sparsely measured HRTF using an extrapolated feature extraction network.
[0040] Specifically, extrapolated feature encoders were designed for binaural ITD and binaural HRTF amplitude spectra, respectively. Both networks consist of convolutional layers, batch normalization layers, and activation functions. The extrapolated features characterize the relationship between the low- and high-resolution HRTF distributions. The number of channels in the input layer of the extrapolated feature extraction network is equal to the density of the HRTF sampling grid.
[0041] Step S103: extracting universal features from the universal HRTF using a universal feature extraction network.
[0042] Specifically, for the dataset used to train the proposed method, the mean ITD and HRTF amplitude spectrum from the training set are extracted as the universal HRTF. Using the same encoder as in step S102, the universal ITD and amplitude spectrum are encoded into corresponding universal features, simply by changing the number of input layer channels. The number of input layer channels of the universal feature extraction network is equal to the density of the HRTF target sampling grid. The universal HRTF can be the average of multiple subjects' HRTFs at high spatial resolution or the HRTF of an artificial head.
[0043] Step S104 , after fusing the extrapolated features with the general features, the fused features are mapped to the HRTF amplitude spectrum and ITD under the target sampling grid using a prediction neural network.
[0044] Specifically, assume that the extrapolated features are given by The general features are represented by Represented as, then the fused hidden layer features are expressed as Then the prediction network will Mapped to the target HRTF amplitude spectrum or target ITD. The prediction neural network consists of convolutional layers, batch normalization layers, and activation layers.
[0045] Step S105 , using the minimum phase to reconstruct the HRIR at the minimum phase and splicing the ITD to reconstruct the upsampled HRIR.
[0046] Specifically, it is assumed that the HRTF after upsampling of the prediction network output is given by It means that the ITD after the prediction network output is upsampled is given by Using the following formula, we can get Obtain the minimum phase of HRTF and reconstruct the HRIR at the minimum phase, that is,
[0047]
[0048] Among them, θ represents the azimuth angle, φ represents the elevation angle, f represents the frequency, and ψ min That is the minimum phase obtained, using this minimum phase to Restore to Afterwards, Splicing The phase information can be completely reconstructed, and finally the upsampled binaural HRIR is obtained.
[0049] In step S106, given a small amount of measured HRTFs of a new subject, the proposed method can be used to upsample and obtain a personalized HRTF of the target spatial resolution.
[0050] In this application, the extrapolated feature extraction network, the general feature extraction network and the prediction network are trained in a joint training manner, that is, data is input into the extrapolated feature extraction network and the general feature extraction network at the same time, and their respective outputs are merged and input into the prediction network. Then, the output of the prediction network is evaluated for loss, and according to the evaluation results, the gradients of the extrapolated feature extraction network, the general feature extraction network and the prediction network are simultaneously returned.
[0051] Example 2
[0052] The present application also provides a personalized head-related transfer function upsampling system, which is implemented based on the above method, and the system includes:
[0053] An upsampling module is used to input the sparsely measured head-related transfer function amplitude spectrum and interaural time difference of the subject into the trained upsampling model, and output the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution;
[0054] The reconstruction module is used to reconstruct the head-related transfer function amplitude spectrum under the minimum phase using the minimum phase, and to obtain the upsampled head-related transfer function by splicing it with the interaural time difference of the target spatial resolution.
[0055] The present invention extracts personalized extrapolated features that characterize the subject's physiological structure from sparsely measured HRTFs through a deep learning method, and compensates for the information loss in the extrapolated features under extremely sparse conditions through universal features, thereby establishing a spatial distribution relationship between the sparsely measured HRTFs and the personalized HRTFs under the target spatial resolution, so that the full-space personalized HRTF can be predicted given the measured HRTFs of the subject in a few directions.
[0056] The present application may also provide a computer device comprising: at least one processor, memory, at least one network interface, and a user interface. The various components in the device are coupled together via a bus system. It will be understood that the bus system is used to enable communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0057] The user interface may include a display, a keyboard, or a pointing device, such as a mouse, a trackball, a touchpad, or a touch screen.
[0058] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0059] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset or an extension thereof: an operating system and applications.
[0060] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and handle hardware-based tasks. Application programs include various application programs, such as media players and browsers, which are used to implement various application services. The program that implements the method of the embodiment of the present disclosure can be included in the application program.
[0061] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in the application program, to:
[0062] Perform the steps of the above method.
[0063] The above method can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The above-disclosed methods, steps, and logic block diagrams can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the above-disclosed method can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0064] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units or combinations thereof for performing the functions described herein.
[0065] For software implementation, the technology of the present application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0066] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0067] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.
Claims
1. A personalized head-related transfer function upsampling method, comprising: The sparsely measured head-related transfer function amplitude spectrum and interaural time difference of the subject are input into the trained upsampling model, and the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution are output; The head-related transfer function amplitude spectrum at the minimum phase is reconstructed using the minimum phase, and the upsampled head-related transfer function is obtained by splicing it with the interaural time difference of the target spatial resolution; The upsampling model includes an extrapolated feature extraction network, a general feature extraction network and a prediction network; wherein, The extrapolation feature extraction network is used to obtain corresponding extrapolation features from the sparsely measured head-related transfer function; The universal feature extraction network is used to obtain corresponding universal features from the universal head-related transfer function; The prediction network is used to input the fused extrapolated features and universal features, and output the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution.
2. The personalized head-related transfer function upsampling method according to claim 1, characterized in that: The method for obtaining the head-related transfer function amplitude spectrum and interaural time difference includes: The envelope is extracted by low-frequency filtering the original head-related transfer function, and the interaural time difference (ITD) (θ) is obtained by calculating the interaural cross-correlation peak delay: Where θ represents the sampling azimuth; τ represents the target interaural time difference; p L represents the left ear head-related transfer function envelope after low-pass filtering; p R represents the envelope of the right ear head-related transfer function after low-pass filtering; T represents the sampling time of the head-related transfer function; IACC(θ,τ) represents the intermediate variable; The original head-related transfer function is subjected to fast Fourier transform to obtain the head-related transfer function amplitude spectrum of N sampling frequency points.
3. The personalized head-related transfer function upsampling method according to claim 1, characterized in that: The extrapolation feature extraction network includes: Two extrapolation encoders take the sparsely measured head-related transfer function amplitude spectrum and interaural time difference as input, respectively, and output corresponding extrapolated features; the extrapolation encoders include a convolutional layer, a batch normalization layer, and an activation layer; the extrapolated features characterize the relationship between the low spatial resolution head-related transfer function distribution and the high spatial resolution head-related transfer function distribution; the number of input layers of the extrapolation encoders is the density of the head-related transfer function sampling grid.
4. The personalized head-related transfer function upsampling method according to claim 1, characterized in that: The general feature extraction network includes: Two universal encoders take a universally measured HRT amplitude spectrum and interaural time difference as input, respectively, and output corresponding universal features; the universal encoder comprises a convolutional layer, a batch normalization layer, and an activation layer; the universal features characterize the distribution of high-resolution HRT functions common to different subjects; and the number of input layers of the universal encoder is the density of the HRT target sampling grid.
5. The personalized head-related transfer function upsampling method according to claim 1, characterized in that: The prediction network includes a convolutional layer, a batch normalization layer and an activation layer.
6. The personalized head-related transfer function upsampling method according to claim 1, characterized in that: The method of reconstructing the head-related transfer function amplitude spectrum at the minimum phase by using the minimum phase includes: Where θ is the azimuth angle; φ is the elevation angle; f is the frequency; ψ min (θ, φ, f) represents the minimum phase obtained; Head-related transfer function magnitude spectrum representing the target spatial resolution; The minimum phase is used to restore the head-related transfer function amplitude spectrum of the target spatial resolution to the head-related transfer function amplitude spectrum under the minimum phase.
7. A personalized head-related transfer function upsampling system, implemented based on the method according to any one of claims 1 to 6, characterized in that: The system comprises: an upsampling module, configured to input the sparsely measured head-related transfer function amplitude spectrum and interaural time difference of the subject into a trained upsampling model and output the head-related transfer function amplitude spectrum and interaural time difference of the target spatial resolution; and The reconstruction module is used to reconstruct the head-related transfer function amplitude spectrum under the minimum phase using the minimum phase, and to obtain the upsampled head-related transfer function by splicing it with the interaural time difference of the target spatial resolution.
Citation Information
Patent Citations
Personalized head-related transfer function prediction method and device based on sparse measurement
CN116506795A
HRTF data fusion and sparse prediction method based on binaural factors
CN119052713A