A head-related transfer function data set fusion method and system
Through deep neural network training, the measurement features of the data set are digested, the binaural time difference and auditory positioning characteristics are extracted, and the loss function is constructed, which solves the HRTF fusion problem of different data sets, and the consistency of the positioning perception results and the expansion of the data set are achieved.
Patent Information
- Application Number
- CN202510355674.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Due to the differences in measurement equipment and methods, HRTFs in different data sets have obvious differences in data distribution, which is difficult to directly integrate. It is difficult for existing methods to effectively eliminate measurement attributes and keep the positioning perception results unchanged.
Through deep neural network training, the head-related impulse response that eliminates the measurement features of the data set is generated, the binaural time difference, sound level difference and auditory positioning characteristics are extracted, and the loss function is constructed for training to achieve the fusion of multiple data sets.
The fusion of multiple data sets is achieved, the consistency of positioning and perceived results is maintained, and a broader data base is provided for HRTF modeling.
Smart Images

Figure CN120217298B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of spatial audio technology, and specifically relates to a head-related transfer function dataset fusion method and system. Background Art
[0002] The head-related transfer function (HRTF) plays a crucial role in spatial audio technology. HRTF characterizes the process by which sound waves travel from a sound source through physiological structures such as the head, auricle, and torso to the eardrum. Differences in individual physiological parameters lead to distinct HRTFs. Using personalized HRTFs in spatial audio playback can effectively improve playback quality. However, personalized HRTFs typically require measurement using specialized equipment in an anechoic environment, which is time-consuming and hinders their practical application. Currently, synthesizing HRTF amplitude spectra from physiological parameters is widely used to derive personalized HRTFs. However, due to the complex HRTF generation process and limited dataset size, existing personalized HRTF prediction methods struggle to effectively characterize the physical processes between sound source propagation from different directions and various physiological parameters of the human body. Therefore, some methods attempt to fuse multiple datasets to expand the data size available for model training. However, due to differences in measurement equipment, measurement methods, and post-processing methods, the data distributions of different datasets vary significantly, making direct fusion difficult. Therefore, how to eliminate the measurement attributes carried by HRTF in each dataset and keep the positioning perception results unchanged is a key issue that needs to be solved in fusing multiple datasets. Summary of the Invention
[0003] To overcome the above-mentioned drawbacks, this application proposes a head-related transfer function dataset fusion method, including:
[0004] Input the sampled data of different subjects in multiple datasets into the trained deep neural network to generate a head-related impulse response that eliminates the measurement features of the dataset;
[0005] The training process of the deep neural network includes:
[0006] The head-related impulse response in the original dataset is used to extract interaural time difference, interaural level difference and auditory localization features;
[0007] Use deep neural networks to generate head-related impulse responses that eliminate the measurement features of the dataset, and extract binaural level difference and auditory localization features;
[0008] The loss function is constructed using the binaural sound level difference between the original head-related impulse response and the head-related impulse response output by the deep neural network, the difference in auditory localization features, and the dataset classification accuracy between the network output head-related transfer functions, and the feedback is used to train the deep neural network.
[0009] As an improvement to the above method, the method of extracting interaural time difference, interaural level difference and auditory localization features by using the head-related impulse response in the original data set includes:
[0010] The HRIs of different subjects in different data sets were resampled with a unified sampling rate and length. The resampled HRIs were filtered using a low-pass filter and the envelope was extracted. The binaural cross-correlation method was used to calculate the peak delay and obtain the interaural time difference (ITD) (θ):
[0011]
[0012] Among them, θ represents the sampling direction; τ represents the target interaural time difference; p L represents the left ear head-related impulse response envelope after low-pass filtering; p R represents the right ear head-related impulse response envelope after low-pass filtering; T represents the head-related impulse response sampling time; t represents time;
[0013] By calculating the energy difference of binaural head-related impulse responses, the binaural level difference ILD(θ) is obtained:
[0014]
[0015] Among them, HRIR L Represents the left ear head-related impulse response after resampling; HRIR R represents the head-related impulse response of the right ear after resampling;
[0016] The resampled head-related impulse response is subjected to fast Fourier transform to obtain the head-related transfer function amplitude spectrum at N sampling frequency points; the head-related transfer function amplitude spectrum is auditory filtered using gammatone auditory filters with M equivalent rectangular bandwidths to obtain auditory localization features related to auditory localization perception.
[0017] As an improvement to the above method, the input of the deep neural network is a two-channel random noise sampling, and the output is a target two-channel head-related impulse response;
[0018] The deep neural network includes three intermediate layers and a fully connected output layer;
[0019] The intermediate layer includes a fully connected layer, batch normalization and an activation function.
[0020] As an improvement to the above method, the loss function of the deep neural network is:
[0021] L=L ild +αL loc +βL cls
[0022] Among them, α and β represent the proportion of perceptual positioning feature loss and classification loss in the total loss respectively; L ild Represents binaural level difference loss:
[0023]
[0024] Where ILD represents the binaural level difference of the original input, Represents the binaural level difference output by the deep neural network;
[0025] L loc Represents the perceptual positioning feature loss:
[0026]
[0027] Where F represents the gammatone auditory filter; HRTF represents the head-related transfer function of the original input; represents the head-related transfer function output by the deep neural network; M represents the number of equivalent rectangular bandwidths;
[0028] L cls Denotes the classification loss:
[0029]
[0030] Among them, c i and c j represents the mean of the head-related transfer function under different data sets; i and j represent the indexes of different data sets respectively; N represents the total number of data sets; ||·|| 2 represents the 2-norm.
[0031] As an improvement to the above method, the classification accuracy is characterized by calculating the mean of the head-related transfer function amplitude spectrum under different data sets as the inter-class distance between the class centers.
[0032] The present application also provides a head-related transfer function dataset fusion system, which is implemented based on the above method, and includes:
[0033] The dataset fusion module is used to input the sampled data of different subjects in multiple datasets into the trained deep neural network to generate a head-related impulse response that eliminates the measurement features of the dataset;
[0034] Network training module, used to train deep neural networks.
[0035] Compared with the existing technology, the advantages of this application are:
[0036] 1. This application proposes a method and system for fusing head-related transfer function (HRTF) datasets. Given HRIRs (head-related impulse responses) from different subjects across multiple datasets, the proposed network can be trained to generate HRIRs that eliminate the measurement features of the datasets. This HRIR retains the localization-sensing features of the original HRFs while making it difficult to discern their dataset sources, thereby enabling the fusion of multiple datasets.
[0037] 2. Due to differences in measurement equipment, measurement methods, post-processing methods, etc. used in different data sets, the data distribution differences between different data sets are very obvious, making it difficult to directly fuse them. The method provided in this application can effectively eliminate the measurement attributes carried by HRTF in different data sets and keep the positioning perception results unchanged, thereby enabling different data sets to be fused, providing a broader data foundation for data-driven HRTF modeling methods. Previously, there was no relevant technology to solve this problem. The method provided in this application is the first method to eliminate the measurement attributes of HRTF in different data sets while keeping the positioning perception results unchanged. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Shown is a flowchart of the head-related transfer function dataset fusion method. DETAILED DESCRIPTION
[0039] The technical solution of this application is described in detail below with reference to the accompanying drawings.
[0040] The present invention provides a method and system for fusing head-related transfer function datasets. By extracting perceptual positioning features from the original HRIR dataset and constructing a loss function feedback training deep neural network based on this, the HRTF generated by the network has the same perceptual positioning results as the original head-related transfer function but does not contain the dataset measurement features, thereby realizing the fusion of multiple datasets.
[0041] Example 1
[0042] like Figure 1 FIG2 is a flow chart of a head-related transfer function dataset fusion method provided by an embodiment of the present application, the method comprising:
[0043] Step S101 : extracting interaural time difference, interaural level difference and auditory localization features using the head-related impulse response (HRIR) in the original data set.
[0044] Specifically, the HRIR of different subjects in different data sets are resampled with a unified sampling rate and sampling length. Then, a Butterworth low-pass filter with a cutoff frequency of 3 kHz is designed to filter the resampled HRIR and extract the envelope. Finally, the binaural cross-correlation method is used to calculate the peak delay to obtain the interaural time difference (ITD), as shown in the following formula.
[0045]
[0046] Among them, θ represents the sampling direction, τ represents the target ITD, and p L represents the left ear HRIR envelope after low-pass filtering, p R represents the right ear HRIR envelope after low-pass filtering, and T represents the HRIR sampling time.
[0047] The binaural level difference (ILD) information can then be extracted by calculating the energy difference between the binaural HRIRs, as shown in the following formula.
[0048]
[0049] Among them, HRIR L / R Represents the resampled left / right ear HRIR.
[0050] By performing a Fast Fourier Transform on the resampled HRIR, we can obtain the HRTF amplitude spectrum at N sampling frequencies. After designing M gammatone auditory filters with equivalent rectangular bandwidths (ERBs), we can use them to perform auditory filtering on the HRTF amplitude spectrum, thereby obtaining auditory localization features related to auditory localization perception.
[0051] Step S102: Using a deep neural network, generate HRIR that eliminates the measurement features of the data set, and extract binaural level difference and auditory localization features.
[0052] Specifically, because eliminating the measurement attributes inherent in the original HRTF data is a no-reference problem, a neural network model is constructed with two-channel random noise samples as input and the target two-channel HRIR as output. This model consists of three modules and a fully connected output layer. Each module includes a fully connected layer, batch normalization, and an activation function. By performing the same feature extraction process as step S101 on the binaural HRIR output from the network, the corresponding ITD, ILD, and auditory localization features are obtained.
[0053] Step S103 , a loss function is constructed using the binaural level difference between the original HRIR and the network output HRIR, the difference in auditory localization features, and the dataset classification accuracy between the network output head-related transfer functions (HRTFs), and the feedback is used to train the deep neural network.
[0054] Specifically, because ITD is less affected by factors such as measurement equipment and methods across different datasets, the ITD in the predicted HRIR can be aligned using the ITD information extracted from the original HRIR. Then, based on the differences between the ILD and auditory localization features before and after processing, the following loss function can be constructed.
[0055]
[0056] in, represents the gammatone auditory filter, ILD represents the ILD of the original input, Represents the ILD, HRTF and denote the original and model output HRTFs respectively.
[0057] A key indicator of the degree of resolution of the original HRIR measurement attributes is the difference in data distribution across different datasets. Specifically, the classification accuracy of the HRTF spectra output by the model across different datasets is determined. By calculating the mean HRTF value across each dataset, the distance between data distributions can be calculated, as shown in the following formula.
[0058]
[0059] Among them, c i / j represents the mean HRTF value under different data sets, i and j represent the index of different data sets, and N represents the total number of data sets; ||·|| 2 represents the 2-norm.
[0060] By combining the perceptual localization loss and classification loss designed above, the total loss function for model training can be constructed, as shown in the following formula.
[0061] L=L ild +αL loc +βL cls
[0062] Here, α and β represent the proportion of perceptual positioning feature loss and classification loss in the total loss, respectively, and their typical values are both 1. The total loss function L can be used to perform feedback training on the network proposed in step S102.
[0063] The classification accuracy is characterized by calculating the mean of the head-related transfer function amplitude spectrum under different data sets as the inter-class distance between the class centers.
[0064] In step S104, given the HRIRs of different subjects in multiple data sets, the proposed network can be trained to generate HRIRs that eliminate the measurement features of the data sets, so that the HRIRs have the positioning perception features of the original head-related transfer function while being difficult to identify the source of the data sets, thereby enabling the fusion of multiple data sets.
[0065] The present invention extracts perceptual positioning features from the original HRIR dataset and constructs a loss function feedback training deep neural network based on it, so that the HRTF generated by the network has the same perceptual positioning results as the original head-related transfer function but does not contain the measurement features of the dataset, thereby realizing the fusion of multiple datasets.
[0066] Example 2
[0067] The present application also provides a head-related transfer function dataset fusion system, which is implemented based on the above method, and includes:
[0068] The dataset fusion module is used to input the sampled data of different subjects in multiple datasets into the trained deep neural network to generate a head-related impulse response that eliminates the measurement features of the dataset;
[0069] Network training module, used to train deep neural networks.
[0070] The present application may also provide a computer device comprising: at least one processor, memory, at least one network interface, and a user interface. The various components in the device are coupled together via a bus system. It will be understood that the bus system is used to enable communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0071] The user interface may include a display, a keyboard, or a pointing device, such as a mouse, a trackball, a touchpad, or a touch screen.
[0072] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0073] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset or an extension thereof: an operating system and applications.
[0074] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and handle hardware-based tasks. Application programs include various application programs, such as media players and browsers, which are used to implement various application services. The program that implements the method of the embodiment of the present disclosure can be included in the application program.
[0075] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in the application program, to:
[0076] Perform the steps of the above method.
[0077] The above method can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The above-disclosed methods, steps, and logic block diagrams can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the above-disclosed method can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0078] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units or combinations thereof for performing the functions described herein.
[0079] For software implementation, the technology of the present application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0080] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0081] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.
Claims
1. A method for fusing head-related transfer function datasets, comprising: Input the sampled data of different subjects in multiple datasets into the trained deep neural network to generate a head-related impulse response that eliminates the measurement features of the dataset; The training process of the deep neural network includes: The head-related impulse response in the original dataset is used to extract interaural time difference, interaural level difference and auditory localization features; Use deep neural networks to generate head-related impulse responses that eliminate the measurement features of the dataset, and extract binaural level difference and auditory localization features; The loss function is constructed using the binaural sound level difference between the original head-related impulse response and the head-related impulse response output by the deep neural network, the difference in auditory localization features, and the dataset classification accuracy between the network output head-related transfer function, and the feedback is used to train the deep neural network. Among them, the auditory localization features are extracted using the head-related impulse response in the original data set, including: performing fast Fourier transform on the resampled head-related impulse response to obtain the head-related transfer function amplitude spectrum at N sampling frequency points; using gammatone auditory filters under M equivalent rectangular bandwidths to perform auditory filtering on the head-related transfer function amplitude spectrum to obtain auditory localization features related to auditory localization perception.
2. The head-related transfer function dataset fusion method according to claim 1, characterized in that: The head-related impulse response in the original data set is used to extract the interaural time difference and interaural sound level difference, including: The HRIs of different subjects in different data sets were resampled with a unified sampling rate and length. The resampled HRIs were filtered using a low-pass filter and the envelope was extracted. The binaural cross-correlation method was used to calculate the peak delay and obtain the interaural time difference (ITD) (θ): Among them, θ represents the sampling direction; τ represents the target interaural time difference; p L represents the left ear head-related impulse response envelope after low-pass filtering; p R represents the right ear head-related impulse response envelope after low-pass filtering; T represents the head-related impulse response sampling time; t represents time; By calculating the energy difference of binaural head-related impulse responses, the binaural level difference ILD(θ) is obtained: Among them, HRIR L Represents the left ear head-related impulse response after resampling; HRIR R represents the resampled head-related impulse response of the right ear.
3. The head-related transfer function dataset fusion method according to claim 1, characterized in that: The input of the deep neural network is a two-channel random noise sampling, and the output is a target two-channel head-related impulse response; The deep neural network includes three intermediate layers and a fully connected output layer; The intermediate layer includes a fully connected layer, batch normalization and an activation function.
4. The head-related transfer function dataset fusion method according to claim 1, characterized in that: The loss function of the deep neural network is: L=L ild +αL loc +βL cls Among them, α and β represent the proportion of perceptual positioning feature loss and classification loss in the total loss respectively; L ild Represents binaural level difference loss: Where ILD represents the binaural level difference of the original input, Represents the binaural level difference output by the deep neural network; L loc Represents the perceptual positioning feature loss: in, represents the gammatone auditory filter; HRTF represents the head-related transfer function of the original input; represents the head-related transfer function output by the deep neural network; M represents the number of equivalent rectangular bandwidths; L cls Denotes the classification loss: Among them, c i and c j represents the mean of the head-related transfer function under different data sets; i and j represent the indexes of different data sets respectively; N represents the total number of data sets; ||·|| 2 represents the 2-norm.
5. The head-related transfer function dataset fusion method according to claim 1, characterized in that: The classification accuracy is characterized by calculating the mean of the head-related transfer function amplitude spectrum under different data sets as the class-heart-time inter-class distance.
6. A head-related transfer function dataset fusion system, implemented based on the method according to any one of claims 1 to 5, characterized in that: The system comprises: The dataset fusion module is used to input the sampled data of different subjects under multiple datasets into the trained deep neural network to generate a head-related impulse response that eliminates the measurement characteristics of the dataset; and Network training module, used to train deep neural networks.
Citation Information
Patent Citations
Head-related transfer function objective evaluation method and system based on auditory perception characteristics
CN117979218A
HRTF data fusion and sparse prediction method based on binaural factors
CN119052713A