A head-related transfer function generation method based on ear-enhanced multi-view reconstruction
By using an enhanced multi-view reconstruction method for the ear, combined with a 3D Gaussian model and camera parameters, a watertight mesh suitable for boundary element acoustics is generated, solving the problem of reconstructing the thin structure of the auricle under consumer-grade image acquisition and realizing high-precision personalized HRTF acoustic simulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to stably reconstruct personalized head correlation transfer functions (HRTFs) under consumer-grade image acquisition conditions, particularly in the geometric restoration of thin auricular structures and the generation of watertight meshes, resulting in insufficient acoustic accuracy in personalized spatial audio systems.
The ear-enhanced multi-view reconstruction method uses a camera to acquire basic and supplementary view images covering the head. By combining a 3D Gaussian model and camera intrinsic and extrinsic parameters, a pseudo-depth map is generated and a watertight mesh is reconstructed. Head coordinates are standardized and ear canal inlet surface is labeled. Finally, boundary element acoustics is used to calculate the personalized HRTF.
It improves the accuracy of auricular geometry reconstruction, generates watertight meshes suitable for acoustic simulation, reduces operational complexity, and enhances the acoustic fidelity of personalized HRTFs, especially in the recovery of spatial cues in the high-frequency band.
Smart Images

Figure CN122492983A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of 3D audio, spatial audio, three-dimensional visual reconstruction and computational acoustics. Specifically, it relates to a method for generating head-related transfer functions based on ear-enhanced multi-view reconstruction. In particular, it relates to constructing a head-related transfer function suitable for boundary element acoustic solutions by acquiring ear-enhanced multi-view images, photogrammetric calibration, three-dimensional Gaussian scatter reconstruction, pseudo-depth derivation, truncated symbolic distance field fusion, watertight mesh generation, and head coordinate standardization. Background Technology
[0002] With the rapid development of virtual reality, augmented reality, immersive games, remote pre-context communication, smart terminals, and hearing aids, users have placed higher demands on the realism, externalization, and directional positioning accuracy of spatial audio systems. The head-related transfer function (HRTF) describes the directional correlation filtering characteristics generated when sound propagates from different directions in space to the listener's ears, after being processed by the head, torso, and auricle. It is a key fundamental data point for achieving personalized spatial audio rendering.
[0003] Existing methods for obtaining personalized HRTF mainly include direct acoustic measurement, data-driven prediction, and physical simulation based on geometric models. Direct acoustic measurement typically requires anechoic chambers, multi-speaker arrays, and specialized measurement equipment, making the process complex, time-consuming, and costly, hindering large-scale deployment in ordinary user scenarios. While learning-based prediction methods based on anthropometric parameters, head and face images, or ear images are easy to deploy, they often lack explicit 3D geometric and acoustic physical constraints, making it difficult to accurately represent the fine spectral cues generated by individual auricular differences in the high-frequency band. Physical simulation methods based on 3D geometric models and the boundary element method can obtain high-fidelity HRTFs when the geometric model is accurate, watertight, and scaled correctly, but this requires the stable acquisition of individual head and auricular meshes that meet the acoustic simulation requirements.
[0004] The auricle is a key structure for high-frequency spectral shaping in HRTF (High-Temperature Formatting). Regions such as the helix, concha, cymba conchae, and tragus exhibit significant variations in texture and thin edges, which substantially influence sound diffraction, reflection, and interference modes. However, the ear region is characterized by severe self-occlusion, weak surface texture, numerous thin structures, and large local curvature variations. When using ordinary consumer-grade cameras for multi-view reconstruction, traditional photogrammetry and multi-view stereo methods easily produce holes, breaks, edge collapses, and topological errors in areas such as the conchae and helix edges. While neural implicit reconstruction or novel perspective synthesis methods offer advantages in visual rendering, their outputs typically do not directly meet the requirements of boundary element acoustics (BEA) solutions for watertightness, normal consistency, scale, and ear canal boundary annotation.
[0005] Furthermore, existing visual reconstruction workflows typically focus only on appearance restoration, lacking a complete processing chain that connects to subsequent acoustic calculations. For example, current methods often fail to specifically acquire trajectories for ear self-occlusion design, nor do they establish a unified head coordinate system, restore physical scale, annotate the ear canal entrance patch, or generate a watertight mesh suitable for boundary element method (BEM) solutions after reconstruction. Therefore, current techniques still struggle to stably obtain high-quality geometric models from ordinary multi-view images that can be directly used for personalized HRTF calculations.
[0006] In summary, the field of personalized spatial audio urgently needs a head reconstruction method that can enhance the acquisition and high-fidelity reconstruction of the complex geometry of the auricle under consumer-grade image acquisition conditions, and can convert the reconstruction results into a watertight mesh suitable for boundary element acoustic simulation, thereby providing a geometric basis for personalized HRTF calculation. Summary of the Invention
[0007] Existing personalized HRTF acquisition technologies suffer from several drawbacks: direct measurement methods are costly and difficult to implement widely; image- or parameter-based prediction methods lack explicit physical constraints; and traditional 3D reconstruction methods struggle to accurately recover the thin auricle structure and generate watertight meshes suitable for boundary element methods. To address these issues, this invention proposes a head-related transfer function generation method based on enhanced multi-view ear reconstruction. This method is used to stably reconstruct an individual head geometry model suitable for subsequent acoustic simulation from ordinary multi-view images.
[0008] The technical solution adopted in this invention is: A method for generating head-related transfer functions based on ear-enhanced multi-view reconstruction, comprising the following steps: A basic viewpoint image covering the overall shape of the subject's head was acquired using a camera, and supplementary viewpoint images were acquired for the left and right ears respectively; based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and a sparse 3D point cloud of the head were obtained. A 3D Gaussian model is generated by training based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. Using the aforementioned 3D Gaussian model and camera intrinsic and extrinsic parameters, the acquired image under each calibrated viewpoint is rendered to obtain a pseudo-depth map of the corresponding viewpoint; a closed watertight head mesh is generated based on the pseudo-depth maps under multiple calibrated viewpoints. A standardized head model is generated based on the closed watertight head mesh, and the ear canal entrance surface is annotated. The boundary element acoustic solver calculates the subject's personalized binaural head correlation transfer function based on the labeled standardized head model boundary element acoustics.
[0009] Preferably, each three-dimensional Gaussian in the set of three-dimensional Gaussians includes at least a center position, anisotropic covariance, color, and opacity parameters.
[0010] Preferably, the method for training and generating a 3D Gaussian model is as follows: A 3D Gaussian set is initialized based on the sparse 3D point cloud of the head; using camera intrinsic and extrinsic parameters, the 3D Gaussian set is projected onto the image planes corresponding to each base view image and supplementary view image; a predicted image is obtained through differentiable rasterization rendering; the predicted image is compared with the corresponding acquired image to obtain the image reconstruction error; depth distortion constraints and normal consistency constraints are added to the image reconstruction error to form a training loss; the center position, anisotropic covariance, color, and opacity parameters of the 3D Gaussian model are iteratively updated based on the training loss to generate the 3D Gaussian model.
[0011] Preferably, the depth distortion constraint is used to constrain the depth distribution corresponding to the Gaussian contribution in the same line of sight, so as to reduce geometric floating layers and depth discrepancies; the normal consistency constraint is used to constrain the normal direction obtained by pseudo-depth or local surface estimation, so as to reduce abrupt changes in local surface normals; during training, adaptive density control is performed on the three-dimensional Gaussian based on image reconstruction error, gradient magnitude or Gaussian coverage.
[0012] Preferably, the pseudo-depth map is a depth result obtained by accumulating the three-dimensional Gaussian depth contribution of each line of sight under the viewpoint according to the transmittance weight; each pseudo-depth map is back-projected to a unified three-dimensional voxel space according to its corresponding camera extrinsic parameters to obtain the corresponding depth observation; the depth observations of each viewpoint are fused to obtain a truncated symbolic distance field; the truncated symbolic distance field is subjected to zero level set extraction, or the truncated symbolic distance field is subjected to surface extraction to obtain a closed watertight head mesh.
[0013] Preferably, the method for generating a standardized head model and annotating the ear canal entrance patches is as follows: The nose tip, left ear key points, and right ear key points are determined based on the closed watertight head mesh; the midpoint of the left and right ear key points is used as the origin of the coordinate system, the direction from the right ear to the left ear is used as the horizontal axis, the direction towards the nose tip is used as the forward axis, and the vertical axis is determined according to the right-hand rule to obtain the standardized head coordinate system; the watertight head mesh is scaled based on the prior distance between the two ears, a reference scale, or known geometric dimensions to obtain the standardized head model; and ear canal entrance patches are annotated at the left and right ear canal entrances of the standardized head model.
[0014] Preferably, the method for obtaining the subject's personalized binaural head correlation transfer function is as follows: the labeled standardized head model and its external normal are input into the boundary element acoustic solver; acoustic hard boundary conditions are applied to the head surface except for the ear canal inlet surface, and normal velocity boundary conditions are applied at the ear canal inlet surface on the side to be determined, and the external sound field is solved in a reciprocal manner; the left and right ear responses are calculated according to preset frequency sampling and spherical direction sampling respectively, to obtain the subject's personalized binaural head correlation transfer function.
[0015] A head-related transfer function generation system based on ear-enhanced multi-view reconstruction, characterized in that it includes: The data acquisition and processing module is used to acquire basic viewpoint images covering the overall shape of the subject's head using a camera, and to acquire supplementary viewpoint images for the left and right ears respectively; based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and sparse three-dimensional point cloud of the head are obtained. The 3D Gaussian model generation module is used to train and generate a 3D Gaussian model based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. The closed watertight head mesh generation module is used to render the acquired image under each calibrated viewpoint using the three-dimensional Gaussian model and camera intrinsic and extrinsic parameters to obtain the pseudo-depth map of the corresponding viewpoint; and to generate a closed watertight head mesh based on the pseudo-depth maps under multiple calibrated viewpoints. The head model generation and annotation module is used to generate a standardized head model based on the closed watertight head mesh and to annotate the ear canal inlet surface. The head-related transfer function generation module is used to calculate the subject's personalized binaural head-related transfer function based on the boundary element acoustics of the labeled standardized head model using a boundary element acoustics solver.
[0016] A computing device, characterized in that it comprises: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above.
[0017] A computer-readable storage medium, characterized in that it stores instructions that, when executed on a computer, cause the computer to perform the above-described method.
[0018] Experiments have shown that the method of the present invention has the following advantages: The method achieves higher precision in ear geometry reconstruction. By supplementing the ear's viewing angle and using directional illumination, the observability of the auricular recess area is enhanced. Combined with adaptive density control of three-dimensional Gaussian scattering, it effectively improves the geometric restoration quality of key high-frequency acoustic structures such as the concha and the helix edge.
[0019] This method can generate watertight meshes suitable for acoustic simulation. Instead of simply outputting visual rendering results, it obtains a watertight, scale-correct, and normal-consistent head mesh through pseudo-depth fusion, truncated symbolic distance field reconstruction, and mesh post-processing, which can be directly connected to boundary element acoustic solutions.
[0020] The personalized HRTF calculation process is complete. This method unifies consumer-grade image acquisition, 3D reconstruction, coordinate standardization, ear canal entrance annotation, and boundary element method solving into a single process, reducing the cost and operational complexity for ordinary users to obtain personalized HRTFs.
[0021] Higher acoustic fidelity. This method improves the geometric accuracy of the auricle, enabling the HRTF obtained from boundary element simulation to outperform various image-based personalized reconstruction baselines in terms of full-band spectral distortion, and is particularly beneficial for preserving high-frequency spatial cues determined by the fine structure of the auricle. Attached Figure Description
[0022] Figure 1 This is a flowchart of the method of the present invention.
[0023] Figure 2 This is a flowchart of the reconstruction and watertight mesh generation based on three-dimensional Gaussian scatter points provided in the embodiments of the present invention.
[0024] Figure 3 This is a schematic diagram of the head-standardized coordinate system provided in an embodiment of the present invention.
[0025] Figure 4 This is a system diagram of the present invention.
[0026] Figure 5 This is a comparison chart of the HRTF average spectral distortion results of the embodiments of the present invention and various comparison methods.
[0027] Figure 6 This is a graph showing the HRTF spectral distortion analysis results at different frequencies according to an embodiment of the present invention.
[0028] Figure 7 This is a diagram showing the effect of pose perturbation on HRTF spectral distortion in an embodiment of the present invention. Detailed Implementation
[0029] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0030] like Figure 1 As shown, an optional embodiment of the present invention provides a method for generating head-related transfer functions based on ear-enhanced multi-view reconstruction, the steps of which include: A basic viewpoint image covering the overall shape of the subject's head was acquired using a camera, and supplementary viewpoint images were acquired for the left and right ears respectively; based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and a sparse 3D point cloud of the head were obtained. A 3D Gaussian model is generated by training based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. Using the aforementioned 3D Gaussian model and camera intrinsic and extrinsic parameters, the acquired image under each calibrated viewpoint is rendered to obtain a pseudo-depth map of the corresponding viewpoint; a closed watertight head mesh is generated based on the pseudo-depth maps under multiple calibrated viewpoints. A standardized head model is generated based on the closed watertight head mesh, and the ear canal entrance surface is annotated. The boundary element acoustic solver calculates the subject's personalized binaural head correlation transfer function based on the labeled standardized head model boundary element acoustics.
[0031] The head-related transfer function generation method based on ear-enhanced multi-view reconstruction provided in a specific embodiment of the present invention has the following steps: 1) Enhanced Multi-View Image Acquisition and Photogrammetric Calibration of the Ear: A horizontal surround image acquisition was performed around the subject's head using a camera to obtain a base view image covering the overall shape of the head. Supplementary view images with an elevation angle were acquired for the left and right ears respectively. During the acquisition of these supplementary view images, a directional light source was placed near the camera axis to enhance the light and shadow gradients in concave and thin structural areas such as the concha, cymba conchae, and helix edges. The multi-view image, composed of the base and supplementary view images, underwent distortion correction, feature matching, registration, and photogrammetric processing to obtain camera intrinsic parameters, camera extrinsic parameters, and a sparse 3D point cloud of the head. The sparse 3D point cloud of the head was used to initialize subsequent 3D Gaussian assemblies, while the camera intrinsic and extrinsic parameters were used for subsequent 3D Gaussian projection, differentiable rendering training, and pseudo-depth map generation.
[0032] 2) Head and Auricle Reconstruction Based on 3D Gaussian Scatter Points: A 3D Gaussian set is initialized based on the 3D spatial points in the sparse 3D point cloud of the head obtained in step 1), ensuring that each 3D Gaussian set includes at least a center position, anisotropic covariance, color, and opacity parameters. Using the camera intrinsic and extrinsic parameters obtained in step 1), the 3D Gaussian set is projected onto the image planes corresponding to the base view image and the supplementary view image. A predicted image is obtained through differentiable rasterization rendering. The predicted image is compared with the corresponding acquired image to obtain the image reconstruction error. Depth distortion constraints and normal consistency constraints are added to the image reconstruction error to form the training loss. The center position, anisotropic covariance, color, and opacity parameters of the 3D Gaussian set are iteratively updated based on the training loss to obtain the trained 3D Gaussian model.
[0033] Among them, the depth distortion constraint is used to constrain the depth distribution corresponding to the 3D Gaussian contribution in the same line of sight, so as to reduce geometric floating and depth discrepancies; the normal consistency constraint is used to constrain the normal direction estimated by pseudo depth or local surface, so as to reduce abrupt changes in local surface normal.
[0034] During training, adaptive density control is performed on the 3D Gaussian based on whether the image reconstruction error, gradient magnitude, or the coverage area of the 3D Gaussian projection meets the preset density control conditions. Specifically, every preset number of iterations, the projection coverage area of each 3D Gaussian in the training viewpoint, the image reconstruction error of the corresponding image region, and the position gradient magnitude are statistically analyzed. If the projection coverage area of the 3D Gaussian is greater than the preset coverage area threshold, or is in the top 5% to 20% of the 3D Gaussian projection coverage area in the current iteration, and the image reconstruction error or gradient magnitude of the corresponding image region is greater than the preset error threshold or preset gradient threshold, or is in the top 5% to 20% of the image reconstruction error or gradient magnitude in the current iteration, then it is determined that it meets the preset density control conditions.
[0035] For a 3D Gaussian model that meets the preset density control conditions and has a large projection coverage, it is split into multiple smaller-scale 3D Gaussian models. For regions with insufficient local detail representation, such as the helix, concha boundary, cymba conchae, and tragus, 3D Gaussian models that meet the preset density control conditions are cloned along the error or gradient direction to increase local representation density. Through the above processing, the trained 3D Gaussian model can simultaneously represent the overall shape of the head and the thin local structures of the auricle. The output of the above training processing is the trained 3D Gaussian model, which is used for subsequent color map and pseudo-depth map rendering.
[0036] 3) Pseudo-depth derivation, truncated symbolic distance field fusion, and watertight mesh generation: Using the trained 3D Gaussian model obtained in step 2), and combining it with the camera intrinsic and extrinsic parameters obtained in step 1), the 3D Gaussian model is rendered under the camera intrinsic and extrinsic parameters corresponding to each calibrated viewpoint, resulting in a color map and a pseudo-depth map corresponding to that calibrated viewpoint. The color map is used to check the consistency between the reconstructed appearance and the corresponding acquired image; the pseudo-depth map is used for subsequent multi-view depth fusion. The pseudo-depth map is the depth result obtained by accumulating the 3D Gaussian depth contribution in each line of sight under the calibrated viewpoint according to the transmittance weight.
[0037] Multiple pseudo-depth maps obtained from calibrated viewpoints are back-projected into a unified 3D voxel space according to their corresponding camera intrinsic and extrinsic parameters to obtain depth observations for each viewpoint. Depth observations from different viewpoints are then fused to obtain a truncated symbolic distance field. This truncated symbolic distance field represents the signed distance from each voxel in 3D space to the surface to be reconstructed. Zero-level set extraction is performed on the truncated symbolic distance field, or the Marching Cubes algorithm is used to extract the surface from the truncated symbolic distance field to obtain a closed watertight head mesh.
[0038] Then, the closed watertight head mesh is post-processed. This post-processing includes removing isolated connected components whose area, volume, or number of faces is below a preset threshold, removing degenerate triangular faces, repairing small holes, performing light smoothing, and unifying the outward normal direction of the closed watertight head mesh, resulting in a closed watertight head mesh that meets the geometric requirements for boundary element acoustics. This closed watertight head mesh is used for subsequent head coordinate standardization, ear canal inlet face annotation, and boundary element acoustics calculations.
[0039] 4) Head Coordinate Standardization, Scale Restoration, and Ear Canal Entrance Labeling: For the closed watertight head mesh obtained in step 3) that meets the geometric requirements of boundary element acoustic solution, determine the nose tip, left ear key point, and right ear key point. Using the midpoint of the left and right ear key points as the origin, the direction from the right ear to the left ear as the horizontal axis, and the direction towards the nose tip as the forward axis, and determining the vertical axis according to the right-hand rule, a standardized head coordinate system is obtained. Based on the prior information of the inter-ear distance, a reference scale, or known geometric dimensions, the closed watertight head mesh is scaled to obtain a standardized head model. Furthermore, ear canal entrance patches are labeled at the left and right ear canal entrances of the standardized head model. These ear canal entrance patches are used for setting boundary conditions in subsequent boundary element acoustic solutions.
[0040] 5) Personalized HRTF Calculation Based on Boundary Element Method: The standardized head model obtained in step 4), the external normal of the standardized head model, and the ear canal inlet surface are input into the boundary element acoustic solver. Hard acoustic boundary conditions are applied to the head surface except for the ear canal inlet surface, and normal velocity boundary conditions are applied at the ear canal inlet surface on the side to be calculated. The external sound field is solved in a reciprocal manner. The left and right ear responses are calculated separately according to preset frequency sampling and spherical direction sampling, and the personalized binaural HRTF of the subject is output.
[0041] In one embodiment, the method for acquiring multi-view images of the head and ears in step one is as follows: To obtain individualized geometric structures of the subject's head and auricle, a basic perspective image covering the overall shape of the head is first captured using a standard camera, mobile phone camera, industrial camera, or other imaging device, taking horizontal, surround-view images of the subject's head. In one embodiment, the number of basic perspective images is 30 to 40. To address the issue of self-occlusion in the auricle region, several supplementary perspective images with an elevation angle are captured for both the left and right ears, allowing the camera to observe areas difficult to cover by a standard horizontal perspective, such as the area above the concha, the cymba conchae, and the inner side of the helix. In one embodiment, the number of supplementary perspective images for each ear is 10 to 15, with the supplementary perspective having an elevation angle of ±20 to 40 degrees relative to the horizontal line of sight, preferably approximately ±30 degrees.
[0042] When acquiring supplementary viewpoint images, small directional light sources are placed near the camera axis to create stable light and shadow variations in the ear's recessed structure, thereby enhancing the local visual cues required for feature matching and 3D reconstruction. After acquisition, distortion correction, feature matching, registration, and photogrammetric processing are performed on the multi-view image composed of the base viewpoint image and the supplementary viewpoint image to obtain camera intrinsic parameters, camera extrinsic parameters, and a sparse 3D point cloud of the head. The camera intrinsic parameters describe the camera imaging model, the camera extrinsic parameters describe the camera pose corresponding to each viewpoint, and the sparse 3D point cloud of the head represents the initial 3D spatial points on the head surface.
[0043] In one embodiment, the head and auricle reconstruction method based on three-dimensional Gaussian scatter points in step two is as follows: A 3D Gaussian set is initialized based on the sparse 3D point cloud of the head obtained in Step 1. Specifically, the 3D spatial points in the sparse 3D point cloud of the head are used as the initial values for the center positions of the 3D Gaussians, and anisotropic covariance, color, and opacity parameters are set for each 3D Gaussian. Using the camera intrinsic and extrinsic parameters obtained in Step 1, the 3D Gaussian set is projected onto the image planes corresponding to the base view image and the supplementary view image, and the predicted image is obtained through differentiable rasterization rendering.
[0044] The predicted image is compared with the corresponding acquired image to obtain the image reconstruction error. Depth distortion constraints are constructed based on the depth distribution obtained during training, and normal consistency constraints are constructed based on the local surface normal or the normal estimated from the depth gradient. The image reconstruction error, depth distortion constraints, and normal consistency constraints are weighted and combined to form the training loss. The training loss is used to iteratively update the center position, anisotropic covariance, color, and opacity parameters of the 3D Gaussian model, resulting in the trained 3D Gaussian model.
[0045] During training, adaptive density control is performed on the 3D Gaussian model based on image reconstruction error, gradient magnitude, or Gaussian coverage. For 3D Gaussians with large coverage or significant errors, they are split into multiple smaller-scale 3D Gaussians. For regions with insufficient representation of local details, such as the helix, concha boundary, cymba conchae, and tragus, 3D Gaussians are cloned along the error or gradient direction to increase local representation density. Through these processes, the trained 3D Gaussian model can simultaneously represent the overall shape of the head and the thin local structures of the auricle.
[0046] In one embodiment, the pseudo-depth derivation and watertight mesh generation method in step three is as follows: Using the trained 3D Gaussian model obtained in step two, and combining it with the camera intrinsic and extrinsic parameters obtained in step one, rendering is performed at each calibrated viewpoint to obtain a color map and a pseudo-depth map corresponding to that viewpoint. The color map is used to check the consistency between the reconstruction result and the input image, and the pseudo-depth map is used for subsequent multi-view depth fusion. The pseudo-depth map is the depth result obtained by accumulating the 3D Gaussian depth contribution of each line of sight in that viewpoint according to the transmittance weight.
[0047] Multiple pseudo-depth maps obtained from calibrated viewpoints are back-projected into a unified 3D voxel space according to their corresponding camera extrinsic parameters. Depth observations from different viewpoints are then fused to obtain a truncated symbolic distance field. This truncated symbolic distance field represents the signed distance from each voxel in 3D space to the surface to be reconstructed. Zero-level set extraction is performed on the truncated symbolic distance field, or the Marching Cubes algorithm is used to extract the surface, resulting in a closed watertight head mesh, such as... Figure 2 As shown.
[0048] Subsequently, post-mesh processing is performed on the closed watertight head mesh. This post-mesh processing includes removing isolated connected components whose area, volume, or number of faces is below a preset threshold, removing degenerate triangular faces, repairing small holes, performing light smoothing, and unifying the outward normal direction. After post-mesh processing, a closed watertight head mesh that meets the geometric requirements for boundary element acoustic solution is obtained.
[0049] In one embodiment, the method for head coordinate standardization, scale restoration, and ear canal entrance annotation in step four is as follows: For the watertight head mesh obtained in step three that meets the geometric requirements of boundary element acoustic solution, determine the nose tip, left ear keypoint, and right ear keypoint. Establish a standardized head coordinate system by using the midpoint between the left and right ear keypoints as the origin, the direction from the right ear to the left ear as the horizontal axis, and the direction towards the nose tip orthogonal to the horizontal axis as the forward axis. Determine the vertical axis according to the right-hand rule. Figure 3 As shown.
[0050] Subsequently, based on the prior information regarding the interauricular distance, calibration scale, or known geometric dimensions, the closed watertight head mesh is scaled back to possess true physical length units, resulting in a standardized head model. Furthermore, facets are selected near the left and right ear canal entrances of the standardized head model and labeled to obtain ear canal entrance facets. These ear canal entrance facets are used for setting boundary conditions in subsequent boundary element acoustic solutions.
[0051] In one embodiment, the personalized HRTF calculation method based on the boundary element method in step five is as follows: The standardized head model, external normal, and ear canal inlet surface obtained in step four are input into the boundary element acoustic solver. Hard acoustic boundary conditions are applied to the head surface except for the ear canal inlet surface, and a preset normal velocity boundary condition is applied at the ear canal inlet surface on the side to be solved, to construct a reciprocal sound source. The external sound field is solved for preset frequency points and spherical direction sampling points to obtain the transfer function values at the corresponding frequencies and directions.
[0052] In one embodiment, the frequency range covers 100Hz to 20kHz, and the spatial orientation uses a spherical sampling grid composed of azimuth and elevation angles. By performing the above solution on the left and right ears respectively, a personalized binaural HRTF for the subject can be obtained. This HRTF can be used for applications such as binaural rendering, virtual sound source localization, personalized filter design, and spatial audio playback.
[0053] like Figure 4 As shown, an optional embodiment of the present invention provides a head-related transfer function generation system based on ear-enhanced multi-view reconstruction, characterized in that it includes... The data acquisition and processing module is used to acquire basic viewpoint images covering the overall shape of the subject's head using a camera, and to acquire supplementary viewpoint images for the left and right ears respectively; based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and sparse three-dimensional point cloud of the head are obtained. The 3D Gaussian model generation module is used to train and generate a 3D Gaussian model based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. The closed watertight head mesh generation module is used to render the acquired image under each calibrated viewpoint using the three-dimensional Gaussian model and camera intrinsic and extrinsic parameters to obtain the pseudo-depth map of the corresponding viewpoint; and to generate a closed watertight head mesh based on the pseudo-depth maps under multiple calibrated viewpoints. The head model generation and annotation module is used to generate a standardized head model based on the closed watertight head mesh and to annotate the ear canal inlet surface. The head-related transfer function generation module is used to calculate the subject's personalized binaural head-related transfer function based on the boundary element acoustics of the labeled standardized head model using a boundary element acoustics solver.
[0054] The present invention provides a computing device, characterized in that it includes: a processor and a memory storing a computer program, wherein the computer program is executed by the processor to perform the above-described method.
[0055] The present invention provides a computer-readable storage medium, characterized in that it stores instructions that, when executed on a computer, cause the computer to perform the above-described method.
[0056] Method evaluation experiment The method of this invention was validated on a publicly available high-precision head dataset. In the experiment, high-precision head scan data was used as the reference geometry to generate multi-view image inputs. The method of this invention was then compared with traditional multi-view stereo reconstruction, neural implicit reconstruction, and existing ear reconstruction methods. Geometric evaluation used the symmetric nearest neighbor distance, Chamfer distance, and Hausdorff distance of the ear region; acoustic evaluation used the HRTF spectral distortion index calculated based on the boundary element method.
[0057] Experimental results show that the method of the present invention can achieve lower geometric errors in the ear region and achieve better results in terms of HRTF spectral distortion across the entire frequency band, such as... Figure 5 , Figure 6 and Figure 7 As shown. Compared to acquisition methods without supplementary ear viewpoints and directional illumination, the complete solution of this invention significantly improves both ear geometric errors and HRTF errors, demonstrating that ear-enhanced acquisition design plays a crucial role in restoring key acoustic geometries.
[0058] In summary, this invention proposes a method for generating head-related transfer functions (HRTFs) based on ear-enhanced multi-view reconstruction. By unifying the design of ear enhancement acquisition, 3D Gaussian scatter reconstruction, pseudo-depth fusion, watertight mesh generation, head coordinate standardization, ear canal entrance annotation, and boundary element acoustic solution, a complete computational process from consumer-grade multi-view images to personalized HRTFs is achieved. This method can improve the accuracy of auricular geometry reconstruction and the acoustic fidelity of personalized spatial audio, demonstrating significant engineering application value.
[0059] Although specific embodiments and accompanying drawings of the invention have been disclosed for illustrative purposes to aid in understanding and implementing the invention, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. Therefore, the invention should not be limited to the content disclosed in the preferred embodiments and accompanying drawings.
Claims
1. A method for generating head-related transfer functions based on ear-enhanced multi-view reconstruction, comprising the following steps: A basic viewpoint image covering the overall shape of the subject's head was acquired using a camera, and supplementary viewpoint images were acquired for the left and right ears respectively. Based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head are obtained; A 3D Gaussian model is generated by training based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. Using the aforementioned 3D Gaussian model and camera intrinsic and extrinsic parameters, the acquired image under each calibrated viewpoint is rendered to obtain a pseudo-depth map of the corresponding viewpoint; a closed watertight head mesh is generated based on the pseudo-depth maps under multiple calibrated viewpoints. A standardized head model is generated based on the closed watertight head mesh, and the ear canal entrance surface is annotated. The boundary element acoustic solver calculates the subject's personalized binaural head correlation transfer function based on the labeled standardized head model boundary element acoustics.
2. The method of claim 1, wherein, Each of the three-dimensional Gaussians in the set of three-dimensional Gaussians includes at least the center position, anisotropic covariance, color, and opacity parameters.
3. The method according to claim 2, characterized in that, The method for training and generating a 3D Gaussian model is as follows: Initialize a 3D Gaussian set based on the sparse 3D point cloud of the head; Project the 3D Gaussian set onto the image plane corresponding to each base view image and supplementary view image using camera intrinsic and extrinsic parameters; Obtain the predicted image through differentiable rasterization rendering; Compare the predicted image with the corresponding acquired image to obtain the image reconstruction error; Add depth distortion constraints and normal consistency constraints to the image reconstruction error to form the training loss; Iteratively update the center position, anisotropic covariance, color, and opacity parameters of the 3D Gaussian model based on the training loss to generate the 3D Gaussian model.
4. The method according to claim 3, characterized in that, The depth distortion constraint is used to constrain the depth distribution corresponding to the Gaussian contribution in the same line of sight, so as to reduce geometric floating layers and depth discrepancies; the normal consistency constraint is used to constrain the normal direction estimated by pseudo-depth or local surface, so as to reduce abrupt changes in local surface normals; during training, adaptive density control is performed on the 3D Gaussian based on image reconstruction error, gradient magnitude or Gaussian coverage.
5. The method according to claim 1, characterized in that, The pseudo-depth map is the depth result obtained by accumulating the three-dimensional Gaussian depth contribution of each line of sight under the viewpoint according to the transmittance weight; each pseudo-depth map is back-projected to a unified three-dimensional voxel space according to its corresponding camera extrinsic parameters to obtain the corresponding depth observation; the depth observations of each viewpoint are fused to obtain the truncated symbolic distance field. Zero-level set extraction is performed on the truncated symbolic distance field, or surface extraction is performed on the truncated symbolic distance field to obtain a closed watertight head mesh.
6. The method according to claim 1, characterized in that, The method for generating a standardized head model and annotating the ear canal entrance patches is as follows: The nose tip, left ear key points, and right ear key points are determined based on the closed watertight head mesh; the midpoint of the left and right ear key points is used as the origin of the coordinate system, the direction from the right ear to the left ear is used as the horizontal axis, the direction towards the nose tip is used as the forward axis, and the vertical axis is determined according to the right-hand rule to obtain the standardized head coordinate system; the watertight head mesh is scaled based on the prior distance between the two ears, a reference scale, or known geometric dimensions to obtain the standardized head model; and ear canal entrance patches are annotated at the left and right ear canal entrances of the standardized head model.
7. The method according to claim 1, characterized in that, The method for obtaining the subject's personalized binaural head correlation transfer function is as follows: the labeled standardized head model and its external normal are input into the boundary element acoustic solver; acoustic hard boundary conditions are applied to the head surface except for the ear canal inlet surface, and normal velocity boundary conditions are applied at the ear canal inlet surface on the side to be determined, and the external sound field is solved in a reciprocal manner; the left and right ear responses are calculated according to preset frequency sampling and spherical direction sampling respectively, and the subject's personalized binaural head correlation transfer function is obtained.
8. A head-related transfer function generation system based on ear-enhanced multi-view reconstruction, characterized in that, include The data acquisition and processing module is used to acquire basic viewpoint images covering the overall shape of the subject's head using a camera, and to acquire supplementary viewpoint images for the left and right ears respectively; based on the acquired images, camera intrinsic parameters, camera extrinsic parameters, and sparse three-dimensional point cloud of the head are obtained. The 3D Gaussian model generation module is used to train and generate a 3D Gaussian model based on camera intrinsic parameters, camera extrinsic parameters, and sparse 3D point cloud of the head. The closed watertight head mesh generation module is used to render the acquired image under each calibrated viewpoint using the three-dimensional Gaussian model and camera intrinsic and extrinsic parameters to obtain the pseudo-depth map of the corresponding viewpoint; and to generate a closed watertight head mesh based on the pseudo-depth maps under multiple calibrated viewpoints. The head model generation and annotation module is used to generate a standardized head model based on the closed watertight head mesh and to annotate the ear canal inlet surface. The head-related transfer function generation module is used to calculate the subject's personalized binaural head-related transfer function based on the boundary element acoustics of the labeled standardized head model using a boundary element acoustics solver.
9. A computing device, characterized in that, include: A processor, a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A storage instruction that, when executed on a computer, causes the computer to perform the method as described in any one of claims 1 to 7.