Head-related transfer function generating device, program, and head-related transfer function generating method
The head-related transfer function generating device estimates and converts head-related transfer functions to generate them in any direction in a three-dimensional space, addressing the limitation of existing devices and enhancing sound localization accuracy.
Patent Information
- Application Number
- JP2022024205
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-11-17
- Estimated Expiration
- 2042-02-18
AI Technical Summary
Existing head-related transfer function selection devices are unable to generate head-related transfer functions in any direction in the entire sky three-dimensional space.
A head-related transfer function generating device that estimates notches and peaks of transfer functions for each angle in the median plane of the listener, acquires interaural difference information, and generates a three-dimensional head-related transfer function by performing time-frequency conversion on the head-related transfer response, allowing for head-related transfer function generation in any direction in a three-dimensional space.
Enables the generation of head-related transfer functions in any direction in the entire sky three-dimensional space, improving sound localization accuracy.
Smart Images

Figure 0007770680000006 
Figure 0007770680000007 
Figure 0007770680000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to a head-related transfer function generating device, a program, and a head-related transfer function generating method. [Background technology]
[0002] Research and development has been ongoing for some time, with the aim of putting into practical use three-dimensional sound systems, sound virtual reality (VR), and the like. To put these technologies into practical use, it is necessary to reproduce the head-related transfer function for each listener. One example of a technology for reproducing the head-related transfer function for each listener is the head-related transfer function selection device disclosed in Patent Document 1. This head-related transfer function selection device includes a measurement unit, a feature extraction unit, and a characteristic selection unit. The measurement unit acquires a user's head impulse response based on an audio signal picked up by a microphone attached to the user's ear while a predetermined sound is generated as a measurement signal from a speaker. The feature extraction unit extracts a feature of a frequency characteristic corresponding to the head impulse response. The characteristic selection unit selects one of the head-related transfer functions from a database that associates the head-related transfer functions of a plurality of people with the feature of the head-related transfer function based on the extracted feature. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2016-201723 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the above-described head-related transfer function selection device cannot generate head-related transfer functions in any direction in the entire sky three-dimensional space.
[0005] In view of the above-mentioned problems, the present invention aims to provide a head-related transfer function generating device, a program, and a head-related transfer function generating method that can generate a head-related transfer function in any direction in the entire sky three-dimensional space. [Means for solving the problem]
[0006] an estimation unit that estimates notches and peaks of the head-related transfer functions for each angle in the median plane of the listener based on parameter information of the notches and peaks of the head-related transfer functions acquired by the head-related transfer function acquisition unit; an interaural difference information acquisition unit that acquires interaural difference information of the listener; a head-related transfer function generation unit that generates a head-related transfer function for any direction in a three-dimensional space centered on the listener based on the estimation results of the notches and peaks for each angle in the median plane estimated by the estimation unit and the interaural difference information acquired by the interaural difference information acquisition unit; and a three-dimensional head-related transfer function generation unit that generates a three-dimensional head-related transfer function of the listener by performing a time-frequency conversion on the head-related transfer response generated by the head-related transfer response generation unit.
[0007] Moreover, according to one embodiment of the present invention, in the above-described head-related transfer function generating device, the three directions are a direction in front of the listener, a direction directly behind the listener, and a zenith direction.
[0008] Furthermore, one embodiment of the present invention is the head-related transfer function generation device described above, wherein the estimation unit obtains a regression equation for the notch and peak parameters with the elevation angle in the median plane of the listener as an independent variable, and estimates the notch and peak parameters for each elevation angle based on the regression equation, and the head-related transfer function generation unit calculates a head-related impulse response in the median plane of the listener based on the notch and peak parameters in the median plane of the listener calculated by the estimation unit, and generates a head-related impulse response in any direction in a three-dimensional space centered on the listener by adding at least one of an interaural time difference and an interaural level difference for each lateral angle of the listener based on the calculated head-related impulse response in the median plane of the listener and the interaural difference information.
[0009] Furthermore, one embodiment of the present invention relates to the above-described head-related transfer function generation device, and further includes a hearing device inverse transfer function convolution unit that performs a convolution operation of the inverse transfer function of the hearing device used by the listener with the head impulse response in any direction in three-dimensional space centered on the listener, generated by the head impulse response generation unit.
[0010] Moreover, according to one embodiment of the present invention, the above-mentioned head-related transfer function generation device further comprises a sound source signal convolution section that performs a convolution operation of a sound source signal on the result of the convolution operation by the hearing device inverse transfer function convolution section.
[0011] Moreover, one embodiment of the present invention is a program for causing a computer to execute the following steps: acquire head-related transfer functions in at least three directions centered on the listener within the median plane of the listener; estimate notches and peaks of the head-related transfer functions for each angle within the median plane of the listener based on parameter information indicating notches and peaks of the acquired head-related transfer functions; acquire interaural difference information of the listener; generate a head-related impulse response for any direction in a three-dimensional space centered on the listener based on the estimated results of the notches and peaks for each angle within the median plane and the acquired interaural difference information; and generate a three-dimensional head-related transfer function for the listener by performing a time-frequency conversion on the generated head-related impulse response.
[0012] Moreover, one embodiment of the present invention is a head-related transfer function generation method comprising: acquiring head-related transfer functions in at least three directions centered on the listener within the median plane of the listener; estimating notches and peaks of the head-related transfer functions for each angle within the median plane of the listener based on parameter information indicating notches and peaks of the acquired head-related transfer functions; acquiring interaural difference information of the listener; generating a head-related transfer function for any direction in a three-dimensional space centered on the listener based on the estimated results of the notches and peaks for each angle within the median plane and the acquired interaural difference information; and generating a three-dimensional head-related transfer function for the listener by performing a time-frequency transform on the generated head-related transfer response. [Effects of the Invention]
[0013] According to the present invention, it is possible to generate a head-related transfer function in any direction in the whole sky three-dimensional space. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram showing a listener according to the present embodiment, and a horizontal plane, a median plane, a sagittal plane, an ear axis, a lateral angle, and an elevation angle relative to the listener. [Figure 2]FIG. 1 shows an example of a parametric notch-peak HRTF model generation algorithm. [Figure 3] FIG. 1 shows an example of a user interface for HRTF personal adaptation software. [Figure 4] FIG. 10 is a diagram showing an example of the processing flow of HRTF personal adaptation software. [Figure 5] FIG. 1 is a diagram illustrating an example of hardware constituting a head-related transfer function generating device according to an embodiment. [Figure 6] FIG. 1 is a diagram illustrating an example of a functional configuration of a head-related transfer function generating device according to an embodiment. [Figure 7] FIG. 10 is a diagram showing an example of a software screen that allows a listener to perform HRTF personalization while listening to the sound using a PNP model. [Figure 8] FIG. 10 is a diagram showing an example of the distribution of N1 frequencies in the median plane. [Figure 9] FIG. 1 is a diagram showing an example of N / P parameters (frequency of N / P) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject. [Figure 10] FIG. 1 is a diagram showing an example of N / P parameters (levels) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject. [Figure 11] FIG. 1 is a diagram showing an example of the N / P parameter (Q) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject. [Figure 12] FIG. 10 is a diagram showing an example of ITD distribution according to an actual head shape. [Figure 13] FIG. 10 is a diagram showing an example of a procedure for generating a personalized head impulse response according to the present embodiment. [Figure 14] FIG. 2 is a diagram showing an example of a setting screen of a 3D rendering toolkit according to the present embodiment. [Figure 15] FIG. 10 is a diagram showing an example of responses from each subject regarding azimuth angles. [Figure 16] FIG. 10 is a diagram showing an example of responses from each subject regarding the angle of elevation. [Figure 17] FIG. 10 is a diagram showing an example of an average lateral angle error for evaluating the accuracy of sound image localization in the left-right direction. [Figure 18] FIG. 10 is a diagram showing an example of a front-rear erroneous determination rate for evaluating the accuracy of sound image localization in the front-rear direction. [Figure 19] FIG. 10 is a diagram showing an example of an average elevation angle error for evaluating the accuracy of sound image localization in the vertical direction. [Figure 20] FIG. 10 is a diagram showing an example of an intra-head localization rate. [Figure 21] FIG. 10 is a diagram showing an example of the results of converting the azimuth angles answered by the test subjects into lateral angles. [Figure 22] FIG. 10 is a diagram showing an example of the result of converting the elevation angle answered by the subject into an ascent. [Figure 23] FIG. 10 is a diagram showing an example of an average lateral angle error for evaluating the accuracy of sound image localization in the left-right direction. [Figure 24] FIG. 10 is a diagram showing an example of a front-rear erroneous determination rate for evaluating the accuracy of sound image localization in the front-rear direction. [Figure 25] FIG. 10 is a diagram showing an example of an average elevation angle error for evaluating the accuracy of sound image localization in the vertical direction. [Figure 26] FIG. 10 is a diagram showing an example of an intra-head localization rate. [Figure 27] FIG. 10 is a diagram showing an example of the results of converting the azimuth angles answered by the test subjects into lateral angles. [Figure 28] FIG. 10 is a diagram showing an example of the result of converting the elevation angle answered by the subject into an ascent. [Figure 29] FIG. 10 is a diagram showing an example of an average lateral angle error for evaluating the accuracy of sound image localization in the left-right direction. [Figure 30] FIG. 10 is a diagram showing an example of an average elevation angle error for evaluating the accuracy of sound image localization in the vertical direction. [Figure 31] FIG. 10 is a diagram showing an example of the localization accuracy using the subject's personalized HRTF when the headphone transfer function is corrected. DETAILED DESCRIPTION OF THE INVENTION
[0015] [Embodiment] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0016] [Outline of head-related transfer function generator] First, an ear axis coordinate system used in explaining the head-related transfer function generating device 1 according to the embodiment will be described with reference to FIG.
[0017] FIG. 1 is a diagram showing a listener according to this embodiment, and the horizontal plane, median plane, sagittal plane, ear axis, lateral angle, and elevation angle relative to the listener. The ear axis coordinate system shown in Figure 1 is defined as follows: Ear axis A is a line connecting the left and right ear canal entrances of listener P. The origin is the midpoint of the line segment that connects the left and right ear canal entrances of listener P and is located on ear axis A. Horizontal plane H is a plane that connects the right orbital point and the left and right tragus. Median plane M is a plane that is perpendicular to the horizontal plane and bisects listener P into left and right halves. Sagittal plane S is an arbitrary plane parallel to median plane M. In addition, the ear axis coordinate system represents the direction in which the sound source is located by a lateral angle α and an elevation angle β. The lateral angle α is the complement of the angle between the ear axis A and the line connecting the point where the sound source is located and the origin. The lateral angle α is 0 degrees in the direction directly in front of listener P within the horizontal plane H, and 180 degrees in the direction behind listener P within the horizontal plane H. The elevation angle β is the elevation angle in the sagittal plane S passing through the point where the sound source is located.
[0018] The head-related transfer function generating device 1 of this embodiment generates a personalized head-related transfer function (HRTF). An HRTF is a frequency domain representation of a change in physical characteristics caused by a sound wave reaching the entrance of a listener's ear canal from a sound source being affected by the listener's head and its surroundings, and includes peaks and notches. A peak refers to an upwardly convex portion of a head-related transfer function. A notch refers to a downwardly convex portion of a head-related transfer function. The measured head-related transfer function is a head-related transfer function generated by actually measuring sound waves. The predetermined direction mentioned above is, for example, the front of each listener.
[0019] Figure 2 shows an example of a parametric notch-peak HRTF model generation algorithm. Each notch peak is represented by three parameters: 1) frequency, 2) level, and 3) sharpness, and a parametric notch peak HRTF model is generated using the algorithm shown in Figure 2. In the following explanation, the three parameters of frequency, level, and sharpness (Q) that represent a particular notch or peak are also referred to as the notch peak FLQ parameters (or simply the notch peak parameters).
[0020] (Steps S1 to S3) Obtain notch peak parameters using multiple HRTFs. (Step S4) Using the obtained parameters, a regression equation is obtained for the frequency of a specific notch or peak (e.g., the first notch) for each notch / peak frequency. (Steps S5 to S6) A parametric notch-peak HRTF model is obtained using the frequency of a specific notch or peak (for example, the first notch) as a parameter.
[0021] To achieve three-dimensional sound localization, it is necessary to provide HRTFs that are specific to the listener or that are adapted to the listener. This process is called HRTF personalization. HRTFs personalized through this process are also called personalized HRTFs.
[0022] FIG. 3 shows an example of a user interface for the HRTF personal adaptation software. FIG. 4 is a diagram showing an example of the processing flow of the HRTF personal adaptation software. The HRTF personal adaptation software is provided by the head-related transfer function generating device 1.
[0023] The HRTF personal adaptation software generates a sound convoluted with HRTFs from the parametric notch-peak HRTF model and the sound source signal, and presents the generated sound to the listener via headphones (not shown). The listener changes the parameters (e.g., the first notch frequency) of the parametric notch-peak HRTF model using the slider SL in the user interface shown in Figure 3 to search for parameter settings that produce sound in the target direction. When the listener hears the sound presented by the headphones in the target direction, they press the save button SV. When the save button SV is pressed, the HRTF personal adaptation software generates a personalized HRTF using the set parameters and stores the generated personalized HRTF in a memory unit (not shown).
[0024] The head-related transfer function generation device 1 of this embodiment generates HRTFs for the upper direction (zenith) in addition to the front (front) and rear (directly behind). From the parameter information of the notches and peaks of the HRTFs generated in these three directions, parameter information of the notches and peaks of the HRTFs in any direction such as the median plane is estimated, and by adding interaural difference information to this, a head-related transfer function generation device 1 generates a head-related transfer function for any direction in the three-dimensional space of the entire sky. The head-related transfer function generation device 1 converts the generated head-related transfer response into a head-related transfer function by Fourier transforming it.
[0025] [Hardware configuration of head-related transfer function generator] Next, the hardware constituting the head-related transfer function generating device 1 according to the embodiment will be described with reference to FIG.
[0026] 5 is a diagram showing an example of hardware constituting the head-related transfer function generation device 1 according to the embodiment. As shown in FIG. 5, the head-related transfer function generation device 1 includes a processor 11, a main memory device 12, a communication interface 13, an auxiliary memory device 14, an input / output device 15, and a bus 16.
[0027] The processor 11 is, for example, a CPU (Central Processing Unit), which reads and executes HRTF personal adaptation software (hereinafter also referred to as a program) to realize each function of the head-related transfer function generation device 1.
[0028] The main storage device 12 is, for example, a RAM (Random Access Memory), and stores in advance programs that are read and executed by the processor 11.
[0029] The communication interface 13 is an interface circuit for communicating with other devices via a network, such as the Internet, an intranet, a wide area network (WAN), or a local area network (LAN).
[0030] The auxiliary storage device 14 is, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, or a read only memory (ROM).
[0031] The input / output device 15 is, for example, an input / output port. To the input / output device 15, for example, a mouse 151, a keyboard 152, and a display 153 shown in FIG. 5 are connected. The mouse 151 and the keyboard 152 are used, for example, for inputting data required to operate the head related transfer function generation device 1. The display 153 is, for example, a liquid crystal display. The display 153 displays, for example, a graphical user interface (GUI) used by a user of the head related transfer function generation device 1. Details of the graphical user interface will be described later.
[0032] The bus 16 connects the processor 11, the main memory device 12, the communication interface 13, the auxiliary memory device 14, and the input / output device 15 so that data can be transmitted and received among them.
[0033] [Functional configuration of the head-related transfer function generator] Next, the functional configuration of the head-related transfer function generating device 1 according to the embodiment will be described with reference to FIG.
[0034] FIG. 6 is a diagram showing an example of the functional configuration of the head-related transfer function generating device 1 according to the embodiment. The head-related transfer function generation device 1 includes a head-related transfer function acquisition unit 111, an estimation unit 112, an interaural difference information acquisition unit 113, a head impulse response generation unit 114, a three-dimensional head-related transfer function generation unit 115, a hearing device inverse transfer function acquisition unit 116, a hearing device inverse transfer function convolution unit 117, a sound source signal acquisition unit 118, and a sound source signal convolution unit 119, which are provided as software or hardware functional units of the processor 11.
[0035] The head-related transfer function acquisition unit 111 acquires personalized HRTFs for the listener. These HRTFs include three types of HRTFs: an HRTF in the front direction of the listener (direction D1 in FIG. 1, also simply referred to as "front"), an HRTF in the direction directly behind the listener (direction D2 in FIG. 1, also simply referred to as "back"), and an HRTF in the zenith direction of the listener (direction D3 in FIG. 1, also simply referred to as "zenith direction" or "vertex") within the listener's median plane. In other words, the HRTFs acquired by the head-related transfer function acquisition unit 111 include head-related transfer functions in at least three directions centered around the listener within the listener's median plane.
[0036] In this example, the HRTFs acquired by the head-related transfer function acquisition unit 111 are described as HRTFs in the direction in front of the listener, the direction directly behind the listener, and the zenith direction, but this is not limiting. The HRTFs acquired by the head-related transfer function acquisition unit 111 may be HRTFs in at least three different directions centered on the listener within the median plane of the listener. For example, an HRTF in the direction in front of the listener (direction D1 in FIG. 1) may have an elevation angle of approximately 0 degrees. Similarly, an HRTF in the direction directly behind the listener (direction D2 in FIG. 1) may have an elevation angle of approximately 180 degrees. An HRTF in the zenith direction of the listener (direction D3 in FIG. 1) may have an elevation angle of approximately 90 degrees. That is, the head-related transfer function acquisition unit 111 acquires HRTFs in at least three directions around the listener within the median plane of the listener.
[0037] The estimation unit 112 estimates the notches and peaks of the head-related transfer functions for each angle in the median plane of the listener based on the parameter information of the notches and peaks of the head-related transfer functions acquired by the head-related transfer function acquisition unit 111.
[0038] The interaural difference information acquisition unit 113 acquires the interaural difference information of the listener.
[0039] The head impulse response generating unit 114 generates a head impulse response in any direction in three-dimensional space centered on the listener, based on the estimation results of the notches and peaks for each angle in the median plane estimated by the estimation unit 112 and the interaural difference information acquired by the interaural difference information acquiring unit 113.
[0040] More specifically, the estimation unit 112 obtains a regression equation for the notch and peak parameters with the elevation angle in the listener's median plane as an independent variable, and estimates the notches and peaks in the listener's median plane by calculating the notch and peak parameters for each elevation angle based on the regression equation. The head impulse response generation unit 114 calculates a head impulse response in the median plane of the listener based on the notch and peak parameters in the median plane of the listener calculated by the estimation unit 112. The head impulse response generation unit 114 generates a head impulse response in any direction in a three-dimensional space centered on the listener by adding at least one of the interaural time difference and the interaural level difference for each lateral angle of the listener based on the calculated head impulse response in the median plane of the listener and the interaural difference information.
[0041] The three-dimensional head-related transfer function generating unit 115 generates a three-dimensional head-related transfer function of the listener by performing time-frequency conversion on the head-related impulse response generated by the head-related impulse response generating unit 114.
[0042] The hearing device inverse transfer function acquisition unit 116 acquires the inverse transfer function of the hearing device used by the listener. The hearing device inverse transfer function convolution unit 117 performs a convolution operation of the inverse transfer function of the hearing device used by the listener, acquired by the hearing device inverse transfer function acquisition unit 116, with the head impulse response in any direction in three-dimensional space centered on the listener, generated by the head impulse response generation unit 114.
[0043] The sound source signal acquisition unit 118 acquires a sound source signal. The sound source signal convolution unit 119 performs a convolution operation on the sound source signal acquired by the sound source signal acquisition unit 118 with the result of the convolution operation by the hearing device inverse transfer function convolution unit 117 . The sound source signal convolution unit 119 outputs a signal after the convolution operation of the sound source signal (that is, a sound source signal after sound image localization).
[0044] [Generating personalized HRTFs for any direction in the whole sky] The procedure for generating personalized HRTFs for any direction in the whole sky, which is performed by the processor 11 of the head-related transfer function generation device 1 described above, will be described in detail with reference to FIGS.
[0045] (Algorithm Overview) As mentioned above, in the median plane, sound image localization accuracy equivalent to that of actually measured HRTFs can be achieved by reproducing only N1, N2, P1, and P2 out of the N (notch; the same applies in the following explanation) / P (peak; the same applies in the following explanation) contained in the HRTF.In addition, by adding interaural difference information (ITD, ILD) to the HRTF in the median plane, it is possible to control the sound image in any three-dimensional direction.We attempted to generate personalized HRTFs for any direction in the whole sky using the following algorithm.
[0046] 1) Personalized HRTFs for elevation angles of the median plane of 0° (front), 90° (zenith), and 180° (back) are generated using the PNP model. 2) The N / P parameters of HRTFs in any direction in the median plane are estimated by linear regression of the parameters of personalized HRTFs in the front, zenith, and rear directions. 3) The HRTF generated using the estimated value in 2) is added with ITD and ILD according to the lateral angle to generate a personalized HRTF for any direction in the whole sky. The details of the algorithm are given below.
[0047] (Generating personalized HRTFs for front, zenith, and rear) The PNP model is an HRTF model that uses the N2 frequency as an independent variable and expresses the other N / P parameters as dependent variables. Of the N / P parameters, frequency is calculated using a regression equation, and level and Q are treated as constants. The front PNP model and rear PNP model are calculated by known conventional methods. In addition to these, a PNP model in the zenith direction is constructed. Furthermore, we developed software that allows listeners to personalize HRTFs while listening to the sound using the PNP model (Figure 7). FIG. 7 shows an example of a software screen that allows a listener to personalize HRTFs while listening to the sound using the PNP model.
[0048] (Generation of personalized HRTFs for any direction in the median plane) FIG. 8 is a diagram showing an example of the distribution of N1 frequencies in the median plane. Using the personalized HRTFs for the front, zenith, and rear, we generated personalized HRTFs for any direction in the median plane of the upper hemisphere. The N1 frequency in the median plane increases as the sound source moves from front to top and decreases as the sound source moves from top to back. On the other hand, the N2 frequency increases as the sound source moves from front to top, but the change between top and back is small (Figure 8). From these relationships, the N / P parameters of the HRTFs in any direction in the anterior half of the median plane of the upper hemisphere (i.e., between an elevation angle of 0 degrees and 90 degrees) were calculated by linear regression using the N / P parameters of the personalized HRTFs in the front and zenith directions, with the elevation angle as the explanatory variable. For the posterior half of the median plane (i.e., between an elevation angle of 90 degrees and 180 degrees), the N / P parameters of the personalized HRTFs in the zenith and rear directions were calculated in a similar manner. Figures 9 to 11 show examples of the N / P frequency, level, and Q for a given subject's median plane of the upper hemisphere at any elevation angle. FIG. 9 is a diagram showing an example of N / P parameters (frequency of N / P) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject. FIG. 10 is a diagram showing an example of the N / P parameter (level) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject. FIG. 11 is a diagram showing an example of the N / P parameter (Q) at an arbitrary elevation angle of the median plane of the upper hemisphere of a certain subject.
[0049] (Generation of personalized HRTFs in any direction across the entire sky) Personalized HRTFs for any direction in the whole sky are generated by adding interaural difference information to personalized HRTFs for the median plane of the upper hemisphere.
[0050] Regarding ITD, if the head is considered as a sphere, it can be geometrically expressed by equation (1).
[0051]
number
[0052] Here, φ is the azimuth angle [rad] and D is the distance between the ears [m]. To reflect ITD due to actual head shape, ITD was calculated for the front half of the head in four directions in the horizontal plane (lateral angle α: 0, 30, 60, 90°) from the measured HRIRs of 18 Japanese adult subjects (Fig. 12).
[0053] FIG. 12 is a diagram showing an example of the distribution of ITDs according to actual head shapes. The ITD increased almost linearly with the lateral angle, so we added the ITD using equation (2).
[0054]
number
[0055] ILD varies depending on the frequency even if the lateral angle is the same. When the sound source is broadband white noise, the sound image direction is in front when ILD is 0 dB and to the side when ILD is ±10 dB, and changes in an almost linear relationship between them. ILD is added using equation (3). However, preliminary experiments have determined that the maximum ILD value is 9 dB.
[0056]
number
[0057] [3D Rendering Toolkit] 13 is a diagram showing an example of the procedure for generating a personalized head-impulse response according to this embodiment. In the following description, the HRTF personal adaptation software executed by the processor 11 according to this embodiment is also referred to as a 3D rendering toolkit. The 3D rendering toolkit generates personalized head-based impulse responses (HRIRs) for any direction in the sky and convolves them with up to 48 sound source signals. The processing steps are as follows:
[0058] (Step S101) Generate personalized HRTFs for the front, zenith, and rear using the PNP model. The processor 11 acquires the parameters of the personalized HRTFs for the front, zenith, and rear.
[0059] (Steps S102 to S105) The processor 11 generates a regression equation for the N / P parameters using the elevation angle of the median plane as an explanatory variable from the generated N / P parameter information of the personalized HRTFs for the three directions.
[0060] (Step S106) Processor 11 acquires the target direction (azimuth angle φ, elevation angle θ) and relative level of each sound source. Processor 11 converts the azimuth angle φ and elevation angle θ into a lateral angle α and an elevation angle β using equations (4) and (5).
[0061]
number
[0062]
number
[0063] The processor 11 calculates the N / P parameter at the target climb angle using the regression equation obtained in steps S104 and S105.
[0064] (Step S107) Processor 11 sets the N / P parameter at the calculated target climb angle in the peaking filter to generate an HRIR at the target climb angle.
[0065] (Step S108) The processor 11 adds an ITD and an ILD corresponding to the target lateral angle, and generates a personalized HRIR for an arbitrary direction in a three-dimensional space centered on the listener.
[0066] (Step S109) The processor 11 acquires the inverse transfer function of the hearing device (e.g., headphones) used by the listener. The processor 11 performs an inverse filter convolution operation of the inverse transfer function of the hearing device on the head-impulse response generated in step S108 to generate a personalized HRIR.
[0067] (Step S110) The processor 11 acquires a sound source signal. The processor 11 convolves the acquired sound source signal with the personalized HRIR generated in step S109 to generate a post-sound image localization sound source signal. The processor 11 outputs the generated post-sound image localization sound source signal (for example, to headphones) to reproduce the post-sound image localization sound source signal.
[0068] 14 is a diagram showing an example of a setting screen of the 3D rendering toolkit of this embodiment. In addition to the functions described above, the 3D rendering toolkit has the following functions. - Output a file of the signal convolved with the source signal and personalized HRIR. - Personalized HRIR file output for each sound source direction. -File output of project information such as N / P parameters of personalized HRTFs for the front, zenith, and rear, and sound source information. Graphic display of sound source location.
[0069] [Verification of sound localization accuracy of personalized HRTF] We verified the sound localization accuracy of personalized HRTFs in the horizontal, median, and transverse planes generated by a 3D rendering toolkit.
[0070] (Sound image localization in the horizontal plane) The experiment was conducted in a soundproof room. HRTFs personalized by the subject using the PNP model and the subject's own HRTF were used. The target directions were seven directions in the horizontal plane on the right half (azimuth angle: 0-180°, 30° intervals). The sound source signal was broadband white noise (200Hz-17kHz), and the presentation time was 1.2 seconds (including 0.1 seconds for both the rise and fall). Stimuli, in which the HRIRs were convolved with the sound source signal, were presented to the subjects via FEC (Free Air Equivalent Coupling to the Ear) headphones. The headphone transfer function was not corrected. The presented sound pressure was 63 dB, and each stimulus was presented 10 times in random order. The subjects were two men in their twenties with normal hearing, who responded to the azimuth and elevation angles of the sound image using a mapping method.
[0071] (1) Answer distribution FIG. 15 is a diagram showing an example of the responses of each subject regarding the azimuth angle. FIG. 16 is a diagram showing an example of the responses of each subject regarding the angle of elevation. Regarding azimuth, subject 1's personal HRTF generally gave answers near the target direction, but at a target azimuth angle of 60°, he gave answers near 90-120°. The answers from the personalized HRTF were nearly identical to those from the personal HRTF, but he made one incorrect front / rear judgment at 0° and two at 30°. Subject 2's personal HRTF showed an inverted S-shaped curve, with answers at target azimuth angles of 60, 90, and 120° distributed near 90°. The personalized HRTF generally gave answers closer to the target direction than the personal HRTF. However, at 30°, he sometimes gave answers near the target direction and sometimes to the side. Regarding elevation angle, subject 1 responded that both his own HRTF and personalized HRTF were near 0°, i.e., near the horizontal plane, at all target azimuth angles. Subject 2 responded that both his own HRTF and personalized HRTF were near 0° at all target azimuth angles, but with personalized HRTF the sound image rose slightly at 0° and 30°.
[0072] (2) Mean lateral angle error Figure 17 shows an example of the average lateral angle error used to evaluate the accuracy of sound image localization in the left-right direction. The lateral angle was calculated using equation (4). The average values for seven directions were 7.9° for subject 1's own HRTF and 10.2° for the personalized HRTF, and 17.6° for subject 2's own HRTF and 12.4° for the personalized HRTF. No significant difference was observed between subject 1's own HRTF and the personalized HRTF for either subject 1 or 2.
[0073] (3) Misjudgment rate before and after Figure 18 shows an example of the front-to-back misidentification rate used to evaluate the accuracy of sound image localization in the front-to-back direction. The average value for seven directions was 0.08 for both subject 1's own HRTF and personalized HRTF, while subject 2's own HRTF was 0.23 and the personalized HRTF was 0.05. The front-to-back misidentification rate for personalized HRTF was equal to or lower than that for subject 1's own HRTF.
[0074] (4) Average elevation angle error Figure 19 shows an example of the average elevation angle error used to evaluate sound image localization accuracy in the vertical direction. The average values of the seven directions of the personalized HRTFs are larger than those of the subject's own HRTF, at 1.2° for subject 1 and 3.2° for subject 2. However, the errors at azimuth angles of 0 and 30° for the personalized HRTF of subject 2 were larger than those in the other directions, at 9.5° and 8.0°, respectively.
[0075] (5) In-head localization rate Figure 20 shows an example of the in-head localization rate. For subject 1, in-head localization did not occur at any target azimuth angle for either the subject's own HRTF or the personalized HRTF. For subject 2, the rate was 0.20 for the subject's own HRTF at a target azimuth angle of 180°. For the personalized HRTF, the rates were 0.10, 0.20, and 0.30 at target azimuth angles of 0, 30, and 180°, respectively.
[0076] (Sound image localization in the median plane) The experimental method was the same as that in the horizontal plane, except that the target direction was seven directions in the median plane of the upper hemisphere (elevation angle: 0-180°, 30° intervals).
[0077] (1) Answer distribution FIG. 21 is a diagram showing an example of the results of converting the azimuth angles answered by the subjects into lateral angles. FIG. 22 is a diagram showing an example of the result of converting the elevation angle answered by the subject into an ascent. Regarding the lateral angle, regardless of the subject or type of HRTF, the responses were near the target lateral angle of 0°, i.e., near the median plane, for all elevation angles. Regarding the angle of elevation, subject 1, using his own HRTF, answered near the target direction at target elevation angles of 0, 30, and 180°, but answered near 90° at 60 and 150°. At 90 and 120°, he answered 90-150°. With the personalized HRTF, he answered near the target direction at 0 and 180°. With subject 2, using his own HRTF, he answered near the target direction at target elevation angles of 60 and 180°. With the personalized HRTF, there was overall variation in answers, but the answers were generally near the target direction.
[0078] (2) Mean lateral angle error Figure 23 shows an example of the average lateral angle error used to evaluate the accuracy of sound image localization in the left-right direction. The average value for seven directions was 0.1° for both subject 1's own HRTF and personalized HRTF, and 0.7° for subject 2's own HRTF and 0.6° for personalized HRTF. Both subjects 1 and 2 localized the sound almost within the median plane.
[0079] (3) Misjudgment rate before and after Figure 24 shows an example of the front-to-back misjudgment rate for evaluating the accuracy of sound image localization in the front-to-back direction. The average values for seven directions were 0.08 for subject 1's own HRTF and 0.03 for the personalized HRTF, while subject 2's own HRTF was 0.18 and 0.17 for the personalized HRTF. No significant difference was observed between subject 1's own HRTF and the personalized HRTF for either subject. However, a difference of about 0.1 was observed between the subjects.
[0080] (4) Average elevation angle error Figure 25 shows an example of the average elevation angle error used to evaluate sound image localization accuracy in the vertical direction. The average values for seven directions were 21.0° for subject 1's own HRTF and 26.0° for the personalized HRTF, and 23.5° for subject 2's own HRTF and 22.3° for the personalized HRTF. No significant difference was observed between subject 1's own HRTF and the personalized HRTF for either subject 1 or 2. Looking at each direction, there was a tendency for the error to be larger in the upward direction than in the other directions.
[0081] (5) In-head localization rate Figure 26 shows an example of in-head localization rate. For subject 1, in-head localization did not occur with either the subject's own HRTF or the personalized HRTF at any target elevation angle. For subject 2, in-head localization did not occur with the subject's own HRTF, but the in-head localization rates for the personalized HRTF were 0.10, 0.10, 0.40, and 0.50 at target azimuth angles of 0, 120, 150, and 180°, respectively.
[0082] (Sound image localization in the cross section) The experimental method was the same as for the horizontal plane, except that the target directions were seven directions in the transverse plane of the upper hemisphere (lateral angles: -90 to +90°, 30° intervals). However, the subject's HRTFs were only three directions (lateral angles: -90, 0, 90°).
[0083] (1) Answer distribution FIG. 27 shows an example of the results of converting the azimuth angles answered by the subjects into lateral angles. FIG. 28 is a diagram showing an example of the result of converting the elevation angle answered by the subject into an ascent. Regarding lateral angles, both subjects 1 and 2 perceived the direction of the target for both their own HRTF and personalized HRTF. However, the variance of subject 2's personalized HRTF responses was greater than that of their own HRTF. For the elevation angle, because it is not possible to define a target lateral angle of ±90°, the results are shown for 0° for the subject's own HRTF and 0, ±30, and ±60° for the personalized HRTF. In response to a stimulus with a target lateral angle of 0°, subject 1 responded with 135-165° for his own HRTF. Subject 2 responded with around 90° and 160-180°. The response distribution of the personalized HRTFs was nearly the same as that of the subject's own HRTF for both subjects 1 and 2.
[0084] (2) Mean lateral angle error Figure 29 shows an example of the average lateral angle error for evaluating the accuracy of sound image localization in the left-right direction. The average values for three directions were 6.3° for subject 1's own HRTF and 12.0° for the personalized HRTF, and 5.8° for subject 2's own HRTF and 17.0° for the personalized HRTF. The average lateral angle error for the personalized HRTFs was larger than that for both subjects 1 and 2. The average lateral angle errors for the seven directions for the personalized HRTFs were 15.8° and 17.2° for subjects 1 and 2, respectively.
[0085] (3) Average elevation angle error Figure 30 shows an example of the average elevation angle error used to evaluate sound image localization accuracy in the vertical direction. The average values for three directions were 25.6° for subject 1's own HRTF and 33.6° for the personalized HRTF, while the average values for subject 2's own HRTF were 14.4° and 22.7° for the personalized HRTF. For both subjects 1 and 2, the average elevation angle error for the personalized HRTF was approximately 8° larger than the own HRTF. The average elevation angle errors for the personalized HRTF in seven directions were 28.5° and 35.3° for subjects 1 and 2, respectively.
[0086] (4) Intrahead localization rate In neither case did intrahead localization occur.
[0087] So far, we have explained the localization accuracy using the subject's own personalized HRTF without headphone transfer function correction. Next, we will explain the localization accuracy using the subject's own personalized HRTF with headphone transfer function correction.
[0088] FIG. 31 is a diagram showing an example of the localization accuracy using the subject's personalized HRTF when the headphone transfer function is corrected. As shown in the figure, when the headphone transfer function was corrected, an improvement in the localization accuracy was observed for subject 2 regarding the elevation angle of the cross section.
[0089] The head-related transfer function generation device 1 of this embodiment constructs a PNP model for the zenith direction in addition to the front and rear directions, and generates personalized HRTFs for any direction in the median plane. Sound image localization experiments using personalized HRTFs generated by the head-related transfer function generation device 1 have shown that accurate sound image localization is possible in any direction in the horizontal plane. The head-related transfer function generation device 1 of this embodiment can accurately generate head-related transfer functions for any direction in the entire sky three-dimensional space.
[0090] Although the embodiments of the present invention have been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and appropriate modifications can be made without departing from the spirit of the present invention. The configurations described in the above-described embodiments may also be combined.
[0091] Each unit included in the head-related transfer function generating device 1 in the above embodiment may be realized by dedicated hardware, or may be realized by a memory and a microprocessor.
[0092] In addition, each part of the head-related transfer function generating device 1 may be composed of a memory and a CPU (central processing unit), and the functions may be realized by loading a program for realizing the functions of each part of the head-related transfer function generating device 1 into the memory and executing the program.
[0093] Furthermore, a program for realizing the functions of each unit of the head-related transfer function generation device 1 may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform processing by each unit of the control unit. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.
[0094] Furthermore, if a WWW system is used, the "computer system" also includes the homepage provision environment (or display environment). "Computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording media" also includes devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or over communication lines like telephone lines, and devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients. Furthermore, the programs may be programs that implement some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system. [Explanation of symbols]
[0095] 1...head-related transfer function generation device, 111...head-related transfer function acquisition unit, 112...estimation unit, 113...interaural difference information acquisition unit, 114...head impulse response generation unit, 115...three-dimensional head-related transfer function generation unit, 116...hearing device inverse transfer function acquisition unit, 117...hearing device inverse transfer function convolution unit, 118...sound source signal acquisition unit, 119...sound source signal convolution unit
Claims
1. a head-related transfer function acquisition unit that acquires head-related transfer functions in at least three directions centered on the listener within a median plane of the listener; an estimation unit that estimates notches and peaks of the head-related transfer functions for each angle in the median plane of the listener based on parameter information of the notches and peaks of the head-related transfer functions acquired by the head-related transfer function acquisition unit; a binaural difference information acquisition unit that acquires binaural difference information of the listener; a head-impulse response generating unit that generates a head-impulse response in an arbitrary direction in a three-dimensional space centered on the listener, based on the estimation results of the notch and the peak for each angle in the median plane estimated by the estimation unit and the interaural difference information acquired by the interaural difference information acquisition unit; a three-dimensional head-related transfer function generation unit that generates a three-dimensional head-related transfer function of the listener by performing a time-frequency transform on the head-related impulse response generated by the head-related impulse response generation unit; A head-related transfer function generating device comprising:
2. The three directions are the front direction, the rear direction, and the zenith direction of the listener. The head-related transfer function generating device according to claim 1 .
3. The estimation unit determining a regression equation for the notch and peak parameters using the elevation angle in the median plane of the listener as an independent variable; estimating the notch and peak in the median plane of the listener by calculating parameters of the notch and peak for each elevation angle based on the regression equation; The head impulse response generation unit calculating a head impulse response in the median plane of the listener based on the parameters of the notch and peak in the median plane of the listener calculated by the estimation unit; Based on the calculated head impulse response in the median plane of the listener and the interaural difference information, at least one of the interaural time difference and the interaural level difference for each lateral angle of the listener is added to generate a head impulse response in any direction in a three-dimensional space centered on the listener. The head-related transfer function generating device according to claim 1 or 2.
4. a hearing device inverse transfer function convolution unit that performs a convolution operation of the inverse transfer function of a hearing device used by the listener on the head impulse response in any direction in a three-dimensional space centered on the listener, the head impulse response being generated by the head impulse response generation unit; The head-related transfer function generating device according to claim 1 , further comprising:
5. a sound source signal convolution unit that performs a convolution operation of a sound source signal on the result of the convolution operation by the hearing device inverse transfer function convolution unit; The head-related transfer function generating device according to claim 4 , further comprising:
6. On the computer, Acquiring head-related transfer functions in at least three directions centered on the listener within the median plane of the listener; estimating notches and peaks of the head-related transfer functions for each angle in the median plane of the listener based on parameter information indicating notches and peaks of the acquired head-related transfer functions; obtaining interaural difference information of the listener; generating a head-impulse response for any direction in a three-dimensional space centered on the listener based on the estimated notch and peak for each angle in the median plane and the acquired interaural difference information; generating a three-dimensional head-related transfer function of the listener by performing a time-frequency transform on the generated head-related impulse response; A program to execute.
7. Acquiring head-related transfer functions in at least three directions centered on the listener within the median plane of the listener; estimating notches and peaks of the head-related transfer functions for each angle in the median plane of the listener based on parameter information indicating notches and peaks of the acquired head-related transfer functions; obtaining interaural difference information of the listener; generating a head-impulse response for any direction in a three-dimensional space centered on the listener based on the estimated notch and peak for each angle in the median plane and the acquired interaural difference information; generating a three-dimensional head-related transfer function of the listener by performing a time-frequency transform on the generated head-related impulse response; A head-related transfer function generation method having the following.
Citation Information
Patent Citations
Sound image localizing device
JP2002218598A
Sound image locating device
JP2006203850A
Head-related transfer function interpolation device
JP2008312113A
Head transfer function selection apparatus, head transfer function selection method, head transfer function selection program and voice reproducer
JP2016201723A
Signal processor, acoustic processing system, and program
JP2020088632A