Head-to-head transfer function generation device, head-to-head transfer function generation method, head-to-head transfer function generation program
By modeling critical sound frequencies and reducing unnecessary calculations, the HRTF generation device addresses the computational burden of high-precision HRTF generation, achieving faster and effective spatial audio reproduction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-24
AI Technical Summary
High-precision HRTF generation for spatial audio requires extensive computation due to the high-bandwidth acoustic characteristics needed for 360-degree audio, leading to lengthy computation times.
The HRTF generation device models sound frequencies or frequency bands that significantly affect spatial sound perception, simplifying calculations by focusing on these critical bands and reducing the number of observation points and frequency ranges analyzed.
This approach significantly reduces computation time by approximately 60% while maintaining the quality of spatial audio perception, enabling efficient HRTF generation and three-dimensional sound reproduction.
Smart Images

Figure 2026052462000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a head-transfer function generation device, a head-transfer function generation method, and a head-transfer function generation program for generating a user's head-transfer function. [Background technology]
[0002] To achieve spatial audio using headphones, a method that generates a virtual sound image using the Head-Related Transfer Function (HRTF) is commonly employed. However, measuring the HRTF requires the subject to be confined to an anechoic chamber for an extended period and to undergo long-term measurements using speakers and earphones. As an alternative to actual measurements, there is a technique that uses a smartphone app, such as Non-Patent Document 1, to photograph the shape of the ear and analyze auditory characteristics. Furthermore, methods for acquiring the ear and head using a 3D scanner and generating the HRTF from three-dimensional data are also being researched. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] SONY, “360 Reality Audio: ‘A new experience where sound rains down from all directions.’” [Retrieved July 2, 2024], Internet<https: / / www.sony.jp / headphone / special / 360_Reality_Audio / > . [Overview of the project] [Problems that the invention aims to solve]
[0004] However, in high-precision HRTF generation using acoustic simulation, the acoustic characteristics required for 360-degree and spatial audio are high-bandwidth (8kHz-16kHz). Therefore, an enormous amount of computation is required to complete the estimation of the HRTF. Accordingly, the present invention aims to shorten the computation time for generating an HRTF that provides users with spatial audio perception. [Means for solving the problem]
[0005] The head-related transfer function (HRTF) generation device of the present invention generates the user's head-related transfer function. The head-related transfer function generation device comprises a modeling means for modeling at a predetermined frequency or frequency band, and an HRTF generation unit for determining the HRTF based on the model. The predetermined frequency or frequency band is a frequency or frequency band in which changes in sound within that frequency range affect the user's spatial sound perception. [Effects of the Invention]
[0006] The head-related transfer function (HRTF) generation device of the present invention models the sound using frequencies or frequency bands in which changes in sound affect the user's spatial sound perception. Since it is possible to simplify the calculation of frequencies that have little effect on spatial sound perception, the computation time for generating the HRTF can be shortened. [Brief explanation of the drawing]
[0007] [Figure 1] A diagram showing an example of the functional configuration of the head-related transfer function generation device and content device of the present invention. [Figure 2] A diagram showing an example of the processing flow of a head-related transfer function generator. [Figure 3] A diagram showing an example of the processing flow of a content device. [Figure 4] A diagram showing the relationship between three-dimensional head data and observation points. [Figure 5] This figure shows an example of the results of simulating the HRTF across all frequency bands being analyzed. [Figure 6] This figure shows the notch frequencies and approximate curves of two notches spaced 20 degrees apart relative to the left ear. [Figure 7]A diagram showing simulation results of three points near the peak frequency and the frequency band near the notch frequency for the left ear. [Figure 8] A diagram showing the HRTF of the left ear with (α,β)=(-60,10). [Figure 9] A diagram showing an example of the functional configuration of a computer.
Mode for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present invention will be described in detail. Components having the same function are denoted by the same reference numerals, and redundant descriptions are omitted.
Examples
[0009] FIG. 1 shows an example of the functional configuration of the head-related transfer function generation device and the content device of the present invention. FIG. 2 shows an example of the processing flow of the head-related transfer function generation device, and FIG. 3 shows an example of the processing flow of the content device. The head-related transfer function generation device 100 generates a user's head-related transfer function (HRTF: Head-Related Transfer Function). The head-related transfer function generation device 100 includes a modeling means 135 that performs modeling at a predetermined frequency or frequency band, an HRTF generation unit 170 that obtains an HRTF based on the model, and a recording unit 190. The predetermined frequency or frequency band is a frequency or frequency band in which a change in sound in the frequency region affects the user's stereo perception. The "frequency or frequency band in which a change in sound affects the user's stereo perception" may be, for example, the frequency of the peak of the HRTF, the frequency of the edge, or the frequency band near the peak, the frequency band near the edge. However, it is not limited to the peak and the edge, and any frequency or frequency band that is likely to affect the user's stereo perception may be used.
[0010] The following provides a more detailed explanation. The head transfer function generation device 100 may also include a three-dimensional data acquisition unit 110 and a data formatting unit 120. The three-dimensional data acquisition unit 110 acquires three-dimensional ear data, which is three-dimensional data of the shape of the user's ear (S110). The three-dimensional data of the shape of the user's ear can be measured using a 3D scanner or the like. The data formatting unit 120 obtains three-dimensional head data 121 by combining the three-dimensional ear data and a dummy head (S120). A pre-selected dummy head can be used as the dummy head. When obtaining the three-dimensional head data 121, the data formatting unit 120 may also perform processing such as supplementing data if there is any missing data in the three-dimensional ear data.
[0011] The modeling means 135 includes a condition setting unit 130, a simulation unit 140, an extraction unit 150, and a notch frequency range calculation unit 160. The condition setting unit 130 sets the observation points and frequencies or frequency bands to be simulated. In the initial state, the condition setting unit 130 sets the pre-selected observation points and the entire frequency band to be analyzed as initial settings to the simulation unit 140 (S131). For example, the "pre-selected observation points" in the initial settings may include multiple observation points on the horizontal plane front side and multiple observation points in the upward direction from one of the selected observation points on the horizontal plane front side.
[0012] Figure 4 shows the relationship between the three-dimensional head data and the observation points. The observation points are located on a sphere centered on the three-dimensional head data 121. The black circles (filled black) and white circles (circles with only outlines) in Figure 4 are the observation points. Hereafter, the filled black circles will be called "black circles" and the circles with only outlines will be called "white circles". In Figure 4, only the upper hemisphere is shown, but observation points may also be placed on the lower sphere. The observation points are identified by the angle viewed from the three-dimensional head data 121. The horizontal angle is denoted as the lateral angle α, and the front direction of the three-dimensional head data 121 is set to α = 0°. There are positions α = -90° and α = 90° directly to the side of the three-dimensional head data 121. The angle viewed from the lateral axis on a plane perpendicular to the lateral axis of the three-dimensional head data 121 is denoted as the ascending angle β, and the front direction of the three-dimensional head data 121 is set to β = 0°. In other words, the angle of the observation point in front of the three-dimensional head data 121 is (α,β)=(0°,0°). Below, the "°" symbol will be omitted and expressed as (α,β)=(0,0). The observation point directly above the three-dimensional head data 121 is (α,β)=(0,90). Also, the observation point directly behind the three-dimensional head data 121 is (α,β)=(0,180). "Multiple observation points on the horizontal front side" are observation points where -90≦α≦90 and β=0. "Multiple observation points in the upward direction" from one of these observation points are observation points where α is constant and 0<β≦180.
[0013] The black circles in Figure 4 show an example where, as the initial "pre-selected observation points," multiple observation points are selected on the front side of the horizontal plane, and multiple observation points are selected in the upward direction from one selected observation point on the front side of the horizontal plane (α=0, β=0). The white circles represent observation points other than those pre-selected. In the example in Figure 4, all observation points on the front side of the horizontal plane are selected, but only some may be selected. Similarly, all observation points in the upward direction are selected, but only some may be selected. In Figure 4, observation points are shown at 10-degree intervals, but they may be at different angles. For example, observation points may be set at 5-degree intervals. If observation points are set at 5-degree intervals in the lateral angle α and at 5-degree intervals in the upward angle β, in the case of a hemisphere like in Figure 4, there will be a total of 1297 (=35×37+2) observation points. Furthermore, if the initial "pre-selected observation points" include all observation points on the horizontal plane front side (37 points) and all observation points in the upward direction from one selected observation point on the horizontal plane front side (α=0, β=0) (36 points), then the "pre-selected observation points" will be 73 points.
[0014] The simulation unit 140 uses three-dimensional head data 121, which combines the shape of the user's ear with a dummy head, to calculate the HRTF for a set frequency or frequency band at a set observation point, and outputs the simulation results. In other words, in the initial settings, the simulation unit 140 calculates the HRTF for all frequency bands of the analysis target at pre-selected observation points (S141). Figure 5 shows an example of the simulation results for the HRTF of all frequency bands of the analysis target. The horizontal axis is frequency (Hz), and the vertical axis is amplitude (dB). The simulation is performed with a sound source placed in the ear of the three-dimensional head data 121 and sound collected at the observation point.
[0015] The extraction unit 150 extracts peaks and notches from the simulation results (S150). A peak refers to a frequency where the amplitude of the HRTF is locally large. A notch refers to a frequency where the amplitude of the HRTF is locally small. When referring to them with frequency in mind, they are sometimes called "peak frequencies" and "notch frequencies." For example, two frequencies can be extracted as peaks and two frequencies as notches. In Figure 5, black circles represent peaks and white circles represent notches.
[0016] The notch frequency range calculation unit 160 determines the frequency range in which the notch exists from the simulation results (S160). The notch frequency range calculation unit 160 predicts the change in notch frequency dependent on the upward angle from the notches extracted at one selected observation point on the front side of the horizontal plane (for example, observation point (α=0,β=0)) and multiple observation points in the upward direction (for example, observation points (α=0,β=0),..., observation point (α=0,β=180)), and determines the frequency range in which the notch exists. Figure 6 shows the notch frequencies and approximate curves of two notches spaced 20 degrees apart for the left ear, such as observation point (α=-5,β=0), observation point (α=-5,β=20), observation point (α=-5,β=40),..., observation point (α=-5,β=180). The first notch is called "notch 1" and is indicated by a black circle. The second notch is called "Notch 2" and is indicated by a white circle. The notch frequency changes depending on the angle of ascent.
[0017] The curves shown in Figure 6 are approximation curves for the notch frequency of notch 1 and notch frequency of notch 2. The approximation curves can be expressed, for example, as a quartic equation in terms of the rise angle β. However, different approximation equations are also acceptable; any appropriate equation should be chosen. An example of a quartic equation obtained by simulation is as follows: Left ear f N1 (β) = 8.650 × 10 -6 ×β 4 -5.967 × 10 -3 ×β 3 +1.181 × β 2 -6.857 × 10 × β +1.486 × 102 +f N1 (0) f N2 (β) = 3.187×10 -6 ×β 4 +1.439×10 -3 ×β 3 -5.742×10 -1 ×β 2 +4.486×10×β +9.615×10 + f N2 (0) Right ear f N1 (β) = 5.463×10 -6 ×β 4 -4.856×10 -3 ×β 3 +1.115×β 2 -7.115×10×β +1.748×10 2 +f N1 (0) f N2 (β) = -2.276×10 -6 ×β 4 +9.919×10 -3 ×β 3 -1.461×β 2 +7.427×10×β -1.399×10 + f N2 (0) Here, f N1 (β) is the approximation formula of notch 1 at the rising angle β, f N2 (β) is the approximation formula of notch 2 at the rising angle β, f N1 (0) is the frequency of notch 1 on the front side of the horizontal plane (rising angle β = 0), f N2 (0) is the frequency of notch 2 on the front side of the horizontal plane (rising angle β = 0).
[0018] Figure 7 shows the simulation results for three points near the peak frequency and the frequency band near the notch frequency for the left ear. The frequency band near the notch frequency can be defined as a frequency band of 0.2 octaves in the lower frequency direction and 0.2 octaves in the higher frequency direction. However, it is sufficient if the vicinity of the notch can be simulated even with frequency bands of different widths. If simulation results exist in the initial settings, the condition setting unit 130 sets the simulation unit 140 with observation points selected from observation points other than those selected in advance, frequencies based on the peaks extracted by the extraction unit 150 based on the simulation results in the initial settings, and frequency bands based on the range of frequencies in which the notch exists, as determined by the notch frequency range calculation unit 160 (S132).
[0019] Figure 7 shows the selection of multiple frequencies (three frequencies) based on the peaks at multiple observation points (observation point (α=-5, β=0)) on the front of the selected horizontal plane, as "peak-based frequencies." Since the peak frequency does not change significantly depending on the rise angle, there is no need to derive an approximation formula like that for the notch frequency; for the same lateral angle, the peak frequency extracted from the simulation results at the multiple observation points on the front of the selected horizontal plane can be used. In addition, as "frequency bands based on the range of frequencies in which notches exist, as determined by the notch frequency range calculation unit," frequency bands of 0.2 octaves in the lower and higher directions are set based on the approximation curve (approximation formula) shown in Figure 6.
[0020] The simulation unit 140 performs simulations for all observation points other than the pre-selected observation points at the frequencies and frequency bands set by the condition setting unit 130, and outputs the simulation results (S142). As a result, for the initial setting of "pre-selected observation points," simulation results are obtained for all frequency bands of the analysis target. For observation points selected from observation points other than the pre-selected observation points, simulation results are obtained for frequency bands based on the frequencies extracted by the extraction unit 150 and the frequency ranges in which notches exist, as determined by the notch frequency range calculation unit 160.
[0021] The HRTF generation unit 170 calculates the HRTF based on the simulation results and records it in the recording unit 190 (S170). For example, as shown in Figure 7, the HRTF generation unit 170 calculates the HRTF based on the simulation results for three frequencies based on the peak and a frequency band of 0.2 octaves in the lower and higher directions based on the frequency range in which the notch exists.
[0022] More specifically, for example, the HRTF generation unit 170 substitutes the amplitudes at three frequencies as "peak-based frequencies" and the amplitudes at a set frequency band as "frequency band based on the range of frequencies in which notches exist, as determined by the notch frequency range calculation unit," and converts them into an impulse response using an IIR filter (Infinite Impulse Response filter) to obtain a head-related impulse response (HRIR). For example, a 16th-order IIR filter can be used. Then, the HRIR is Fourier transformed to obtain the head-related transfer function (HRTF). Note that HRIR is a time-domain representation of the HRTF.
[0023] Figure 8 shows the HRTF for the left ear at (α,β)=(-60,10). The white circles represent the simulation results for the frequency and frequency band set by the condition setting unit 130, as determined by the simulation unit 140. Based on these simulation results, the HRTF shown by the thick solid line in Figure 8 is obtained. The thin solid line represents the HRTF obtained when simulating across the entire frequency band of the analysis target. There are differences, but they do not affect the user's perception of spatial sound.
[0024] The head-related transfer function (HRTF) generation device of the present invention models using frequencies or frequency bands in which changes in sound affect the user's spatial sound perception. Since it is possible to simplify the calculation of frequencies that have little effect on spatial sound perception, the computation time for generating the HRTF can be shortened. For example, consider a case where observation points are set every 5 degrees in a hemisphere as shown in Figure 4, and all observation points on the front of the horizontal plane (37 points) and all observation points in the upward direction (36 points) from one selected observation point on the front of the horizontal plane (α=0, β=0) are selected as the initial "pre-selected observation points". In this case, the computational amount of the present invention is a ratio of the computational amount when simulating the HRTF of all frequency bands being analyzed for all observation points (1297 points) as shown in the following equation. The percentage to calculate = (73 / 1297 + (range of frequencies where notches exist) / (total frequency band to be analyzed) × (1297 - 73) / 1297) × 100 (%) The calculation ratio depends on the initial setting, specifically the number of "pre-selected observation points," and the width of the frequency range in which the notch exists. Simulation results showed that the ratio could be reduced to approximately 42% for both ears. Therefore, the simulation results indicate that the calculation time can be reduced by about 60%.
[0025] Next, an example of the processing flow of the content device will be explained using Figure 3. The content device 200 includes a content input unit 210, a sound source direction extraction unit 220, an acoustic content generation unit 230, and a playback unit 240. The content input unit 210 acquires a combination of sound signal and direction for each sound source as content input (S210). There can be any number of sound sources, and the direction of the sound sources may move. For example, if it is the sound of birdsong, and there are several birds, the sound and direction for each bird will be acquired as content input. The sound source direction extraction unit 220 extracts the direction for each sound source (S220) and acquires the information necessary for processing in the acoustic content generation unit 230 from the HRTF recorded in the recording unit 190 of the head-related transfer function generator 100 (S221). The acoustic content generation unit 230 generates acoustic content for playback using the HRTF based on the direction of each sound source (S230). The playback unit 240 plays the acoustic content (S240). Since the audio content is generated based on the user's HRTF, it can reproduce three-dimensional sound. [Example 1]
[0026] In Example 1, the HRTF for the entire sphere in Figure 4 was determined. However, if the sound sources acquired by the content input unit 210 exist only in a specific direction, then sound content can be generated even without the HRTF for the entire sphere, at least for that content input. A modified example will be explained using Figure 2. In this modified example, a content device 200 is also used. The content input unit 210 acquires a combination of sound signal and direction for each sound source as content input (S210). The sound source direction extraction unit 220 extracts the direction for each sound source (S220).
[0027] The condition setting unit 130 selects the initial "pre-selected observation points" based on the direction of the sound source of the content, for which the direction of each sound source has been identified. In other words, it selects them based on the direction of each sound source extracted by the sound source direction extraction unit 220. The condition setting unit 130 also selects "observation points selected from observation points other than the pre-selected observation points" as observation points that were not among the "pre-selected observation points" among all the observation points necessary to generate audio content from the content. For example, in step S131, the condition setting unit 130 selects observation points within the range where HRTF is required so that the extraction unit 150 can extract peaks and notches. Then, in step S132, the condition setting unit 130 selects observation points within the range where HRTF is required that were not selected in step S131.
[0028] Even in the case of modified images, the computation time for generating the necessary HRTF can be reduced. Furthermore, it enables the reproduction of three-dimensional sound.
[0029] [Processor, program, recording medium] The functions realized by the components described herein may be implemented in a circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the functions described herein. A processor includes transistors and other circuits and is considered a circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in memory.
[0030] In this specification, circuitry, unit, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0031] If the hardware is a processor that is considered to be a type of circuitry, then the circuitry, means, or unit is a combination of hardware and software used to constitute the hardware and / or processor.
[0032] The various processes described above can be carried out by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 9, and then causing the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc. to operate.
[0033] The program describing this process can be recorded on a computer-readable recording medium. Any computer-readable recording medium can be used, such as a magnetic recording device, optical disc, magneto-optical recording medium, or semiconductor memory.
[0034] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0035] A computer executing such a program may, for example, first store the program recorded on a portable storage medium or a program transferred from a server computer in its own storage device. Then, when processing is to be executed, the computer reads the program stored on its own storage medium and executes the processing according to the read program. Alternatively, the computer may directly read the program from the portable storage medium and execute the processing according to that program, or it may sequentially execute the processing according to the received program each time a program is transferred to it from a server computer. Furthermore, the processing may be executed by a so-called ASP (Application Service Provider) type service, where the processing function is realized only by execution instructions and result acquisition, without transferring the program from the server computer to this computer. Furthermore, the processing may be executed using a so-called SaaS (Software as a Service) type service, where a part of the server computer is made available to the user along with the program. In this form, the program includes information used for processing by an electronic computer that is equivalent to a program (data that is not a direct instruction to the computer but has the property of defining the computer's processing).
[0036] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware. [Explanation of Symbols]
[0037] 100 Head Transfer Function Generator 110 Three-Dimensional Data Acquisition Unit 120 Data Shaping Section 121 Three-Dimensional Head Data 130 Condition setting unit 135 Modeling means 140 Simulation Unit 150 Extraction Unit 160 Notch frequency range calculation unit 170 HRTF generation unit 190 Recording unit 200 Content device 210 Content input section 220 Sound source direction extraction section 230 Audio content generation unit 240 Playback unit
Claims
1. A head-transfer function generator that generates a user's head-transfer function, A modeling means for performing modeling at a predetermined frequency or frequency band, An HRTF generation unit that determines HRTF based on the aforementioned model. Equipped with, The predetermined frequency or frequency band is a frequency or frequency band in which a change in sound within that frequency range affects the user's perception of spatial sound. A head-related transfer function generation device characterized by the following features.
2. A condition setting unit for setting the observation points and frequencies or frequency bands to be simulated, A simulation unit that uses three-dimensional head data combining the shape of the user's ear and a dummy head to calculate the HRTF at a set frequency or frequency band at a set observation point and outputs the simulation results. An extraction unit that extracts peaks and notches from the simulation results, A notch frequency range calculation unit that determines the frequency range in which a notch exists from the simulation results, An HRTF generation unit that calculates HRTF based on simulation results and Equipped with, In its initial state, the condition setting unit sets the pre-selected observation points and the entire frequency band of the analysis target as initial settings. The condition setting unit, when simulation results exist in the initial settings, sets an observation point selected from among the pre-selected observation points, a frequency based on the peak extracted by the extraction unit based on the simulation results in the initial settings, and a frequency band based on the frequency range in which the notch exists, as determined by the notch frequency range calculation unit. A head-related transfer function generation device characterized by the following features.
3. A head-of-transfer function generation device according to claim 2, The initial settings of the condition setting unit include a plurality of observation points selected in advance on the horizontal plane, and a plurality of observation points in the upward direction from one of the selected observation points on the horizontal plane. A head-related transfer function generation device characterized by the following features.
4. A head transfer function generation device according to claim 3, The notch frequency range calculation unit predicts the change in notch frequency dependent on the upward angle from the notches extracted from one selected observation point on the horizontal plane front and multiple observation points in the upward direction, and determines the frequency range in which the notches exist. A head-related transfer function generation device characterized by the following features.
5. A head-of-transfer function generation device according to claim 4, The condition setting unit, when simulation results exist in the initial settings, sets all observation points other than the pre-selected observation points, multiple frequencies based on the peaks at the multiple selected observation points in front of the horizontal plane, and a frequency band based on the frequency range in which the notch found by the notch frequency range calculation unit exists. The HRTF generation unit determines the HRTF based on simulation results of multiple frequencies based on peaks and frequency bands based on the range of frequencies in which notches exist. A head-related transfer function generation device characterized by the following features.
6. A head-of-transfer function generation device according to claim 2, The pre-selected observation points in the initial settings of the condition setting unit are selected based on the direction of the sound source of the content, whose direction is specified for each sound source. Observation points selected from locations other than the aforementioned pre-selected observation points are observation points necessary for generating audio content from the aforementioned content. A head-related transfer function generation device characterized by the following features.
7. Modeling is performed in a specified frequency domain, The head-transfer function is determined based on the aforementioned model. A method for generating a head transfer function for a user, The predetermined frequency range is a frequency range in which changes in sound within that frequency range affect the user's perception of spatial sound. A method for generating head transfer functions characterized by the following features.
8. A head-transfer function generation program for causing a computer to function as a head-transfer function generation device according to any one of claims 1 to 6.